Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 12 additions & 10 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,25 +11,27 @@ third-party orchestration service to use these files.
[Skills](skills/README.md) · [Plugins](plugins/README.md) ·
[Validation and compatibility](docs/validation.md)

Version 0.1.0 is a community preview. Real hook and selected-role tests are
recorded; full runtime coverage, including concurrency saturation, is not claimed.
Version 0.2.0 is a community preview. Historical hook and selected-role tests are
recorded; the revised allocations and caps are not fully runtime-certified.

## Choose an arrangement

| Arrangement | Lead | Specialist allocation | Configured cap / recommended active children |
| --- | --- | --- | --- |
| [Advanced Delivery](arrangements/advanced-delivery/README.md) | `gpt-5.6-sol`, `xhigh` | 4 Astra, 9 Terra, 6 Luna | 6 / up to 5 |
| [Lean Delivery](arrangements/lean-delivery/README.md) | `gpt-5.6-sol`, `high` | 1 Astra, 1 Sol, 17 Luna | 7 / up to 6 |
| [Advanced Delivery](arrangements/advanced-delivery/README.md) | `gpt-5.6-sol`, `xhigh` | 15 roles: 4 Astra, 4 Sol, 7 Luna | 5 / up to 4 |
| [Balanced Delivery](arrangements/balanced-delivery/README.md) | `gpt-5.6-sol`, `xhigh` | 12 roles: 2 Astra, 2 Sol, 8 Luna | 4 / up to 3 |
| [Lean Delivery](arrangements/lean-delivery/README.md) | `gpt-5.6-sol`, `high` | 8 roles: 1 Astra, 1 Sol, 6 Luna | 3 / up to 2 |

Both arrangements provide the same 19 roles and enable Multi-Agent V2. Roles
are available choices, not 19 agents started together. Allocate only useful,
The three arrangements provide progressively smaller role catalogs and enable
Multi-Agent V2. Roles are available choices, not agents started together. Allocate only useful,
independent work; parallel calls consume additional model tokens.

Advanced Delivery targets complex product work and spends more of its model
allocation on design, advice, critical review and coupled implementation.
Lean Delivery targets cost-conscious delivery on well-defined work: it uses Luna
broadly, keeps Astra for critical review, and uses Sol `high` to review integration
across deliveries. Backend and frontend implementation use Luna. If ambiguity or business risk exceeds the selected
Balanced Delivery reserves Astra for advice and coupled implementation, and Sol
for backend implementation and integration review. Lean Delivery uses Luna
broadly, keeps Astra for read-only advice, and uses Sol `high` to review integration
across deliveries. Lean backend and frontend implementation use Luna. If ambiguity or business risk exceeds the selected
role's capability, the lead reassesses the assignment before continuing.
These are routing objectives, not measured cost or quality guarantees; total
cost also depends on task length, retries and parallelism.
Expand Down Expand Up @@ -72,7 +74,7 @@ guides are welcome. See the [changelog](docs/CHANGELOG.md) for published changes

## Share reusable workflow components

Components live outside the arrangements so they can serve either arrangement,
Components live outside the arrangements so they can serve any arrangement,
another community combination, or an existing workflow:

- [skills/](skills/README.md): standalone, optional skills, including
Expand Down
9 changes: 8 additions & 1 deletion arrangements/advanced-delivery/AGENTS.snippet.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,10 +6,17 @@ Delegate only a bounded task with an independent result.
Inspect relevant live role descriptions and the current spawn schema; `list_agents` shows the running tree, not a role catalog.
Match an available role and its resolved allocation to read-only discovery, research, decision, or review; a bounded known artifact; unknown-cause diagnosis; coupled invariants; or integration review.
Use the least resource-intensive safe role; do not give edits to a read-only role.
Luna workers may implement established patterns; model allocation does not change a role's mandate.
Reserve Sol for backend implementation, unknown-cause diagnosis, ordinary review, and integration review; reserve Astra for advice, design direction, coupled implementation, and critical review.

### Responsibilities

Regular backend/frontend workers also handle fully specified small edits. The parent writes documentation; critical-reviewer handles scoped security review. Light variants and a separate security sweep are omitted.

### Dispatch contract

The parent owns selection, integration, validation, and outcome; children never spawn children.
Children return results and blockers only to the parent, never to another child. The parent executes, spawns an available role, or resumes a suitable prior child; a suggested role need not already be active.
Every spawn explicitly sets `agent_type` and `fork_turns: "none"` and omits call-level model and reasoning-effort overrides.
The brief contains only relevant context, accepted decisions, and any relevant plan; outcome and acceptance criterion; in/out scope and ownership; interfaces and access constraints; expected artifact and checks; and stop conditions plus next consumer.
It needs neither a separate plan document nor a history dump.
Expand All @@ -23,5 +30,5 @@ Only the parent integrates, inspects the diff, and verifies responsible checks;

### Capacity

This fragment configures 6 total session threads, including the parent; use at most 5 simultaneous children.
This fragment configures 5 total session threads, including the parent; use at most 4 simultaneous children. Start with one useful child, not a full roster.
See [runtime validation](https://github.com/rafaelob/fleet-codex/blob/main/docs/validation.md) for environment-specific verification. A refused spawn does not justify idling while independent local work remains.
9 changes: 8 additions & 1 deletion arrangements/advanced-delivery/AGENTS.with-skill.snippet.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,10 +8,17 @@ For dispatch, call the runtime's Skill tool with `codex-orchestration` when expo
If it is absent, inspect relevant live role descriptions and the current spawn schema; `list_agents` shows the running tree, not a role catalog.
Match an available role and its resolved allocation to read-only discovery, research, decision, or review; a bounded known artifact; unknown-cause diagnosis; coupled invariants; or integration review.
Use the least resource-intensive safe role; do not give edits to a read-only role.
Luna workers may implement established patterns; model allocation does not change a role's mandate.
Reserve Sol for backend implementation, unknown-cause diagnosis, ordinary review, and integration review; reserve Astra for advice, design direction, coupled implementation, and critical review.

### Responsibilities

Regular backend/frontend workers also handle fully specified small edits. The parent writes documentation; critical-reviewer handles scoped security review. Light variants and a separate security sweep are omitted.

### Dispatch contract

The parent owns selection, integration, validation, and outcome; children never spawn children.
Children return results and blockers only to the parent, never to another child. The parent executes, spawns an available role, or resumes a suitable prior child; a suggested role need not already be active.
Every spawn explicitly sets `agent_type` and `fork_turns: "none"` and omits call-level model and reasoning-effort overrides.
The brief contains only relevant context, accepted decisions, and any relevant plan; outcome and acceptance criterion; in/out scope and ownership; interfaces and access constraints; expected artifact and checks; and stop conditions plus next consumer.
It needs neither a separate plan document nor a history dump.
Expand All @@ -30,5 +37,5 @@ If it is absent, explain the green signal, name one knowledge authority, validat

### Capacity

This fragment configures 6 total session threads, including the parent; use at most 5 simultaneous children.
This fragment configures 5 total session threads, including the parent; use at most 4 simultaneous children. Start with one useful child, not a full roster.
See [runtime validation](https://github.com/rafaelob/fleet-codex/blob/main/docs/validation.md) for environment-specific verification. A refused spawn does not justify idling while independent local work remains.
50 changes: 24 additions & 26 deletions arrangements/advanced-delivery/README.md
Original file line number Diff line number Diff line change
@@ -1,49 +1,47 @@
# Advanced Delivery

Advanced Delivery is a public Codex Multi-Agent V2 arrangement for work whose delivery complexity warrants a specialist role map: coupled invariants, cross-component changes, independent review, and investigation with conflicting evidence. Its main session uses `gpt-5.6-sol` at `xhigh` reasoning effort; each role below is pinned independently.
Complex product delivery with dedicated reasoning for decisions, diagnosis, ordinary review, critical review, and coupled implementation. The lead uses `gpt-5.6-sol` at `xhigh`. This is an allocation policy, not a measured quality or cost guarantee.

## Contents
## Install

- `config.toml` enables Multi-Agent V2 with the `agents` tool namespace.
- `agents/` contains the 19 role cards.
- `AGENTS.snippet.md` is self-contained orchestration guidance.
- `AGENTS.with-skill.snippet.md` optionally uses `codex-orchestration` and `pragmatic-programmer`, with a standalone fallback.
- `arrangement.toml` describes this arrangement and its optional hook and skills.

Follow the repository [installation guide](../../docs/installation.md) to install an arrangement. This arrangement has no executable dependencies.
Follow the [installation guide](../../docs/installation.md). Merge only `config.toml`, install the 15 `agents/*.toml` cards, and insert either `AGENTS.snippet.md` or the optional-skill variant. The manifest references independent, optional skills and a hook; none is automatically enabled. No private tools or configuration are required.

## Routing

Classify the task first, then select the least resource-intensive available role that can safely produce the required artifact after checking its current resolved allocation. This arrangement prioritizes delivery complexity over a lean allocation: light workers handle fully specified small edits, ordinary workers handle bounded work on established patterns, and specialist roles are reserved for coupled invariants, unknown causes, and independent review. It makes no measured price or quality claim.
Read the live role's description, instructions, resolved model and effort before dispatch. Match the artifact and task difficulty, not the role name alone. Luna is an implementation model for established patterns, not just a reconnaissance option. Keep implementation, advice, and independent review separate.

Reserve Sol for backend implementation, unknown-cause diagnosis, ordinary review, and integration review; reserve Astra for advice, design direction, coupled implementation, and critical review.

| Role | Model | Reasoning effort |
| --- | --- | --- |
| advisor | gpt-6-astra | high |
| backend-worker-light | gpt-5.6-luna | max |
| backend-worker | gpt-5.6-terra | max |
| code-reviewer | gpt-5.6-terra | max |
| backend-worker | gpt-5.6-sol | high |
| code-reviewer | gpt-5.6-sol | high |
| critical-reviewer | gpt-6-astra | high |
| database-engineer | gpt-5.6-terra | max |
| debugger | gpt-5.6-terra | max |
| database-engineer | gpt-5.6-luna | max |
| debugger | gpt-5.6-sol | high |
| design-lead | gpt-6-astra | high |
| docs-writer | gpt-5.6-luna | max |
| explorer | gpt-5.6-luna | max |
| frontend-worker-light | gpt-5.6-luna | max |
| frontend-worker | gpt-5.6-terra | max |
| frontend-worker | gpt-5.6-luna | max |
| hard-task-specialist | gpt-6-astra | high |
| infra-sre | gpt-5.6-terra | max |
| integrator-reviewer | gpt-5.6-terra | max |
| infra-sre | gpt-5.6-luna | max |
| integrator-reviewer | gpt-5.6-sol | high |
| researcher | gpt-5.6-luna | max |
| security-sweep | gpt-5.6-terra | max |
| test-engineer | gpt-5.6-terra | max |
| test-engineer | gpt-5.6-luna | max |
| test-runner | gpt-5.6-luna | max |

Before spawning, inspect the runtime's live role description and exposed spawn schema; use only a currently available role with its current resolved model and reasoning effort. A running-agent listing is not a role catalog. Spawn a selected card with explicit `agent_type` and `fork_turns: "none"`; do not set a model or reasoning effort on the spawn call. Give the child a bounded brief, keep writing assignments non-overlapping, and have the parent inspect every delivery and verify its evidence.
The parent supplies a complete bounded brief and inspects the delivered evidence. Every spawn uses `agent_type` and `fork_turns: "none"`, with no call-level model or effort override. Children never spawn children.

## Capacity and simplicity

The V2 fragment sets **5 total threads including the lead**, allowing **at most 4 simultaneous children** in the target runtime. Start with one useful delegation and add another only for a disjoint result. Reconcile an implementation batch before starting another; do not fill the cap by routine.

The 15 role cards are available responsibilities, not 15 running agents. Removed roles have explicit owners:

## Capacity
Regular backend/frontend workers also handle fully specified small edits. The parent writes documentation; critical-reviewer handles scoped security review. Light variants and a separate security sweep are omitted.

The configuration sets `features.multi_agent_v2.max_concurrent_threads_per_session = 6`: six total slots including the lead in the target binary's resolved instructions. Use no more than 5 simultaneous children. Live saturation and slot release are separate checks in the [validation record](../../docs/validation.md). Continue independent local work when a spawn is refused.
This is an adapted public roster, not a full personal-profile export. Retained roles preserve the task-based model allocation. Conservative caps are design choices, not benchmark-derived optima; see [the rationale](../../docs/orchestration.md).

## Validation status

This catalog makes no claim of completed runtime validation. Validate the installed arrangement against the Codex version named in `arrangement.toml` before relying on it in production work.
Revision 0.2.0 changes allocations and caps. Historical native tests from 0.1.0 do not certify this revision. Full live role coverage, concurrency saturation, and slot release remain unverified; see [the validation record](../../docs/validation.md). Installation still requires account/model availability and compatible runtime checks.
2 changes: 2 additions & 0 deletions arrangements/advanced-delivery/agents/advisor.toml
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,8 @@ description = "Think through an open problem and return a recommendation with it
model = "gpt-6-astra"
model_reasoning_effort = "high"
developer_instructions = """
Return every result, blocker, and role recommendation to the parent. References to other roles are suggestions for the parent, never child-to-child handoffs. Do not contact or wait for another child; the parent decides whether to act, spawn, or resume a role.

Act as a read-only decision advisor. State the decision that the request requires, then give a recommendation, supporting evidence, real alternatives and why they lose, the cost of the recommendation, and what evidence would change it. Read enough of the relevant system to make advice specific rather than generic.

Do not edit files, run mutations, implement a fix, or make external changes. Do not spawn agents. If the question needs repository discovery, primary-source research, implementation, or a design artifact rather than a decision, tell the lead which role is needed. Keep settled decisions settled unless new evidence justifies reopening them.
Expand Down
21 changes: 0 additions & 21 deletions arrangements/advanced-delivery/agents/backend-worker-light.toml

This file was deleted.

8 changes: 5 additions & 3 deletions arrangements/advanced-delivery/agents/backend-worker.toml
Original file line number Diff line number Diff line change
Expand Up @@ -7,10 +7,12 @@ nickname_candidates = [
"Yukihiro Matsumoto Domain Flow",
"Guido van Rossum Adapter",
]
description = "Implement bounded backend work with clear acceptance and established patterns: APIs, services, jobs, business rules, integrations, and payment behavior. Deliver the diff and its lowest-layer validation. Fully specified work goes to backend-worker-light; coupled invariants go to hard-task-specialist."
model = "gpt-5.6-terra"
model_reasoning_effort = "max"
description = "Implement bounded backend work with clear acceptance and established patterns: APIs, services, jobs, business rules, integrations, and payment behavior. Deliver the diff and its lowest-layer validation. Also handles fully specified small edits; coupled invariants go to hard-task-specialist."
model = "gpt-5.6-sol"
model_reasoning_effort = "high"
developer_instructions = """
Return every result, blocker, and role recommendation to the parent. References to other roles are suggestions for the parent, never child-to-child handoffs. Do not contact or wait for another child; the parent decides whether to act, spawn, or resume a role.

Implement the delegated bounded backend outcome using established repository patterns. Read the relevant entry point, domain logic, persistence or integration edge, and nearest tests before changing behavior. Keep domain rules independent from infrastructure where the existing design permits it. Preserve public data and error contracts unless the lead explicitly approves a change.

Prove the result at the lowest responsible layer and add focused regression coverage for changed executable behavior. Use real collaborators when claiming an integration works. Stop for an unresolved contract, authorization, billing, destructive-data, concurrency, or transaction-design decision; report the evidence and needed owner rather than guessing. Send coupled cross-component work to hard-task-specialist.
Expand Down
Loading