diff --git a/.claude-plugin/marketplace.json b/.claude-plugin/marketplace.json index c9f77ae..f903788 100644 --- a/.claude-plugin/marketplace.json +++ b/.claude-plugin/marketplace.json @@ -1,7 +1,7 @@ { "$schema": "https://anthropic.com/claude-code/marketplace.schema.json", "name": "bro-code", - "description": "droniu's Claude Code toolkit — the bro persona for your pair programmer, plus a multi-model council, end-to-end ship workflow, and CodeRabbit triage", + "description": "droniu's Claude Code toolkit — the bro persona for your pair programmer, plus a multi-model council, single-model consults, end-to-end ship workflow, and CodeRabbit triage", "owner": { "name": "Droniu", "email": "droniu@droniu.dev" @@ -12,6 +12,11 @@ "description": "Multi-model council: fan a question or diff out to Codex, Grok, and a sandboxed Claude seat, run an anonymized rebuttal round, synthesize with disagreements surfaced", "source": "./plugins/council" }, + { + "name": "consult", + "description": "Consult one other model — Codex, Grok, or a Claude model — on the issue at hand, with the same access as the current session", + "source": "./plugins/consult" + }, { "name": "ship", "description": "Ship current work end-to-end: branch off the correct base, run the project's verify gate, commit per repo conventions, push, open a PR", diff --git a/.github/workflows/validate.yml b/.github/workflows/validate.yml index 3258fba..f20011c 100644 --- a/.github/workflows/validate.yml +++ b/.github/workflows/validate.yml @@ -21,14 +21,36 @@ jobs: - name: Validate marketplace and plugin manifests run: | claude plugin validate .claude-plugin/marketplace.json --strict - for p in council ship coderabbit-triage bro-mode; do - claude plugin validate "plugins/$p" --strict - done - - name: Portability scan — no machine-specific facts in published plugins + # Validation targets come from the marketplace itself: a plugin can + # never be registered and then silently left unvalidated. Marketplace + # --strict passes over a missing source dir and never opens skills, + # so each plugin directory is validated on its own. + jq -r '.plugins[].source' .claude-plugin/marketplace.json \ + | sed 's#^\./##;s#/$##' | sort > /tmp/listed.txt + while read -r src; do + if [ ! -d "$src" ]; then + echo "marketplace.json lists '$src' but that directory does not exist" >&2 + exit 1 + fi + claude plugin validate "$src" --strict + done < /tmp/listed.txt + + # ...and the reverse: a plugin directory nobody registered. + ls -d plugins/*/ | sed 's#/$##' | sort > /tmp/present.txt + if ! diff -u /tmp/listed.txt /tmp/present.txt; then + echo "plugins/ and marketplace.json disagree (-marketplace +on disk)" >&2 + exit 1 + fi + + - name: Portability scan — no machine-specific facts anywhere in the repo run: | - if grep -rEn '/Users/|/home/[a-z]|[A-Za-z]:\\\\Users|linear\.app/[a-z]|atlassian\.net' plugins/; then - echo "Machine- or workspace-specific facts found in published plugin content" >&2 + # Scans the whole repo, not just plugins/: README, CHANGELOG and docs + # leak machine paths just as easily. This workflow is excluded because + # it necessarily contains the patterns it searches for. + if grep -rEn '/Users/|/home/[a-z]|[A-Za-z]:\\Users|/private/tmp|claude-501|linear\.app/[a-z]|atlassian\.net' \ + --exclude-dir=.git --exclude=validate.yml .; then + echo "Machine- or workspace-specific facts found in repo content" >&2 exit 1 fi diff --git a/CHANGELOG.md b/CHANGELOG.md index e6e7d92..5418d93 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -6,6 +6,50 @@ conventions. Versions are per-plugin semver from each plugin's `plugin.json`. ## council +### 0.2.0 — 2026-09-12 + +#### Added + +- Per-seat preflight: a one-turn smoke call decides the roster, so a broken + codex install or an exhausted grok quota is caught before three seats get + announced and two get retracted mid-run. +- Nickname resolution for seat selection ("just fable and astra"), with an + alias table mapping nicknames to provider and model. +- Seat budgets: turn caps, a codex wall-clock timeout, and a watchdog on the + Claude seat, so one runaway seat cannot swallow a run. +- Wall-clock expectations up front, plus a spend line drawn from the provider + envelopes. +- An evidence gate: a seat whose answer carries no `file:line` citation is + reported degraded instead of counted as a position. +- A pointer to `consult` for the single-model case. + +#### Fixed + +- Codex seats no longer use `codex exec review`, which rejects `-C`, refuses + `--base` alongside a prompt, and ignores `--output-schema` (it returns + prose). The diff target moved into the brief. +- `find-schema.json` now lists every property in `required`, as OpenAI strict + structured output demands; the old schema was rejected outright. +- Grok failures are diagnosed as quota (`429`, `free-usage-exhausted`, + `reauthable: false`) instead of advising a pointless `grok login`, and grok + is always invoked by name so a shell wrapper's API key and `GROK_HOME` + still apply. +- Dropped `--permission-mode plan` from the Grok seat: the flag only applies + `default` and `bypassPermissions`, so `plan` was accepted and silently + ignored. `--sandbox read-only` was and remains the actual enforcement. +- Corrected the Grok model and effort claims: `grok models` lists the catalog + and `-m` pins a seat model (new `--grok-model` flag), and + `--reasoning-effort` is silently ignored when the model's catalog entry says + `supports_reasoning_effort: false`, so the roster now reports "effort: n/a" + instead of a level that was never applied. +- Grok auth detection no longer probes a hardcoded `~/.grok/auth.json`, which + missed wrapper setups that relocate `GROK_HOME`. +- The Claude seat's brief is passed as literal text and asserted non-empty, + since workflow args can arrive stringified or empty; seats also write results + to disk so notification truncation is off the critical path. +- Documented that a headless Grok seat ends its whole run + (`stopReason: "Cancelled"`) at the first command needing approval. + ### 0.1.0 — 2026-07-19 #### Added @@ -15,6 +59,18 @@ conventions. Versions are per-plugin semver from each plugin's `plugin.json`. seat separation, capability discovery, full repo read access for every seat (enforced read-only), degraded-roster reporting. +## consult + +### 0.1.0 — 2026-09-10 + +#### Added + +- Initial release: one-on-one consultation with a single named model — + Codex, Grok, or a Claude subagent. The consultant's access mirrors the + parent session's permission mode, it gets a self-contained brief, changes + it makes to the working tree are reported, and follow-ups resume the same + consultant session. + ## ship ### 0.1.0 — 2026-07-19 diff --git a/README.md b/README.md index 5430c78..8ce6a1f 100644 --- a/README.md +++ b/README.md @@ -6,7 +6,7 @@ Claude Code plugins by [Droniu](https://github.com/Droniu). The headliner is **bro-mode** — an output style that makes your AI pair programmer talk like -your best bro. Three workflow plugins ride along, built and dogfooded daily. +your best bro. Four workflow plugins ride along, built and dogfooded daily. > **Alpha honesty:** everything here is `0.x`. Interfaces, flags, and skill > names may change between minor versions. Pin nothing to muscle memory yet. @@ -104,6 +104,9 @@ What the side-by-side "ask three models" tools don't do, this does: - **Refutation by default** — in review mode, findings go to a *different* model instructed to refute them, with `refuted: true` as the default, because AI review's dominant failure mode is confident false positives. +- **Preflighted seats** — every provider gets a one-turn smoke call before the + roster is announced, so a broken CLI or an exhausted quota is caught up + front instead of halfway through a long run. - **Honest degradation** — no provider CLIs installed? It says "degraded council, single model family" instead of pretending. @@ -119,6 +122,31 @@ accounts. Every council seat has full read access to your repository and sends what it reads to its provider — read the skill's data-exposure section before pointing it at anything sensitive. +### consult — one model, your session's access + +When you want one specific model's take instead of a whole council: + +```text +/plugin install consult@bro-code +/consult gpt-6 astra (codex) on this issue +/consult grok on why the build broke — let it try fixing it +/consult sonnet on whether SSE or websockets fits here +``` + +The consultant — Codex, Grok, or a Claude subagent — gets a self-contained +brief (it never saw your conversation) and the **same access as your +session**: same repo, its own full config and MCP servers, network, and your +current permission mode mapped onto its CLI. Plan mode stays read-only, auto +mode auto-approves inside a workspace sandbox, bypass stays bypass. Where your +session would ask you first, a headless consultant is denied rather than +silently upgraded. + +Unlike council's read-only seats, a consultant in a write-capable mode *can* +change files; its brief says not to unless you asked ("let it try fixing +it"), and anything it changes in the working tree is reported back. +Follow-ups resume the same consultant session. Same data-exposure rule as +council: Codex and Grok send what they read to their providers. + ### ship End-to-end "take my working state to an open PR": classifies your branch diff --git a/plugins/consult/.claude-plugin/plugin.json b/plugins/consult/.claude-plugin/plugin.json new file mode 100644 index 0000000..bfb93c0 --- /dev/null +++ b/plugins/consult/.claude-plugin/plugin.json @@ -0,0 +1,11 @@ +{ + "name": "consult", + "version": "0.1.0", + "description": "Consult one other model — Codex, Grok, or a Claude model — on the issue at hand, with the same access as the current session", + "author": { + "name": "Droniu", + "email": "droniu@droniu.dev" + }, + "license": "MIT", + "homepage": "https://github.com/Droniu/bro-code" +} diff --git a/plugins/consult/skills/consult/SKILL.md b/plugins/consult/skills/consult/SKILL.md new file mode 100644 index 0000000..d6dc2f6 --- /dev/null +++ b/plugins/consult/skills/consult/SKILL.md @@ -0,0 +1,168 @@ +--- +name: consult +description: Consult one other model — Codex (GPT), Grok, or a Claude model — on the issue at hand, with the same access as the current session. Use when the user types /consult, or asks for one named model's take on the current problem ("ask gpt-6 about this", "what does grok think", "have sonnet look at it"). For a multi-model fan-out with rebuttal, use council instead. +argument-hint: " [(codex|grok|claude)] [on ] [--effort ] [--mode ]" +--- + +# Consult + +One outside model, one issue, **the same access as this session**. You stay the lead: the consultant investigates and answers, you relay that answer faithfully, check it against the code, and keep driving the task. + +## Usage + +```text +/consult gpt-6 astra (codex) on this issue +/consult grok on why the build broke — let it try fixing it +/consult sonnet on whether SSE or websockets fits here +``` + +Flags: `--effort ` (default: this session's effort) · `--mode ` (default: this session's permission mode). + +Follow-ups — "ask it why", "tell codex the test passes now" — go to the same consultant session (Step 5). + +## Step 1 — Resolve the consultant + +| Signal in the arguments | Provider | +|---|---| +| `(codex)` / `(grok)` / `(claude)` | that provider — wins over everything else | +| `gpt-*`, `o*`, `codex` | codex | +| `grok*` | grok | +| `claude`, `opus`, `sonnet`, `haiku`, `fable` | claude | +| no model or provider named | ask the user, offering installed providers only | + +Turn spoken names into slugs (`gpt-6 astra` → `gpt-6-astra`) and **check the slug against the provider's own catalog**: + +- **codex** — slugs in `~/.codex/models_cache.json` (`.models[].slug`, with `supported_reasoning_levels` per model). No model named → the `model` key in `~/.codex/config.toml`, else the CLI default. +- **grok** — `grok models`. No model named → its listed default. +- **claude** — the Agent tool takes `opus | sonnet | haiku | fable`. No model named → `opus`. A pinned version ("opus 4.8") can't be guaranteed through an alias — say so. + +An unknown slug gets the catalog shown and a question; a CLI rejecting a model gets reported. Neither is a cue to drop `-m` and quietly run the default model instead. + +`command -v codex` / `command -v grok` must succeed. Invoke them by name, never by a resolved binary path — `grok` in particular may be a shell function that sets its home and credentials. Auth failures surface on the first call — report them; never swap in a different provider. + +## Step 2 — Resolve the session's access + +The consultant runs in the same working directory, loads its own full user config (instructions, MCP servers, plugins, rules — Grok additionally reads Claude's `settings.json`, `CLAUDE.md`, hooks, and plugins), has network, and gets **this session's permission mode** mapped onto the provider. + +**Mode** — first match wins; name the source in the roster: + +1. `--mode` flag +2. The latest system reminder says plan mode is active → `plan` +3. `permissions.defaultMode` in managed settings (`/Library/Application Support/ClaudeCode/managed-settings.json` on macOS, `/etc/claude-code/managed-settings.json` on Linux) +4. A flag on this session's process: `ps -o command= -p "$CLAUDE_PID"` (`--permission-mode `; `--dangerously-skip-permissions` → `bypassPermissions`) +5. `permissions.defaultMode` in `.claude/settings.local.json`, then `.claude/settings.json`, then `~/.claude/settings.json` +6. `default` + +A mid-session Shift+Tab switch (other than into plan mode) is invisible to all of these, which is why the roster states the mode and where it came from. + +**Effort** — `--effort`, else `$CLAUDE_EFFORT`, else `high`. Clamp to the provider and note any clamp: + +- **codex** — the model's `supported_reasoning_levels`. +- **grok** — read `models_cache.json` next to the `user` layer path in `grok inspect --json` → `.configSources.layers` (it follows wrappers that set `GROK_HOME`; default `~/.grok`). When `.models[].info.supports_reasoning_effort` is false the flag is silently ignored: drop it and put "effort: n/a" in the roster; otherwise clamp to `.info.reasoning_efforts[].value`. + +| Mode | codex flags | grok flags | claude | +|---|---|---|---| +| `plan` | `-s read-only` | `--permission-mode default --sandbox read-only` | inherits | +| `default`, `dontAsk` | `-s read-only` | `--permission-mode default` | inherits | +| `acceptEdits` | `-s workspace-write` | `--permission-mode default --allow Edit --allow Write --sandbox workspace` | inherits | +| `auto` | `--approve-for-me -c sandbox_workspace_write.network_access=true` | `--permission-mode bypassPermissions --sandbox workspace` | inherits | +| `bypassPermissions` | `--dangerously-bypass-approvals-and-sandbox` | `--permission-mode bypassPermissions` | inherits | + +Verified 2026-09-10 against codex-cli 0.153.2, grok 0.2.103, Claude Code 2.1.267: + +- A headless consultant has nobody to ask. Where this session would prompt the user, the consultant is denied — never silently upgraded. Codex can't split edits from commands, so `acceptEdits` confines both to the workspace with no network. +- Codex `workspace-write` blocks network unless `network_access=true`. +- Grok's `--permission-mode` flag only applies `default` and `bypassPermissions`; `plan`, `acceptEdits`, and `auto` are accepted and ignored, so those rows are rebuilt from allow rules plus a kernel sandbox profile. +- Grok headless **ends the whole run** when a call needs approval (`stopReason: "Cancelled"`) — observed, although grok's docs say the call is reported back to the model. Report "stopped at the permission wall", offer a `--mode`, don't retry silently. +- Grok runs Claude's hooks but feeds them its own payload (`toolName`, `toolInput`). A guard hook that reads Claude's `tool_name` / `tool_input` sees nothing and allows the call, since hooks fail open — don't count on it for a grok consultant that can write. +- Claude consultants are `general-purpose` Agent-tool subagents. With no `permissionMode` of their own they inherit this session's mode in every row, along with its MCP tools, permission rules, and hooks. In the background they keep the working built-ins but lose interactive ones like AskUserQuestion; their permission prompts surface in this session. + +**The mapping is the access — never narrow it.** The reflex is to lock a consultant down; `/council` does that on purpose, this skill does not. + +| Tempting move | Why it's wrong here | +|---|---| +| read-only sandbox "since it's only advice" | Advice vs. changes is the brief's mandate (Step 3), not a sandbox | +| `--ignore-user-config`, banning MCP/web tools, `Plan`/`Explore` agent type | Strips tools this session has; the consultant can't reproduce, run, or look anything up | +| no network "for safety" | This session has network | +| `fork` subagent for a Claude consultant | Ignores the model override and inherits your anchoring — not a second opinion | + +## Step 3 — Write the brief + +The consultant has seen none of this conversation. Write `$SP/brief.md` with these parts, in order — `$SP` is a fresh directory per consultation, `/consult/-`, so a later consult never overwrites an earlier one's session id. In plan mode, where you can't write files, pass the same text inline: a quoted heredoc on codex's stdin (`- <<'BRIEF'`), `-p "$(cat <<'BRIEF' … BRIEF)"` for grok, the Agent `prompt` for claude. + +1. **Role** — "You are being consulted by another coding agent (Claude Code) working with a user. Investigate independently; the hypotheses below are unverified — try to refute them. Respond in a neutral technical register." (Provider configs can set a persona.) +2. **The ask** — the user's words, verbatim, plus one line saying what "this" refers to when they point at the conversation. +3. **Context** — the problem; exact errors and output, verbatim; what was tried and what happened; the current hypothesis, marked unverified; relevant paths; `git status --short`. +4. **Project rules the CLI won't load** — for codex, which reads `AGENTS.md` but not `CLAUDE.md`: copy the rules that matter (package manager, verify commands). Claude and Grok consultants load `CLAUDE.md` themselves. +5. **Mandate** — by default: "Investigate and advise. Run whatever helps (tests, builds, repro scripts), but leave project files as you found them." When the user said fix / implement / let it try: "You may change files in the workspace to fix this." Either way: no commits, pushes, or other outward-facing actions. +6. **Access** — one line on what the mapped access allows and what happens at its edge, so the consultant doesn't spend its run discovering it. Codex: blocked commands fail back to it. Grok in `plan`, `default`, or `acceptEdits`: a command that needs approval ends the run — name the verification commands instead of running them. +7. **Response contract** — "Reply with exactly these sections: Answer (first line is the verdict) · Evidence (file:line; commands run with key output) · Key assumption · Confidence (high/medium/low) · Changes made (paths, or none) · Open questions." + +## Step 4 — Run + +Announce the roster in one line before running: +`Consulting codex gpt-6-astra @xhigh · access: auto (~/.claude/settings.json) → approve-for-me + network · reads go to OpenAI` + +`$SP`, `$REPO`, `$MODEL`, `$EFFORT`, and `$ACCESS` below stand for concrete values — shell state doesn't survive between Bash calls, so write them into every command. + +When the mapped access can write, snapshot the working tree — tracked and untracked — into a throwaway index; your real index and files stay untouched: + +```bash +SP="/consult/-"; mkdir -p "$SP"; REPO=$(git rev-parse --show-toplevel) +cp "$(git -C "$REPO" rev-parse --path-format=absolute --git-path index)" "$SP/index" +GIT_INDEX_FILE="$SP/index" git -C "$REPO" add -A && GIT_INDEX_FILE="$SP/index" git -C "$REPO" write-tree > "$SP/tree.before" +``` + +Run with Bash `run_in_background: true` — consultations routinely outlast the foreground timeout. Wait for the completion notification; never predict the answer. Leave the working tree alone until it returns: anything you edit meanwhile shows up as the consultant's change. + +```bash +# codex — $ACCESS from the table; drop -m to use the config default. Answer lands +# in -o; session id is the first JSONL event: {"type":"thread.started","thread_id":"..."} +codex exec --json -C "$REPO" -m "$MODEL" -c model_reasoning_effort="$EFFORT" $ACCESS \ + -o "$SP/codex.md" - < "$SP/brief.md" > "$SP/codex.jsonl" 2> "$SP/codex.err" +``` + +```bash +# grok — drop -m for its default, and --reasoning-effort when effort is n/a. +# Answer at .text, session id at .sessionId; .stopReason must be "EndTurn" +grok --prompt-file "$SP/brief.md" --cwd "$REPO" -m "$MODEL" --reasoning-effort "$EFFORT" $ACCESS \ + --output-format json > "$SP/grok.json" 2> "$SP/grok.err" +``` + +```text +claude — Agent tool (subagents run in the background): + subagent_type: "general-purpose" (all tools and MCP servers, this session's mode) + model: opus | sonnet | haiku | fable + prompt: the brief +Keep the agent id it returns — follow-ups go there. +No effort parameter: roster says "effort: session default"; a --effort flag is ignored — say so. +Data stays with Anthropic. When the model matches this session's, say so: fresh context, same family. +``` + +## Step 5 — Report, verify, follow up + +1. **Failure first.** Non-zero exit, an empty answer, no `turn.completed` event in codex's JSONL, or a grok `stopReason` other than `EndTurn` → report the consultation as failed or cut short, quoting the relevant stderr lines. CLIs log noise (MCP auth errors and the like) on every run, so stderr alone isn't failure. Never write its answer for it. +2. **What it changed** — whenever you snapshotted, failed and cut-short runs included: rerun the `add -A` / `write-tree` line into `$SP/tree.after`, then `git -C "$REPO" diff --stat `. Gitignored paths (build output, generated code) aren't covered. A mismatch with its "Changes made" is itself a finding. Undoing exactly its edits, when the user asks: `git -C "$REPO" diff --binary | git -C "$REPO" apply`. +3. **Present** it attributed and intact — including where it disagrees with you: + +```text +gpt-6-astra (codex) · auto access · session + +Evidence: · Key assumption: <…> · Confidence: <…> +Changed: · Open questions: <…> +My read: +``` + +Add grok's `total_cost_usd` to the header when it's reported. Read the code its answer rests on before repeating any claim as fact. Neither the consultant nor you is the authority — the code is. Don't adopt its fix or keep its changes on its say-so; the user decides, then the project's verify gate runs. + +**Follow-ups** go to the same session, run in the background from `$REPO`, with the follow-up as a new file in `$SP` (inline in plan mode). Re-snapshot first when access can write: + +| Provider | Follow-up | +|---|---| +| codex | `codex exec $ACCESS resume --json -c model_reasoning_effort="$EFFORT" -o "$SP/codex-2.md" - < "$SP/followup.md" > "$SP/codex-2.jsonl" 2> "$SP/codex-2.err"` — repeat the first call's `$ACCESS` *before* `resume`: without it a resumed session silently drops `--approve-for-me` to approval policy `never` and loses `-c` overrides like network. Resume has no `-C`, so run it from `$REPO` | +| grok | `grok --prompt-file "$SP/followup.md" --resume --cwd "$REPO" -m "$MODEL" --reasoning-effort "$EFFORT" $ACCESS --output-format json > "$SP/grok-2.json" 2> "$SP/grok-2.err"` — the sandbox profile is fixed per session, but permission flags are read per process: repeat `$ACCESS` or the follow-up silently loses `--allow` rules and `bypassPermissions` | +| claude | `SendMessage` to the agent id the Agent call returned | + +## Cost and data exposure + +A consultation is a real session on that provider's quota; follow-ups add turns to it. Codex and Grok send what they read to OpenAI and xAI — no sandbox mode changes that. Don't consult an external provider on code you may not share with it. A first Codex run in a directory also adds a `[projects.""]` trust entry to `~/.codex/config.toml`. diff --git a/plugins/council/.claude-plugin/plugin.json b/plugins/council/.claude-plugin/plugin.json index a617c62..254590f 100644 --- a/plugins/council/.claude-plugin/plugin.json +++ b/plugins/council/.claude-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "council", - "version": "0.1.0", + "version": "0.2.0", "description": "Multi-model council: fan a question or diff out to Codex, Grok, and a sandboxed Claude seat, run an anonymized rebuttal round, synthesize with disagreements surfaced", "author": { "name": "Droniu", diff --git a/plugins/council/skills/council/SKILL.md b/plugins/council/skills/council/SKILL.md index 44f2c7d..4a34bc6 100644 --- a/plugins/council/skills/council/SKILL.md +++ b/plugins/council/skills/council/SKILL.md @@ -1,6 +1,6 @@ --- name: council -description: Convene a multi-model council for adversarial code review or architecture brainstorming. Fans a question or diff out to every available model CLI (Codex, Grok) plus a sandboxed Claude seat in parallel, runs an anonymized rebuttal or refutation round, and synthesizes with disagreements surfaced. Use when the user types /council, or asks for a second opinion, multi-model review, architecture critique, or "what would other models say". +description: Convene a multi-model council for adversarial code review or architecture brainstorming. Fans a question or diff out to every available model CLI (Codex, Grok) plus a sandboxed Claude seat in parallel, runs an anonymized rebuttal or refutation round, and synthesizes with disagreements surfaced. Use when the user types /council, or asks for a multi-model review, cross-model disagreement, or an architecture critique from several models at once. For one named model, use consult instead. --- # Council @@ -17,10 +17,22 @@ Claude Code chairs. Every model — including Claude — argues from a seat with /council review --base main review against a base branch ``` -Flags: `--only codex,claude` select seats — the chair always remains; omitting `claude` from the list drops Claude's seat too · `--model ` override the Codex model · `--claude-model ` override the Claude seat model (default `opus`) · `--effort ` override fan-out effort (clamped per provider). +Flags: `--only codex,claude` select seats — the chair always remains; omitting `claude` from the list drops Claude's seat too · `--model ` override the Codex model · `--claude-model ` override the Claude seat model (default `opus`) · `--grok-model ` pin the Grok seat model (default: the CLI's own default) · `--effort ` override fan-out effort (clamped per provider). When the Claude seat runs the **same model as the chair**, say so in the roster — a sibling instance holds no stake in the outcome, but the chair reasons the way its sibling argues; weigh cross-family agreement (codex+grok) a notch higher that run. +**Users name seats by model nickname**, not by provider — "drop grok, just fable and astra". Resolve nicknames before anything else: + +| The user says | Seat | Resolve against | +|---|---|---| +| `opus`, `sonnet`, `haiku`, `fable` | claude | the `--claude-model` value | +| `astra`, `sol`, `terra`, `luna`, `gpt-*`, `codex` | codex | `.models[].slug` in `~/.codex/models_cache.json` (`gpt-6 astra` → `gpt-6-astra`) | +| `grok*` | grok | `grok models` | + +A nickname you cannot resolve is a question for the user, never a silent substitution into somebody else's default. + +**One model is not a council — that is `/consult`.** A single named model with your session's access and no rebuttal round is the lighter tool; reach for the council when you want cross-family disagreement. + ## Prerequisites, cost, and data exposure - **Provider CLIs are all optional.** The council runs with whoever is present: [Codex CLI](https://github.com/openai/codex) (`codex`), [Grok CLI](https://docs.x.ai) (`grok`). Each must be installed and authenticated by you, on your account, at your cost. The Claude seat needs no extra install. @@ -36,29 +48,41 @@ When the Claude seat runs the **same model as the chair**, say so in the roster Discover capabilities instead of assuming them: - **Codex model**: read the `model` key from `~/.codex/config.toml` if the file exists; otherwise omit `-m` and let the CLI use its default. Report whichever applies in the roster. -- **Grok model**: the CLI offers no model flag on all accounts; confirm the served model post-hoc from `.modelUsage` in the response envelope and report it. +- **Grok model**: `grok models` prints the catalog and marks the default; pin one with `-m `. Confirm the served model post-hoc from `.modelUsage` in the response envelope and report that. - **Claude seat**: default `opus`, overridden only by `--claude-model`. **Never let the seat inherit the session model.** Workflow's `agent()` defaults to the main-loop model when `model:` is omitted — pass `model:` explicitly on every seat call, both modes. "The user is running X, so the seat should be X" is exactly the drift this rule exists to stop. -**Effort is set per stage, not per mode**: the fan-out (where quality is decided) runs one notch above rebuttal/refutation (many small judgment calls; diminishing returns). Known-accepted ranges as of the verified CLI versions below: codex `low…xhigh`, grok `low|medium|high`, claude `low…max`. If a CLI rejects an effort value, step down one notch and note the clamp in the roster. **Never leave effort unset** — unset effort means unpredictable cost and non-comparable answers. +**Effort is set per stage, not per mode**: the fan-out (where quality is decided) runs one notch above rebuttal/refutation (many small judgment calls; diminishing returns). Known-accepted ranges as of the verified CLI versions below: codex `low…ultra` (per-model `supported_reasoning_levels`), grok `none…xhigh` **only where the model supports effort at all**, claude `low…max`. Grok silently ignores `--reasoning-effort` when its catalog marks the model `supports_reasoning_effort: false` — true of API-key accounts today. Check `models_cache.json` beside the `user` layer path in `grok inspect --json` → `.configSources.layers`, and report "effort: n/a" rather than claiming a level you did not get. If a CLI rejects an effort value, step down one notch and note the clamp in the roster. **Never leave effort unset** — unset effort means unpredictable cost and non-comparable answers. | Stage | Codex | Grok | Claude | |---|---|---|---| | Fan-out (positions / findings) | `xhigh` | `high` | `xhigh` | | Rebuttal / refutation | `high` | `medium` | `high` | -Ceiling asymmetry is real (claude reaches `max`, codex `xhigh`, grok `high`): when providers disagree, some of that gap is intensity rather than judgment. Say so instead of scoring it as pure signal. +Ceiling asymmetry is real (claude reaches `max`, codex `xhigh`, grok `xhigh` where effort applies at all): when providers disagree, some of that gap is intensity rather than judgment. Say so instead of scoring it as pure signal. -## Step 1 — Detect providers +## Step 1 — Detect providers, then smoke-test them -Always run this first. The council runs with whoever is present; never fail because someone is missing. +`command -v` plus an auth file proves nothing: a CLI can be installed, authenticated, and still fail every call — a broken helper binary, an exhausted quota, a revoked token. **Preflight every seat with a real call, and let the answer decide the roster.** Two one-turn `low`-effort sessions buy you a roster that is actually true. + +Run each block as its **own** Bash call — one long chained command trips the permission splitter and stalls the whole detection step waiting for approval. ```bash -SP="/council"; mkdir -p "$SP" -echo "codex: $(command -v codex >/dev/null && echo yes || echo no) auth=$(python3 -c "import json;print(json.load(open('$HOME/.codex/auth.json'))['auth_mode'])" 2>/dev/null || echo none)" -echo "grok: $(command -v grok >/dev/null && echo yes || echo no) auth=$([ -f "$HOME/.grok/auth.json" ] && echo yes || echo no)" +SP="/council"; mkdir -p "$SP"; command -v codex; command -v grok ``` -Claude always chairs. Its seat runs unless `--only` excludes `claude`. Report the roster with model and effort before doing work. +```bash +timeout 90 codex exec --ephemeral -s read-only -c model_reasoning_effort=low \ + -o "$SP/smoke-codex.txt" < /dev/null "Reply with exactly: OK"; echo "codex exit=$?"; cat "$SP/smoke-codex.txt" +``` + +```bash +timeout 90 grok -p "Reply with exactly: OK" --output-format json --max-turns 1 > "$SP/smoke-grok.json"; echo "grok exit=$?" +python3 -c "import json;d=json.load(open('$SP/smoke-grok.json'));print(d.get('stopReason'),'|',(d.get('text') or d.get('message') or '')[:120])" +``` + +A seat that does not come back `OK` **is not on the roster**. Name it and its reason in one line and convene without it. Two seats announced honestly beat three announced and two retracted mid-run. + +Claude chairs and needs no smoke test. Its seat runs unless `--only` excludes `claude`. ## Step 2 — Sandboxing (non-negotiable) @@ -67,15 +91,15 @@ Every seat has **full repository read access**, enforced read-only: reads everyt | Provider | Enforcement | |---|---| | **codex** | `-s read-only` | -| **grok** | `--sandbox read-only` (dedicated filesystem/network sandbox profile) plus `--permission-mode plan` | +| **grok** | `--sandbox read-only` (dedicated filesystem/network sandbox profile) | -`--permission-mode plan` on Grok is a **permission gate, not a sandbox** — on its own it does not enforce read-only filesystem access. The `--sandbox read-only` profile is the enforcement; plan mode just suppresses write-tool approval churn on top of it. Never describe plan mode alone as sandboxing. +Grok's `--permission-mode` flag only applies `default` and `bypassPermissions`; `plan` is accepted and silently ignored (verified on grok 0.2.103), so it is no longer passed — it never gated anything. `--sandbox read-only` is the entire enforcement. Note also that a headless grok seat **ends its whole run** (`stopReason: "Cancelled"`) the first time it reaches for a command needing approval: report that seat as degraded rather than empty. A read-only sandbox stops the model *writing to your disk*. It does **not** stop a vendor transmitting what it reads — see the data-exposure note above. Different threat models; never conflate them. ## Parsing provider output — the shapes differ -Envelope shapes verified 2026-07-19 against codex-cli 0.144.2 and grok 0.2.103. These CLIs version independently of this skill — treat envelope drift as an expected failure mode, not a surprise. +Envelope shapes verified 2026-09-12 against codex-cli 0.153.2 and grok 0.2.103. These CLIs version independently of this skill — treat envelope drift as an expected failure mode, not a surprise. | Provider | Where the object actually is | |---|---| @@ -137,9 +161,11 @@ codex exec --ephemeral -s read-only -C "$REPO" $CODEX_MODEL_ARGS \ ``` ```bash -# grok — repo access, read-only sandbox. Real object lands at .structuredOutput -grok -p "$(cat "$SP/brief.md")" --cwd "$REPO" \ - --sandbox read-only --permission-mode plan \ +# grok — repo access, read-only sandbox. $GROK_MODEL_ARGS is "-m " when a +# model is pinned, else empty. Drop --reasoning-effort when effort is n/a. +# Real object lands at .structuredOutput +grok -p "$(cat "$SP/brief.md")" --cwd "$REPO" $GROK_MODEL_ARGS \ + --sandbox read-only \ --reasoning-effort "${GROK_EFFORT:-high}" \ --output-format json --json-schema "$(cat "$SP/arch-schema.json")" \ --max-turns 15 > "$SP/grok.json" @@ -149,10 +175,17 @@ grok -p "$(cat "$SP/brief.md")" --cwd "$REPO" \ Claude's seat — spawn via the Workflow tool when available: agent() declares both model and effort, and its schema option returns a validated object: - agent(, {model: 'opus', effort: 'xhigh', schema: }) + agent(, {model: 'opus', effort: 'xhigh', schema: }) -Inline the brief text into the script (or guard args: workflow args can arrive -JSON-stringified — use `typeof args === 'string' ? args : args.brief`). +Read the brief and inline its text when you author the script. NEVER pass it +through workflow args: they can arrive JSON-stringified or empty, and a seat +handed an empty brief burns a full run refusing to invent a position from +nothing. Assert the brief is non-empty before launching, and treat a seat that +reports an empty task as a re-launch, not as a position. + +Tell the seat to write its answer to $SP/claude.json as well as returning it. +Task notifications truncate long results, and digging the full object back out +of the workflow journal costs a round trip you do not need. Fallback when Workflow is unavailable: the Agent tool. It has NO effort parameter — flag the seat as "effort: session default (uncontrolled)" in the @@ -217,7 +250,7 @@ You are the validation layer and the chair, not a vote counter and not a competi "claim":{"type":"string"}, "failure_scenario":{"type":"string"}, "confidence":{"type":"string","enum":["high","medium","low"]} - },"required":["file","severity","claim","failure_scenario","confidence"], + },"required":["file","line","severity","category","claim","failure_scenario","confidence"], "additionalProperties":false}}}, "required":["findings"],"additionalProperties":false} ``` @@ -227,19 +260,20 @@ You are the validation layer and the chair, not a vote counter and not a competi ### 3b. Fan out — each in its native idiom ```bash -# codex — native review command (-m IS supported on the review subcommand) -codex exec review --uncommitted -C "$REPO" $CODEX_MODEL_ARGS \ +# codex — plain `exec`, never `exec review`: that subcommand rejects -C, refuses +# --base alongside a prompt, and ignores --output-schema (it answers in prose). +codex exec --ephemeral -s read-only -C "$REPO" $CODEX_MODEL_ARGS \ -c model_reasoning_effort="${EFFORT:-xhigh}" \ --output-schema "$SP/find-schema.json" -o "$SP/codex-find.json" \ - < /dev/null "Report findings matching the schema. failure_scenario must be concrete." + < "$SP/review-brief.md" ``` -Use `--base ` or `--commit ` instead of `--uncommitted` when the user specified a target. +The diff target belongs in the brief, not in flags: *"Review the uncommitted changes (`git diff HEAD`)"*, or `git diff main...HEAD` when the user named a base. Append the schema to the brief too — a seat that loses the flag still answers in shape. ```bash # grok — reads the diff itself, read-only sandbox grok -p "Review the uncommitted changes in this repository. Report findings matching the schema. failure_scenario must be concrete." \ - --cwd "$REPO" --sandbox read-only --permission-mode plan \ + --cwd "$REPO" $GROK_MODEL_ARGS --sandbox read-only \ --reasoning-effort "${GROK_EFFORT:-high}" \ --output-format json --json-schema "$(cat "$SP/find-schema.json")" \ --max-turns 20 > "$SP/grok-find.json" @@ -276,6 +310,8 @@ Rank: co-discovered + survived refutation → single-source + survived → refut You have the final say. If a finding survived every model but you can read the code and see it is wrong, kill it and explain. Three models agreeing is not evidence; the code is evidence. +**Keep the process out of the deliverable.** The roster, the degraded seats, and the retry story belong in the chat with the user — never in a PR description, commit message, or review comment. When the user does want provenance in a shipped artifact, name the models and nothing else. + ## Cost Every consultation is a real session doing real work against each provider's quota. @@ -290,8 +326,10 @@ Full roster is 3 seats — codex, grok, claude. Sessions per run: | `review --deep` | 3 + one per critical/high finding | - `--deep` is the expensive path — never make it the default -- `--max-turns` on Grok bounds runaway exploration; Codex is bounded by effort -- Announce before an expensive fan-out: *"3 seats, repo access, ~2 min, 6 sessions. Go?"* +- **Cap every seat.** Grok takes `--max-turns`; wrap codex in `timeout 900`; give the Claude seat an explicit turn budget in its brief plus a wall-clock watchdog. An uncapped seat can run for hours and return nothing, turning a three-seat council into one. +- **Say how long it will take.** A seat with repo access runs for minutes, not seconds, and a full architect run with the rebuttal round can take half an hour. Quote that up front — half an hour is a different product from "a couple of minutes". +- Announce before an expensive fan-out: *"3 seats, repo access, ~30 min, 6 sessions. Go?"* +- **Report what it cost.** Grok's envelope carries `total_cost_usd` and `usage`, codex reports tokens. Add one spend line to the synthesis for a metered run, and write "cost not reported" rather than implying it was free. - Narrow with `--only` or `--effort medium` when the question doesn't warrant full depth ## Failure handling @@ -300,7 +338,10 @@ Providers fail independently and that is fine. A dead provider is a smaller coun - Missing/unauthenticated → skip, name it in the roster - Non-JSON despite schema → parse what you can (see the defensive-parsing rule), report the provider as degraded, never fabricate its position -- Grok auth errors like `invalid_grant: Refresh token has been revoked` can be triggered by token rotation racing under parallel fan-out → tell the user to re-run `grok login` -- Timeout → report partial council; do not silently drop a member +- **Grok quota, not auth.** `429`, `subscription:free-usage-exhausted`, `You've reached your free Grok Build usage limit`, or `reauthable: false` all mean the seat is out of budget. `grok login` fixes none of them — report "grok seat out of quota", give the reset window if the error carries one, and convene without it. One repo-reading seat can eat most of a free daily allowance in a single run. +- **Grok auth routing.** Invoke `grok` by name so a shell function that exports `XAI_API_KEY` or relocates `GROK_HOME` still applies; calling a resolved binary path silently drops to the free tier. `grok models` prints the route in use, and `.modelUsage` names the model actually served — `grok-4.5` served as `grok-4.5-build-free` is a degraded seat, not the model you asked for. +- Genuine auth errors (`invalid_grant: Refresh token has been revoked`) can be triggered by token rotation racing a parallel fan-out → re-run `grok login` +- **Schema-valid is not evidence of work.** A seat can return well-formed JSON without ever opening the repo. No `file:line` citation, or a single-turn finish, means it did not read your code — mark it degraded and say which. +- Timeout → report partial council; do not silently drop a member. A late seat that lands after you called it dead reopens the synthesis; never publish a verdict a background seat is still contradicting Never invent a provider's opinion. If a provider did not answer, it has no position — say that.