From 2835951e4b15ac6bb105718b688ec348982a5841 Mon Sep 17 00:00:00 2001 From: Gabe Borges Date: Tue, 22 Sep 2026 13:02:04 -0400 Subject: [PATCH] feat(agents): pin opus-medium and opus-xhigh to Opus 5.5 Anthropic shipped Opus 5.5 on 2026-09-22 as the successor to Opus 5 at 4 / 0.20 / 20 dollars per million tokens against Opus 5's 5 / 0.50 / 25. The two Opus agents now pin the id 'claude-opus-5-5' instead of the 'opus' alias, so a machine on an older Claude Code build runs the same model as every other. Efforts stay at medium and xhigh. The README table, docs/subagent-routing.md, and docs/model-routing.md carry the new prices and the redone ratios. Sonnet is now 50% of Opus on input and output, Fable is 250%, and Sol on the Codex side matches Opus 5.5's price exactly. The published score rows keep their Opus 5 numbers, labeled as such, because no independent Opus 5.5 score exists yet. Plugin version 1.33.0 to 1.34.0, rev 39 to 40. --- README.md | 6 +- docs/model-routing.md | 111 ++++++++++++++---- docs/subagent-routing.md | 76 +++++++----- .../.claude-plugin/plugin.json | 2 +- .../.codex-plugin/plugin.json | 2 +- .../gborges-standard/agents/opus-medium.md | 4 +- plugins/gborges-standard/agents/opus-xhigh.md | 4 +- scripts/cloud-bootstrap.sh | 2 +- tests/test_hooks.py | 2 +- 9 files changed, 147 insertions(+), 62 deletions(-) diff --git a/README.md b/README.md index f7fce39..b202856 100644 --- a/README.md +++ b/README.md @@ -40,7 +40,7 @@ Paste this loader rather than the body of `scripts/cloud-bootstrap.sh`, so the l ```bash #!/bin/bash -# rev: 39 +# rev: 40 curl -fsSL https://raw.githubusercontent.com/gborges0727/claude-plugins/main/scripts/cloud-bootstrap.sh | bash || true exit 0 ``` @@ -89,8 +89,8 @@ Every PR bumps the `rev`, in the snippet above and in `scripts/cloud-bootstrap.s | `setup-repo` | Command | Writes `.claude/gborges-standard.json` at a repo's root, the per-repo `docs` folder and `tracker` choice the document-writing skills read. Wraps `scripts/setup-repo.sh` | | `setup` | Command | Writes `~/.claude/gborges-standard.json`, the per-machine switches for Fable access and Codex delegation, the `attribution` key in `~/.claude/settings.json` that stops Claude Code asking for AI attribution, and the Codex CLI's own model, subagent, and status line config under `~/.codex`. Wraps `scripts/setup.sh`, which does the same with no model turn | | `sonnet-medium` | Agent | Sonnet 5 at medium effort. Edits and runs with a command check in the brief, parallel copies of one such task, and fetching a named doc page | -| `opus-medium` | Agent | Opus 5 at medium effort. The default, and the floor for anything that reads code to reach a conclusion | -| `opus-xhigh` | Agent | Opus 5 at xhigh effort. One escalation step for a task that failed below it, and the stand-in for Fable on an account without it | +| `opus-medium` | Agent | Opus 5.5 at medium effort. The default, and the floor for anything that reads code to reach a conclusion | +| `opus-xhigh` | Agent | Opus 5.5 at xhigh effort. One escalation step for a task that failed below it, and the stand-in for Fable on an account without it | | `fable-xhigh` | Agent | Fable 5.1 at xhigh effort. Runs only when the user's message names Fable. See [docs/subagent-routing.md](docs/subagent-routing.md) for the routing rule and the cost reasoning | | `frontend-design` | Dependency | From `claude-plugins-official` | | `context7` | Dependency | From `claude-plugins-official` | diff --git a/docs/model-routing.md b/docs/model-routing.md index cd28d67..3e6d6a5 100644 --- a/docs/model-routing.md +++ b/docs/model-routing.md @@ -10,7 +10,9 @@ The `model-routing-review` skill rebuilds this file when either vendor ships a model or moves a price. Every number carries its date and source so the next review can tell what moved. -Last reviewed 2026-09-04 against Codex CLI 0.153.3. +Last reviewed 2026-09-22 against Codex CLI 0.153.3. The Claude side +moved to Opus 5.5 that day, and the OpenAI side was last checked on +2026-09-04. ## The two ladders @@ -42,9 +44,11 @@ says Codex is on, two questions pick the host: on what was said in the session stays on Claude. 2. Does a command check the result? A test run, a build, or a diff that applies catches a failure at no cost in judgment. A task whose result - is a conclusion nobody downstream checks stays on Claude, where Opus 5 - holds the better accuracy record (63.0 against Sol's 58.9 on the - Intelligence Index, 79.2 against 64.6 on SWE-Bench Pro). + is a conclusion nobody downstream checks stays on Claude, where Opus + holds the better accuracy record. The published Opus rows are Opus 5's + (63.0 against Sol's 58.9 on the Intelligence Index, 79.2 against 64.6 + on SWE-Bench Pro), and Anthropic's testing puts Opus 5.5 at `medium` + at or above Opus 5 at `high`. Two yeses send the task to the Codex rung that mirrors the Claude rung it would have taken. A failure escalates inside the host that ran the task, @@ -71,7 +75,8 @@ the API price is the plan burn rate too. | GPT-5.6 Sol | 4 | 0.40 | 20 | 8 / 0.80 / 30 | 2026-08-21 cut, promised through 2026-11-21, then 5 / 0.50 / 30 | | GPT-6 Astra | 10 | 1 | 50 | 20 / 2 / 75 | 2026-08 launch, checked 2026-09-04 | | Claude Sonnet 5 | 2 | 0.20 | 10 | n/a | checked 2026-09-01 | -| Claude Opus 5 | 5 | 0.50 | 25 | n/a | checked 2026-09-01 | +| Claude Opus 5.5 | 4 | 0.20 | 20 | n/a | launched 2026-09-22 | +| Claude Opus 5 | 5 | 0.50 | 25 | n/a | checked 2026-09-01, off the ladder since 2026-09-22 | | Claude Fable 5.1 | 10 | 0.25 | 50 | n/a | checked 2026-09-01 | The OpenAI long-context row applies to the whole request once its input @@ -80,24 +85,31 @@ GPT-5.6 model and Astra list a 1,050,000-token window and a 128,000-token output cap on the API. The Codex CLI catalog reports a 272,000-token working window and an 872,000-token maximum for the same models. -Rung for rung, the Codex side is cheaper than the Claude side on the two -lower rungs and dearer on the escalation rung: +Rung for rung, the Codex side is cheaper than the Claude side on the +mechanical rung, level on the default rung, and dearer on the escalation +rung: | Rung | Claude, in / out | Codex, in / out | Codex as a share of Claude | |---|---|---|---| | mechanical | Sonnet 5, 2 / 10 | Luna, 0.20 / 1.20 | 10% / 12% | -| default | Opus 5, 5 / 25 | Sol, 4 / 20 | 80% / 80% | -| escalation | Opus 5, 5 / 25 | Astra, 10 / 50 | 200% / 200% | +| default | Opus 5.5, 4 / 20 | Sol, 4 / 20 | 100% / 100% | +| escalation | Opus 5.5, 4 / 20 | Astra, 10 / 50 | 250% / 250% | | summoned | Fable 5.1, 10 / 50 | Astra, 10 / 50 | 100% / 100% | -The escalation rung pays double per token and gets it back in tokens. +Opus 5.5 took the default rung at Sol's exact list price, and Sol's price +is a cut that OpenAI promised only through 2026-11-21. So the 20% price +edge Sol held over Opus 5 on the default rung is gone, and the choice +between the two hosts rests on the two questions above alone. + +The escalation rung pays two and a half times per token and gets some +of it back in tokens. Artificial Analysis measured Astra at max using a third of Sol's tokens on its coding harness, and on this repo's three briefs Astra at medium spent 29% fewer tokens than Sol at max. -Astra's cache hit costs 1.00 against Fable's 0.25, so a long -many-turn Astra session pays four times Fable's rate on the tokens it -resends. Per-token price is an input to the ranking, not the ranking. The +Astra's cache hit costs 1.00 against Fable's 0.25 and Opus 5.5's 0.20, +so a long many-turn Astra session pays four times Fable's rate and five +times Opus 5.5's on the tokens it resends. Per-token price is an input to the ranking, not the ranking. The ranking is cost per finished task, which counts the retry a cheap failure causes and the orchestrator tokens spent writing the brief again. @@ -106,7 +118,10 @@ causes and the orchestrator tokens spent writing the brief again. All OpenAI numbers ran at the model's maximum effort unless noted. Scores compare safely only when the benchmark, the harness, and the effort match, and the vendor tables mix all three, so treat a gap under three points as -noise. +noise. Every Claude row below is from before 2026-09-22. No independent +Opus 5.5 score had been published on that date, so the Opus 5 rows stay as +the Claude side's record, and Anthropic's own Opus 5.5 claims sit in the +"Why the ladder changed on 2026-09-22" section. | Benchmark | Luna | Terra | Sol | Astra | Claude | Source | |---|---|---|---|---|---|---| @@ -158,7 +173,8 @@ What the rows say, one model at a time: $11.84 at the same rate. Opus 5 leads it on the Intelligence Index (63.0 against 58.9), SWE-Bench Pro (79.2 against 64.6), and Terminal-Bench 4.0 (52.3 against 37.3), which is why the Claude escalation rung stays - on Opus rather than moving to Sol. + on Opus rather than moving to Sol. Opus 5.5 succeeds Opus 5 on both + Claude rungs at Sol's price, so that record now costs no premium. - Astra leads every OpenAI benchmark it appears on and beats Fable 5.1 on the agentic coding and math rows, while Fable 5.1 leads on Humanity's Last Exam with tools. On Terminal-Bench 4.0 OpenAI estimates Astra's API @@ -213,7 +229,7 @@ stays. The same three briefs then ran on Luna, Sol, Sol at max, and Astra at medium through Codex, and on Opus 5 through the plugin's `opus-medium` -and `opus-xhigh` agents. All twenty-four runs passed. Token totals are what each host billed for the three runs, so +and `opus-xhigh` agents, which pinned Opus 5 until 2026-09-22. All twenty-four runs passed. Token totals are what each host billed for the three runs, so the Codex figures include Codex's own system prompt and the Opus figures include Claude Code's plus the writing rules the spawn hook appends. The API-equivalent cost takes 80% of tokens at the input price and 20% at @@ -234,11 +250,16 @@ the output price. Every model finished every brief, so these briefs cannot rank the models on accuracy. They rank them on cost and speed for work all of them can -do: Luna costs 3% of Opus at medium, Terra at xhigh costs 29%, Sol costs -60%, and Opus finishes fastest. The Codex costs land on the ChatGPT plan allowance, -not the API bill. Ranking on the hard tenth of tasks still rests on the -published scores above, where Opus 5 leads Sol by 15 points on -Terminal-Bench 4.0. +do: Luna costs 3% of Opus 5 at medium, Terra at xhigh costs 29%, Sol costs +60%, and Opus finishes fastest. + +At Opus 5.5's prices the same token counts +would cost $0.92 at medium and $0.98 at xhigh, so Sol's share rises to +75%, and Anthropic reports Opus 5.5 finishing agentic coding tasks in +about half the tokens, which would cut those figures again. The Codex +costs land on the ChatGPT plan allowance, not the API bill. Ranking on the +hard tenth of tasks still rests on the published scores above, where Opus +5 leads Sol by 15 points on Terminal-Bench 4.0. The orchestrator never sets max or ultra. Ultra spawns subagents inside the call, which multiplies the allowance one call spends. Only Pro and @@ -246,6 +267,39 @@ Business Premium plans draw Astra from the full Codex allowance. Plus holds a limited Astra allowance, so on Plus an escalation can hit that cap before the 5-hour limit does. +## Why the ladder changed on 2026-09-22 + +Anthropic shipped Opus 5.5 (`claude-opus-5-5`) on 2026-09-22 as the +successor to Opus 5, and it took both Claude rungs Opus 5 held, at the +same efforts. Rule 1 of the review decides it, because the vendor's named +successor takes the rung. Rule 3 agrees on its own, since the price fell +and the vendor's numbers rose. The facts: + +- Price. Opus 5.5 lists at 4 / 0.20 / 20 (input, cache hit, output) + against Opus 5's 5 / 0.50 / 25, so 20% less on input and output and 60% + less on cache hits. Fast mode is 8 / 40. Batch is 2 / 10. +- Effort. The API default is `medium` where Opus 5 defaulted to `high`, + and the plugin's agents set the effort explicitly, so the change + reaches neither rung. Anthropic's testing has Opus 5.5 at `medium` + matching or beating Opus 5 at `high` on multistep coding in a real + codebase, in fewer steps and with about half the tokens, and `low` + close behind on several coding evaluations. At a given effort it thinks + more per turn than Opus 5, most of all at `xhigh` and `max`, so + `opus-xhigh` turns run longer. +- Behavior. Thinking cannot be turned off, forced tool choice returns an + error, thinking blocks are bound to the model and the conversation, and + computer use runs only through the 2026-08-01 toolset. Claude Code + keeps the conversation prefix intact, so none of those reach a spawned + agent. The agent files pin `claude-opus-5-5` by id rather than the + `opus` alias, so a machine on an older Claude Code build that still + resolves the alias to Opus 5 runs the same model as every other. +- Scores. No independent Opus 5.5 row was published on 2026-09-22. The + claims above are Anthropic's migration guide's, run in its own harness. + +The OpenAI side did not move. Sol at 4 / 20 now matches Opus 5.5's +price exactly, so the default rung's host is picked by the two questions +above and not by price. + ## Why the ladder changed on 2026-09-04 Before this review the Codex side ran Luna at xhigh as the default worker @@ -271,12 +325,18 @@ medium, Astra at xhigh. The facts that moved it: ## Open questions +- Opus 5.5 has no independent score on any row of the table above and no + point-by-point effort curve. Anthropic's claim that `medium` beats Opus + 5 at `high` in half the tokens is the whole case for keeping `medium`. + The three briefs above, rerun on Opus 5.5 at `medium` and at `xhigh`, + would give the first measured token counts on this repo's work. - No effort curve is published for any GPT-5.6 model or for Astra. The - Claude side runs Opus at medium because Anthropic published that curve. + Claude side ran Opus 5 at medium because Anthropic published that curve, + and Opus 5.5 keeps it on the vendor's word. The three-brief Terra measurement above is one run per cell. A repeat on ten briefs with three runs each would give a curve worth acting on. - Sol's price reverts on or after 2026-11-21 unless OpenAI extends it. At - 5 / 30 Sol still sits under Opus. + 5 / 30 Sol would then cost more than Opus 5.5's 4 / 20 on both counts. - Astra at medium has no published score on the agentic rows. The case for it rests on Astra's four-point effort spread against its 20-point lead at max. Ten hard briefs that Sol at xhigh fails, rerun on Astra at @@ -304,7 +364,10 @@ medium, Astra at xhigh. The facts that moved it: read the same way - July price cut, https://community.openai.com/t/announcing-a-major-price-drop-for-5-6-terra-and-luna-and-fast-mode-for-5-6-sol/1388484 - August Sol cut, https://www.explainx.ai/blog/openai-gpt-5-6-sol-api-price-cut-20-percent-august-2026 -- Anthropic prices, `docs/subagent-routing.md`, checked 2026-09-01 +- Anthropic prices, `docs/subagent-routing.md`, checked 2026-09-22 +- Opus 5.5 prices, effort default, and the `medium` against `high` + claims, Anthropic's Opus 5.5 migration guide as shipped in Claude + Code's `claude-api` skill, read 2026-09-22 - The local catalog, `codex debug models` - CodeRabbit run, https://www.coderabbit.ai/blog/gpt-5-6-sol-and-terra-benchmark - Sonar run, https://www.sonarsource.com/blog/openai-gpt-5-6-sol-and-terra/ diff --git a/docs/subagent-routing.md b/docs/subagent-routing.md index d0aed07..c071c3b 100644 --- a/docs/subagent-routing.md +++ b/docs/subagent-routing.md @@ -4,16 +4,17 @@ How the orchestrating session picks a subagent, and the cost reasoning behind the ladder. The rule itself lives in the output style's Subagents section, `plugins/gborges-standard/output-styles/plain-english.md`. This file holds the why, so the rule can stay short. Prices and published -numbers below are from Anthropic's API pricing, its cost guidance, and its -Fable 5.1 prompting guide as of 2026-09-01. +numbers below are from Anthropic's API pricing, its cost guidance, its +Fable 5.1 prompting guide, and its Opus 5.5 migration guide as of +2026-09-22. ## The four agents | Agent | Model | Effort | Takes | |---|---|---|---| | `sonnet-medium` | Sonnet 5 | medium | An edit or a run whose brief names the exact change and a command that checks it. Parallel copies of one such task across files. Fetching a named doc page outside the codebase | -| `opus-medium` | Opus 5 | medium | The default. The floor for any task that reads code to reach a conclusion | -| `opus-xhigh` | Opus 5 | xhigh | A task that failed once below it. A task that is one dependent chain the orchestrator cannot split. The stand-in for Fable on an account without it | +| `opus-medium` | Opus 5.5 | medium | The default. The floor for any task that reads code to reach a conclusion | +| `opus-xhigh` | Opus 5.5 | xhigh | A task that failed once below it. A task that is one dependent chain the orchestrator cannot split. The stand-in for Fable on an account without it | | `fable-xhigh` | Fable 5.1 | xhigh | Only when the user's message names Fable | Each name states its model and effort so the orchestrator sees the cost of @@ -21,20 +22,26 @@ a dispatch in the name it types. ## Prices -| Model | Input, $ per million tokens | Cache hit, $ per million tokens | Output, $ per million tokens | Against Opus 5 (input, cache hit, output) | +| Model | Input, $ per million tokens | Cache hit, $ per million tokens | Output, $ per million tokens | Against Opus 5.5 (input, cache hit, output) | |---|---|---|---|---| -| Sonnet 5 | 2 | 0.20 | 10 | 40%, 40%, 40% | -| Opus 5 | 5 | 0.50 | 25 | 100%, 100%, 100% | -| Fable 5.1 | 10 | 0.25 | 50 | 200%, 50%, 200% | - -Fable 5.1 prices a cache hit at 2.5% of its input price, where every other -model uses 10%, so a Fable 5.1 cache hit costs half of an Opus 5 cache hit. -A subagent that reads many files sends its -whole context back on every turn, and after the first turn most of that -input is cache hits, so Fable 5.1 pays less than Opus 5 for those tokens. -Uncached input and output stay at double. Nobody has measured what share -of a real dispatch's tokens are cache hits, so 200% is the ceiling on the -Fable premium and 50% is the floor. +| Sonnet 5 | 2 | 0.20 | 10 | 50%, 100%, 50% | +| Opus 5.5 | 4 | 0.20 | 20 | 100%, 100%, 100% | +| Fable 5.1 | 10 | 0.25 | 50 | 250%, 125%, 250% | + +Opus 5.5 shipped on 2026-09-22 at 20% under Opus 5 on input and output +(Opus 5 was 5 / 0.50 / 25) and at 60% under it on cache hits. Opus 5.5 +prices a cache hit at 5% of its input price and Fable 5.1 at 2.5%, so a +Fable 5.1 cache hit now costs a quarter more than an Opus 5.5 cache hit, +where against Opus 5 it cost half. + +A subagent that reads many files sends its whole context back on every turn, and after the first turn most of +that input is cache hits, so the Fable premium on a long dispatch sits +nearer 125% than 250%. Nobody has measured what share of a real +dispatch's tokens are cache hits, so 250% is the ceiling on the Fable +premium and 125% is the floor. + +Sonnet 5 and Opus 5.5 charge the same for a cache hit, so on a long +dispatch Sonnet's saving comes only from uncached input and output. Per-token price is an input to the analysis, not the ranking. The ranking is cost per finished task, which counts the retry a cheap failure causes @@ -43,14 +50,26 @@ and the orchestrator tokens spent writing the brief again. ## Why Opus at medium is the default Anthropic's coding runs on Opus 5 put the effort curve like this. At -`medium`, Opus 5 gives up about 2 points of pass rate for half the cost of -the default (`high`). At `low` it gives up about 8 points for a quarter of +`medium`, Opus 5 gave up about 2 points of pass rate for half the cost of +its default (`high`). At `low` it gave up about 8 points for a quarter of the cost. Two points is the price of halving the bill, and eight points is -not, so `medium` is the default and `low` appears nowhere in the ladder. +not, so `medium` became the default and `low` appears nowhere in the +ladder. -On research and knowledge work the curve is nearly flat. `medium` matched +On research and knowledge work the curve was nearly flat. `medium` matched the default's accuracy at 70% to 85% of its cost across four benchmarks, -so nothing there argues for a higher default either. +so nothing there argued for a higher default either. + +Opus 5.5 keeps `medium` for a stronger reason. Anthropic's migration guide +sets `medium` as the API default on Opus 5.5 (Opus 5 defaulted to `high`), +and its testing has Opus 5.5 at `medium` matching or beating Opus 5 at +`high` on multistep coding in a real codebase, in fewer steps and with +about half the tokens. The same guide says `low` comes close on several +coding evaluations at much lower cost, and that at a given level Opus 5.5 +thinks more per turn than Opus 5, most of all at `xhigh` and `max`. So +`opus-xhigh` turns run longer than they did on Opus 5. Anthropic has +published no point-by-point effort curve for Opus 5.5, so the Opus 5 curve +above is still the only one with numbers. ## Why Sonnet stays out of code investigation @@ -98,17 +117,20 @@ Anthropic's guidance is to sweep effort on the current model before dropping a tier, and it reports that the larger model at lower effort often wins on cost per task. Fable 5 at `low` beat Sonnet 5 on a deep-research benchmark while costing about 10% less per task. Nobody has measured Sonnet -5 at `xhigh` against Opus 5 at `medium` on this user's work, and a fourth +5 at `xhigh` against Opus 5.5 at `medium` on this user's work, and a fourth agent makes every dispatch a harder choice. The model step is the one the published numbers say moves accuracy, so a failed Sonnet task goes to Opus at `medium`, not to Sonnet at a higher effort. ## Why Fable is user-only -Fable costs double Opus on uncached input and on output, and on a coding -subset Opus 5 matched Fable 5 (91.7% against 91.3%) at about 60% of its -cost. Those are Fable 5 numbers. Anthropic's Fable 5.1 prompting guide -says the 5.1 gains over Fable 5 are largest at the higher effort levels, +Fable costs two and a half times Opus 5.5 on uncached input and on +output, and on a coding subset Opus 5 matched Fable 5 (91.7% against +91.3%) at about 60% of its cost. Those are Opus 5 and Fable 5 numbers, and +Opus 5.5 at `medium` beats Opus 5 at `high` in Anthropic's own testing at +a lower price, so the gap Fable has to justify grew. + +Anthropic's Fable 5.1 prompting guide says the 5.1 gains over Fable 5 are largest at the higher effort levels, that 5.1 at `medium` roughly matches Fable 5 at lower cost, and that 5.1 at `low` is often competitive with Opus and Sonnet on cost per task while scoring higher. None of that is measured on this user's briefs. diff --git a/plugins/gborges-standard/.claude-plugin/plugin.json b/plugins/gborges-standard/.claude-plugin/plugin.json index c99115f..6c6ce62 100644 --- a/plugins/gborges-standard/.claude-plugin/plugin.json +++ b/plugins/gborges-standard/.claude-plugin/plugin.json @@ -2,7 +2,7 @@ "name": "gborges-standard", "displayName": "Gabe's standard set", "description": "Standard Claude Code setup used across all of Gabe Borges' repos. Applies the Plain English output style everywhere, strips AI-attribution footers from GitHub writes, appends the writing rules to subagent prompts, keeps Codex calls off GPT-6 Astra unless the user named it, re-injects the writing rules reminder with every user message, and bundles the writing-voice, read-aloud-prep, bear-notes, codex-delegate, model-routing-review, and pair-debate skills, the grill, build, spec, routine, wayfinder, handoff, research, wait-what, teach, writing-for-agents, domain-modeling, investigate, and resolving-merge-conflicts skills rewritten in plain English from Matt Pocock's set, four routed subagents (sonnet-medium, opus-medium, opus-xhigh, and the user-summoned fable-xhigh), the add-to-git, setup, and setup-repo commands, and the frontend-design and context7 plugins.", - "version": "1.33.0", + "version": "1.34.0", "author": { "name": "Gabe Borges", "email": "gbborges@proton.me" diff --git a/plugins/gborges-standard/.codex-plugin/plugin.json b/plugins/gborges-standard/.codex-plugin/plugin.json index b07ddcd..b0c645b 100644 --- a/plugins/gborges-standard/.codex-plugin/plugin.json +++ b/plugins/gborges-standard/.codex-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "gborges-standard", - "version": "1.33.0", + "version": "1.34.0", "description": "Standard Codex setup used across all of Gabe Borges' repos. Injects the Plain English rules at session start and into every subagent, strips AI-attribution footers from GitHub writes, re-injects the writing rules reminder with every user message, and ships the writing-voice, read-aloud-prep, bear-notes, codex-delegate, model-routing-review, and pair-debate skills, plus grill, build, spec, routine, wayfinder, handoff, research, wait-what, teach, writing-for-agents, domain-modeling, investigate, and resolving-merge-conflicts.", "author": { "name": "Gabe Borges", diff --git a/plugins/gborges-standard/agents/opus-medium.md b/plugins/gborges-standard/agents/opus-medium.md index ed39461..cc77f6f 100644 --- a/plugins/gborges-standard/agents/opus-medium.md +++ b/plugins/gborges-standard/agents/opus-medium.md @@ -1,7 +1,7 @@ --- name: opus-medium -description: Opus 5 at medium effort. The default for delegated work, and the floor for any task that reads code to reach a conclusion (an investigation, a diagnosis, a review, a design choice). -model: opus +description: Opus 5.5 at medium effort. The default for delegated work, and the floor for any task that reads code to reach a conclusion (an investigation, a diagnosis, a review, a design choice). +model: claude-opus-5-5 effort: medium --- You are a general-purpose worker handling a delegated unit of work from the orchestrating session. When the brief names a check, run it before you report. If the check fails or you cannot finish, say so plainly and quote the failing output. diff --git a/plugins/gborges-standard/agents/opus-xhigh.md b/plugins/gborges-standard/agents/opus-xhigh.md index 762117a..7cc345d 100644 --- a/plugins/gborges-standard/agents/opus-xhigh.md +++ b/plugins/gborges-standard/agents/opus-xhigh.md @@ -1,7 +1,7 @@ --- name: opus-xhigh -description: Opus 5 at xhigh effort. Use for a task that already failed once on opus-medium or sonnet-medium, or for a task that is one long dependent chain the orchestrator cannot split into parallel pieces. Also the stand-in for fable-xhigh when the account cannot run Fable. -model: opus +description: Opus 5.5 at xhigh effort. Use for a task that already failed once on opus-medium or sonnet-medium, or for a task that is one long dependent chain the orchestrator cannot split into parallel pieces. Also the stand-in for fable-xhigh when the account cannot run Fable. +model: claude-opus-5-5 effort: xhigh --- You are a worker taking a hard unit of work from the orchestrating session. The brief may carry the exact output of a failed earlier attempt. Start from that output, not from the earlier approach. When the brief names a check, run it before you report, and quote the failing output if it still fails. diff --git a/scripts/cloud-bootstrap.sh b/scripts/cloud-bootstrap.sh index cfad617..b4fe865 100755 --- a/scripts/cloud-bootstrap.sh +++ b/scripts/cloud-bootstrap.sh @@ -11,7 +11,7 @@ # missing or the two numbers differ. After the merge, paste the new rev into # each environment's Setup script field when the change cannot wait out the # snapshot expiry. -# rev: 39 +# rev: 40 set -u diff --git a/tests/test_hooks.py b/tests/test_hooks.py index 873f1d9..046dbcb 100644 --- a/tests/test_hooks.py +++ b/tests/test_hooks.py @@ -238,7 +238,7 @@ def test_a_fork_on_a_fable_session_is_denied_and_pointed_at_opus_medium(self): self.assertNotIn("updatedInput", block) def test_a_fork_on_an_opus_session_passes_untouched(self): - transcript = self.write_transcript("claude-opus-5") + transcript = self.write_transcript("claude-opus-5-5") self.assertIsNone(self.spawn("fork", transcript=transcript)) def test_the_newest_reply_decides_the_model(self):