Skip to content

feat(plugins): add Claude Opus 5.5 and GPT-6 support - #14

Merged
SpiGAndromeda merged 3 commits into
mainfrom
feat/opus-5-5-gpt-6-support
Sep 24, 2026
Merged

SpiGAndromeda merged 3 commits into
mainfrom
feat/opus-5-5-gpt-6-support

Conversation

@SpiGAndromeda

@SpiGAndromeda SpiGAndromeda commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Add prompting and routing guidance for Claude Opus 5.5 and the GPT-6 models (Astra, Sol, Luna) to llm-author and work-orchestrator. Recalibrate work-orchestrator's per-model routing against a local replication, because five per-model GPT-5.6 observations the routing table relied on did not reproduce.

Changes

Prompting guidance (llm-author 3.13.0)

  • Add an Opus 5.5 section to claude-5-guide.md: default effort medium instead of Opus 5's high, adaptive thinking always on, and a 400 error for a forced tool_choice.
  • Rename the OpenAI guides from gpt-56-* to openai-gpt-*, target GPT-6 while keeping the GPT-5.6 guidance, and archive the superseded vendor docs verbatim.

Effort and refusals (work-orchestrator 4.5.0)

  • State effort defaults per model, so an Opus 5.5 definition's rung is not read as Opus 5's.
  • Re-route a safeguard refusal from a claude worker instead of accepting it as a verdict.
  • Note GPT-6 in the GPT-5.6 docs and correct the GPT-5.6 prices.

Routing recalibration (work-orchestrator 4.6.0)

  • Re-rank every codex model's severity labels before triage instead of sol's alone. Six repeated broad reviews each of gpt-5.6-sol, gpt-6-sol, and gpt-5.6-terra labeled a planted non-blocking defect as blocking in 17 of 18 runs.
  • Drop the sol exclusion from security-flavored review. Four security reviews, sol included, met no safeguard friction.
  • Limit the optional luna pass to adding a third model's coverage.
  • Extend refusal re-routing to codex workers: a different codex model, a claude definition, or the session.
  • State the sample size of every capability statement in docs/, record each routing row's basis there, and phrase per-model behavior in references/ and SKILL.md as working defaults. The measurements come from one small Python fixture, run through a proxy whose routing to the named models is unverified.

@SpiGAndromeda SpiGAndromeda changed the title Support Claude Opus 5.5 and GPT-6 in llm-author and work-orchestrator feat(plugins): add Claude Opus 5.5 and GPT-6 support Sep 24, 2026
@SpiGAndromeda
SpiGAndromeda force-pushed the feat/opus-5-5-gpt-6-support branch from 3d1bf72 to cf5cec8 Compare September 24, 2026 22:28
@SpiGAndromeda
SpiGAndromeda changed the base branch from main to feat/work-orchestrator-nested-subagents September 24, 2026 22:28
@SpiGAndromeda
SpiGAndromeda force-pushed the feat/work-orchestrator-nested-subagents branch from ead7b35 to 589b4d0 Compare September 24, 2026 22:32
@SpiGAndromeda
SpiGAndromeda force-pushed the feat/opus-5-5-gpt-6-support branch from cf5cec8 to b540d8c Compare September 24, 2026 22:35
@SpiGAndromeda
SpiGAndromeda changed the base branch from feat/work-orchestrator-nested-subagents to main September 24, 2026 22:35
Opus 5.5 is Anthropic's new recommended starting model, defaults to `medium` effort, and rejects forced `tool_choice` with a 400. The Claude 5 guide still named Opus 5 / Fable 5 as the current lineup and `high` as the default on every model, so prompts migrated with it would copy the wrong effort level.

GPT-6 ships as Astra, Sol, and Luna with no Terra tier, and GPT-6 Sol sits below Astra, so GPT-5.6 tier names do not map to GPT-6 by name. The OpenAI guide and its adaptation example move to generation-neutral names (`openai-gpt-guide.md`, `openai-gpt-adaptation.md`), make GPT-6 the current target, and keep the GPT-5.6 content labeled. GPT-6 behavior guidance is labeled as observed on Astra, since OpenAI publishes it for Astra and applies it to the family as a starting point.

All five skills move to 3.13.0, which clears the earlier skew between skill and manifest versions.
Claude Code 2.1.280 resolves the `opus` alias to `claude-opus-5-5`, so every opus agent definition now runs on Opus 5.5. Its safety classifiers can decline life-sciences biology work and high-risk dual-use cybersecurity activity, and the routing rules had no handling for a declined dispatch. The new rule treats a refusal as off-profile output that re-routes the checkpoint to the table's codex actor, another claude model, or the session itself, never as a verdict.

Opus 5.5 also defaults to `medium` effort, and Anthropic states that effort names do not correspond to the same amount of thinking across models. The opus definitions therefore stop claiming a model-default or best-measured rung and state their depth instead.

The codex rows keep the GPT-5.6 IDs. GPT-6 has no Terra tier and GPT-6 Sol sits below Astra, so no row moves to a GPT-6 model before a re-validation run.
Five per-model GPT-5.6 observations behind the routing table did not reproduce on a fresh planted-defect fixture, and repeated runs showed label errors on non-blocking items across all three tested models. Every codex model's severity labels are now re-ranked before triage instead of sol's alone, security-flavored review no longer excludes sol, and the optional luna pass claims only a third model's coverage.

A safeguard refusal from a codex worker is now re-routed to a different codex model, a claude definition, or the session, the same as a claude refusal.
@SpiGAndromeda
SpiGAndromeda force-pushed the feat/opus-5-5-gpt-6-support branch from b540d8c to cc0feca Compare September 24, 2026 22:37
@SpiGAndromeda
SpiGAndromeda merged commit c3dd8ce into main Sep 24, 2026
7 checks passed
@SpiGAndromeda
SpiGAndromeda deleted the feat/opus-5-5-gpt-6-support branch September 24, 2026 22:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant