Skip to content

fix(anthropic): correct Opus 5.5 output budget - #1854

Open
jaszhix wants to merge 1 commit into
Zoo-Code-Org:mainfrom
jaszhix:fix/opus55-output-budget
Open

jaszhix wants to merge 1 commit into
Zoo-Code-Org:mainfrom
jaszhix:fix/opus55-output-budget

Conversation

@jaszhix

@jaszhix jaszhix commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Related GitHub Issue

Closes: #1853

Description

Claude Opus 5.5's direct Anthropic model entry declares a 128K output limit and uses adaptive thinking, but it also declares supportsReasoningBudget: true. The output-limit resolver interprets that capability as legacy hybrid-budget handling, so Zoo Code replaces the declared model limit with 16,384 tokens when reasoning is enabled and 8,192 otherwise.

For Opus 5.5, supportsReasoningBinary already controls adaptive-thinking request construction. This PR removes only the incorrect supportsReasoningBudget capability and updates its stale output-limit comment. The existing maxTokens: 128_000 value then becomes the effective request limit.

Regression coverage verifies that:

  • Opus 5.5 still uses { type: "adaptive" } when reasoning is enabled;
  • temperature remains omitted;
  • the resolved and requested output limit is 128000;
  • supportsReasoningBudget is absent from the resolved model info.

The change is limited to Claude Opus 5.5 on the direct Anthropic provider path.

Test Procedure

  1. Run the focused Anthropic provider suite:

    pnpm --dir src exec vitest run api/providers/__tests__/anthropic.spec.ts
  2. Run affected typechecks:

    pnpm --dir packages/types check-types
    pnpm --dir src check-types
  3. Run lint for the changed source and test files.

  4. Verify the Opus 5.5 request-shape regression:

    • thinking remains { type: "adaptive" }
    • temperature remains omitted
    • default max_tokens is 128000
    • getModel().info.supportsReasoningBudget is absent
  5. Optional manual verification: run a complex Opus 5.5 agent task that previously exhausted 16K tokens during thinking and confirm it progresses to text or tool use.

Pre-Submission Checklist

  • Issue Linked: This PR is linked to the GitHub issue above.
  • Scope: My changes are focused on the linked issue (one major fix per PR).
  • Self-Review: I have performed a thorough self-review of my code.
  • Testing: Updated tests cover model resolution and request construction.
  • Visual Snapshot: No UI or rendered-state change.
  • Documentation Impact: I have considered whether documentation updates are required.
  • Contribution Guidelines: I have read and agree to the Contributor Guidelines.

Visual Snapshots

N/A. This changes provider metadata and request construction only.

Videos (interaction / animation only)

N/A.

Documentation Updates

  • No documentation updates are required.
  • Yes, documentation updates are required.

@coderabbitai

coderabbitai Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: Zoo-Code-Org/Zoo-Code/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 425c62f3-8d71-4040-b6f6-f69f5257b26b

📥 Commits

Reviewing files that changed from the base of the PR and between 778ad3e and 9f7e474.

📒 Files selected for processing (2)
  • packages/types/src/providers/anthropic.ts
  • src/api/providers/__tests__/anthropic.spec.ts

Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 3 remain after this review.

📜 Recent review details
🧰 Additional context used
📓 Path-based instructions (6)
Treat model, provider, MCP, path, command, and tool data as untrusted.

⚙️ CodeRabbit configuration file

Files:

  • src/api/providers/__tests__/anthropic.spec.ts
For persisted settings, verify the complete schema/storage/runtime/webview round trip, shared default semantics, and focused true plus false/unset tests.

⚙️ CodeRabbit configuration file

Files:

  • packages/types/src/providers/anthropic.ts
Require regression coverage at the lowest valid harness with behavior-focused assertions, including relevant negative, error, false/unset, and boundary cases.

⚙️ CodeRabbit configuration file

Files:

  • src/api/providers/__tests__/anthropic.spec.ts
Check strict typing and exhaustive behavior across normal, boundary, error, cancellation, retry, and compatibility paths.

⚙️ CodeRabbit configuration file

Files:

  • packages/types/src/providers/anthropic.ts
  • src/api/providers/__tests__/anthropic.spec.ts
Verify extension/webview contracts, cancellation and error propagation, VS Code lifecycle correctness, and behavior under retries and partial failure.

⚙️ CodeRabbit configuration file

Files:

  • src/api/providers/__tests__/anthropic.spec.ts
Act as an adversarial second-opinion reviewer.

⚙️ CodeRabbit configuration file

Files:

  • packages/types/src/providers/anthropic.ts
  • src/api/providers/__tests__/anthropic.spec.ts
🔇 Additional comments (2)
packages/types/src/providers/anthropic.ts (1)

168-168: LGTM!

Also applies to: 176-176

src/api/providers/__tests__/anthropic.spec.ts (1)

504-504: LGTM!

Also applies to: 757-757, 759-759


📝 Summary

Summary by CodeRabbit

  • Bug Fixes
    • Claude Opus 5.5 requests now use a 128,000-token limit by default.
    • Updated model information to reflect that manual reasoning-budget settings are not supported.

Walkthrough

Claude Opus 5.5 metadata and Anthropic provider tests now reflect its declared 128,000-token output limit. The tests also check that the model does not expose reasoning-budget support.

Changes

Claude Opus 5.5 output limit

Layer / File(s) Summary
Model limit and provider expectations
packages/types/src/providers/anthropic.ts, src/api/providers/__tests__/anthropic.spec.ts
The model comment now states that Claude Opus 5.5 uses adaptive thinking and rejects manual budget_tokens. Tests no longer set a 32,768-token override and expect a 128,000-token request limit. Model assertions expect maxTokens to be 128,000 and supportsReasoningBudget to be undefined.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix · Severity of issue fixed: Medium

Merge Risk: ⚪ Minimal · up to 9f7e4

Claude Opus 5.5’s direct Anthropic request is expected to retain its declared 128,000-token limit with adaptive thinking. The supplied changes and regression test reveal no actionable merge-blocking risk.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 9f7e4

The change raises the per-request output allowance on an existing provider path without evident new access or authority. The larger allowance could increase resource use, while the surrounding access and spending controls were not fully evidenced.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The established effect is a higher bounded per-request output allowance on the direct Anthropic path, not a demonstrated expansion of independently attackable entrypoints or privileges.

Trust Boundaries and Controls

  • observed — The request still goes through the existing provider branch, and central resolution retains a configured output ceiling; the supplied evidence does not characterize controls before that branch.
🚥 Pre-merge checks | ✅ 7 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Regression Evidence ⚠️ Warning The explicit disabled branch lacks focused coverage. The change removes supportsReasoningBudget for claude-opus-5-5, so the resolver now retains maxTokens: 128_000 when reasoning is disabled, an… Add an Anthropic handler regression test for claude-opus-5-5 with enableReasoningEffort: false. Collect the mocked request and assert thinking is undefined and max_tokens is 128000; retain the existing enabled-reasoning assertions…
✅ Passed checks (7 passed)
Check name Status Explanation
Linked Issues check ✅ Passed The change satisfies issue #1853. It removes supportsReasoningBudget from the direct claude-opus-5-5 Anthropic model entry while retaining supportsReasoningBinary: true. The model keeps `maxToke…
Out of Scope Changes check ✅ Passed The changes stay within issue #1853. The model declaration, its explanatory comment, and the related Anthropic regression tests directly address the output-limit resolution and adaptive-thinking behav…
Security Boundaries ✅ Passed PASS. The changed paths only adjust Claude Opus 5.5 metadata and regression expectations. They remove supportsReasoningBudget, retain supportsReasoningBinary, and change the expected max_tokens …
Persistence Integrity ✅ Passed PASS — The PR changes only Anthropic model metadata and related test expectations. The changed files contain no persistence, storage write, rollback, atomic-write, or partial-failure path. The updated…
Lifecycle Resource Cleanup ✅ Passed PASS — The pull request changes only Claude Opus 5.5 model metadata and related test expectations. The patch adds no listener, watcher, timer, task, provider, cancellation, disposal, or restart lifecy…
Title check ✅ Passed The title clearly identifies the main change: correcting Claude Opus 5.5 output budget handling for Anthropic.
Description check ✅ Passed The description covers the linked issue, implementation details, scope, regression tests, test commands, and checklist. The optional Additional Notes and Get in Touch sections are not included, but th…
Full details: Regression Evidence

Explanation

The explicit disabled branch lacks focused coverage. The change removes supportsReasoningBudget for claude-opus-5-5, so the resolver now retains maxTokens: 128_000 when reasoning is disabled, and createMessage should send max_tokens: 128000 without thinking. The changed tests cover enabled reasoning in the request and an unset setting through getModel(), but no Opus 5.5 test sets enableReasoningEffort: false or checks this request shape. The issue and existing tests identify disabled reasoning as a distinct path.

Resolution

Add an Anthropic handler regression test for claude-opus-5-5 with enableReasoningEffort: false. Collect the mocked request and assert thinking is undefined and max_tokens is 128000; retain the existing enabled-reasoning assertions.

  • Fix all pre-merge checks with AI
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Review status

Thanks for contributing. This comment tracks the review sequence and the next action.

Current step: Awaiting fresh human maintainer or CODEOWNER approval.

Automated review is complete for the latest commit but does not replace human approval.

Review-state labels are managed by this workflow; do not edit them manually. community-approved is managed the same way — do not add or remove it manually. It signals a fresh community code approval for the current head as an advisory priority only; maintainer review is still required.

@codecov

codecov Bot commented Sep 29, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@github-actions github-actions Bot added coderabbit-review-active Required CI passed; CodeRabbit review is active awaiting-coderabbit Waiting for CodeRabbit to approve the latest commit labels Sep 29, 2026
@github-actions github-actions Bot added awaiting-maintainer CodeRabbit approved; waiting for a human maintainer and removed coderabbit-review-active Required CI passed; CodeRabbit review is active awaiting-coderabbit Waiting for CodeRabbit to approve the latest commit labels Sep 29, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting-maintainer CodeRabbit approved; waiting for a human maintainer

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[BUG] Claude Opus 5.5 defaults to legacy 8K/16K output budgets on Anthropic

1 participant