Skip to content

fix(catalog): Codewhale corrections hold online, not only offline (#6396) - #6620

Merged
4 commits merged into
mainfrom
fix/6396-catalog-corrections
Sep 28, 2026
Merged

4 commits merged into
mainfrom
fix/6396-catalog-corrections

Conversation

@Hmbown

@Hmbown Hmbown commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Stacked on #6618. The base is that branch, so this diff shows only this slice. GitHub retargets it to main once #6618 merges.

What was wrong

Codewhale's deliberate catalog holds existed only as hand edits to the offline seed (models_dev.bundled.json). They covered prices left out where a flat rate would mislead, and DeepSeek's output limit kept at its published 384K. A live Models.dev refresh replaces seed rows wholesale, so the holds disappeared on every networked install. Upstream today:

  • MiMo and DeepSeek: priced.
  • Alibaba Token and Coding Plan rows: listed at $0, which reads as free.
  • DeepSeek output limit: 393216.

The "honesty" rule in the seed's _meta was prose that the runtime never read.

The fix

The holds now live in crates/config/assets/catalog_corrections.json as field patches in the cloud-facts ModelFact shape. The signed layer's patch code applies them to every Models.dev row as it is hydrated, whether it comes from the bundled seed (bundled_catalog_offerings) or a live refresh (live_offerings_from_models_dev). A correction therefore ranks above both Models.dev layers and below signed cloud facts and provider rosters. The same correction holds online and offline.

  • ModelFact.pricing_withheld: Option<String> (a reason) clears the price, so the route reports PricingSku::UnknownOrStale. The field is optional, and payloads without it parse unchanged. Per-provider rules expand to one withheld-price patch per row.
  • One patch path. catalog_patch::apply_patches is shared. Signed facts may create rows; corrections may not. The loader refuses:
    • hide, deprecate and allow_unlisted;
    • labels;
    • a patch that changes nothing;
    • a limit patch without a reason.
  • A corrected row keeps its own source. The price a correction owns reports CatalogSource::CodewhaleBundled through cost_source. See the deviations section for why.
  • Dead field removed. CatalogCompiler.codewhale_bundled was declared, but nothing had ever set it. The layer docs in catalog.rs, CATALOG_REFRESH.md and CLOUD_FACTS.md now describe the rank where corrections actually apply.

Reasoning controls (second commit)

Regenerating the seed in #6612 turned up a second kind of hold that works only offline: reasoning controls. Before this change, a live refresh replaced all of these, and the picker, /effort and the effort clamp all read them:

ModelFact.reasoning_options (optional) replaces a row's controls and keeps any signed annotation already on the row. These 18 rows get corrections, each with the source it cites.

Two groups were left out:

  • GLM-5.3's high/max ladder. It was only a placeholder inherited from GLM-5.2, and upstream now publishes low/high/max.
  • The Coding Plan qwen budget_tokens entries. They were copies of Token Plan rows, not a Codewhale decision.

Each hold, checked again on its merits

Hold Decision Why
DeepSeek native pricing kept (provider rule) billed by time of day; crates/tui/src/pricing.rs has the time-aware table
xAI pricing kept (provider rule) upstream's own tiers field shows rates doubling past 200K
Xiaomi MiMo pricing kept (provider rule) PAYG and Token Plan keys cannot be told apart
Model Studio Token/Coding Plan pricing (4 providers) kept (provider rules) plan quota, not per-token; upstream 0 would read as free
StepFun pricing kept (provider rule) the same ids serve Step Plan and PAYG; pricing.rs has endpoint-scoped rates
DeepSeek native max_output 384000 kept, marked unverified DeepSeek publishes 384K; upstream 393216 = 384×1024. Which value the API accepts is unverified, and the lower bound is always accepted. Settling it takes one real provider call (provider spend, so a human runs it)
MiniMax-M3 pricing added upstream tiers the price at 512K, the same rule as Grok. The seed already listed it unpriced without saying why
aggregator-hosted DeepSeek pricing (9 rows) kept (restored in the third commit) I dropped this first, then #6612's generated seed broke hosted_flash_and_pro_routes_price_from_bundled_docs_rates. crates/tui/src/pricing.rs prices these routes from a reviewed provider-docs table (Fireworks V4 Pro: 1.74 vs Models.dev 1.20), and a catalog rate outranks it
Novita DeepSeek max_output 384000 added same published-384K question as native DeepSeek (upstream 393216)
grok-4.3 reasoning_options: [] added xAI documents no effort control for 4.3 (#6501); upstream lists a ladder
OpenRouter dots-3-note-preview:free pricing dropped a free model's price is 0
zai GLM-5.3 pricing kept (restored in the fourth commit) I first dropped it, reasoning that zai was the pay-as-you-go API. That was wrong: DEFAULT_ZAI_BASE_URL is the Coding Plan (api/coding/paas/v4), which bills credits
MiMo 2.6 text-only image input dropped #6618 made a stale text-only row harmless; holding it down would strip images
MiMo 1,000,000 context dropped the only citation was our own rows; upstream says 1,048,576

Deviations from the approved design

  • Where corrections apply. The design put corrections inside CatalogCompiler::compile. But CatalogCompiler is only used in tests; the runtime merge is crates/tui/src/provider_lake.rs::compute_merged_snapshot. Applying corrections at hydration, the single funnel both Models.dev layers pass through, reaches the resolver, the provider lake, pricing fallbacks and the export path without adding a second merge.
  • Row source. The design relabelled corrected rows CodewhaleBundled. That moved a live Models.dev row across a layer boundary: CatalogCompiler::with_live sorts rows by source, so a relabelled row ranked above signed facts, and signed_facts_correct_a_live_models_dev_row_but_never_the_provider_roster failed. Rows now keep their source, and only the owned price is attributed to CodewhaleBundled.

Not in this PR

  • Tiered pricing is inconsistent. Upstream tiers 569 rows by prompt length. The seed prices the OpenAI gpt-5.5/5.6 rows (tier at 272K) and three OpenRouter Qwen rows (256K) at their base rate while withholding Grok and MiniMax-M3. The real fix is tier-aware pricing, not more holds, and it needs its own issue.
  • A third offline store still overlaps. The legacy crates/models/assets/model_catalog.bundled.json duplicates some of these facts. Its module doc already schedules removal once the models/pricing compat callers migrate.
  • Upstream providers are matched by id. A new upstream MiMo or Model Studio id is covered by the provider rules, but a new upstream provider id is not.

Verification

  • CARGO_BUILD_JOBS=3 scripts/dev-cargo.sh test -p codewhale-config --lib: 707 passed, 0 failed, 1 ignored (after the second commit)
  • CARGO_BUILD_JOBS=3 scripts/dev-cargo.sh test -p codewhale-config --test configured_models: 9 passed
  • CARGO_BUILD_JOBS=3 scripts/dev-cargo.sh test -p codewhale-tui --lib -- vision image badge capabilit pricing provider_lake models_dev_live model_picker provider_catalog_live catalog deepseek xiaomi mimo modelstudio grok stepfun minimax effort reasoning thinking muse: 1169 passed, 0 failed, 2 ignored (after the second commit)
    • One earlier run failed route_budget::tests::uncatalogued_remote_model_keeps_a_conservative_ceiling. It passes alone and on rerun. It reads output-limit env vars without taking the env lock.
  • cargo fmt --all -- --check: clean
  • clippy on -p codewhale-config -p codewhale-tui --all-targets --all-features with CI's flags: clean
  • check-blocking-calls-budget.py, check-dead-code-budget.py, check-command-crate-boundaries.py, split/module_graph.py --check: pass
  • Changelog chain (sync, derive-*, check-versions.sh --range-audit-advisory, contributor credit, public-copy.test.ts 6/6): pass

New tests cover:

  • a live MiMo, Alibaba-plan or DeepSeek row reads as unknown price (never free);
  • DeepSeek output clamped live and offline;
  • a signed fact still overrides a correction;
  • corrections never add or hide rows, and the loader refuses the forbidden shapes;
  • pricing_withheld round-trips through serde, and older payloads still parse;
  • every committed correction names a row the seed carries.

Refs #6396

🤖 Generated with Claude Code

@Hmbown Hmbown added this to the v0.10.1 milestone Sep 26, 2026
Copilot AI lite review requested due to automatic review settings September 26, 2026 09:01

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Hmbown pushed a commit that referenced this pull request Sep 26, 2026
# Conflicts:
#	crates/config/src/catalog/tests.rs
Hmbown pushed a commit that referenced this pull request Sep 26, 2026
Hmbown pushed a commit that referenced this pull request Sep 26, 2026
Hmbown pushed a commit that referenced this pull request Sep 26, 2026
Hmbown pushed a commit that referenced this pull request Sep 26, 2026
`seed lock` then `seed render`. The seed now matches what a networked
install already sees from live refresh, instead of a hand-kept copy:
- image input on the Claude, GPT-5.5/5.3-codex, Kimi K2.x, MiMo, Grok and
  OpenRouter Qwen/MiniMax rows;
- upstream limits (for example step-5-preview output 65536, glm-5.1 context
  200000) and reasoning ladders;
- prices where upstream has them. Held prices are withheld by
  corrections, which already applied to live rows.

Running the TUI tests against the generated seed found four things the old
seed handled only offline. They are fixed where they apply online too:
- Arcee trinity-large-thinking: Arcee publishes no cache rate, so cache
  hits bill at the input rate (reviewed table in tui pricing.rs); Models.dev
  says 0.06. Added as a pricing correction.
- `ModelsDevCatalog` readers bypass corrections: xAI effort dispatch read
  the raw seed, so grok-4.3 would have sent `reasoning_effort`. It now reads
  the corrected bundled row.
- A Models.dev `toggle` beside an effort ladder means reasoning can be
  switched off, but the effort clamp and the picker ignored it and rounded
  Off up to High. Both now keep Off (DeepSeek Ctrl+T and hotbar, and a
  neighbouring K3 gateway).
- Three earlier holds were restored as corrections in #6620: hosted
  DeepSeek pricing, Novita and native DeepSeek output 384000, GLM-5.3 on the
  Coding Plan.

Test expectations that pinned hand-seed values move to upstream's, each
with a comment: kimi-k3 output, glm-5.1 context, step-5 and step-3.5
limits, grok pdf input, qwen3.8-flash price, GLM-5.3's own ladder, and
DeepSeek V4 Pro's published effort ladder in the runtime API. The PR 1 and
#6612 tests that assumed a text-only seed row now build a stale row
themselves. MiMo's text-only assertion is removed, since that hold is
dropped.

Tests (focused):
- codewhale-config --lib: 709 passed, 0 failed, 1 ignored;
  --test configured_models: 9 passed
- codewhale-models: 44 passed
- codewhale-tui `-- vision image badge capabilit pricing provider_lake
  models_dev_live model_picker provider_catalog_live catalog deepseek xiaomi
  mimo modelstudio grok stepfun minimax effort reasoning thinking muse kimi
  glm qwen claude anthropic openai openrouter route_budget limit scorecard
  cost_status runtime_api provider_capability`: 1840 passed, 0 failed,
  2 ignored
- python3 -m unittest scripts/catalog_models_dev_test.py: 16 passed;
  seed render --check: ok
- cargo fmt --check clean; clippy -p codewhale-config -p codewhale-tui
  -p codewhale-models --all-targets --all-features with CI flags clean;
  blocking-calls, dead-code, command-crate-boundaries, module_graph and
  provider-registry checks pass; changelog chain passes (public-copy 6/6)

Refs #6396

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Base automatically changed from fix/6396-bundled-image-unknown to main September 26, 2026 14:30
CodeWhale Bot and others added 4 commits September 26, 2026 19:44
)

The offline seed carried Codewhale's deliberate holds as hand edits: prices
left out where a flat rate would mislead, DeepSeek's output limit kept at its
published 384K. A live Models.dev refresh replaces seed rows wholesale, so
every hold vanished on a networked install: live rows priced MiMo and
DeepSeek, listed Alibaba plan rows at $0 (read as free), and raised DeepSeek's
output limit to 393216.

The holds now live in `crates/config/assets/catalog_corrections.json` as field
patches in the cloud-facts `ModelFact` shape. The signed layer's patch code
applies them to every Models.dev row as it is hydrated, bundled seed and live
refresh alike. So a correction ranks above both Models.dev layers and below
signed cloud facts (which may still correct it) and provider rosters.

- `ModelFact.pricing_withheld` (optional reason) clears the row's price so the
  route reports `PricingSku::UnknownOrStale`. It is an optional field;
  payloads without it parse unchanged.
- `catalog_patch::apply_patches` is shared: the signed layer calls it with
  row materialization, corrections without, so they never add a row. The
  loader also refuses hide, deprecate, `allow_unlisted`, labels, a patch that
  changes nothing, and any limit patch without a reason.
- A corrected row keeps its own `source`. `CatalogCompiler::with_live` and
  other code sort rows by source, and a relabelled live row was placed above
  signed facts (an existing cloud-facts test caught it). The price a
  correction owns reports `CatalogSource::CodewhaleBundled` through
  `cost_source`.
- The declared-but-unused `CatalogCompiler.codewhale_bundled` field is
  removed; the layer docs (catalog.rs, CATALOG_REFRESH.md, CLOUD_FACTS.md)
  now describe the rank where corrections actually apply.

Each hold was checked again on its merits:
- kept: DeepSeek native pricing (time-aware table in tui pricing.rs), xAI
  (rates double past 200K), MiMo (PAYG and Token Plan keys look the same),
  the four Model Studio plan providers (quota), StepFun (plan and PAYG share
  ids), DeepSeek native output 384000 (published 384K; which value the API
  accepts is unverified, and the lower bound is always accepted);
- added: MiniMax-M3 (upstream price tiers at 512K, same rule as Grok; the
  seed already listed it unpriced without saying why);
- dropped: aggregator-hosted DeepSeek (an aggregator's published rate is its
  own price, as it is for every other aggregator row we price), OpenRouter
  dots :free (a free model is free), zai GLM-5.3 (the zai route is the
  pay-as-you-go API, not the Coding Plan the hold cited), MiMo's text-only
  image hold (made harmless by the previous commit) and MiMo's 1,000,000
  context (the only citation was our own rows).

Tests (focused):
- codewhale-config --lib: 706 passed, 0 failed, 1 ignored;
  --test configured_models: 9 passed
- codewhale-tui `-- vision image badge capabilit pricing provider_lake
  models_dev_live model_picker provider_catalog_live catalog deepseek xiaomi
  mimo modelstudio grok stepfun minimax`: 872 passed, 0 failed, 2 ignored
  (one earlier run failed
  route_budget::uncatalogued_remote_model_keeps_a_conservative_ceiling; it
  passes alone and on rerun, and reads output-limit env vars without the
  env lock)
- cargo fmt --check clean; clippy -p codewhale-config -p codewhale-tui
  --all-targets --all-features with CI flags clean; budget, boundary and
  module-graph gates pass; changelog chain passes (public-copy 6/6)

Refs #6396

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…6396)

Regenerating the offline seed from Models.dev (#6612) showed a second kind
of offline-only hold. The seed carried the reasoning controls Codewhale
relies on, but live Models.dev rows replace them:
- Grok: xAI documents `high` as the default effort (6be755f / #6501).
  Upstream lists the ladder with no default.
- MiniMax: M2.7 always thinks; M3 is adaptive or disabled, with a different
  default per dialect (47163da). Upstream lists a toggle or nothing.
- qwen3.8-max: thinking-only (the owner's console, ec4a5ae). Upstream
  lists a toggle.
- Muse Spark: also takes the `none` and `ultra` efforts (a50b653).
- Step 3.5: base model uses adaptive reasoning (StepFun guide, recorded
  2026-09-19). Upstream lists low/high.

The picker, `/effort` and the effort clamp read these, so today a networked
install loses them.

`ModelFact.reasoning_options` (optional) replaces a row's reasoning
controls and keeps any signed annotation already on the row. The 18 rows
above get it as a correction, each with its reason. GLM-5.3's high/max
ladder was left out: the seed only inherited it from GLM-5.2 as a
placeholder, and upstream now publishes low/high/max for 5.3. The Coding
Plan qwen `budget_tokens` entries were left out too: they were copies of
Token Plan rows, not a Codewhale decision.

Tests (focused):
- codewhale-config --lib: 707 passed, 0 failed, 1 ignored;
  --test configured_models: 9 passed
- codewhale-tui `-- vision image badge capabilit pricing provider_lake
  models_dev_live model_picker provider_catalog_live catalog deepseek xiaomi
  mimo modelstudio grok stepfun minimax effort reasoning thinking muse`:
  1169 passed, 0 failed, 2 ignored
- cargo fmt --check clean; clippy -p codewhale-config --all-targets
  --all-features with CI flags clean; changelog chain passes
  (public-copy 6/6)

Refs #6396

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The first commit dropped the aggregator-hosted DeepSeek price hold on the
grounds that an aggregator's published rate is its own price. Running the
TUI pricing tests against a seed generated from Models.dev (#6612) showed
the hold's real purpose: `crates/tui/src/pricing.rs` prices these routes
from a reviewed provider-docs table, and a catalog rate outranks that table.
For Fireworks V4 Pro, for example, the table has 1.74 per 1M input and
Models.dev has 1.20. The hold is restored for the nine hosted DeepSeek rows,
with that reason.

Also from the generated seed:
- Novita DeepSeek rows: max_output 384000. Models.dev lists 393216; this is
  the same published-384K question as native DeepSeek.
- grok-4.3: reasoning_options []. xAI documents no effort control for it,
  so Codewhale sends none (#6501); Models.dev lists a ladder.

All three are no-ops against the current hand seed. They take effect on
live refresh today, and offline once the seed is generated.

Tests (focused): codewhale-config --lib 707 passed, 0 failed, 1 ignored.
Changelog chain: public-copy 6/6.

Refs #6396

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The first commit dropped the GLM-5.3 price hold. Its stated reason was that
the `zai` route is the pay-as-you-go API, and that was wrong:
DEFAULT_ZAI_BASE_URL is api.z.ai/api/coding/paas/v4, the GLM Coding Plan,
which bills credit multipliers. It failed against a seed generated from
Models.dev (#6612): Models.dev lists 1.40/4.40, and
provider_cost_does_not_fabricate_price_for_costless_catalog_route caught
it. The hold is restored as a correction, so it now also holds online.

Tests (focused): codewhale-config --lib 707 passed, 0 failed, 1 ignored.
Changelog chain: public-copy 6/6.

Refs #6396

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@Hmbown
Hmbown force-pushed the fix/6396-catalog-corrections branch from bb5e68f to 56941d8 Compare September 27, 2026 02:55
Hmbown pushed a commit that referenced this pull request Sep 27, 2026
…receipt test

#6612 regenerated the offline seed from Models.dev, where Z.ai's GLM-5.3
lists low/high/max, and updated the routing tests on purpose (a Low
preference now reaches the wire as Low). The Auto dispatch receipt test in
tui::ui::tests still expected the old high/max ladder (low->high); #6612's
CI never ran the Rust suite because the PR targeted #6620's branch, so the
integration run was the first to see it. Also adds the #6612 release-note
receipt that Version drift requires.

Checks: cargo test -p codewhale-tui --lib -- auto_dispatch effort catalog
reasoning zai glm: 559 passed; 0 failed; check-versions.sh
--range-audit-advisory: receipts OK.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@Hmbown Hmbown closed this pull request by merging all changes into main in 0bfe04e Sep 28, 2026
@Hmbown
Hmbown deleted the fix/6396-catalog-corrections branch September 28, 2026 08:40
pull Bot pushed a commit to soitun/CodeWhale that referenced this pull request Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants