fix(catalog): Codewhale corrections hold online, not only offline (#6396) - #6620
Merged
4 commits merged intoSep 28, 2026
Merged
4 commits merged into
4 commits merged into
Conversation
Hmbown
pushed a commit
that referenced
this pull request
Sep 26, 2026
Hmbown
pushed a commit
that referenced
this pull request
Sep 26, 2026
Hmbown
pushed a commit
that referenced
this pull request
Sep 26, 2026
Hmbown
pushed a commit
that referenced
this pull request
Sep 26, 2026
Hmbown
pushed a commit
that referenced
this pull request
Sep 26, 2026
`seed lock` then `seed render`. The seed now matches what a networked install already sees from live refresh, instead of a hand-kept copy: - image input on the Claude, GPT-5.5/5.3-codex, Kimi K2.x, MiMo, Grok and OpenRouter Qwen/MiniMax rows; - upstream limits (for example step-5-preview output 65536, glm-5.1 context 200000) and reasoning ladders; - prices where upstream has them. Held prices are withheld by corrections, which already applied to live rows. Running the TUI tests against the generated seed found four things the old seed handled only offline. They are fixed where they apply online too: - Arcee trinity-large-thinking: Arcee publishes no cache rate, so cache hits bill at the input rate (reviewed table in tui pricing.rs); Models.dev says 0.06. Added as a pricing correction. - `ModelsDevCatalog` readers bypass corrections: xAI effort dispatch read the raw seed, so grok-4.3 would have sent `reasoning_effort`. It now reads the corrected bundled row. - A Models.dev `toggle` beside an effort ladder means reasoning can be switched off, but the effort clamp and the picker ignored it and rounded Off up to High. Both now keep Off (DeepSeek Ctrl+T and hotbar, and a neighbouring K3 gateway). - Three earlier holds were restored as corrections in #6620: hosted DeepSeek pricing, Novita and native DeepSeek output 384000, GLM-5.3 on the Coding Plan. Test expectations that pinned hand-seed values move to upstream's, each with a comment: kimi-k3 output, glm-5.1 context, step-5 and step-3.5 limits, grok pdf input, qwen3.8-flash price, GLM-5.3's own ladder, and DeepSeek V4 Pro's published effort ladder in the runtime API. The PR 1 and #6612 tests that assumed a text-only seed row now build a stale row themselves. MiMo's text-only assertion is removed, since that hold is dropped. Tests (focused): - codewhale-config --lib: 709 passed, 0 failed, 1 ignored; --test configured_models: 9 passed - codewhale-models: 44 passed - codewhale-tui `-- vision image badge capabilit pricing provider_lake models_dev_live model_picker provider_catalog_live catalog deepseek xiaomi mimo modelstudio grok stepfun minimax effort reasoning thinking muse kimi glm qwen claude anthropic openai openrouter route_budget limit scorecard cost_status runtime_api provider_capability`: 1840 passed, 0 failed, 2 ignored - python3 -m unittest scripts/catalog_models_dev_test.py: 16 passed; seed render --check: ok - cargo fmt --check clean; clippy -p codewhale-config -p codewhale-tui -p codewhale-models --all-targets --all-features with CI flags clean; blocking-calls, dead-code, command-crate-boundaries, module_graph and provider-registry checks pass; changelog chain passes (public-copy 6/6) Refs #6396 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
) The offline seed carried Codewhale's deliberate holds as hand edits: prices left out where a flat rate would mislead, DeepSeek's output limit kept at its published 384K. A live Models.dev refresh replaces seed rows wholesale, so every hold vanished on a networked install: live rows priced MiMo and DeepSeek, listed Alibaba plan rows at $0 (read as free), and raised DeepSeek's output limit to 393216. The holds now live in `crates/config/assets/catalog_corrections.json` as field patches in the cloud-facts `ModelFact` shape. The signed layer's patch code applies them to every Models.dev row as it is hydrated, bundled seed and live refresh alike. So a correction ranks above both Models.dev layers and below signed cloud facts (which may still correct it) and provider rosters. - `ModelFact.pricing_withheld` (optional reason) clears the row's price so the route reports `PricingSku::UnknownOrStale`. It is an optional field; payloads without it parse unchanged. - `catalog_patch::apply_patches` is shared: the signed layer calls it with row materialization, corrections without, so they never add a row. The loader also refuses hide, deprecate, `allow_unlisted`, labels, a patch that changes nothing, and any limit patch without a reason. - A corrected row keeps its own `source`. `CatalogCompiler::with_live` and other code sort rows by source, and a relabelled live row was placed above signed facts (an existing cloud-facts test caught it). The price a correction owns reports `CatalogSource::CodewhaleBundled` through `cost_source`. - The declared-but-unused `CatalogCompiler.codewhale_bundled` field is removed; the layer docs (catalog.rs, CATALOG_REFRESH.md, CLOUD_FACTS.md) now describe the rank where corrections actually apply. Each hold was checked again on its merits: - kept: DeepSeek native pricing (time-aware table in tui pricing.rs), xAI (rates double past 200K), MiMo (PAYG and Token Plan keys look the same), the four Model Studio plan providers (quota), StepFun (plan and PAYG share ids), DeepSeek native output 384000 (published 384K; which value the API accepts is unverified, and the lower bound is always accepted); - added: MiniMax-M3 (upstream price tiers at 512K, same rule as Grok; the seed already listed it unpriced without saying why); - dropped: aggregator-hosted DeepSeek (an aggregator's published rate is its own price, as it is for every other aggregator row we price), OpenRouter dots :free (a free model is free), zai GLM-5.3 (the zai route is the pay-as-you-go API, not the Coding Plan the hold cited), MiMo's text-only image hold (made harmless by the previous commit) and MiMo's 1,000,000 context (the only citation was our own rows). Tests (focused): - codewhale-config --lib: 706 passed, 0 failed, 1 ignored; --test configured_models: 9 passed - codewhale-tui `-- vision image badge capabilit pricing provider_lake models_dev_live model_picker provider_catalog_live catalog deepseek xiaomi mimo modelstudio grok stepfun minimax`: 872 passed, 0 failed, 2 ignored (one earlier run failed route_budget::uncatalogued_remote_model_keeps_a_conservative_ceiling; it passes alone and on rerun, and reads output-limit env vars without the env lock) - cargo fmt --check clean; clippy -p codewhale-config -p codewhale-tui --all-targets --all-features with CI flags clean; budget, boundary and module-graph gates pass; changelog chain passes (public-copy 6/6) Refs #6396 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…6396) Regenerating the offline seed from Models.dev (#6612) showed a second kind of offline-only hold. The seed carried the reasoning controls Codewhale relies on, but live Models.dev rows replace them: - Grok: xAI documents `high` as the default effort (6be755f / #6501). Upstream lists the ladder with no default. - MiniMax: M2.7 always thinks; M3 is adaptive or disabled, with a different default per dialect (47163da). Upstream lists a toggle or nothing. - qwen3.8-max: thinking-only (the owner's console, ec4a5ae). Upstream lists a toggle. - Muse Spark: also takes the `none` and `ultra` efforts (a50b653). - Step 3.5: base model uses adaptive reasoning (StepFun guide, recorded 2026-09-19). Upstream lists low/high. The picker, `/effort` and the effort clamp read these, so today a networked install loses them. `ModelFact.reasoning_options` (optional) replaces a row's reasoning controls and keeps any signed annotation already on the row. The 18 rows above get it as a correction, each with its reason. GLM-5.3's high/max ladder was left out: the seed only inherited it from GLM-5.2 as a placeholder, and upstream now publishes low/high/max for 5.3. The Coding Plan qwen `budget_tokens` entries were left out too: they were copies of Token Plan rows, not a Codewhale decision. Tests (focused): - codewhale-config --lib: 707 passed, 0 failed, 1 ignored; --test configured_models: 9 passed - codewhale-tui `-- vision image badge capabilit pricing provider_lake models_dev_live model_picker provider_catalog_live catalog deepseek xiaomi mimo modelstudio grok stepfun minimax effort reasoning thinking muse`: 1169 passed, 0 failed, 2 ignored - cargo fmt --check clean; clippy -p codewhale-config --all-targets --all-features with CI flags clean; changelog chain passes (public-copy 6/6) Refs #6396 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The first commit dropped the aggregator-hosted DeepSeek price hold on the grounds that an aggregator's published rate is its own price. Running the TUI pricing tests against a seed generated from Models.dev (#6612) showed the hold's real purpose: `crates/tui/src/pricing.rs` prices these routes from a reviewed provider-docs table, and a catalog rate outranks that table. For Fireworks V4 Pro, for example, the table has 1.74 per 1M input and Models.dev has 1.20. The hold is restored for the nine hosted DeepSeek rows, with that reason. Also from the generated seed: - Novita DeepSeek rows: max_output 384000. Models.dev lists 393216; this is the same published-384K question as native DeepSeek. - grok-4.3: reasoning_options []. xAI documents no effort control for it, so Codewhale sends none (#6501); Models.dev lists a ladder. All three are no-ops against the current hand seed. They take effect on live refresh today, and offline once the seed is generated. Tests (focused): codewhale-config --lib 707 passed, 0 failed, 1 ignored. Changelog chain: public-copy 6/6. Refs #6396 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The first commit dropped the GLM-5.3 price hold. Its stated reason was that the `zai` route is the pay-as-you-go API, and that was wrong: DEFAULT_ZAI_BASE_URL is api.z.ai/api/coding/paas/v4, the GLM Coding Plan, which bills credit multipliers. It failed against a seed generated from Models.dev (#6612): Models.dev lists 1.40/4.40, and provider_cost_does_not_fabricate_price_for_costless_catalog_route caught it. The hold is restored as a correction, so it now also holds online. Tests (focused): codewhale-config --lib 707 passed, 0 failed, 1 ignored. Changelog chain: public-copy 6/6. Refs #6396 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Hmbown
force-pushed
the
fix/6396-catalog-corrections
branch
from
September 27, 2026 02:55
bb5e68f to
56941d8
Compare
Hmbown
pushed a commit
that referenced
this pull request
Sep 27, 2026
…receipt test #6612 regenerated the offline seed from Models.dev, where Z.ai's GLM-5.3 lists low/high/max, and updated the routing tests on purpose (a Low preference now reaches the wire as Low). The Auto dispatch receipt test in tui::ui::tests still expected the old high/max ladder (low->high); #6612's CI never ran the Rust suite because the PR targeted #6620's branch, so the integration run was the first to see it. Also adds the #6612 release-note receipt that Version drift requires. Checks: cargo test -p codewhale-tui --lib -- auto_dispatch effort catalog reasoning zai glm: 559 passed; 0 failed; check-versions.sh --range-audit-advisory: receipts OK. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #6618. The base is that branch, so this diff shows only this slice. GitHub retargets it to
mainonce #6618 merges.What was wrong
Codewhale's deliberate catalog holds existed only as hand edits to the offline seed (
models_dev.bundled.json). They covered prices left out where a flat rate would mislead, and DeepSeek's output limit kept at its published 384K. A live Models.dev refresh replaces seed rows wholesale, so the holds disappeared on every networked install. Upstream today:$0, which reads as free.393216.The "honesty" rule in the seed's
_metawas prose that the runtime never read.The fix
The holds now live in
crates/config/assets/catalog_corrections.jsonas field patches in the cloud-factsModelFactshape. The signed layer's patch code applies them to every Models.dev row as it is hydrated, whether it comes from the bundled seed (bundled_catalog_offerings) or a live refresh (live_offerings_from_models_dev). A correction therefore ranks above both Models.dev layers and below signed cloud facts and provider rosters. The same correction holds online and offline.ModelFact.pricing_withheld: Option<String>(a reason) clears the price, so the route reportsPricingSku::UnknownOrStale. The field is optional, and payloads without it parse unchanged. Per-provider rules expand to one withheld-price patch per row.catalog_patch::apply_patchesis shared. Signed facts may create rows; corrections may not. The loader refuses:allow_unlisted;source. The price a correction owns reportsCatalogSource::CodewhaleBundledthroughcost_source. See the deviations section for why.CatalogCompiler.codewhale_bundledwas declared, but nothing had ever set it. The layer docs incatalog.rs,CATALOG_REFRESH.mdandCLOUD_FACTS.mdnow describe the rank where corrections actually apply.Reasoning controls (second commit)
Regenerating the seed in #6612 turned up a second kind of hold that works only offline: reasoning controls. Before this change, a live refresh replaced all of these, and the picker,
/effortand the effort clamp all read them:high(xAI docs, Thinking display: Grok 4.7 and Xiaomi 2.6 render incorrectly; audit all supported reasoning formats #6501).noneandultraefforts.ModelFact.reasoning_options(optional) replaces a row's controls and keeps any signed annotation already on the row. These 18 rows get corrections, each with the source it cites.Two groups were left out:
budget_tokensentries. They were copies of Token Plan rows, not a Codewhale decision.Each hold, checked again on its merits
crates/tui/src/pricing.rshas the time-aware tabletiersfield shows rates doubling past 200K0would read as freepricing.rshas endpoint-scoped ratesmax_output384000hosted_flash_and_pro_routes_price_from_bundled_docs_rates.crates/tui/src/pricing.rsprices these routes from a reviewed provider-docs table (Fireworks V4 Pro: 1.74 vs Models.dev 1.20), and a catalog rate outranks itmax_output384000reasoning_options: []dots-3-note-preview:freepricing0GLM-5.3pricingzaiwas the pay-as-you-go API. That was wrong:DEFAULT_ZAI_BASE_URLis the Coding Plan (api/coding/paas/v4), which bills creditsDeviations from the approved design
CatalogCompiler::compile. ButCatalogCompileris only used in tests; the runtime merge iscrates/tui/src/provider_lake.rs::compute_merged_snapshot. Applying corrections at hydration, the single funnel both Models.dev layers pass through, reaches the resolver, the provider lake, pricing fallbacks and the export path without adding a second merge.CodewhaleBundled. That moved a live Models.dev row across a layer boundary:CatalogCompiler::with_livesorts rows by source, so a relabelled row ranked above signed facts, andsigned_facts_correct_a_live_models_dev_row_but_never_the_provider_rosterfailed. Rows now keep their source, and only the owned price is attributed toCodewhaleBundled.Not in this PR
crates/models/assets/model_catalog.bundled.jsonduplicates some of these facts. Its module doc already schedules removal once themodels/pricingcompat callers migrate.Verification
CARGO_BUILD_JOBS=3 scripts/dev-cargo.sh test -p codewhale-config --lib: 707 passed, 0 failed, 1 ignored (after the second commit)CARGO_BUILD_JOBS=3 scripts/dev-cargo.sh test -p codewhale-config --test configured_models: 9 passedCARGO_BUILD_JOBS=3 scripts/dev-cargo.sh test -p codewhale-tui --lib -- vision image badge capabilit pricing provider_lake models_dev_live model_picker provider_catalog_live catalog deepseek xiaomi mimo modelstudio grok stepfun minimax effort reasoning thinking muse: 1169 passed, 0 failed, 2 ignored (after the second commit)route_budget::tests::uncatalogued_remote_model_keeps_a_conservative_ceiling. It passes alone and on rerun. It reads output-limit env vars without taking the env lock.cargo fmt --all -- --check: clean-p codewhale-config -p codewhale-tui --all-targets --all-featureswith CI's flags: cleancheck-blocking-calls-budget.py,check-dead-code-budget.py,check-command-crate-boundaries.py,split/module_graph.py --check: passcheck-versions.sh --range-audit-advisory, contributor credit,public-copy.test.ts6/6): passNew tests cover:
pricing_withheldround-trips through serde, and older payloads still parse;Refs #6396
🤖 Generated with Claude Code