Skip to content

feat(providers)!: AI SDK-shaped provider factories, request layer and replay tests - #214

Draft
cunninghamcard-bit wants to merge 114 commits into
rfc-0036/provider-result-typesfrom
rfc-0036/provider-factory
Draft

cunninghamcard-bit wants to merge 114 commits into
rfc-0036/provider-result-typesfrom
rfc-0036/provider-factory

Conversation

@cunninghamcard-bit

@cunninghamcard-bit cunninghamcard-bit commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

What

Moves provider construction, the shared request layer and their callers to package factories shaped by the AI SDK. It also adds the fixtures/aisdk samples, the factory replay tests and the migration documentation. Core, provider-utils, vendor packages, FFI, bindings, CLI, web console and replay change together because they share these interfaces. Base rfc-0036/provider-result-types (A); follows A, precedes C. Replaces #222, #223, #224 and the previous content of #214.

Layers

Branch PR Content Diff against its updated base
rfc-0036/factory-test-removal #222 Deletes four hand-written test files; no source or type changes 4 files, 1594 deletions(-)
rfc-0036/upstream-samples #223 Adds the fixtures/aisdk sample set 29 files, 3440 insertions(+)
rfc-0036/provider-factory-source #224 Factories, request layer, callers; restores and migrates upstream-ported tests 435 files, 36665 insertions(+), 24801 deletions(-)
rfc-0036/provider-factory #214 Factory replay tests, helpers and documentation; no production source changes 34 files, 4277 insertions(+), 387 deletions(-)

Before and after

Area Before After Upstream reference
Deleted tests Tests of registry construction and configuration, profile and body overrides, the realtime WebSocket stream, the Node-to-Rust HTTP path Deleted: aimux-providers/tests/config_snapshot_test.rs, reasoning_map_test.rs, openai_transcription_stream_test.rs, bindings/node/__test__/body_overrides.test.ts openai-provider.ts; provider request modules; openai-transcription-model.ts
Samples No fixtures/aisdk Reconstructed JSON inputs, requests, responses and results for Anthropic messages, Google generation and embeddings, OpenAI chat, responses and embeddings, compatible chat and embeddings; not verbatim upstream payloads @ai-sdk/anthropic, google, openai, openai-compatible
Sample versions No manifest VERSIONS.json: ai 7.0.127 with matching @ai-sdk/*; the local reference checkout for ported tests is 7.0.122 Package manifests
Construction Previous model construction paths and provider-specific setup Package create_* factories, settings and default instances create models through Provider. create_provider, provider_names, default_providers assemble factories and compatible presets (an aimux extension) createOpenAI, createAnthropic and other package provider modules
Docs Config builders and older by-name exports; Chat Completions as default; base-URL variables unread; reasoning option as reasoning_format Package settings and factories, create_provider(name, PresetSettings), provider_names(), default_providers(), Provider::discovery() (discovery is an aimux extension); option is provider_options.groq.reasoningFormat, converted to reasoning_format groq-chat-language-model-options.ts
OpenAI default FFI, Node, Python chose chat; Rust factory chose Responses Provider::language_model and call select Responses; chat selects Chat Completions explicitly; Responses factory supplies the file- prefix openai-provider.ts
Settings Extra resolvable keys or request transforms; supported hooks missing OpenAI, Anthropic, Google, Mistral, Cohere, Groq, DeepSeek, xAI and OpenAI-compatible use static Option<String> keys and header maps; env-key fallback loads at request time; Azure and PresetSettings keep resolvable keys, Azure also a token provider; Mistral, Cohere, Anthropic expose generateId Provider settings modules
Base URL, auth Docs implied every package loaded an env key and sent an empty key OpenAI, Anthropic: explicit base URL, then OPENAI_BASE_URL or ANTHROPIC_BASE_URL, then default; an explicit empty key does not fall back to env. OpenAI-compatible reads no env key and sends no authorization for a missing or empty key openai-compatible-provider.ts
Compatible requests URL support and usage conversion not suppliable; body transform reached other modalities OpenAICompatibleProviderSettings has supported_urls, convert_usage; transform_request_body applies to chat only; Groq and DeepSeek have own settings and conversion openai-compatible-provider.ts
Request execution Transport and retry mixed into providers provider-utils holds injectable fetch, resolvable values, response handlers, media-type detection; Core owns call retry; redirects followed only within the issuing origin; User-Agent suffixes append withUserAgentSuffix, postToApi
Azure auth Token resolved before header merge A caller Authorization header suppresses token resolution azure-openai-provider.ts
Registry, middleware No caller-assembled registry; incomplete wrapping create_provider_registry splits IDs at the first separator and applies language and image middleware; wrappers accept ID overrides; language hooks get generate and stream callbacks; NoSuchProvider carries provider ID, model ID, model type, available providers createProviderRegistry, wrapLanguageModel, wrapImageModel, NoSuchProviderError
Conversions Anthropic iteration metadata lost fields; Mistral merged choices; fallback file representations Anthropic keeps iteration model and cache fields, custom-name metadata conditional on options; Mistral reads the first choice and assembles tool-call fragments; Cohere and Mistral detect bare image media types; Bedrock converter rejects unsupported tool-content files anthropic-language-model.ts, mistral-chat-language-model.ts
Replay headers Replays removed user-agent Replays expect the suffixes with-user-agent-suffix.ts

Vendor option schemas live in each package: the shared OpenAI-style parser takes its schema from the caller; OpenAI, Azure, Groq, DeepSeek own theirs in options.rs; Anthropic spells its namespace and skill type once there. The files model exposes no specification version, as scripts/check_provider_boundaries.sh requires. TypeScript-only constraints (JSONSchema7, dotted IDs, URL, Date, numbers) are not enforced upstream at runtime; Rust keeps Value, String, u32.

Tests

  • Delete the four hand-written files above; no replacements. Cassette replays are kept.
  • Migrate the sixteen upstream-ported embedding, image, speech, transcription, files and video test files to the factory API. Restore the DeepSeek strict-option cases. Delete hand-written tests in anthropic/options.rs and provider-utils fetch.rs, and the hand-written DeepSeek namespace and env-key cases; no bedrock_factory_test.rs or ws_connector_test.rs remains.
  • The OpenAI, OpenAI-compatible, Anthropic and Google factory suites read fixtures/aisdk through the shared mock-fetch helper. The samples are recordings: scripts/aisdk-fixtures/record.mjs runs ai and @ai-sdk/* over a mocked fetch and writes the request sent and the normalised result. Requests and results are upstream behaviour; response bodies are fixed inputs. Replays check agreement with these records, not upstream test payloads.
  • The DeepSeek cassette replay is kept. The fetch-injection suite keeps its guarded-download test: validated downloads use the fixed guarded transport, a security decision. Hand-written preset and Node provider-config tests remain.
  • Cassettes are unchanged. Not every remaining hand-written test is removed or every upstream case ported.
  • create_provider with an unknown name reports every built-in name in NoSuchProvider.available_providers. Before dfe9a453 the miss fell through to preset::create, which lists only the registry rows, so a misspelled openai got a list without any of the 45 vendor packages. Unit test: default_providers::tests::an_unknown_name_lists_every_built_in_provider.

Not in this pull request

  • Serde naming, the by-name entries in aimux_providers::provider, and the call-layer ModelMessage tool-result shape: covered by D.
  • Samples stay as recorded. Aligning the recorder pin, covering the other eight packages and scheduled re-recording are follow-up work.
  • Google speech, transcription and evaluation; xAI speech, transcription and video; Groq transcription; Mistral speech and transcription; registry evaluationModel and customProvider: covered by C. xAI image and files, and OpenAI skills, are not ported.
  • Vertex application default credentials, Azure/Vertex transcription routing, Azure Foundry Responses explicit message-item typing, Groq first-SSE-error handling: covered by C.
  • specificationVersion stays omitted by the public-surface decision. Search, discovery and built-in assembly are aimux extensions, not parity claims.

Relation to the maintainer's design

Reference order: the pinned AI SDK source first; where it has no counterpart (C ABI, bindings, recording, code generation), the maintainer's RFCs 0039 (ops and bindings, #209), 0040 (provider construction and auth, #210), 0041 (replay governance, #211) and 0042 (V4 type boundaries, #212), which are Drafts that refine the #200 design doc docs/aisdk-architecture-alignment.md; that design doc only where those RFCs defer to it or are silent. Where an RFC restates upstream behaviour differently from the pinned source, upstream wins and the RFC gets an amendment request. The stack targets rfc-0036/integration; RFC-0039 §11 and RFC-0040 §7.2 require the whole closure to switch to master at once, so items marked "before master" must land on integration first. "Branch E" is the follow-up branch stacked on D, in progress.

  • Recording schema 3 (implemented in branch E). This PR sets RECORDING_SCHEMA to 3 with input, provider: ProviderRecord { provider_id, provider, model_id }, exchanges, outcome, complete, transport_closed. It drops the configuration snapshot, but it is a subset of RFC-0041 §4's schema 3: target (model reference and identity), operation, capture_policy and Captured values, model_calls, exchange id/phase/boundary/timing, ws_sessions and replayability are missing, and it identifies the provider through ProviderRecord instead of §4.1's target; RFC-0041 §2 drops ProviderRecord as a replay mechanism. Mock replay still decodes recorded OpenAI bodies in aimux-core/src/replay.rs (rebuild_generate_result) instead of RFC-0041 §3's ReplaySession with replay_fetch / replay_web_socket running the real provider converters. Branch E replaces it with RFC-0041's schema 3 and replay, so no file with this subset ships under the number 3.
  • Package descriptors and provider registration (implemented in branch E). RFC-0040 §2.1–2.2: package descriptor, ProviderRegistration (provider + extensions + metadata), pure-assembly providers_from_config(document). This PR has create_<vendor> factories, default_providers and create_provider(name, PresetSettings), which returns Arc<dyn Provider>; named methods such as chat, responses and completion exist only on the concrete provider types and are lost through it.
  • Fixture governance (with the schema 3 work in branch E). RFC-0041 §5 (Draft): fixture manifest with source=synthetic|approved_capture, redaction policy and review status. The 28 fixtures/aisdk samples carry only VERSIONS.json; the 2,799 cassettes keep their format.
  • One-shot breaking change (agrees). No forwarding, old-config or old-recording reading (RFC-0039 §0.1.9, RFC-0040 §7.2, RFC-0041 §2).

Verification

Gate on the head 419ac220: cargo fmt --check, cargo clippy --workspace --all-targets -D warnings, cargo doc with -D warnings, tests of core, provider-utils, providers, ffi, web, replay and cli: 2,437 passed, 0 failed, 6 ignored; Node and Python bindings cargo check; provider boundary script; gen_ts_types.py --check; gen_providers_doc.py --check. Cassettes unchanged; the 29 files under fixtures/aisdk are added. Diff against the base: 502 files changed, 44382 insertions(+), 26782 deletions(-). After the restack onto the current master, only this head was tested; the intermediate layers were checked with fmt and clippy only.

Restacked head 78c9a41f. Restack (2026-10-09): merged master (#213, SSE) and the rewritten #204 tracker down the stack by merge, no history rewrite. This layer ports the groq, deepseek, mistral and openai_compatible call sites to the new tracker API (the openai_compatible thought signature now travels as StreamingToolCallDelta::provider_metadata). Checked: cargo fmt --check, cargo clippy --workspace --all-targets -D warnings, tests of provider-utils and providers (107 test binaries, 0 failed). cargo doc, binding checks and the generator --check scripts were not rerun.

Fix dfe9a453 (merged into head 9bd3da58): the new unit test passed on provider-factory-source, on this head, and after merging into wire-format. presets_test and cargo clippy -p aimux-providers --all-targets -D warnings were run on provider-factory-source. The full gate above was not rerun.

Not verified

No GitHub check runs were attached at the last refresh. Go, Java, Kotlin, Swift and Flutter were not compiled locally in that refresh; Node npm test and Python pytest were not run. Live vendor calls and unported upstream cases are outside validation.

🤖 Generated with Claude Code

claude and others added 30 commits October 3, 2026 11:39
The AI SDK alignment design (docs/aisdk-architecture-alignment.md §9.3)
requires validating aimux providers against what the AI SDK sends and
returns for the same input under a mocked transport. This adds the SDK
half: scripts/aisdk-fixtures runs the official @ai-sdk packages
(ai 7.0.127, provider 4.0.21, provider-utils 5.0.53, openai 4.0.83,
openai-compatible 3.0.62, anthropic 4.0.71, google 4.0.87, pinned exactly)
with a recording fetch and writes fixtures/aisdk/<package>/<case>.json:
request (URL, headers with credentials redacted, body), canned response,
and the SDK's normalized result or thrown error.

28 cases across openai, openai-compatible (name "groq"), anthropic and
google pin the behaviours the Rust provider rewrite must match, among them:
an explicit empty apiKey is sent verbatim and never falls back to the env
var; a missing key fails at call time with AI_LoadAPIKeyError and no
request; createAnthropic({ name }) sets model.provider to the name verbatim
(default "anthropic.messages") and reads providerOptions under both the
canonical and the custom key, custom winning; openai-compatible derives
its providerOptions key from the first segment of model.provider and
spreads unknown fields under that key into the request body; header
merging is case-insensitive with call-level headers overriding provider
headers.

Rust tests consume these in the provider factory rewrite; regenerate with
`cd scripts/aisdk-fixtures && npm ci && node record.mjs`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
… call-level retry

AI SDK alignment design (docs/aisdk-architecture-alignment.md §3.2 / §4.5,
impact map S1-1, S1-2, S1-3, S1-5). First commit group
(A0) of the provider-factory rewrite: the primitives every createXxx(settings)
factory needs, with no change to the factories themselves yet.

Transport (S1-1)
- `Fetch` trait with `FetchRequest` / `FetchResponse` / `FetchError`;
  `ReqwestFetch` carries the former `shared_client()` machinery (per-runtime
  client sharding, pool settings, proxy config); `default_fetch()` is resolved
  per request, never frozen at factory time.
- `HttpRequest.fetch: Option<FetchFunction>` threads an injected transport
  through `post_*_to_api` / `get_from_api`; `send_request_once` sends through
  it. The SSRF download guard keeps its pinned-DNS client as `PinnedFetch`
  and ignores the injected fetch when `validate_url` is set (D26).
- `SigV4Fetch` decorator signs the final request bytes (same algorithm and
  test vectors as `bedrock/sigv4.rs`; providers switch to it in A4).
- `WsConnector` injection point on `WebSocketRequest`; `TungsteniteConnector`
  is today's behaviour.
- `catalogue.rs` no longer builds its own reqwest client.

Settings (S1-2, S1-3)
- `Resolvable<T>` (Value / Fn / AsyncFn / Future) and `HeadersFn`;
  `combine_headers` (lower-cased names, later layers override, `None`
  removes) and `normalize_headers`; requests insert headers instead of
  appending, so a credential header is never sent twice.
- `load_api_key(Some(s))` returns `s` verbatim, including the empty string;
  only `None` falls back to the environment (pinned by the AI SDK fixture
  `openai/empty-api-key`). Missing values are `AiMuxError::LoadApiKey` /
  `LoadSetting`; `load_setting` / `load_optional_setting` added. The FFI,
  Node and Python mappings route the two new variants to the existing
  invalid-argument code until A5 adds dedicated codes.

Retry (S1-5, D14)
- `prepare_retries(max_retries, abort)` with the constant defaults 2 /
  2000 ms / x2; `RetryConfig` is deleted along with every provider config
  field, `with_retry_config` builder and `retry_config()` trait method.
  Retry is a call-level concern: the nine core operations read only the
  caller's `max_retries`, provider-internal list/files/poll sites use the
  default budget for now (A4 applies the §4.5 table).

Not in this commit: RecordingFetch (S1-4), the Provider trait reshape (A1),
provider factories (A2–A4), FFI/binding error codes (A5).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…rding, and the boundary gate

Second commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.1, §3.3, §4.3, §7).

Provider trait (aimux-core/src/provider.rs)
- `Provider` now has the `ProviderV4` shape: language / embedding / image
  are required and return `NoSuchModel { model_type }` when a vendor has
  no such modality; transcription / speech / reranking / files are
  optional (`None` = not offered); video and search stay as aimux
  extensions. All constructors return `Arc<dyn Model>`.
- `Provider::name()` is gone: the registered name belongs to whoever
  holds the registry, and each model reports `provider()` itself.
- `specification_version()` is removed from the nine model traits and
  from `TraceLayer`; `LanguageModel::config_snapshot()` is removed and
  `supported_urls()` added.
- `list_models` moves to a separate `ProviderDiscovery` trait (no
  `ProviderV4` equivalent; most vendors cannot list). `delegate_list_models!`
  emits that impl; `provider_discovery(name, ..)` resolves a discovery
  handle for FFI / Node / Python.

Recording (RECORDING_SCHEMA = 3)
- `ProviderRecord` is identity only: `provider_id`, `provider`, `model_id`.
  Configuration (base URL, key source, profile, provider options, retry)
  is no longer recorded; schema-2 files are rejected on deserialization.
- `rebuild_provider` rebuilds from `provider_id` + `model_id` through the
  registry; `is_openai_compatible_provider` is deleted. Native-protocol
  packages return `NoSuchProvider` until they join the registry (A4).

Call-level overrides
- `CallOptions.body_overrides` is deleted everywhere (core, openai and
  anthropic converters, FFI, Node, Python, aimux-web playground, wire
  fixture). Provider-level overrides are applied to the finished body
  until `transform_request_body` replaces them (A3).
- `ProviderOptions` / `config_json` / Node `ProviderConfig` reject
  `max_retries` and `body_overrides` with `InvalidArgument` instead of
  silently ignoring them.

Boundary gate
- `scripts/check_provider_boundaries.sh` (wired into the contract-tests
  CI job) fails on provider code reading `options.max_retries` /
  `timeout` / `session_id` / `call_id` outside `HttpRequest::new`, and on
  any removed name coming back. Eleven hand-built `HttpRequest` literals
  now go through `HttpRequest::new`; the eight poll-loop providers use
  `prepare_retries(None, ..)` constants for their status GETs.

Tests: new `provider_trait_test.rs` in core and providers (shape, error
mapping, discovery); `config_snapshot_test.rs` deleted; overlay test now
proves the override through a wiremock request; body-override tests
exercise the provider-level path; Node `provider_config.test.ts` covers
the rejection.

Not yet done here (later groups): `api_key_source` fields and
`ExternalProviderEntry.max_retries` / `body_overrides` still exist
unread (A3); `provider_id` is the first segment of `provider()` until the
registry name fills it in (A3); Go / Java / Kotlin / Swift / Flutter still
declare a call-level `body_overrides` field that core ignores (A5).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Third commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.2, §3.3, §4.1, §7), modelled on
`createOpenAI` in @ai-sdk/openai 4.0.83 and checked against the recorded
fixtures under fixtures/aisdk/openai/.

Native OpenAI package
- `OpenAIProviderSettings { base_url, api_key: Option<Resolvable<String>>,
  organization, project, headers: Option<HeaderMapOpt>, name, fetch,
  transform_request_body }` and `create_openai(settings)`. The factory
  only validates `base_url` and fixes `name` (default "openai"); the key
  is loaded on every request: `None` reads OPENAI_API_KEY, `Some("")` is
  sent verbatim, a missing key fails the call with `LoadApiKey`, not the
  factory. `openai()` is the infallible `OnceLock` default instance.
- `OpenAIModelConfig` (crate-private, no getters, no snapshot) carries
  `provider`, a `url()` closure, an async `headers` resolver (credential,
  then organization/project, then user headers via `combine_headers`),
  `fetch`, `supported_urls`, `transform_request_body` and the base URL
  for the credentialed-origin guard. Chat, Responses, embedding, image,
  speech, transcription and files models read only this config.
- `provider()` is `"{name}.{method}"` (`openai.chat`, `openai.responses`,
  `proxy.chat` for `name: "proxy"`); the providerOptions namespace stays
  the fixed `openai` key, as upstream.
- `transform_request_body` is the only provider-level body hook; it runs
  once on the finished JSON body (multipart untouched). `OpenAIConfig`'s
  old `body_overrides` becomes such a closure at the compat boundary.
- `Resolvable::Future` is awaited once, `Resolvable::AsyncFn` on every
  request (counter tests).

Compat consumers (transitional until A3)
- `OpenAIConfig` and the new `OpenAIConfigProvider` move to
  `aimux-providers/src/openai_legacy.rs`; they keep the builder API for
  the 34 thin wrappers, the registry (`provider.rs`), Codex, xAI and
  Hugging Face, and feed the same `OpenAIModelConfig` through
  `into_model_config`. Removed from `OpenAIConfig`: getters, `from_env`,
  `with_api_key_source`, `with_body_overrides`, the `api_key_source`
  field. `apply_body_overrides` / `deep_merge_json` move to
  `body_merge.rs` for Anthropic until A4.
- Compat `provider()` strings are unchanged (`"groq"`), except files,
  which now report `"{provider}.files"`.

Constructor sites: the twelve `aimux_openai_*` FFI symbols, the Node and
Python OpenAI constructors, aimux-cli and aimux-web probes go through
`create_openai`; symbol names are unchanged. CLI / web cache probes
filter traces by `model.provider()`.

Boundary gate: rule 3 keeps `config.api_key` / `config.base_url` /
`body_overrides` / `api_key_source` / `from_env` out of
`aimux-providers/src/openai/`.

Also: an invalid header value no longer echoes the value in the error
(it could contain an API key).

Tests: new `openai_factory_test.rs` asserts every fixture in
fixtures/aisdk/openai/ (URL, method, headers, body, `provider()`, stream)
and fails when a fixture has no case; eleven openai test files moved to
`create_openai`; `body_overrides_test` covers `transform_request_body`.

Left for later groups: delete `OpenAIConfig` / `OpenAICompatProfile` /
`OpenAIConfigProvider` and the thin wrappers, move compat identity to
`"{name}.{method}"` with an explicit dialect flag (A3); Anthropic
`transform_request_body`, Azure `fetch` / `credentialed_origin`, Codex /
xAI / Hugging Face off `OpenAIConfig` (A4); `OPENAI_BASE_URL`, structured
env references, the recorded `"openai"` → `"openai.chat"` consumers and
the CHANGELOG entry (A5).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…and registry-generated presets

Fourth commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.1, D9, D12, D20), modelled
on @ai-sdk/openai-compatible 3.0.62 and checked against the recorded
fixtures under fixtures/aisdk/openai-compatible/.

openai-compatible (aimux-providers/src/openai_compatible/)
- `OpenAICompatibleProviderSettings { name, base_url, api_key, headers,
  query_params, fetch, include_usage, supports_structured_outputs,
  supports_multi_part_tool_content, transform_request_body }` and
  `create_openai_compatible`. `name` and `base_url` are required and
  fixed at the factory; the key resolves per request; `None` sends no
  `Authorization` (local servers).
- Chat, embedding and image models (the upstream set) with
  `provider() = "{name}.{method}"`. providerOptions are read from
  `openaiCompatible`, then the name, then its camelCase form; unknown
  fields under the provider's own namespace pass through to the body in
  the upstream position; metadata uses the name as key.
- Dialect hooks (`pub(crate)`) replace `OpenAICompatProfile`: usage
  conversion, tool and response-format preparation, file-part policy,
  `max_tokens_key`, stream usage key. No `Box::leak`.
- `list_models` is one request, no retry.

Groq and DeepSeek (groq/, deepseek/)
- Own packages on the compat internals: `create_groq` / `groq()` with
  `x_groq` stream usage, no `top_k`, `max_completion_tokens`, Groq
  structured-output rules, browser_search and the Groq file-part
  rejection; `create_deepseek` / `deepseek()` with cache-read usage.
  `provider()` is `groq.chat` / `deepseek.chat`; providerOptions and
  metadata keys are `groq` / `deepseek`. The eight `== "groq"` branches
  in openai/convert.rs are gone; the native OpenAI package knows nothing
  about other vendors.

Presets (scripts/gen_presets.py → aimux-providers/src/presets/, committed)
- `provider_registry.json` now has 283 rows (251 + the 32 former wrapper
  vendors) with `auth: api_key | none`, `base_url_env`, `params` with
  `{param}` templates (env / default / derived host maps for Vertex
  locations) and `family`. Each row generates a `PresetDescriptor`,
  `create_<name>(PresetSettings)` and an infallible `<name>()`.
- `auth: none` resolves no key and injects no placeholder; there is no
  `PLACEHOLDER_API_KEY` any more. Template parameters are only the
  declared ones, host parameters reject `/`, `@`, `:`, `?`, and an
  unexpanded placeholder is an error. Unknown preset names are
  `NoSuchProvider`; nothing falls back to OpenAI.
- `provider.rs` resolves names through the preset table and the overlay
  table; `ExternalProviderEntry` rejects `max_retries` and
  `body_overrides`. `gen_presets.py --check` runs in the contract-tests
  CI job; `gen_providers_doc.py` reads the new columns.

Deleted: `openai_legacy.rs` (`OpenAIConfig`, `OpenAIConfigProvider`),
`OpenAICompatProfile`, the 22 local/cloud thin wrappers (incl.
bedrock_mantle, openrouter) and the 10 `vertex_ai_*_models.rs` files.
Codex, xAI, Hugging Face and Azure compile through a crate-private
`StaticBearerConfig` until their own factories (A4).

Behaviour changes recorded here: the native OpenAI package warns on and
drops `top_k` (upstream does not send it); registry and preset providers
expose chat / embedding / image only; compat providers no longer read
the `openai` providerOptions key; `providerOptions.deepseek` is honoured
(`reasoningEffort`, `thinking`); `litellm_proxy` reads
`LITELLM_PROXY_BASE_URL` (the old wrapper read the API-key variable as a
URL).

Boundary gate: rule 4 forbids provider-name comparisons inside the
shared packages, rule 5 keeps the removed names removed.

Tests: `openai_compatible_factory_test.rs` (every compat fixture, plus
a directory-listing guard), `presets_test.rs` (all 283 rows create
without env, keyless rows send no Authorization, template and parameter
errors, Vertex hosts), `deepseek_test.rs`, the Groq package module;
about twenty test files moved from wrapper types to presets. Cassette
recordings are untouched.

Left for later groups: Codex / xAI / Hugging Face / Azure factories and
removal of `StaticBearerConfig` (A4); Anthropic off `body_merge.rs`
(A4a); `params` in the FFI / Node / Python configs, bindings tests on
provider names, docs and the CHANGELOG entry (A5).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…reuse the Anthropic core

Fifth commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.1, D-d, D-h, D-i),
modelled on `createAnthropic` in @ai-sdk/anthropic 4.0.71 and checked
against the recorded fixtures under fixtures/aisdk/anthropic/.

anthropic/
- `AnthropicProviderSettings { base_url, api_key, auth_token, headers,
  name, fetch, transform_request_body }`, `create_anthropic` and the
  infallible `anthropic()`. The factory validates `base_url` and rejects
  `api_key` together with `auth_token`; the credential resolves per
  request (`None` reads ANTHROPIC_API_KEY, `Some("")` is sent verbatim,
  `auth_token` sends only `Authorization: Bearer`). A missing key fails
  the call with `LoadApiKey`.
- Private `AnthropicModelConfig` with hooks (URL, body, headers, error
  handler, feature flags) feeds one `AnthropicMessagesModel`; the files
  model reads the same config.
- `provider()` is the name verbatim: `anthropic.messages` by default,
  `proxy` for `name: "proxy"`; files derive `{name minus .messages}.files`.
- providerOptions are read from the canonical `anthropic` key merged
  with the custom first segment (custom wins); response metadata is
  written under the custom key. The twenty hard-coded `"anthropic"`
  sites go through `anthropic/options.rs`.
- Base URL follows upstream: it includes `/v1` and the endpoint is
  `{base}/messages`; only the bare `https://api.anthropic.com` is
  rewritten. `ANTHROPIC_BASE_URL` is no longer read (no `from_env`).
- Request body: `stream` is omitted on non-streaming calls;
  `providerOptions.anthropic.metadata.userId` maps to `metadata.user_id`;
  `anthropic-beta` is sent on every host; `supported_urls` declares the
  upstream image and PDF patterns.
- `list_models` is one exchange without retry.

anthropic_aws/
- `AnthropicAwsProviderSettings` with `ApiKey(Resolvable)` or
  `SigV4(Resolvable<AwsCredentials>)`; signing is now the `SigV4Fetch`
  transport decorator (service `aws-external-anthropic`) over the final
  bytes, the model no longer signs, and `list_models` works with SigV4.
  The provider string stays `anthropic-aws`.

vertex/anthropic_model.rs
- `VertexAnthropicModel` is `AnthropicMessagesModel` with Vertex hooks
  (rawPredict / streamRawPredict URL, `anthropic_version` in the body,
  bearer or x-goog-api-key headers, Google error handler, empty
  `supported_urls`, structured outputs and strict tools off); 591 → 117
  lines. Provider string `googleVertex.anthropic.messages`.

Deleted: `AnthropicConfig`, `AnthropicConfigBuilder`,
`AnthropicAwsProviderConfig`, `api_key_source`, `from_env`, `with_*`,
and the family's `body_overrides` (replaced by `transform_request_body`).
Constructor sites in aimux-ffi, Node, Python, aimux-cli and aimux-web go
through the factories; symbol names are unchanged. Boundary gate rule 6
covers the Anthropic family.

Tests: `anthropic_factory_test.rs` replays every fixture in
fixtures/aisdk/anthropic/ (URL, method, headers, body, `provider()`,
stream) with a directory-listing guard, plus credential timing, custom
namespace, SigV4-over-final-bytes and Vertex-envelope tests; fifteen
test files moved to the factories. Cassette recordings are untouched.

Left for later groups: Vertex factory with real ADC / express headers
and the `googleVertex` namespace (4b); Bedrock on `SigV4Fetch` and the
`bedrock/sigv4.rs` shim removal (4b); files upload retry (4d);
`body_merge.rs` and its two tests, the base-URL release note and the
result-level providerMetadata (A5).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…on_bedrock factories

Sixth commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.1, D-h, D11), modelled on
@ai-sdk/google 4.0.87, @ai-sdk/google-vertex and @ai-sdk/amazon-bedrock
and checked against the recorded fixtures under fixtures/aisdk/google/.

google/
- `GoogleProviderSettings`, `create_google`, the infallible `google()`;
  chat, embedding, image, video and files models read a private
  per-model config (`shared/exchange.rs`). The key resolves per request
  (`x-goog-api-key`, GOOGLE_GENERATIVE_AI_API_KEY); `provider()` is the
  name (`google.generative-ai`) with `{name}.files` / `{name}.video`.
- Request bodies now match upstream: `generationConfig` is always sent,
  function tools go out as `parametersJsonSchema`, `providerOptions.google`
  (thinkingConfig, responseModalities, audioTimestamp, mediaResolution,
  imageConfig, safetySettings, cachedContent, labels, serviceTier,
  retrievalConfig) is mapped; response `modelId` comes from
  `modelVersion`. providerOptions are read from and written to `google`
  only (`google/options.rs`). `list_models` and the files upload are
  single exchanges.

vertex/
- `VertexProviderSettings { api_key (express), project, location,
  base_url, headers, access_token, fetch, transform_request_body }`,
  `create_google_vertex`, `google_vertex()`. Express versus standard
  mode, project / location (`LoadSetting`) and the bearer token
  (`GOOGLE_VERTEX_ACCESS_TOKEN`, or any `Resolvable`) resolve per
  request; location must be one DNS label and the host follows the
  upstream global / rep / regional rule; Gemini models use `v1beta1`,
  Anthropic-on-Vertex `v1`. `provider()` is `google.vertex` (plus
  `.video`, `.transcription`, and `googleVertex.anthropic.messages`).
  providerOptions: `googleVertex` (canonical only; the legacy `vertex`
  key is neither read nor written), then `google` for the shared Gemini
  model.

bedrock/
- `AmazonBedrockProviderSettings { region, api_key, access_key_id,
  secret_access_key, session_token, base_url, headers, fetch,
  credential_provider }`, `create_amazon_bedrock`, `amazon_bedrock()`.
  AWS_BEARER_TOKEN_BEDROCK selects bearer auth; otherwise `SigV4Fetch`
  signs the final bytes with credentials resolved per request
  (`LoadSetting` for AWS_REGION / AWS_ACCESS_KEY_ID /
  AWS_SECRET_ACCESS_KEY). `provider()` is `amazon-bedrock` for chat,
  embedding, image and reranking; providerOptions use `amazonBedrock`
  only (the legacy `bedrock` key is neither read nor written).
  `bedrock/sigv4.rs` is deleted; `aws_polly` signs through `SigV4Fetch`.

Deleted: `GoogleConfig`, `VertexProviderConfig`, `VertexAuth`,
`BedrockProviderConfig`, `BedrockAuth`, their `from_env` / `with_*` /
`api_key_source`. Constructor sites in aimux-ffi (14), Node, Python,
aimux-cli and aimux-web go through the factories; symbols unchanged.
Boundary gate rule 7 keeps the builder-era names, `bedrock::sigv4` and
stray namespace literals out.

`anthropic_aws` keeps its name: it is Claude Platform on AWS
(`aws-external-anthropic`), not Bedrock InvokeModel; aimux has no
Bedrock-Anthropic InvokeModel provider and this PR does not add one.

Tests: `google_factory_test.rs` (every fixture, directory-listing guard),
`vertex_factory_test.rs` (express and standard paths, three hosts,
invalid location, missing settings), `bedrock_factory_test.rs` (bearer
without signature, SigV4 over the final bytes, provider strings), a
shared scripted `Fetch` mock in tests/common; about twenty-eight test
files moved to the factories. Cassette recordings are untouched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…Face, Codex, open_responses, Voyage and ElevenLabs

Seventh commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.1, D-d, D11, D13), each
package modelled on its @ai-sdk counterpart where one exists.

- Every package gets `XxxProviderSettings`, `create_xxx`, an infallible
  `xxx()` and a private per-model config; credentials resolve per
  request through the shared helpers (`None` reads the package's env
  var, `Some("")` is sent verbatim, a missing key fails the call with
  `LoadApiKey`). `provider()` strings follow upstream: `azure.chat` /
  `azure.responses` / `azure.embeddings` / `azure.image` /
  `azure.transcription` / `azure.speech`, `xai.responses`,
  `mistral.chat` / `mistral.embedding`, `cohere.chat` /
  `cohere.textEmbedding` / `cohere.reranking`, `huggingface.responses`,
  `codex.responses`, `{name}.responses` for open_responses,
  `voyage.embedding` / `voyage.reranking`, `elevenlabs.speech` /
  `elevenlabs.transcription`. All but Azure accept a `name` override.
- Azure reuses the OpenAI chat / Responses / embedding / image /
  transcription / speech models through the injected config: `api-key`
  or bearer token (`token_provider` AsyncFn) per request, resource name
  from AZURE_RESOURCE_NAME (`LoadSetting`), upstream defaults
  (`api_version = "v1"`, `/v1{path}?api-version=`), deployment-based
  URLs on request. The Responses model reads providerOptions `azure`
  then `openai` and writes `azure`.
- xAI and Hugging Face default to Responses (as upstream); their
  chat-completions models remain reachable as the `chat_completions(id)`
  extension. Codex keeps RFC-0018's two modes (`ApiKey` with
  CODEX_API_KEY fallback, `ChatGptAccount { token, account_id }`), the
  stateless refresh, 401 → `TokenExpired`, and `store: false` as a
  package rule. ElevenLabs' WebSocket handshake headers come from the
  resolved headers and the connector is injectable. open_responses takes
  `base_url` and `Resolvable` headers; `api_key: None` sends nothing.
- Deleted: the nine `XxxConfig` types and builders, `AzureAuth`,
  `TokenProvider`, `AzureModel`, `azure/{model,responses}.rs`,
  `api_key_source` / `from_env` in these packages, and the transitional
  `StaticBearerConfig`. `OpenAIModelConfig` now anchors credentialed
  headers to the request URL (Azure's host is per request) and carries
  a `ResponsesProfile`; namespace literals live in each package's
  `options.rs`.
- Constructor sites in aimux-ffi (12), Node (6), Python (6), aimux-cli
  and aimux-web go through the factories; symbols unchanged. Boundary
  gate rule 8 covers these packages.

Not ported from upstream Azure: the deepseek and completion models, the
MAI / speech endpoints and the Foundry item type.

Tests: `vendor_factories_test.rs` (28 tests over a scripted `Fetch`:
credential timing, empty key, default instances, provider strings per
package) and 27 adapted test files. Cassette recordings are untouched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…ge-constant poll loops; uploads are not retried

Eighth commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.5, §9.3).

- Every single-modality package (serper, prodia, deepgram, you_com, luma,
  lmnt, klingai, dataforseo, replicate, fal, recraft, aws_polly, jina_ai,
  gladia, tavily, linkup, tinyfish, black_forest_labs, assemblyai, revai,
  hume, cartesia, searxng, parallel_ai, firecrawl, runwayml, exa_ai,
  stability, google_pse) gets `XxxProviderSettings`, `create_xxx`, an
  infallible `xxx()` and `XxxProvider` on the shared `EndpointConfig`;
  credentials resolve per request (missing key → `LoadApiKey`, missing
  URL or AWS setting → `LoadSetting`). `provider()` is `"{name}.{method}"`
  (`luma.image`, `deepgram.transcription`, `tavily.search`,
  `amazon-polly.speech`, `blackForestLabs.image`, …). providerOptions
  keys live in one `options.rs` per package. The old `XxxConfig`,
  `from_env` and `with_*` are gone; no `*Config` struct remains in
  aimux-providers.
- Poll loops (luma, black_forest_labs, fal, gladia, assemblyai, revai)
  run on package constants through `shared/poll.rs`: a fixed interval
  and attempt cap (previously unbounded loops now stop after 6000 × 100
  ms), a transient poll error spends one attempt, exhaustion returns the
  last error or `Timeout`, cancellation is checked before every wait.
  The job-creating request is sent exactly once; nothing re-submits.
  Only the interval can be overridden, through the package namespace
  (`pollIntervalMillis` for luma as upstream, `pollIntervalMs`
  elsewhere); pacing keys are stripped before the body is forwarded.
  Download stages keep a bounded three-try retry. `aimux_core::retry`
  is no longer used anywhere in aimux-providers (boundary rule 9).
- Files uploads (OpenAI, Anthropic, Google) are single exchanges.
- Constructor sites (`aimux_tavily_search_new*`, Node and Python
  Tavily) go through the factory; symbols unchanged. Boundary rule 10
  covers the 28 packages.

Behaviour changes recorded here: cartesia `version` is a fixed header
(override via `headers`); runway poll pacing is constant; google_pse
`cx` precedence is settings, providerOptions, then GOOGLE_CSE_ID per
request; DataForSEO missing credentials report `LoadApiKey`; searxng
accepts an optional bearer token.

Tests: `single_modality_factories_test.rs` (defaults, missing key,
provider strings), `poll_stage_test.rs` (exhaustion never re-submits per
package, interval override, abort, ignored attempt keys, stripped pacing
keys), upload-not-retried tests in the three files packages, unit tests
in `shared/poll.rs`; thirty-five test files converted. Cassette
recordings are untouched.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…binding configs and error classes, CHANGELOG and docs

Ninth and last commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §5, §6.2, §8, §9.2, §9.4).

C ABI
- `AIMUX_E_LOAD_API_KEY` (18) and `AIMUX_E_LOAD_SETTING` (19) for
  `AiMuxError::LoadApiKey` / `LoadSetting`; the env var travels in
  `aimux_error_provider_code` and the description or setting name in
  `aimux_error_provider_message`. `aimux-error.h` / `aimux-ffi.h` synced,
  with a header-versus-source parity test. `provider_handle_new` and
  `register_providers` reject `max_retries` / `body_overrides` in
  `config_json`.

Node / Python
- `LoadAPIKeyError` and `LoadSettingError` classes (`envVar` /
  `description` / `settingName`; snake_case in Python). Every Node native
  constructor goes through one `native_config()` that rejects
  `maxRetries` / `bodyOverrides` and applies `headers`;
  `ProviderConfig.params` carries preset template parameters.
  `index.d.ts` regenerated; both bindings are fmt- and clippy-clean; the
  ava and pytest suites pass.

Go / Java / Kotlin / Swift / Dart
- Codes 18 / 19 with an `EnvVar` field; the call-level `body_overrides`
  field (which core no longer reads) is deleted; Go and Dart
  `ProviderConfig` drop `max_retries` / `body_overrides` and gain
  `params`. Go and Java suites run here; Kotlin, Swift and Dart edits
  are symmetrical and unverified in this container.

Residue
- `body_merge.rs` and `body_overrides_test.rs` are gone
  (`transform_request_body_test.rs` replaces them); `hmac` dropped,
  `sha2` / `hex` are dev-dependencies; Codex subscription default URL
  fixed to `…/backend-api/codex`; Bedrock `list_models` host fixed to
  `bedrock.{region}.amazonaws.com`; Anthropic result-level
  `providerMetadata` (usage, stopSequence, container) implemented and
  checked against the fixtures; boundary gate also forbids
  `api_key_source` and the body-merge helpers.

Docs
- CHANGELOG `[Unreleased]` gains the Rust, C ABI, Node / Python and
  other-binding breaking lists for the whole rewrite. README,
  CONTRIBUTING, docs/api/*, docs/API.md, the provider-config manual, the
  Codex guide, the error model, the request-pipeline doc (retry is a
  call-level constant) and `rfc/0036-aisdk-architecture-alignment.md` §5
  (no `specification_version` / `name()`, `.call()` form, OpenAI default
  stays chat, Bedrock InvokeModel and unported Azure models, video
  `poll_config`, native `rebuild_provider`) are updated.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
… in §5

Same change as the docs/rfc-0036-aisdk-alignment follow-up (no ABI
coexistence promise after the cutover; §3.3 / S4-7 wording for the
xAI and Hugging Face Chat Completions extension), plus the §5 product
difference row for `chat_completions(id)` that this branch introduces.
The shared part becomes a no-op once the docs PR lands on master.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…s models

The AI SDK's xAI package (5.0.12) and Hugging Face package serve the
Responses API only. The Chat Completions models that 7c54cfd kept as a
`chat_completions(id)` extension outside the `Provider` trait are removed:
`XaiModel`, its wire types and chat converter, and the Hugging Face chat
constructor. `xai/convert.rs` keeps only the helpers the Responses
converter uses.

- Tests of the removed model go with it (the chat-level modules of
  `xai_test.rs`, the xAI / Hugging Face rows of the OpenAI-compatible
  suites). The xai-error and xai-provider cases now run against
  `responses(id)`. The 16 Hugging Face cassettes are unchanged and replay
  through the OpenAI-compatible package in `conformance_test.rs`.
- The xAI Responses model sends `top_k` and warns for `frequencyPenalty` and
  `presencePenalty`, as upstream does (found while moving the
  unsupported-parameter case).
- docs / ROADMAP: same wording change as the docs PR follow-up (the
  extension is withdrawn; superseded ROADMAP lines rewritten), and the §5
  product-difference row for the extension is dropped.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…in only for API calls

The pooled client followed redirects with reqwest's default policy, which
strips only Authorization-style headers on a cross-host hop, so vendor
credential headers (x-api-key, x-goog-api-key, api-key,
x-amz-security-token, ...) were forwarded to the redirect target, and a
SigV4 request was replayed with its old signature.

Redirect handling now lives in one place, the hop-by-hop loop that the
validated-download path already had in http.rs (#163):

- An ordinary API call follows a redirect only while it stays on the
  origin of the hop that issued it; a cross-origin 3xx is returned as a
  regular non-2xx response. Nothing is sent to the other origin, neither
  the request headers nor what a transport decorator would add.
- Every followed hop goes through the request's transport again, so a
  signing transport signs the URL it actually sends.
- A transport never follows a redirect: the pooled client is built with
  no redirect policy, and RedirectPolicy / FetchRequest.redirect /
  PinnedFetch::unpinned are removed. An injected Fetch gets the same
  protection because the rule is above it.
- Validated downloads behave as before (per-hop validation and pinning,
  caller headers stripped off the credentialed origin).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…pi_key in ExternalProviderEntry Debug

The Bedrock region was interpolated into the request host unchecked, so a
crafted value could change the host. It now must be a single DNS label (the
helper the Vertex location check uses) and fails with InvalidArgument before
any request is sent.

ExternalProviderEntry no longer derives Debug: a manual impl lists the same
fields but reports only whether api_key is set.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…tten factory tests

Kept: the tests that replay recorded AI SDK fixtures (fixtures/aisdk/*) in the
OpenAI, OpenAI-compatible, Anthropic and Google factory tests, with their
fixture-directory guard tests and only the helpers they use; one poll test
proving the job-creating request is sent once when polling is exhausted.

Dropped: the hand-written factory unit tests, the vendor/single-modality/
Vertex/transform-request-body/provider-trait test files, the other poll-stage
tests, and the helpers and imports that became unused. Fixture replays and
end-to-end behavior are the coverage we want; hand-written unit tests for
factories are not.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…istry; drop the generated sources

`aimux-providers/src/presets/` (26 generated files, 12,774 lines) and
`scripts/gen_presets.py` are removed. `provider_registry.json` is embedded
and parsed once into the descriptor table (`preset::entries` / `names` /
`lookup`); a preset is created by name with
`PresetProvider::create(descriptor, settings)`. The per-name Rust
functions (`presets::create_<name>`, `presets::<name>()`) are gone with no
replacement, so the registry data is no longer stored twice.

- Every check the generator made on a row (key whitelist, auth / env_var
  pairing, family and auth values, template parameters matching the URL
  placeholders, duplicate names) now runs when the table is first loaded;
  an invalid registry panics there and fails the table test.
- `presets_test.rs` keeps the table test (all 283 rows load and create
  without reading the environment) and one end-to-end case each for a
  keyless row, an unknown name and a rejected template parameter.
- CI no longer runs `gen_presets.py --check`; docs and CHANGELOG describe
  the runtime table. docs/aisdk-architecture-alignment.md carries the
  same D12 / D20 wording as the docs PR follow-up.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Keeps SigV4 signing of the final request bytes, the bearer-token path and
the region host-label check; the remaining hand-written factory cases are
dropped like the other factory test files.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Takes the design docs as merged in #200 (including the maintainer's review
fixes) and keeps this branch's follow-ups on top of them:

- xAI / Hugging Face expose Responses only; the Chat Completions extension
  recorded in the merged doc is withdrawn (§3.3, S4-7, revision note).
- Presets are a runtime descriptor table (D12, D20, §0.3, §0.7, §4.1, §5,
  §6.5).
- §3.2: redirects are handled by the helper, same-origin only for API
  calls; `FetchRequest` has no redirect field.
- §0.3 drops NDJSON from aimux-stream; the reference-baseline header notes
  that fixtures are pinned by fixtures/aisdk/VERSIONS.json.
- ROADMAP: the lines that still carried superseded commitments are
  rewritten (C2 shim, 0.7 / 0.8 version semantics, C1 coexistence,
  #174 / #175, #167, C4). Where the merged text already annotated a line
  (B1 + C2, #185, the Part II ts-rs sentence) the merged wording is kept.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: `create_provider_registry(providers, options)` returns a
`ProviderRegistry` over a caller-assembled map of provider ids to
providers. A model is addressed as "{provider}:{model}"; the id is split at
the first separator and the rest is passed to the provider unchanged. One
method per modality of the `Provider` trait, plus `files(provider_id)`.

Why: this is the one by-name mechanism the AI SDK has
(`createProviderRegistry` in `ai/src/registry/provider-registry.ts`). Until
now the workspace answered "name -> provider" in three separate places
that cover different vendors (the `provider(name, ..)` function for registry
presets, hand-written matches in the CLI and web tools for native vendors,
and per-vendor FFI constructors). The following commits move all of them
onto this registry.

Behaviour, as upstream: an id without the separator is `NoSuchModel`, an
unregistered provider is `NoSuchProvider`, a modality the provider does not
offer is `NoSuchModel`, the separator is configurable. The registry knows
no provider on its own and never falls back.

Not ported yet: language / image middleware and the `skills` accessor.

Tests: six cases ported from provider-registry.test.ts (ai@7.0.127).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…, default_providers)

What:
- `create_provider(name, settings)`: a name selects a factory, and the same
  few settings (key, base URL, extra headers, transport) go to whichever
  factory it is. The name is either a vendor package (`openai`, `anthropic`,
  `google`, ... 16 of them) or a row of `provider_registry.json`, which goes
  to the OpenAI-compatible factory. A package wins over a row of the same
  name. An unknown name is `NoSuchProvider`; there is no fallback.
- `default_providers()`: every name with default settings, as the map
  `create_provider_registry` takes. Nothing is read from the environment
  and no entry can fail to be created; keys are evaluated per request.
- `preset::create(name, settings)` returns an ordinary
  `OpenAICompatibleProvider` for a registry row.
- `Provider::discovery()` (aimux extension, default `None`): the
  `/models` side of a provider is reachable from the provider, so a
  registry entry needs no second handle for it. A reference to a provider
  is a provider, so `&'static` default instances register as they are.

Why: the AI SDK ships no vendor list. The reference for a built-in list on
top of it is models.dev (what opencode consumes): one row per vendor naming
the factory to use and what to pass it, native and compatible vendors in
the same table. Until now only the compatible vendors were reachable by
name, so the CLI, the web tool and the FFI each carried their own mapping
for the native ones. This is that mapping, once.

Tests: one end-to-end case, a registry over `default_providers()` serving a
package, a registry row and a name both have.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
… create_provider

What: `build_model` in `aimux-cli` (probe) and in `aimux-web` is now one
call, `create_provider(name, settings)?.language_model(model_id)`. The web
console's provider list is the keys of `default_providers()`.

Why: both tools carried the same hand-written match (openai / anthropic /
google / mistral / xai / cohere each to its own factory, everything else to
the by-name `provider()` function) because the by-name function could not
reach the vendor packages. With one table for every built-in provider the
match has nothing left to do. 164 lines removed, 26 added.

Behaviour change: a missing key for a vendor package is no longer checked
when the model is built. It fails the request that needs it, with
`LoadApiKey` naming the variable, which is when every factory evaluates its
key. The provider list now also contains the vendor packages that the
hand-written list did not name (azure, amazon_bedrock, google_vertex, ...).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: `rebuild_provider(record, api_key)` creates the recorded provider
with `create_provider(provider_id, ..)` and takes its language model. If
that model is not the one the recording was made with (the record's
`provider` string, e.g. `openai.responses` against the default
`openai.chat`), the rebuild is refused with a message pointing at
`replay_with_model`.

Why: it was the last caller in this crate of the by-name `provider()`
function, and it shared that function's blind spot: a recording made with a
vendor package (`anthropic`, `google`, ...) was `NoSuchProvider`. Those now
rebuild. The refusal exists because the identity-only record cannot select
a non-default method, and replaying a Responses recording against the chat
endpoint would be silently wrong.

Behaviour change: a missing key is no longer reported when the model is
rebuilt; it fails the replayed request, like any other call.

Tests: the three cases that registered a runtime overlay are replaced by
one case for a registry row plus a vendor package, one for the refusal and
one for an unknown id.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
The Mistral chat stream built tool-input-start/delta/end and tool-call
parts with its own hand-written logic, assuming every tool call arrives
complete in one chunk. Upstream (@ai-sdk/mistral
mistral-chat-language-model.ts doStream) instead feeds every
choice.delta.tool_calls entry to StreamingToolCallTracker, constructed
with the model's id generator, and flushes it after the text/reasoning
ends and before the finish part.

mistral/model.rs now does the same with the existing
aimux_provider_utils::StreamingToolCallTracker (as the openai and
openai_compatible chat models already do): each delta is forwarded to
process_delta, a malformed delta ends the stream with an
invalid-response-data error, and flush() runs before Finish. The
hand-written per-chunk emission is removed. A small private generate_id
supplies ids for calls the server sends without one, standing in for
upstream's generateId option.

mistral/types.rs: DeltaToolCall now mirrors upstream's chunk schema
(index, id and function.name/arguments are all optional), so partial
deltas can reach the tracker. Module docs no longer claim tool calls
always arrive complete.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
GroqProvider::chat (and call / language_model) used to return the generic
OpenAI-compatible chat model parameterised by a Groq dialect. The AI SDK's
@ai-sdk/groq has its own GroqChatLanguageModel and does not depend on the
compatible package, so the Rust package now has one too, file for file:

- model.rs mirrors groq-chat-language-model.ts: getArgs, doGenerate and
  doStream, with streamed tool calls going through the shared
  StreamingToolCallTracker (type validation "required", as upstream) and
  stream error chunks mapped to a status through getGroqStreamErrorMetadata.
- convert.rs mirrors convert-to-groq-chat-messages.ts.
- prepare_tools.rs mirrors groq-prepare-tools.ts, with
  browser_search_models.rs for groq-browser-search-models.ts.
- usage.rs mirrors convert-groq-usage.ts, finish_reason.rs mirrors
  map-groq-finish-reason.ts, options.rs mirrors
  groq-chat-language-model-options.ts, error.rs mirrors groq-error.ts, and
  types.rs holds the response and chunk shapes the model reads.
- mod.rs builds the provider the way the Mistral package does (an
  EndpointConfig per model, headers evaluated on every request); the public
  surface (create_groq, groq(), GroqProviderSettings, "groq.chat") is
  unchanged and GroqProvider::chat now returns GroqChatLanguageModel.

Where the existing package tests pin behaviour that differs from upstream, the
tests win and the model keeps it: the token limit is sent as
max_completion_tokens, the generic openaiCompatible options namespace is read
under the groq one, unknown fields of the groq options go to the body, and
provider metadata {"groq": {}} is reported.

groq/dialect.rs is left in place: the registry preset still builds Groq
through groq::profile() on the compatible model.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
DeepSeekProvider::chat used to return the generic OpenAI-compatible chat
model parameterised by a profile. The AI SDK's @ai-sdk/deepseek has its own
chat model and does not depend on the OpenAI-compatible package, so the Rust
package now has one too: DeepSeekChatLanguageModel, returned by chat(),
call() and language_model(). The provider keeps its public surface
(create_deepseek, deepseek(), DeepSeekProviderSettings, provider() ==
"deepseek.chat") and is assembled like the Mistral provider: a fixed
EndpointConfig, bearer credential headers loaded per request, the shared
exchange for URL, headers, transport and recording, and list_data_models for
discovery. Streamed tool calls go through the shared StreamingToolCallTracker.

New files under aimux-providers/src/deepseek/, each mirroring one upstream
file of packages/deepseek/src/chat/:
- model.rs      <- deepseek-chat-language-model.ts (getArgs, doGenerate,
                   doStream, stream error classification)
- convert.rs    <- convert-to-deepseek-chat-messages.ts (also carries the
                   file part handling of deepseek-file-part-options.ts)
- usage.rs      <- convert-to-deepseek-usage.ts
- prepare_tools.rs <- deepseek-prepare-tools.ts
- types.rs      <- deepseek-chat-api-types.ts (response and chunk shapes)
- options.rs    <- deepseek-chat-language-model-options.ts (chat, message and
                   file part options)
- finish_reason.rs <- map-deepseek-finish-reason.ts
- is_v4_model.rs   <- is-deepseek-v4-model.ts
The files/ module of the upstream package is not ported.

Where the existing acceptance tests pin behaviour that differs from upstream,
the tests win: thinking and reasoningEffort are sent as given with no
normalisation, unknown providerOptions.deepseek fields (user) reach the body,
topK is sent, reasoning is read as a fallback for reasoning_content, no JSON
system message is injected and a JSON schema becomes response_format
json_schema, strict tool flags are passed through without validation, an
assistant message with only tool calls has null content, and provider
metadata carries no cache counters or response ids. The old
deepseek::profile() stays only for the registry presets that still use it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
There were three ways to turn a provider name into a provider: the old
`provider()` family in provider.rs, `PresetProvider::create`, and the new
`create_provider`. This commit leaves one.

- provider.rs keeps only what the language bindings still call
  (`provider`, `provider_handle`, the runtime overlay for user-registered
  providers). `provider_handle` now checks the overlay and otherwise calls
  `create_provider`. The module is no longer re-exported from the crate
  root; callers name it as `aimux_providers::provider::...`. It goes away
  when the bindings expose the Rust surface directly.
- Removed: `ResolvedProvider`, `resolve_provider`, `provider_discovery`,
  `provider_from_env`, `provider_names`, `provider_registry_entry` and the
  module's unit tests. Model listing is reached through
  `Provider::discovery()`.
- preset.rs: the `PresetProvider` wrapper type is gone. `preset::create`
  returns a plain `OpenAICompatibleProvider`, the same type
  `create_openai_compatible` returns, so a preset is only a row of settings.
- openai_compatible: drop `resolve_base_url`, which had no caller left.

The FFI crate and the Node and Python bindings still import the removed
root exports and are updated in the following commits.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…, moonshotai, fireworks, cerebras, baseten, alibaba, deepinfra)

The vendor test files built their models with the by-name `provider(..)`
and `provider_from_env(..)` functions, which are leaving the crate's
public surface. Their `make_provider` helpers and the individual tests now
call `create_provider(name, PresetSettings { api_key, base_url, .. })` and
`.language_model(model_id)` instead. Assertions are unchanged.

The `from_env_fails_without_env_var` test in each file used to assert an
error at creation time. `create_provider` no longer checks the key, so the
test is now async and asserts that the first `do_generate` call fails. No
tests were deleted.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…e, openai_compatible, thin_wrapper, list_models, openrouter, presets, reasoning_map)

The by-name `provider(..)`, `provider_handle`, `provider_discovery`,
`provider_from_env`, `ProviderOptions`, `provider_registry_entry` and
`PresetProvider::create` are leaving the crate's public surface. These test
files now build their models with `create_provider(name, PresetSettings)` (and
`preset::create(name, settings)` for the OpenAI-compatible presets) followed by
`language_model(..)` / `discovery()`. Local helpers (`registry_model`,
`preset_model`) carry the change so the call sites keep their assertions.

Files with no use of the removed surface (anthropic_prepare_tools, openai_files,
google_files, open_responses) are unchanged.

Deleted:
- list_models_test: `provider_handle_unknown_name` and
  `discovery_for_unknown_provider_is_no_such_provider` only exercised the
  removed by-name entry points (the unknown-name error is covered by
  presets_test through `create_provider`).
- thin_wrapper_config_test: the two `from_env_fails_without_env_var` tests
  (togetherai, vercel) were `#[ignore]`d and asserted a creation-time missing
  key error that no longer exists.

Moved: openrouter `from_env_fails_without_env_var` now asserts `LoadApiKey` on
the first request instead of at creation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
The previous commit gave the Groq package its own chat model but kept four
behaviours that upstream does not have, because hand-written tests pinned them.
Upstream wins, so in groq/model.rs (groq-chat-language-model.ts) and
groq/options.rs (groq-chat-language-model-options.ts):

- the token limit is sent as max_tokens, as getArgs does;
- provider options are read from the groq namespace only; the generic
  openaiCompatible namespace is no longer merged in;
- unknown fields of the groq options are dropped, not copied into the body;
- no provider metadata is reported from doGenerate or doStream, as upstream
  returns none.

A chunk that fails to parse is now reported the way upstream does it: an error
part, finish reason "error", and the stream carries on; only a transport
failure ends it. An error reported by the very first event still rejects the
call, which is the convention of the other own-model packages in this crate
and keeps the failure inside Core's retry boundary.

tests/groq_test.rs now builds every model through create_groq(..).chat(..)
instead of the by-name provider() function, so it exercises the new model. Tests
that expected non-upstream behaviour are corrected to the upstream cases
(reasoning mapped to low/medium/high, none handled per model, file references
raising an unsupported-functionality error, tool_choice "auto", max_tokens, no
metadata, groq namespace only) and upstream stream cases are ported: reasoning
kept active across empty tool_calls, unparsable chunks, raw chunks, error chunks,
and a response without choices.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
The DeepSeek chat model kept a list of behaviours that differed from the
upstream package only because older, hand-translated tests pinned them. The
repo rule is that upstream wins, so the model now does what
deepseek-chat-language-model.ts, convert-to-deepseek-chat-messages.ts,
deepseek-prepare-tools.ts and deepseek-chat-language-model-options.ts do:

- thinking / reasoning: `thinking.type` `adaptive` becomes `enabled` with a
  compatibility warning; top-level `reasoning` derives `thinking`
  (`none` is `disabled`, anything else `enabled`) and `reasoning_effort`
  (minimal/low to low, medium/high to high, xhigh to max, with the same
  warnings); `reasoningEffort` `medium`/`xhigh` map to `high`/`max`; no
  effort is sent when thinking is disabled.
- providerOptions.deepseek is validated like the zod schema: unknown fields
  are dropped, `thinking` and `reasoningEffort` are enums, `userId` errors
  carry the upstream messages.
- `top_k` is warned about and not sent.
- The model reads `reasoning_content` only; the `reasoning` fallback is gone.
- JSON response format injects the "Return JSON." / schema system message
  (with a compatibility warning for a schema) and sends `json_object`.
- Strict tools are rejected unless the base URL ends in `/beta`, and strict
  and non-strict function tools may not be mixed (prepare_tools).
- An assistant message with only tool calls has `content: ""`.
- Provider metadata carries promptCacheHitTokens, promptCacheMissTokens,
  responseObject, choiceIndex, messageRole, toolCallTypes, logprobs and
  systemFingerprint, in generate and in the stream.
- The default base URL is `https://api.deepseek.com` (the request path is
  `/chat/completions`), as in createDeepSeek.
- Stream error events keep the whole error envelope as the error data.
- `request_body` and `RequestBodyResult` are private: they mirror the
  private `getArgs`.

The tests are replaced by direct ports of the upstream describe-blocks of
deepseek-chat-language-model.test.ts and
convert-to-deepseek-chat-messages.test.ts in deepseek_chat_test.rs, built
through create_deepseek(..).chat(..). deepseek_reasoning_test.rs is folded
into it, deepseek_test.rs keeps the provider-level cases, and the DeepSeek
cases of reasoning_map_test.rs that pinned the old behaviour are removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
@cunninghamcard-bit cunninghamcard-bit changed the title feat(providers)!: AI SDK-shaped provider factories (RFC-0036 Part I, Rust) test(providers): add factory replay coverage and migration docs (RFC-0036) Oct 7, 2026
@cunninghamcard-bit cunninghamcard-bit changed the title test(providers): add factory replay coverage and migration docs (RFC-0036) feat(providers)!: AI SDK-shaped provider factories, request layer and replay tests Oct 8, 2026
@cunninghamcard-bit
cunninghamcard-bit changed the base branch from rfc-0036/provider-factory-source to rfc-0036/provider-result-types October 8, 2026 03:10
chenhaonan and others added 4 commits October 8, 2026 17:22
A name that is neither a vendor package nor a registry row fell through to
preset::create, whose NoSuchProvider lists only the 281 registry rows. A
misspelled "openai" therefore got an available-providers list without
openai, anthropic or any other vendor package. create_provider now reports
the full provider_names() list itself.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@cunninghamcard-bit

cunninghamcard-bit commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor Author

[P1] Correct the HMAC derivation in the moved SigV4 signer

This is inherited from the previous Bedrock signer rather than a new regression in the refactor, but the shared SigV4Fetch still produces signatures that do not match AWS SigV4.

hmac_sha256 takes (key, data). The calls at lines 169–178 reverse those arguments at every step, including the final signature, and omit the AWS4 prefix on the secret. With the current helper, the correct order is:

let k_date = hmac_sha256(
    format!("AWS4{}", credentials.secret_access_key).as_bytes(),
    date_stamp.as_bytes(),
);
let k_region = hmac_sha256(&k_date, credentials.region.as_bytes());
let k_service = hmac_sha256(&k_region, service.as_bytes());
let k_signing = hmac_sha256(&k_service, b"aws4_request");
let signature = hex::encode(hmac_sha256(&k_signing, string_to_sign.as_bytes()));

This matches AWS's documented derivation.

I reproduced the existing signer_output_is_unchanged_by_the_move fixture offline with its public example credentials and fixed timestamp. Holding both the canonical request and string-to-sign identical, aws4fetch 1.0.20 gives:

current golden: 50198f4d3c821ae4aa0518965cd3ecaeb81e542a155dc57766e1e28a6ae696b8
AWS-compatible: cd5e8832ab041bc9557910ef3dd094d8eb37b14b0edd903c4b45d83bbda888ca

The comparison explicitly asserts equal canonical requests and strings-to-sign; singleEncode: true keeps the PR's existing URI representation so this isolates the HMAC defect. It does not validate the rest of the current canonicalization. The existing golden test pins the previous algorithm, and the decorator test compares against that same signer, so neither establishes AWS compatibility. Please replace the incorrect golden with an independent AWS-compatible reference and retain the checks that the inner fetch receives the same signed body bytes.

For implementation, the existing SigV4Fetch → inner.fetch boundary can stay. An alternative to maintaining a private SigV4 implementation is AWS's standalone aws-sigv4 HTTP signer; AWS documents using it independently of the SDK. A full aws-sdk-bedrockruntime client is not required for this fix. The pinned AI SDK likewise delegates its signing algorithm to AwsV4Signer from aws4fetch, while keeping its own fetch wrapper.

Clarification: Rig vs. the standalone signer

The Rig Bedrock transport inspected here calls the full aws-sdk-bedrockruntime client via client.converse().customize().send(). The AWS SDK then invokes aws-sigv4 internally before its HTTP connector. Rig does not directly integrate the standalone signer.

For aimux, I recommend using that official signer directly inside the existing decorator:

SigV4Fetch::fetch(final FetchRequest)
  ① resolve credentials
  ② Credentials::new(...).into() → Identity
  ③ v4::SigningParams::builder()
     identity / region / service="bedrock" / clock / settings
  ④ SignableRequest::new(method, URL, headers, SignableBody::Bytes(body))
  ⑤ aws_sigv4::http_request::sign(...)
  ⑥ apply the returned SigningInstructions
  ⑦ existing inner.fetch(signed request)

This keeps aimux's existing request conversion, response parsing, call-level retry and fetch injection. It delegates canonicalization and cryptographic signing without introducing a Bedrock service client. See the official HTTP signing API. Check the signing settings against Bedrock's URI encoding, path normalization and header rules rather than preserving the old sign-every-header behavior.

Dependency compatibility also needs checking: the locally inspected aws-sigv4 1.3.8 declares Rust 1.88, while aimux declares 1.85. Select compatible dependency versions or explicitly raise the MSRV.

…S does

The signing key chain passed every HMAC its key and data swapped and
dropped the AWS4 prefix on the secret, so no signature matched AWS.
The canonical request also left path segments unencoded (a ':' in a
Bedrock model id), kept the query unsorted and did not collapse header
whitespace.

The golden signatures now come from botocore's SigV4Auth instead of
the previous implementation.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A3WxDBWBA4ituJjXsHaEAa
…nd Polly

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01A3WxDBWBA4ituJjXsHaEAa

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants