Repository navigation
feat(providers)!: AI SDK-shaped provider factories, request layer and replay tests - #214
cunninghamcard-bit wants to merge 114 commits into
Conversation
The AI SDK alignment design (docs/aisdk-architecture-alignment.md §9.3)
requires validating aimux providers against what the AI SDK sends and
returns for the same input under a mocked transport. This adds the SDK
half: scripts/aisdk-fixtures runs the official @ai-sdk packages
(ai 7.0.127, provider 4.0.21, provider-utils 5.0.53, openai 4.0.83,
openai-compatible 3.0.62, anthropic 4.0.71, google 4.0.87, pinned exactly)
with a recording fetch and writes fixtures/aisdk/<package>/<case>.json:
request (URL, headers with credentials redacted, body), canned response,
and the SDK's normalized result or thrown error.
28 cases across openai, openai-compatible (name "groq"), anthropic and
google pin the behaviours the Rust provider rewrite must match, among them:
an explicit empty apiKey is sent verbatim and never falls back to the env
var; a missing key fails at call time with AI_LoadAPIKeyError and no
request; createAnthropic({ name }) sets model.provider to the name verbatim
(default "anthropic.messages") and reads providerOptions under both the
canonical and the custom key, custom winning; openai-compatible derives
its providerOptions key from the first segment of model.provider and
spreads unknown fields under that key into the request body; header
merging is case-insensitive with call-level headers overriding provider
headers.
Rust tests consume these in the provider factory rewrite; regenerate with
`cd scripts/aisdk-fixtures && npm ci && node record.mjs`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
… call-level retry AI SDK alignment design (docs/aisdk-architecture-alignment.md §3.2 / §4.5, impact map S1-1, S1-2, S1-3, S1-5). First commit group (A0) of the provider-factory rewrite: the primitives every createXxx(settings) factory needs, with no change to the factories themselves yet. Transport (S1-1) - `Fetch` trait with `FetchRequest` / `FetchResponse` / `FetchError`; `ReqwestFetch` carries the former `shared_client()` machinery (per-runtime client sharding, pool settings, proxy config); `default_fetch()` is resolved per request, never frozen at factory time. - `HttpRequest.fetch: Option<FetchFunction>` threads an injected transport through `post_*_to_api` / `get_from_api`; `send_request_once` sends through it. The SSRF download guard keeps its pinned-DNS client as `PinnedFetch` and ignores the injected fetch when `validate_url` is set (D26). - `SigV4Fetch` decorator signs the final request bytes (same algorithm and test vectors as `bedrock/sigv4.rs`; providers switch to it in A4). - `WsConnector` injection point on `WebSocketRequest`; `TungsteniteConnector` is today's behaviour. - `catalogue.rs` no longer builds its own reqwest client. Settings (S1-2, S1-3) - `Resolvable<T>` (Value / Fn / AsyncFn / Future) and `HeadersFn`; `combine_headers` (lower-cased names, later layers override, `None` removes) and `normalize_headers`; requests insert headers instead of appending, so a credential header is never sent twice. - `load_api_key(Some(s))` returns `s` verbatim, including the empty string; only `None` falls back to the environment (pinned by the AI SDK fixture `openai/empty-api-key`). Missing values are `AiMuxError::LoadApiKey` / `LoadSetting`; `load_setting` / `load_optional_setting` added. The FFI, Node and Python mappings route the two new variants to the existing invalid-argument code until A5 adds dedicated codes. Retry (S1-5, D14) - `prepare_retries(max_retries, abort)` with the constant defaults 2 / 2000 ms / x2; `RetryConfig` is deleted along with every provider config field, `with_retry_config` builder and `retry_config()` trait method. Retry is a call-level concern: the nine core operations read only the caller's `max_retries`, provider-internal list/files/poll sites use the default budget for now (A4 applies the §4.5 table). Not in this commit: RecordingFetch (S1-4), the Provider trait reshape (A1), provider factories (A2–A4), FFI/binding error codes (A5). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…rding, and the boundary gate
Second commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.1, §3.3, §4.3, §7).
Provider trait (aimux-core/src/provider.rs)
- `Provider` now has the `ProviderV4` shape: language / embedding / image
are required and return `NoSuchModel { model_type }` when a vendor has
no such modality; transcription / speech / reranking / files are
optional (`None` = not offered); video and search stay as aimux
extensions. All constructors return `Arc<dyn Model>`.
- `Provider::name()` is gone: the registered name belongs to whoever
holds the registry, and each model reports `provider()` itself.
- `specification_version()` is removed from the nine model traits and
from `TraceLayer`; `LanguageModel::config_snapshot()` is removed and
`supported_urls()` added.
- `list_models` moves to a separate `ProviderDiscovery` trait (no
`ProviderV4` equivalent; most vendors cannot list). `delegate_list_models!`
emits that impl; `provider_discovery(name, ..)` resolves a discovery
handle for FFI / Node / Python.
Recording (RECORDING_SCHEMA = 3)
- `ProviderRecord` is identity only: `provider_id`, `provider`, `model_id`.
Configuration (base URL, key source, profile, provider options, retry)
is no longer recorded; schema-2 files are rejected on deserialization.
- `rebuild_provider` rebuilds from `provider_id` + `model_id` through the
registry; `is_openai_compatible_provider` is deleted. Native-protocol
packages return `NoSuchProvider` until they join the registry (A4).
Call-level overrides
- `CallOptions.body_overrides` is deleted everywhere (core, openai and
anthropic converters, FFI, Node, Python, aimux-web playground, wire
fixture). Provider-level overrides are applied to the finished body
until `transform_request_body` replaces them (A3).
- `ProviderOptions` / `config_json` / Node `ProviderConfig` reject
`max_retries` and `body_overrides` with `InvalidArgument` instead of
silently ignoring them.
Boundary gate
- `scripts/check_provider_boundaries.sh` (wired into the contract-tests
CI job) fails on provider code reading `options.max_retries` /
`timeout` / `session_id` / `call_id` outside `HttpRequest::new`, and on
any removed name coming back. Eleven hand-built `HttpRequest` literals
now go through `HttpRequest::new`; the eight poll-loop providers use
`prepare_retries(None, ..)` constants for their status GETs.
Tests: new `provider_trait_test.rs` in core and providers (shape, error
mapping, discovery); `config_snapshot_test.rs` deleted; overlay test now
proves the override through a wiremock request; body-override tests
exercise the provider-level path; Node `provider_config.test.ts` covers
the rejection.
Not yet done here (later groups): `api_key_source` fields and
`ExternalProviderEntry.max_retries` / `body_overrides` still exist
unread (A3); `provider_id` is the first segment of `provider()` until the
registry name fills it in (A3); Go / Java / Kotlin / Swift / Flutter still
declare a call-level `body_overrides` field that core ignores (A5).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Third commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.2, §3.3, §4.1, §7), modelled on
`createOpenAI` in @ai-sdk/openai 4.0.83 and checked against the recorded
fixtures under fixtures/aisdk/openai/.
Native OpenAI package
- `OpenAIProviderSettings { base_url, api_key: Option<Resolvable<String>>,
organization, project, headers: Option<HeaderMapOpt>, name, fetch,
transform_request_body }` and `create_openai(settings)`. The factory
only validates `base_url` and fixes `name` (default "openai"); the key
is loaded on every request: `None` reads OPENAI_API_KEY, `Some("")` is
sent verbatim, a missing key fails the call with `LoadApiKey`, not the
factory. `openai()` is the infallible `OnceLock` default instance.
- `OpenAIModelConfig` (crate-private, no getters, no snapshot) carries
`provider`, a `url()` closure, an async `headers` resolver (credential,
then organization/project, then user headers via `combine_headers`),
`fetch`, `supported_urls`, `transform_request_body` and the base URL
for the credentialed-origin guard. Chat, Responses, embedding, image,
speech, transcription and files models read only this config.
- `provider()` is `"{name}.{method}"` (`openai.chat`, `openai.responses`,
`proxy.chat` for `name: "proxy"`); the providerOptions namespace stays
the fixed `openai` key, as upstream.
- `transform_request_body` is the only provider-level body hook; it runs
once on the finished JSON body (multipart untouched). `OpenAIConfig`'s
old `body_overrides` becomes such a closure at the compat boundary.
- `Resolvable::Future` is awaited once, `Resolvable::AsyncFn` on every
request (counter tests).
Compat consumers (transitional until A3)
- `OpenAIConfig` and the new `OpenAIConfigProvider` move to
`aimux-providers/src/openai_legacy.rs`; they keep the builder API for
the 34 thin wrappers, the registry (`provider.rs`), Codex, xAI and
Hugging Face, and feed the same `OpenAIModelConfig` through
`into_model_config`. Removed from `OpenAIConfig`: getters, `from_env`,
`with_api_key_source`, `with_body_overrides`, the `api_key_source`
field. `apply_body_overrides` / `deep_merge_json` move to
`body_merge.rs` for Anthropic until A4.
- Compat `provider()` strings are unchanged (`"groq"`), except files,
which now report `"{provider}.files"`.
Constructor sites: the twelve `aimux_openai_*` FFI symbols, the Node and
Python OpenAI constructors, aimux-cli and aimux-web probes go through
`create_openai`; symbol names are unchanged. CLI / web cache probes
filter traces by `model.provider()`.
Boundary gate: rule 3 keeps `config.api_key` / `config.base_url` /
`body_overrides` / `api_key_source` / `from_env` out of
`aimux-providers/src/openai/`.
Also: an invalid header value no longer echoes the value in the error
(it could contain an API key).
Tests: new `openai_factory_test.rs` asserts every fixture in
fixtures/aisdk/openai/ (URL, method, headers, body, `provider()`, stream)
and fails when a fixture has no case; eleven openai test files moved to
`create_openai`; `body_overrides_test` covers `transform_request_body`.
Left for later groups: delete `OpenAIConfig` / `OpenAICompatProfile` /
`OpenAIConfigProvider` and the thin wrappers, move compat identity to
`"{name}.{method}"` with an explicit dialect flag (A3); Anthropic
`transform_request_body`, Azure `fetch` / `credentialed_origin`, Codex /
xAI / Hugging Face off `OpenAIConfig` (A4); `OPENAI_BASE_URL`, structured
env references, the recorded `"openai"` → `"openai.chat"` consumers and
the CHANGELOG entry (A5).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…and registry-generated presets
Fourth commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.1, D9, D12, D20), modelled
on @ai-sdk/openai-compatible 3.0.62 and checked against the recorded
fixtures under fixtures/aisdk/openai-compatible/.
openai-compatible (aimux-providers/src/openai_compatible/)
- `OpenAICompatibleProviderSettings { name, base_url, api_key, headers,
query_params, fetch, include_usage, supports_structured_outputs,
supports_multi_part_tool_content, transform_request_body }` and
`create_openai_compatible`. `name` and `base_url` are required and
fixed at the factory; the key resolves per request; `None` sends no
`Authorization` (local servers).
- Chat, embedding and image models (the upstream set) with
`provider() = "{name}.{method}"`. providerOptions are read from
`openaiCompatible`, then the name, then its camelCase form; unknown
fields under the provider's own namespace pass through to the body in
the upstream position; metadata uses the name as key.
- Dialect hooks (`pub(crate)`) replace `OpenAICompatProfile`: usage
conversion, tool and response-format preparation, file-part policy,
`max_tokens_key`, stream usage key. No `Box::leak`.
- `list_models` is one request, no retry.
Groq and DeepSeek (groq/, deepseek/)
- Own packages on the compat internals: `create_groq` / `groq()` with
`x_groq` stream usage, no `top_k`, `max_completion_tokens`, Groq
structured-output rules, browser_search and the Groq file-part
rejection; `create_deepseek` / `deepseek()` with cache-read usage.
`provider()` is `groq.chat` / `deepseek.chat`; providerOptions and
metadata keys are `groq` / `deepseek`. The eight `== "groq"` branches
in openai/convert.rs are gone; the native OpenAI package knows nothing
about other vendors.
Presets (scripts/gen_presets.py → aimux-providers/src/presets/, committed)
- `provider_registry.json` now has 283 rows (251 + the 32 former wrapper
vendors) with `auth: api_key | none`, `base_url_env`, `params` with
`{param}` templates (env / default / derived host maps for Vertex
locations) and `family`. Each row generates a `PresetDescriptor`,
`create_<name>(PresetSettings)` and an infallible `<name>()`.
- `auth: none` resolves no key and injects no placeholder; there is no
`PLACEHOLDER_API_KEY` any more. Template parameters are only the
declared ones, host parameters reject `/`, `@`, `:`, `?`, and an
unexpanded placeholder is an error. Unknown preset names are
`NoSuchProvider`; nothing falls back to OpenAI.
- `provider.rs` resolves names through the preset table and the overlay
table; `ExternalProviderEntry` rejects `max_retries` and
`body_overrides`. `gen_presets.py --check` runs in the contract-tests
CI job; `gen_providers_doc.py` reads the new columns.
Deleted: `openai_legacy.rs` (`OpenAIConfig`, `OpenAIConfigProvider`),
`OpenAICompatProfile`, the 22 local/cloud thin wrappers (incl.
bedrock_mantle, openrouter) and the 10 `vertex_ai_*_models.rs` files.
Codex, xAI, Hugging Face and Azure compile through a crate-private
`StaticBearerConfig` until their own factories (A4).
Behaviour changes recorded here: the native OpenAI package warns on and
drops `top_k` (upstream does not send it); registry and preset providers
expose chat / embedding / image only; compat providers no longer read
the `openai` providerOptions key; `providerOptions.deepseek` is honoured
(`reasoningEffort`, `thinking`); `litellm_proxy` reads
`LITELLM_PROXY_BASE_URL` (the old wrapper read the API-key variable as a
URL).
Boundary gate: rule 4 forbids provider-name comparisons inside the
shared packages, rule 5 keeps the removed names removed.
Tests: `openai_compatible_factory_test.rs` (every compat fixture, plus
a directory-listing guard), `presets_test.rs` (all 283 rows create
without env, keyless rows send no Authorization, template and parameter
errors, Vertex hosts), `deepseek_test.rs`, the Groq package module;
about twenty test files moved from wrapper types to presets. Cassette
recordings are untouched.
Left for later groups: Codex / xAI / Hugging Face / Azure factories and
removal of `StaticBearerConfig` (A4); Anthropic off `body_merge.rs`
(A4a); `params` in the FFI / Node / Python configs, bindings tests on
provider names, docs and the CHANGELOG entry (A5).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…reuse the Anthropic core
Fifth commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.1, D-d, D-h, D-i),
modelled on `createAnthropic` in @ai-sdk/anthropic 4.0.71 and checked
against the recorded fixtures under fixtures/aisdk/anthropic/.
anthropic/
- `AnthropicProviderSettings { base_url, api_key, auth_token, headers,
name, fetch, transform_request_body }`, `create_anthropic` and the
infallible `anthropic()`. The factory validates `base_url` and rejects
`api_key` together with `auth_token`; the credential resolves per
request (`None` reads ANTHROPIC_API_KEY, `Some("")` is sent verbatim,
`auth_token` sends only `Authorization: Bearer`). A missing key fails
the call with `LoadApiKey`.
- Private `AnthropicModelConfig` with hooks (URL, body, headers, error
handler, feature flags) feeds one `AnthropicMessagesModel`; the files
model reads the same config.
- `provider()` is the name verbatim: `anthropic.messages` by default,
`proxy` for `name: "proxy"`; files derive `{name minus .messages}.files`.
- providerOptions are read from the canonical `anthropic` key merged
with the custom first segment (custom wins); response metadata is
written under the custom key. The twenty hard-coded `"anthropic"`
sites go through `anthropic/options.rs`.
- Base URL follows upstream: it includes `/v1` and the endpoint is
`{base}/messages`; only the bare `https://api.anthropic.com` is
rewritten. `ANTHROPIC_BASE_URL` is no longer read (no `from_env`).
- Request body: `stream` is omitted on non-streaming calls;
`providerOptions.anthropic.metadata.userId` maps to `metadata.user_id`;
`anthropic-beta` is sent on every host; `supported_urls` declares the
upstream image and PDF patterns.
- `list_models` is one exchange without retry.
anthropic_aws/
- `AnthropicAwsProviderSettings` with `ApiKey(Resolvable)` or
`SigV4(Resolvable<AwsCredentials>)`; signing is now the `SigV4Fetch`
transport decorator (service `aws-external-anthropic`) over the final
bytes, the model no longer signs, and `list_models` works with SigV4.
The provider string stays `anthropic-aws`.
vertex/anthropic_model.rs
- `VertexAnthropicModel` is `AnthropicMessagesModel` with Vertex hooks
(rawPredict / streamRawPredict URL, `anthropic_version` in the body,
bearer or x-goog-api-key headers, Google error handler, empty
`supported_urls`, structured outputs and strict tools off); 591 → 117
lines. Provider string `googleVertex.anthropic.messages`.
Deleted: `AnthropicConfig`, `AnthropicConfigBuilder`,
`AnthropicAwsProviderConfig`, `api_key_source`, `from_env`, `with_*`,
and the family's `body_overrides` (replaced by `transform_request_body`).
Constructor sites in aimux-ffi, Node, Python, aimux-cli and aimux-web go
through the factories; symbol names are unchanged. Boundary gate rule 6
covers the Anthropic family.
Tests: `anthropic_factory_test.rs` replays every fixture in
fixtures/aisdk/anthropic/ (URL, method, headers, body, `provider()`,
stream) with a directory-listing guard, plus credential timing, custom
namespace, SigV4-over-final-bytes and Vertex-envelope tests; fifteen
test files moved to the factories. Cassette recordings are untouched.
Left for later groups: Vertex factory with real ADC / express headers
and the `googleVertex` namespace (4b); Bedrock on `SigV4Fetch` and the
`bedrock/sigv4.rs` shim removal (4b); files upload retry (4d);
`body_merge.rs` and its two tests, the base-URL release note and the
result-level providerMetadata (A5).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…on_bedrock factories
Sixth commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.1, D-h, D11), modelled on
@ai-sdk/google 4.0.87, @ai-sdk/google-vertex and @ai-sdk/amazon-bedrock
and checked against the recorded fixtures under fixtures/aisdk/google/.
google/
- `GoogleProviderSettings`, `create_google`, the infallible `google()`;
chat, embedding, image, video and files models read a private
per-model config (`shared/exchange.rs`). The key resolves per request
(`x-goog-api-key`, GOOGLE_GENERATIVE_AI_API_KEY); `provider()` is the
name (`google.generative-ai`) with `{name}.files` / `{name}.video`.
- Request bodies now match upstream: `generationConfig` is always sent,
function tools go out as `parametersJsonSchema`, `providerOptions.google`
(thinkingConfig, responseModalities, audioTimestamp, mediaResolution,
imageConfig, safetySettings, cachedContent, labels, serviceTier,
retrievalConfig) is mapped; response `modelId` comes from
`modelVersion`. providerOptions are read from and written to `google`
only (`google/options.rs`). `list_models` and the files upload are
single exchanges.
vertex/
- `VertexProviderSettings { api_key (express), project, location,
base_url, headers, access_token, fetch, transform_request_body }`,
`create_google_vertex`, `google_vertex()`. Express versus standard
mode, project / location (`LoadSetting`) and the bearer token
(`GOOGLE_VERTEX_ACCESS_TOKEN`, or any `Resolvable`) resolve per
request; location must be one DNS label and the host follows the
upstream global / rep / regional rule; Gemini models use `v1beta1`,
Anthropic-on-Vertex `v1`. `provider()` is `google.vertex` (plus
`.video`, `.transcription`, and `googleVertex.anthropic.messages`).
providerOptions: `googleVertex` (canonical only; the legacy `vertex`
key is neither read nor written), then `google` for the shared Gemini
model.
bedrock/
- `AmazonBedrockProviderSettings { region, api_key, access_key_id,
secret_access_key, session_token, base_url, headers, fetch,
credential_provider }`, `create_amazon_bedrock`, `amazon_bedrock()`.
AWS_BEARER_TOKEN_BEDROCK selects bearer auth; otherwise `SigV4Fetch`
signs the final bytes with credentials resolved per request
(`LoadSetting` for AWS_REGION / AWS_ACCESS_KEY_ID /
AWS_SECRET_ACCESS_KEY). `provider()` is `amazon-bedrock` for chat,
embedding, image and reranking; providerOptions use `amazonBedrock`
only (the legacy `bedrock` key is neither read nor written).
`bedrock/sigv4.rs` is deleted; `aws_polly` signs through `SigV4Fetch`.
Deleted: `GoogleConfig`, `VertexProviderConfig`, `VertexAuth`,
`BedrockProviderConfig`, `BedrockAuth`, their `from_env` / `with_*` /
`api_key_source`. Constructor sites in aimux-ffi (14), Node, Python,
aimux-cli and aimux-web go through the factories; symbols unchanged.
Boundary gate rule 7 keeps the builder-era names, `bedrock::sigv4` and
stray namespace literals out.
`anthropic_aws` keeps its name: it is Claude Platform on AWS
(`aws-external-anthropic`), not Bedrock InvokeModel; aimux has no
Bedrock-Anthropic InvokeModel provider and this PR does not add one.
Tests: `google_factory_test.rs` (every fixture, directory-listing guard),
`vertex_factory_test.rs` (express and standard paths, three hosts,
invalid location, missing settings), `bedrock_factory_test.rs` (bearer
without signature, SigV4 over the final bytes, provider strings), a
shared scripted `Fetch` mock in tests/common; about twenty-eight test
files moved to the factories. Cassette recordings are untouched.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…Face, Codex, open_responses, Voyage and ElevenLabs
Seventh commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.1, D-d, D11, D13), each
package modelled on its @ai-sdk counterpart where one exists.
- Every package gets `XxxProviderSettings`, `create_xxx`, an infallible
`xxx()` and a private per-model config; credentials resolve per
request through the shared helpers (`None` reads the package's env
var, `Some("")` is sent verbatim, a missing key fails the call with
`LoadApiKey`). `provider()` strings follow upstream: `azure.chat` /
`azure.responses` / `azure.embeddings` / `azure.image` /
`azure.transcription` / `azure.speech`, `xai.responses`,
`mistral.chat` / `mistral.embedding`, `cohere.chat` /
`cohere.textEmbedding` / `cohere.reranking`, `huggingface.responses`,
`codex.responses`, `{name}.responses` for open_responses,
`voyage.embedding` / `voyage.reranking`, `elevenlabs.speech` /
`elevenlabs.transcription`. All but Azure accept a `name` override.
- Azure reuses the OpenAI chat / Responses / embedding / image /
transcription / speech models through the injected config: `api-key`
or bearer token (`token_provider` AsyncFn) per request, resource name
from AZURE_RESOURCE_NAME (`LoadSetting`), upstream defaults
(`api_version = "v1"`, `/v1{path}?api-version=`), deployment-based
URLs on request. The Responses model reads providerOptions `azure`
then `openai` and writes `azure`.
- xAI and Hugging Face default to Responses (as upstream); their
chat-completions models remain reachable as the `chat_completions(id)`
extension. Codex keeps RFC-0018's two modes (`ApiKey` with
CODEX_API_KEY fallback, `ChatGptAccount { token, account_id }`), the
stateless refresh, 401 → `TokenExpired`, and `store: false` as a
package rule. ElevenLabs' WebSocket handshake headers come from the
resolved headers and the connector is injectable. open_responses takes
`base_url` and `Resolvable` headers; `api_key: None` sends nothing.
- Deleted: the nine `XxxConfig` types and builders, `AzureAuth`,
`TokenProvider`, `AzureModel`, `azure/{model,responses}.rs`,
`api_key_source` / `from_env` in these packages, and the transitional
`StaticBearerConfig`. `OpenAIModelConfig` now anchors credentialed
headers to the request URL (Azure's host is per request) and carries
a `ResponsesProfile`; namespace literals live in each package's
`options.rs`.
- Constructor sites in aimux-ffi (12), Node (6), Python (6), aimux-cli
and aimux-web go through the factories; symbols unchanged. Boundary
gate rule 8 covers these packages.
Not ported from upstream Azure: the deepseek and completion models, the
MAI / speech endpoints and the Foundry item type.
Tests: `vendor_factories_test.rs` (28 tests over a scripted `Fetch`:
credential timing, empty key, default instances, provider strings per
package) and 27 adapted test files. Cassette recordings are untouched.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…ge-constant poll loops; uploads are not retried
Eighth commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §3.3, §4.5, §9.3).
- Every single-modality package (serper, prodia, deepgram, you_com, luma,
lmnt, klingai, dataforseo, replicate, fal, recraft, aws_polly, jina_ai,
gladia, tavily, linkup, tinyfish, black_forest_labs, assemblyai, revai,
hume, cartesia, searxng, parallel_ai, firecrawl, runwayml, exa_ai,
stability, google_pse) gets `XxxProviderSettings`, `create_xxx`, an
infallible `xxx()` and `XxxProvider` on the shared `EndpointConfig`;
credentials resolve per request (missing key → `LoadApiKey`, missing
URL or AWS setting → `LoadSetting`). `provider()` is `"{name}.{method}"`
(`luma.image`, `deepgram.transcription`, `tavily.search`,
`amazon-polly.speech`, `blackForestLabs.image`, …). providerOptions
keys live in one `options.rs` per package. The old `XxxConfig`,
`from_env` and `with_*` are gone; no `*Config` struct remains in
aimux-providers.
- Poll loops (luma, black_forest_labs, fal, gladia, assemblyai, revai)
run on package constants through `shared/poll.rs`: a fixed interval
and attempt cap (previously unbounded loops now stop after 6000 × 100
ms), a transient poll error spends one attempt, exhaustion returns the
last error or `Timeout`, cancellation is checked before every wait.
The job-creating request is sent exactly once; nothing re-submits.
Only the interval can be overridden, through the package namespace
(`pollIntervalMillis` for luma as upstream, `pollIntervalMs`
elsewhere); pacing keys are stripped before the body is forwarded.
Download stages keep a bounded three-try retry. `aimux_core::retry`
is no longer used anywhere in aimux-providers (boundary rule 9).
- Files uploads (OpenAI, Anthropic, Google) are single exchanges.
- Constructor sites (`aimux_tavily_search_new*`, Node and Python
Tavily) go through the factory; symbols unchanged. Boundary rule 10
covers the 28 packages.
Behaviour changes recorded here: cartesia `version` is a fixed header
(override via `headers`); runway poll pacing is constant; google_pse
`cx` precedence is settings, providerOptions, then GOOGLE_CSE_ID per
request; DataForSEO missing credentials report `LoadApiKey`; searxng
accepts an optional bearer token.
Tests: `single_modality_factories_test.rs` (defaults, missing key,
provider strings), `poll_stage_test.rs` (exhaustion never re-submits per
package, interval override, abort, ignored attempt keys, stripped pacing
keys), upload-not-retried tests in the three files packages, unit tests
in `shared/poll.rs`; thirty-five test files converted. Cassette
recordings are untouched.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…binding configs and error classes, CHANGELOG and docs
Ninth and last commit group of the provider-factory rewrite
(docs/aisdk-architecture-alignment.md §5, §6.2, §8, §9.2, §9.4).
C ABI
- `AIMUX_E_LOAD_API_KEY` (18) and `AIMUX_E_LOAD_SETTING` (19) for
`AiMuxError::LoadApiKey` / `LoadSetting`; the env var travels in
`aimux_error_provider_code` and the description or setting name in
`aimux_error_provider_message`. `aimux-error.h` / `aimux-ffi.h` synced,
with a header-versus-source parity test. `provider_handle_new` and
`register_providers` reject `max_retries` / `body_overrides` in
`config_json`.
Node / Python
- `LoadAPIKeyError` and `LoadSettingError` classes (`envVar` /
`description` / `settingName`; snake_case in Python). Every Node native
constructor goes through one `native_config()` that rejects
`maxRetries` / `bodyOverrides` and applies `headers`;
`ProviderConfig.params` carries preset template parameters.
`index.d.ts` regenerated; both bindings are fmt- and clippy-clean; the
ava and pytest suites pass.
Go / Java / Kotlin / Swift / Dart
- Codes 18 / 19 with an `EnvVar` field; the call-level `body_overrides`
field (which core no longer reads) is deleted; Go and Dart
`ProviderConfig` drop `max_retries` / `body_overrides` and gain
`params`. Go and Java suites run here; Kotlin, Swift and Dart edits
are symmetrical and unverified in this container.
Residue
- `body_merge.rs` and `body_overrides_test.rs` are gone
(`transform_request_body_test.rs` replaces them); `hmac` dropped,
`sha2` / `hex` are dev-dependencies; Codex subscription default URL
fixed to `…/backend-api/codex`; Bedrock `list_models` host fixed to
`bedrock.{region}.amazonaws.com`; Anthropic result-level
`providerMetadata` (usage, stopSequence, container) implemented and
checked against the fixtures; boundary gate also forbids
`api_key_source` and the body-merge helpers.
Docs
- CHANGELOG `[Unreleased]` gains the Rust, C ABI, Node / Python and
other-binding breaking lists for the whole rewrite. README,
CONTRIBUTING, docs/api/*, docs/API.md, the provider-config manual, the
Codex guide, the error model, the request-pipeline doc (retry is a
call-level constant) and `rfc/0036-aisdk-architecture-alignment.md` §5
(no `specification_version` / `name()`, `.call()` form, OpenAI default
stays chat, Bedrock InvokeModel and unported Azure models, video
`poll_config`, native `rebuild_provider`) are updated.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
… in §5 Same change as the docs/rfc-0036-aisdk-alignment follow-up (no ABI coexistence promise after the cutover; §3.3 / S4-7 wording for the xAI and Hugging Face Chat Completions extension), plus the §5 product difference row for `chat_completions(id)` that this branch introduces. The shared part becomes a no-op once the docs PR lands on master. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…s models The AI SDK's xAI package (5.0.12) and Hugging Face package serve the Responses API only. The Chat Completions models that 7c54cfd kept as a `chat_completions(id)` extension outside the `Provider` trait are removed: `XaiModel`, its wire types and chat converter, and the Hugging Face chat constructor. `xai/convert.rs` keeps only the helpers the Responses converter uses. - Tests of the removed model go with it (the chat-level modules of `xai_test.rs`, the xAI / Hugging Face rows of the OpenAI-compatible suites). The xai-error and xai-provider cases now run against `responses(id)`. The 16 Hugging Face cassettes are unchanged and replay through the OpenAI-compatible package in `conformance_test.rs`. - The xAI Responses model sends `top_k` and warns for `frequencyPenalty` and `presencePenalty`, as upstream does (found while moving the unsupported-parameter case). - docs / ROADMAP: same wording change as the docs PR follow-up (the extension is withdrawn; superseded ROADMAP lines rewritten), and the §5 product-difference row for the extension is dropped. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…in only for API calls The pooled client followed redirects with reqwest's default policy, which strips only Authorization-style headers on a cross-host hop, so vendor credential headers (x-api-key, x-goog-api-key, api-key, x-amz-security-token, ...) were forwarded to the redirect target, and a SigV4 request was replayed with its old signature. Redirect handling now lives in one place, the hop-by-hop loop that the validated-download path already had in http.rs (#163): - An ordinary API call follows a redirect only while it stays on the origin of the hop that issued it; a cross-origin 3xx is returned as a regular non-2xx response. Nothing is sent to the other origin, neither the request headers nor what a transport decorator would add. - Every followed hop goes through the request's transport again, so a signing transport signs the URL it actually sends. - A transport never follows a redirect: the pooled client is built with no redirect policy, and RedirectPolicy / FetchRequest.redirect / PinnedFetch::unpinned are removed. An injected Fetch gets the same protection because the rule is above it. - Validated downloads behave as before (per-hop validation and pinning, caller headers stripped off the credentialed origin). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…pi_key in ExternalProviderEntry Debug The Bedrock region was interpolated into the request host unchecked, so a crafted value could change the host. It now must be a single DNS label (the helper the Vertex location check uses) and fails with InvalidArgument before any request is sent. ExternalProviderEntry no longer derives Debug: a manual impl lists the same fields but reports only whether api_key is set. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…tten factory tests Kept: the tests that replay recorded AI SDK fixtures (fixtures/aisdk/*) in the OpenAI, OpenAI-compatible, Anthropic and Google factory tests, with their fixture-directory guard tests and only the helpers they use; one poll test proving the job-creating request is sent once when polling is exhausted. Dropped: the hand-written factory unit tests, the vendor/single-modality/ Vertex/transform-request-body/provider-trait test files, the other poll-stage tests, and the helpers and imports that became unused. Fixture replays and end-to-end behavior are the coverage we want; hand-written unit tests for factories are not. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…istry; drop the generated sources `aimux-providers/src/presets/` (26 generated files, 12,774 lines) and `scripts/gen_presets.py` are removed. `provider_registry.json` is embedded and parsed once into the descriptor table (`preset::entries` / `names` / `lookup`); a preset is created by name with `PresetProvider::create(descriptor, settings)`. The per-name Rust functions (`presets::create_<name>`, `presets::<name>()`) are gone with no replacement, so the registry data is no longer stored twice. - Every check the generator made on a row (key whitelist, auth / env_var pairing, family and auth values, template parameters matching the URL placeholders, duplicate names) now runs when the table is first loaded; an invalid registry panics there and fails the table test. - `presets_test.rs` keeps the table test (all 283 rows load and create without reading the environment) and one end-to-end case each for a keyless row, an unknown name and a rejected template parameter. - CI no longer runs `gen_presets.py --check`; docs and CHANGELOG describe the runtime table. docs/aisdk-architecture-alignment.md carries the same D12 / D20 wording as the docs PR follow-up. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Keeps SigV4 signing of the final request bytes, the bearer-token path and the region host-label check; the remaining hand-written factory cases are dropped like the other factory test files. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
Takes the design docs as merged in #200 (including the maintainer's review fixes) and keeps this branch's follow-ups on top of them: - xAI / Hugging Face expose Responses only; the Chat Completions extension recorded in the merged doc is withdrawn (§3.3, S4-7, revision note). - Presets are a runtime descriptor table (D12, D20, §0.3, §0.7, §4.1, §5, §6.5). - §3.2: redirects are handled by the helper, same-origin only for API calls; `FetchRequest` has no redirect field. - §0.3 drops NDJSON from aimux-stream; the reference-baseline header notes that fixtures are pinned by fixtures/aisdk/VERSIONS.json. - ROADMAP: the lines that still carried superseded commitments are rewritten (C2 shim, 0.7 / 0.8 version semantics, C1 coexistence, #174 / #175, #167, C4). Where the merged text already annotated a line (B1 + C2, #185, the Part II ts-rs sentence) the merged wording is kept. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: `create_provider_registry(providers, options)` returns a
`ProviderRegistry` over a caller-assembled map of provider ids to
providers. A model is addressed as "{provider}:{model}"; the id is split at
the first separator and the rest is passed to the provider unchanged. One
method per modality of the `Provider` trait, plus `files(provider_id)`.
Why: this is the one by-name mechanism the AI SDK has
(`createProviderRegistry` in `ai/src/registry/provider-registry.ts`). Until
now the workspace answered "name -> provider" in three separate places
that cover different vendors (the `provider(name, ..)` function for registry
presets, hand-written matches in the CLI and web tools for native vendors,
and per-vendor FFI constructors). The following commits move all of them
onto this registry.
Behaviour, as upstream: an id without the separator is `NoSuchModel`, an
unregistered provider is `NoSuchProvider`, a modality the provider does not
offer is `NoSuchModel`, the separator is configurable. The registry knows
no provider on its own and never falls back.
Not ported yet: language / image middleware and the `skills` accessor.
Tests: six cases ported from provider-registry.test.ts (ai@7.0.127).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…, default_providers) What: - `create_provider(name, settings)`: a name selects a factory, and the same few settings (key, base URL, extra headers, transport) go to whichever factory it is. The name is either a vendor package (`openai`, `anthropic`, `google`, ... 16 of them) or a row of `provider_registry.json`, which goes to the OpenAI-compatible factory. A package wins over a row of the same name. An unknown name is `NoSuchProvider`; there is no fallback. - `default_providers()`: every name with default settings, as the map `create_provider_registry` takes. Nothing is read from the environment and no entry can fail to be created; keys are evaluated per request. - `preset::create(name, settings)` returns an ordinary `OpenAICompatibleProvider` for a registry row. - `Provider::discovery()` (aimux extension, default `None`): the `/models` side of a provider is reachable from the provider, so a registry entry needs no second handle for it. A reference to a provider is a provider, so `&'static` default instances register as they are. Why: the AI SDK ships no vendor list. The reference for a built-in list on top of it is models.dev (what opencode consumes): one row per vendor naming the factory to use and what to pass it, native and compatible vendors in the same table. Until now only the compatible vendors were reachable by name, so the CLI, the web tool and the FFI each carried their own mapping for the native ones. This is that mapping, once. Tests: one end-to-end case, a registry over `default_providers()` serving a package, a registry row and a name both have. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
… create_provider What: `build_model` in `aimux-cli` (probe) and in `aimux-web` is now one call, `create_provider(name, settings)?.language_model(model_id)`. The web console's provider list is the keys of `default_providers()`. Why: both tools carried the same hand-written match (openai / anthropic / google / mistral / xai / cohere each to its own factory, everything else to the by-name `provider()` function) because the by-name function could not reach the vendor packages. With one table for every built-in provider the match has nothing left to do. 164 lines removed, 26 added. Behaviour change: a missing key for a vendor package is no longer checked when the model is built. It fails the request that needs it, with `LoadApiKey` naming the variable, which is when every factory evaluates its key. The provider list now also contains the vendor packages that the hand-written list did not name (azure, amazon_bedrock, google_vertex, ...). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
What: `rebuild_provider(record, api_key)` creates the recorded provider with `create_provider(provider_id, ..)` and takes its language model. If that model is not the one the recording was made with (the record's `provider` string, e.g. `openai.responses` against the default `openai.chat`), the rebuild is refused with a message pointing at `replay_with_model`. Why: it was the last caller in this crate of the by-name `provider()` function, and it shared that function's blind spot: a recording made with a vendor package (`anthropic`, `google`, ...) was `NoSuchProvider`. Those now rebuild. The refusal exists because the identity-only record cannot select a non-default method, and replaying a Responses recording against the chat endpoint would be silently wrong. Behaviour change: a missing key is no longer reported when the model is rebuilt; it fails the replayed request, like any other call. Tests: the three cases that registered a runtime overlay are replaced by one case for a registry row plus a vendor package, one for the refusal and one for an unknown id. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
The Mistral chat stream built tool-input-start/delta/end and tool-call parts with its own hand-written logic, assuming every tool call arrives complete in one chunk. Upstream (@ai-sdk/mistral mistral-chat-language-model.ts doStream) instead feeds every choice.delta.tool_calls entry to StreamingToolCallTracker, constructed with the model's id generator, and flushes it after the text/reasoning ends and before the finish part. mistral/model.rs now does the same with the existing aimux_provider_utils::StreamingToolCallTracker (as the openai and openai_compatible chat models already do): each delta is forwarded to process_delta, a malformed delta ends the stream with an invalid-response-data error, and flush() runs before Finish. The hand-written per-chunk emission is removed. A small private generate_id supplies ids for calls the server sends without one, standing in for upstream's generateId option. mistral/types.rs: DeltaToolCall now mirrors upstream's chunk schema (index, id and function.name/arguments are all optional), so partial deltas can reach the tracker. Module docs no longer claim tool calls always arrive complete. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
GroqProvider::chat (and call / language_model) used to return the generic
OpenAI-compatible chat model parameterised by a Groq dialect. The AI SDK's
@ai-sdk/groq has its own GroqChatLanguageModel and does not depend on the
compatible package, so the Rust package now has one too, file for file:
- model.rs mirrors groq-chat-language-model.ts: getArgs, doGenerate and
doStream, with streamed tool calls going through the shared
StreamingToolCallTracker (type validation "required", as upstream) and
stream error chunks mapped to a status through getGroqStreamErrorMetadata.
- convert.rs mirrors convert-to-groq-chat-messages.ts.
- prepare_tools.rs mirrors groq-prepare-tools.ts, with
browser_search_models.rs for groq-browser-search-models.ts.
- usage.rs mirrors convert-groq-usage.ts, finish_reason.rs mirrors
map-groq-finish-reason.ts, options.rs mirrors
groq-chat-language-model-options.ts, error.rs mirrors groq-error.ts, and
types.rs holds the response and chunk shapes the model reads.
- mod.rs builds the provider the way the Mistral package does (an
EndpointConfig per model, headers evaluated on every request); the public
surface (create_groq, groq(), GroqProviderSettings, "groq.chat") is
unchanged and GroqProvider::chat now returns GroqChatLanguageModel.
Where the existing package tests pin behaviour that differs from upstream, the
tests win and the model keeps it: the token limit is sent as
max_completion_tokens, the generic openaiCompatible options namespace is read
under the groq one, unknown fields of the groq options go to the body, and
provider metadata {"groq": {}} is reported.
groq/dialect.rs is left in place: the registry preset still builds Groq
through groq::profile() on the compatible model.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
DeepSeekProvider::chat used to return the generic OpenAI-compatible chat
model parameterised by a profile. The AI SDK's @ai-sdk/deepseek has its own
chat model and does not depend on the OpenAI-compatible package, so the Rust
package now has one too: DeepSeekChatLanguageModel, returned by chat(),
call() and language_model(). The provider keeps its public surface
(create_deepseek, deepseek(), DeepSeekProviderSettings, provider() ==
"deepseek.chat") and is assembled like the Mistral provider: a fixed
EndpointConfig, bearer credential headers loaded per request, the shared
exchange for URL, headers, transport and recording, and list_data_models for
discovery. Streamed tool calls go through the shared StreamingToolCallTracker.
New files under aimux-providers/src/deepseek/, each mirroring one upstream
file of packages/deepseek/src/chat/:
- model.rs <- deepseek-chat-language-model.ts (getArgs, doGenerate,
doStream, stream error classification)
- convert.rs <- convert-to-deepseek-chat-messages.ts (also carries the
file part handling of deepseek-file-part-options.ts)
- usage.rs <- convert-to-deepseek-usage.ts
- prepare_tools.rs <- deepseek-prepare-tools.ts
- types.rs <- deepseek-chat-api-types.ts (response and chunk shapes)
- options.rs <- deepseek-chat-language-model-options.ts (chat, message and
file part options)
- finish_reason.rs <- map-deepseek-finish-reason.ts
- is_v4_model.rs <- is-deepseek-v4-model.ts
The files/ module of the upstream package is not ported.
Where the existing acceptance tests pin behaviour that differs from upstream,
the tests win: thinking and reasoningEffort are sent as given with no
normalisation, unknown providerOptions.deepseek fields (user) reach the body,
topK is sent, reasoning is read as a fallback for reasoning_content, no JSON
system message is injected and a JSON schema becomes response_format
json_schema, strict tool flags are passed through without validation, an
assistant message with only tool calls has null content, and provider
metadata carries no cache counters or response ids. The old
deepseek::profile() stays only for the registry presets that still use it.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
There were three ways to turn a provider name into a provider: the old `provider()` family in provider.rs, `PresetProvider::create`, and the new `create_provider`. This commit leaves one. - provider.rs keeps only what the language bindings still call (`provider`, `provider_handle`, the runtime overlay for user-registered providers). `provider_handle` now checks the overlay and otherwise calls `create_provider`. The module is no longer re-exported from the crate root; callers name it as `aimux_providers::provider::...`. It goes away when the bindings expose the Rust surface directly. - Removed: `ResolvedProvider`, `resolve_provider`, `provider_discovery`, `provider_from_env`, `provider_names`, `provider_registry_entry` and the module's unit tests. Model listing is reached through `Provider::discovery()`. - preset.rs: the `PresetProvider` wrapper type is gone. `preset::create` returns a plain `OpenAICompatibleProvider`, the same type `create_openai_compatible` returns, so a preset is only a row of settings. - openai_compatible: drop `resolve_base_url`, which had no caller left. The FFI crate and the Node and Python bindings still import the removed root exports and are updated in the following commits. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…, moonshotai, fireworks, cerebras, baseten, alibaba, deepinfra)
The vendor test files built their models with the by-name `provider(..)`
and `provider_from_env(..)` functions, which are leaving the crate's
public surface. Their `make_provider` helpers and the individual tests now
call `create_provider(name, PresetSettings { api_key, base_url, .. })` and
`.language_model(model_id)` instead. Assertions are unchanged.
The `from_env_fails_without_env_var` test in each file used to assert an
error at creation time. `create_provider` no longer checks the key, so the
test is now async and asserts that the first `do_generate` call fails. No
tests were deleted.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
…e, openai_compatible, thin_wrapper, list_models, openrouter, presets, reasoning_map) The by-name `provider(..)`, `provider_handle`, `provider_discovery`, `provider_from_env`, `ProviderOptions`, `provider_registry_entry` and `PresetProvider::create` are leaving the crate's public surface. These test files now build their models with `create_provider(name, PresetSettings)` (and `preset::create(name, settings)` for the OpenAI-compatible presets) followed by `language_model(..)` / `discovery()`. Local helpers (`registry_model`, `preset_model`) carry the change so the call sites keep their assertions. Files with no use of the removed surface (anthropic_prepare_tools, openai_files, google_files, open_responses) are unchanged. Deleted: - list_models_test: `provider_handle_unknown_name` and `discovery_for_unknown_provider_is_no_such_provider` only exercised the removed by-name entry points (the unknown-name error is covered by presets_test through `create_provider`). - thin_wrapper_config_test: the two `from_env_fails_without_env_var` tests (togetherai, vercel) were `#[ignore]`d and asserted a creation-time missing key error that no longer exists. Moved: openrouter `from_env_fails_without_env_var` now asserts `LoadApiKey` on the first request instead of at creation. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
The previous commit gave the Groq package its own chat model but kept four behaviours that upstream does not have, because hand-written tests pinned them. Upstream wins, so in groq/model.rs (groq-chat-language-model.ts) and groq/options.rs (groq-chat-language-model-options.ts): - the token limit is sent as max_tokens, as getArgs does; - provider options are read from the groq namespace only; the generic openaiCompatible namespace is no longer merged in; - unknown fields of the groq options are dropped, not copied into the body; - no provider metadata is reported from doGenerate or doStream, as upstream returns none. A chunk that fails to parse is now reported the way upstream does it: an error part, finish reason "error", and the stream carries on; only a transport failure ends it. An error reported by the very first event still rejects the call, which is the convention of the other own-model packages in this crate and keeps the failure inside Core's retry boundary. tests/groq_test.rs now builds every model through create_groq(..).chat(..) instead of the by-name provider() function, so it exercises the new model. Tests that expected non-upstream behaviour are corrected to the upstream cases (reasoning mapped to low/medium/high, none handled per model, file references raising an unsupported-functionality error, tool_choice "auto", max_tokens, no metadata, groq namespace only) and upstream stream cases are ported: reasoning kept active across empty tool_calls, unparsable chunks, raw chunks, error chunks, and a response without choices. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
The DeepSeek chat model kept a list of behaviours that differed from the upstream package only because older, hand-translated tests pinned them. The repo rule is that upstream wins, so the model now does what deepseek-chat-language-model.ts, convert-to-deepseek-chat-messages.ts, deepseek-prepare-tools.ts and deepseek-chat-language-model-options.ts do: - thinking / reasoning: `thinking.type` `adaptive` becomes `enabled` with a compatibility warning; top-level `reasoning` derives `thinking` (`none` is `disabled`, anything else `enabled`) and `reasoning_effort` (minimal/low to low, medium/high to high, xhigh to max, with the same warnings); `reasoningEffort` `medium`/`xhigh` map to `high`/`max`; no effort is sent when thinking is disabled. - providerOptions.deepseek is validated like the zod schema: unknown fields are dropped, `thinking` and `reasoningEffort` are enums, `userId` errors carry the upstream messages. - `top_k` is warned about and not sent. - The model reads `reasoning_content` only; the `reasoning` fallback is gone. - JSON response format injects the "Return JSON." / schema system message (with a compatibility warning for a schema) and sends `json_object`. - Strict tools are rejected unless the base URL ends in `/beta`, and strict and non-strict function tools may not be mixed (prepare_tools). - An assistant message with only tool calls has `content: ""`. - Provider metadata carries promptCacheHitTokens, promptCacheMissTokens, responseObject, choiceIndex, messageRole, toolCallTypes, logprobs and systemFingerprint, in generate and in the stream. - The default base URL is `https://api.deepseek.com` (the request path is `/chat/completions`), as in createDeepSeek. - Stream error events keep the whole error envelope as the error data. - `request_body` and `RequestBodyResult` are private: they mirror the private `getArgs`. The tests are replaced by direct ports of the upstream describe-blocks of deepseek-chat-language-model.test.ts and convert-to-deepseek-chat-messages.test.ts in deepseek_chat_test.rs, built through create_deepseek(..).chat(..). deepseek_reasoning_test.rs is folded into it, deepseek_test.rs keeps the provider-level cases, and the DeepSeek cases of reasoning_map_test.rs that pinned the old behaviour are removed. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WGwanvo9WLU9HWD7sRR8TS
A name that is neither a vendor package nor a registry row fell through to preset::create, whose NoSuchProvider lists only the 281 registry rows. A misspelled "openai" therefore got an available-providers list without openai, anthropic or any other vendor package. create_provider now reports the full provider_names() list itself. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…factory # Conflicts: # CHANGELOG.md
[P1] Correct the HMAC derivation in the moved SigV4 signerThis is inherited from the previous Bedrock signer rather than a new regression in the refactor, but the shared
let k_date = hmac_sha256(
format!("AWS4{}", credentials.secret_access_key).as_bytes(),
date_stamp.as_bytes(),
);
let k_region = hmac_sha256(&k_date, credentials.region.as_bytes());
let k_service = hmac_sha256(&k_region, service.as_bytes());
let k_signing = hmac_sha256(&k_service, b"aws4_request");
let signature = hex::encode(hmac_sha256(&k_signing, string_to_sign.as_bytes()));This matches AWS's documented derivation. I reproduced the existing The comparison explicitly asserts equal canonical requests and strings-to-sign; For implementation, the existing Clarification: Rig vs. the standalone signerThe Rig Bedrock transport inspected here calls the full For aimux, I recommend using that official signer directly inside the existing decorator: This keeps aimux's existing request conversion, response parsing, call-level retry and fetch injection. It delegates canonicalization and cryptographic signing without introducing a Bedrock service client. See the official HTTP signing API. Check the signing settings against Bedrock's URI encoding, path normalization and header rules rather than preserving the old sign-every-header behavior. Dependency compatibility also needs checking: the locally inspected |
…S does The signing key chain passed every HMAC its key and data swapped and dropped the AWS4 prefix on the secret, so no signature matched AWS. The canonical request also left path segments unencoded (a ':' in a Bedrock model id), kept the query unsorted and did not collapse header whitespace. The golden signatures now come from botocore's SigV4Auth instead of the previous implementation. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A3WxDBWBA4ituJjXsHaEAa
…nd Polly Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01A3WxDBWBA4ituJjXsHaEAa
What
Moves provider construction, the shared request layer and their callers to package factories shaped by the AI SDK. It also adds the
fixtures/aisdksamples, the factory replay tests and the migration documentation. Core, provider-utils, vendor packages, FFI, bindings, CLI, web console and replay change together because they share these interfaces. Baserfc-0036/provider-result-types(A); follows A, precedes C. Replaces #222, #223, #224 and the previous content of #214.Layers
rfc-0036/factory-test-removalrfc-0036/upstream-samplesfixtures/aisdksample setrfc-0036/provider-factory-sourcerfc-0036/provider-factoryBefore and after
aimux-providers/tests/config_snapshot_test.rs,reasoning_map_test.rs,openai_transcription_stream_test.rs,bindings/node/__test__/body_overrides.test.tsopenai-provider.ts; provider request modules;openai-transcription-model.tsfixtures/aisdk@ai-sdk/anthropic,google,openai,openai-compatibleVERSIONS.json:ai7.0.127 with matching@ai-sdk/*; the local reference checkout for ported tests is 7.0.122create_*factories, settings and default instances create models throughProvider.create_provider,provider_names,default_providersassemble factories and compatible presets (an aimux extension)createOpenAI,createAnthropicand other package provider modulesreasoning_formatcreate_provider(name, PresetSettings),provider_names(),default_providers(),Provider::discovery()(discovery is an aimux extension); option isprovider_options.groq.reasoningFormat, converted toreasoning_formatgroq-chat-language-model-options.tsProvider::language_modelandcallselect Responses;chatselects Chat Completions explicitly; Responses factory supplies thefile-prefixopenai-provider.tsOption<String>keys and header maps; env-key fallback loads at request time; Azure andPresetSettingskeep resolvable keys, Azure also a token provider; Mistral, Cohere, Anthropic exposegenerateIdOPENAI_BASE_URLorANTHROPIC_BASE_URL, then default; an explicit empty key does not fall back to env. OpenAI-compatible reads no env key and sends no authorization for a missing or empty keyopenai-compatible-provider.tsOpenAICompatibleProviderSettingshassupported_urls,convert_usage;transform_request_bodyapplies to chat only; Groq and DeepSeek have own settings and conversionopenai-compatible-provider.tswithUserAgentSuffix,postToApiazure-openai-provider.tscreate_provider_registrysplits IDs at the first separator and applies language and image middleware; wrappers accept ID overrides; language hooks get generate and stream callbacks;NoSuchProvidercarries provider ID, model ID, model type, available providerscreateProviderRegistry,wrapLanguageModel,wrapImageModel,NoSuchProviderErroranthropic-language-model.ts,mistral-chat-language-model.tswith-user-agent-suffix.tsVendor option schemas live in each package: the shared OpenAI-style parser takes its schema from the caller; OpenAI, Azure, Groq, DeepSeek own theirs in
options.rs; Anthropic spells its namespace and skill type once there. The files model exposes no specification version, asscripts/check_provider_boundaries.shrequires. TypeScript-only constraints (JSONSchema7, dotted IDs,URL,Date, numbers) are not enforced upstream at runtime; Rust keepsValue,String,u32.Tests
anthropic/options.rsand provider-utilsfetch.rs, and the hand-written DeepSeek namespace and env-key cases; nobedrock_factory_test.rsorws_connector_test.rsremains.fixtures/aisdkthrough the shared mock-fetch helper. The samples are recordings:scripts/aisdk-fixtures/record.mjsrunsaiand@ai-sdk/*over a mockedfetchand writes the request sent and the normalised result. Requests and results are upstream behaviour; response bodies are fixed inputs. Replays check agreement with these records, not upstream test payloads.create_providerwith an unknown name reports every built-in name inNoSuchProvider.available_providers. Beforedfe9a453the miss fell through topreset::create, which lists only the registry rows, so a misspelledopenaigot a list without any of the 45 vendor packages. Unit test:default_providers::tests::an_unknown_name_lists_every_built_in_provider.Not in this pull request
aimux_providers::provider, and the call-layerModelMessagetool-result shape: covered by D.evaluationModelandcustomProvider: covered by C. xAI image and files, and OpenAI skills, are not ported.specificationVersionstays omitted by the public-surface decision. Search, discovery and built-in assembly are aimux extensions, not parity claims.Relation to the maintainer's design
Reference order: the pinned AI SDK source first; where it has no counterpart (C ABI, bindings, recording, code generation), the maintainer's RFCs 0039 (ops and bindings, #209), 0040 (provider construction and auth, #210), 0041 (replay governance, #211) and 0042 (V4 type boundaries, #212), which are Drafts that refine the #200 design doc
docs/aisdk-architecture-alignment.md; that design doc only where those RFCs defer to it or are silent. Where an RFC restates upstream behaviour differently from the pinned source, upstream wins and the RFC gets an amendment request. The stack targetsrfc-0036/integration; RFC-0039 §11 and RFC-0040 §7.2 require the whole closure to switch to master at once, so items marked "before master" must land on integration first. "Branch E" is the follow-up branch stacked on D, in progress.RECORDING_SCHEMAto 3 withinput,provider: ProviderRecord { provider_id, provider, model_id },exchanges,outcome,complete,transport_closed. It drops the configuration snapshot, but it is a subset of RFC-0041 §4's schema 3:target(model reference and identity),operation,capture_policyandCapturedvalues,model_calls, exchange id/phase/boundary/timing,ws_sessionsandreplayabilityare missing, and it identifies the provider throughProviderRecordinstead of §4.1'starget; RFC-0041 §2 dropsProviderRecordas a replay mechanism. Mock replay still decodes recorded OpenAI bodies inaimux-core/src/replay.rs(rebuild_generate_result) instead of RFC-0041 §3'sReplaySessionwithreplay_fetch/replay_web_socketrunning the real provider converters. Branch E replaces it with RFC-0041's schema 3 and replay, so no file with this subset ships under the number 3.ProviderRegistration(provider + extensions + metadata), pure-assemblyproviders_from_config(document). This PR hascreate_<vendor>factories,default_providersandcreate_provider(name, PresetSettings), which returnsArc<dyn Provider>; named methods such aschat,responsesandcompletionexist only on the concrete provider types and are lost through it.source=synthetic|approved_capture, redaction policy and review status. The 28fixtures/aisdksamples carry onlyVERSIONS.json; the 2,799 cassettes keep their format.Verification
Gate on the head
419ac220:cargo fmt --check,cargo clippy --workspace --all-targets -D warnings,cargo docwith-D warnings, tests of core, provider-utils, providers, ffi, web, replay and cli: 2,437 passed, 0 failed, 6 ignored; Node and Python bindingscargo check; provider boundary script;gen_ts_types.py --check;gen_providers_doc.py --check. Cassettes unchanged; the 29 files underfixtures/aisdkare added. Diff against the base: 502 files changed, 44382 insertions(+), 26782 deletions(-). After the restack onto the currentmaster, only this head was tested; the intermediate layers were checked with fmt and clippy only.Restacked head
78c9a41f. Restack (2026-10-09): mergedmaster(#213, SSE) and the rewritten #204 tracker down the stack by merge, no history rewrite. This layer ports the groq, deepseek, mistral and openai_compatible call sites to the new tracker API (the openai_compatible thought signature now travels asStreamingToolCallDelta::provider_metadata). Checked:cargo fmt --check,cargo clippy --workspace --all-targets -D warnings, tests of provider-utils and providers (107 test binaries, 0 failed).cargo doc, binding checks and the generator--checkscripts were not rerun.Fix
dfe9a453(merged into head9bd3da58): the new unit test passed onprovider-factory-source, on this head, and after merging intowire-format.presets_testandcargo clippy -p aimux-providers --all-targets -D warningswere run onprovider-factory-source. The full gate above was not rerun.Not verified
No GitHub check runs were attached at the last refresh. Go, Java, Kotlin, Swift and Flutter were not compiled locally in that refresh; Node
npm testand Pythonpytestwere not run. Live vendor calls and unported upstream cases are outside validation.🤖 Generated with Claude Code