Repository navigation
Add Fireworks AI, Baseten, Nebius provider entries; refresh Together AI - #22
Merged
Merged
Conversation
Adds primary-source-grounded governance entries for three new inference providers, and refreshes one existing entry whose catalogue had drifted past the freshness SLA. Every entry was independently re-verified by a Cowork browser session after the initial automated-research draft, which surfaced real corrections (documented per-entry in each changelog block and in docs/primary-sources/<id>/browser-verified/). Fireworks AI (AOI 64.8, C): Zero Data Retention and the Response API's default-retention exception are confirmed via readable docs, but the Terms of Service and DPA turned out to be unreachable to a real browser (page-level noindexed) - confidentiality is coded unknown rather than asserted, and sub-processor disclosure corrected to false. Baseten (AOI 73.6, B): a genuinely mutual confidentiality clause was confirmed on two independent live fetches, and a real 99.9% SLA was found where none had been located - but a prior ISO 27001 claim was retracted for lack of any basis anywhere. Nebius / Token Factory (AOI 60.8, C): its central finding - that the binding Terms of Service contradict the marketing claim of no training on customer content - was confirmed verbatim. Three overstatements were corrected downward (EU-region count, a claimed inference SLA that does not exist, and a mis-scoped 'dedicated-only' catalogue claim), while compliance strengthened with a named SOC 2 auditor. Together AI: catalogue/pricing refreshed (a generation behind, and its prior 'embeddings' claim was wrong - none are served via serverless anymore); confidentiality-disclaimed finding re-confirmed unchanged, and zdr_available corrected from yes_default to yes_optional. Validation: python scripts/validate.py -> 33 entries, 0 errors, 0 warnings. cd site && npm run build -> clean, 161 pages. Dash/slop-clean; British spelling.
Deploying aoi with
|
| Latest commit: |
52f0218
|
| Status: | ✅ Deploy successful! |
| Preview URL: | https://f48be5ed.aoi-a63.pages.dev |
| Branch Preview URL: | https://batch-inference-providers-se.aoi-a63.pages.dev |
PR #22's CI surfaced that every model entry (last_verified in late July/early August) had drifted past the 30-day fast SLA. A Cowork browser session re-checked each entry's licence, safety posture and availability against current primary sources - not a bulk date-bump. 20 of 21 re-verified clean with no score change. Several flag newer- generation releases as new entities for a future build pass, deliberately not folded into the existing entries' scores: the DeepSeek V4 line and V3.1/V3.2, a licence-divergent GLM-5 generation (custom licence, MaaS security-review gate), Nemotron 3.5 Lightning 30B-A3B (distinct licence), OLMo 3, EuroLLM-22B-2512, Qwen3.5/3.6/3.8, and Kimi K2.5/K2.6/K2.7-Code. soofi needed actual judgment, not just a re-check: the research surfaced that Soofi-S-Base's model card reads 'closed-beta' and the repo is gated, in tension with the project's own 'open-source' marketing. On review, the entry's 2026-07-25 correction already accounts for exactly this (openness held at the open_weights ceiling, ownership already partial) - no further downgrade warranted. Full research pack archived under docs/freshness-sweeps/2026-09-21/. Note: 8 more entries (3 hosting-providers, 5 inference-providers) are also past the freshness SLA but were out of scope for this pack - python scripts/validate.py still reports 8 errors pending a follow-up pass on those. Validation: python scripts/validate.py models -> 21 entries, 0 errors. cd site && npm run build -> clean, 161 pages. Dash/slop-clean.
…ntries The second half of the PR #22 freshness debt - the same daily-cron finding as the 21-model sweep, out of scope for that pass. Six entries re-verified clean with no score change: Hugging Face (malware/pickle scanning unchanged), Berget (confidentiality/ZDR unchanged, plus a positive delta - a refreshed DPA now names five EEA-only sub-processors), Groq (Sec 10 mutual confidentiality unchanged; the 'Eligible Customers' ZDR gap persists; corrected stale 'very recent' wording on the ~14-month old Helsinki data centre - no score change, since no region-pinning clause exists either way), Infercom (on-prem/air-gap still shipped; SOC 2 badge seen elsewhere confirmed facility-level, not Infercom's own - already correctly uncredited), and two hosting providers where a fresh in-window CVE each (llama.cpp/GGUF: CVE-2026-52131; Ollama: CVE-2026-85180, offset by an in-progress, unmerged signing PR) reinforced rather than moved their existing grade. Two needed real re-grades. DeepInfra's Terms of Service was rewritten (effective 2026-08-17) to add a genuine mutual confidentiality clause (Sec 17) and make Zero Data Retention contractual and controlling with a bounded 30-day exception (Sec 7(b)), replacing the prior open-ended debugging carve-out - data_governance 3 -> 4, headline 52.0 -> 56.8 (D -> C). Runware's current Terms no longer contain the adverse, non-confidentiality clause a prior pass had found - the customer now owns Outputs outright with no broad licence granted to Runware, and a verbatim no-train clause resolves trains_on_inputs from unclear to never - lifting the adverse hard flag and data_governance 2 -> 3, headline 45.6 -> 50.4 (grade stays D; compliance and residency, unrelated to this finding, were not re-examined). Runware's correction carries a load-bearing caveat, documented in the entry: the removal is inferred from the current text's absence of the old clauses, not a verbatim historical diff, since web archives were proxy-blocked this pass. Research pack archived under docs/freshness-sweeps/2026-09-21-providers/. Validation: python scripts/validate.py -> 33 entries, 0 errors, 0 warnings - the full registry, for the first time this session. cd site && npm run build -> clean, 161 pages. Dash/slop-clean.
welsbach
self-requested a review
September 21, 2026 13:14
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds primary-source-grounded governance entries for three new inference providers, and refreshes
one existing entry whose catalogue had drifted past the freshness SLA. Every entry was
independently re-verified by a Cowork browser session after the initial automated-research draft,
which surfaced real corrections - documented per-entry in each
entry.yamlchangelog block andin
docs/primary-sources/<id>/browser-verified/.Fireworks AI (new, AOI 64.8, grade C) - Zero Data Retention and the Response API's
default-retention exception are confirmed via readable docs, but the Terms of Service and DPA
turned out to be unreachable to a real browser (page-level noindexed, ROBOTS_DISALLOWED) -
confidentiality is coded
unknownrather than asserted, and sub-processor disclosure is correctedto
false.Baseten (new, AOI 73.6, grade B) - a genuinely mutual confidentiality clause (Sec 11) was
confirmed on two independent live fetches, and a real 99.9% SLA with service credits was found
where none had been located - but a prior ISO 27001 claim was retracted for lack of any basis
anywhere.
Nebius / Token Factory (new, AOI 60.8, grade C) - its central finding, that the binding Terms
of Service (Sec 7) contradict the marketing claim of no training on customer content, was
confirmed verbatim under independent review. Three overstatements were corrected downward
(EU-region count: four public regions claimed, two actually public; a claimed inference SLA that
does not exist at all; and a "dedicated-only" catalogue claim that turned out to be ordinary
catalogue churn), while compliance strengthened with a named SOC 2 auditor (Deloitte).
Together AI (refreshed) - catalogue/pricing updated (a generation behind; its prior
"embeddings" claim was wrong - none are served via serverless anymore); the confidentiality-
disclaimed finding was re-confirmed unchanged on two independent fetches, and
zdr_availablewascorrected from
yes_defaulttoyes_optional(Zero Data Retention is account-level opt-in, notthe default) - a field-level fix with no score impact, since
data_governancewas already cappedat 3 by the confidentiality disclaimer.
Why this batch went through a second pass
The first draft of these three new entries was built from automated web-fetch/search-agent
research, not a human-supervised browser session reading the actual governing documents. That is
not the AOI process. A Cowork browser session then re-verified everything against the same
primary sources, and the corrections above are the result - some downward (retractions), some
upward (stronger grounding for what held up), one genuinely important non-finding (Fireworks' ToS
and DPA are not independently reachable by a normal browser, which the earlier automated pass had
implied it read directly).
Validation
python scripts/validate.py-> 33 entries, 0 errors, 0 warnings.cd site && npm run build-> clean, 161 pages.docs/primary-sources/<id>/_sources.mdanddocs/primary-sources/<id>/browser-verified/present for all three new providers and updated for Together AI.
Disclosure (ASDD)