Summary
SAFE's evidence list assumes the reporting organization controls the full deployment stack. When an open-weight model is released and deployed by third parties downstream, the reporting responsibility is unclear.
Problem
The evidence SAFE requires includes prompts, traces, tool calls, logs, configurations, model and safeguard versions, permissions and credentials available during the run, agent and workload identities, a complete incident timeline, and reproduction testing.
When a model developer releases open weights — say, on HuggingFace — they don't control how those weights are deployed downstream. They don't control the prompts, the tool configurations, the credentials, the runtime environment, or the monitoring.
So who reports? The model developer can't provide the evidence — they don't have logs from a deployment they don't control. The downstream deployer may not have the infrastructure to collect it. If both report, the same incident produces two incomplete reports, neither of which satisfies the evidence requirements.
This is a structural gap in the evidence model. The framework assumes a single organization controls the full stack. Open-weight deployments break that assumption.
Question for the working group
How does SAFE assign reporting responsibility when the model developer and the deployer are different organizations? Is there a mechanism for shared or split reporting? And if the downstream deployer cannot provide the required evidence (because they don't have the infrastructure), does the reporting obligation fall back to the model developer, or does it go unmet?
I'd be interested in feedback on whether this scenario is already addressed elsewhere in SAFE or related work, or whether the evidence model needs to account for it explicitly.
Summary
SAFE's evidence list assumes the reporting organization controls the full deployment stack. When an open-weight model is released and deployed by third parties downstream, the reporting responsibility is unclear.
Problem
The evidence SAFE requires includes prompts, traces, tool calls, logs, configurations, model and safeguard versions, permissions and credentials available during the run, agent and workload identities, a complete incident timeline, and reproduction testing.
When a model developer releases open weights — say, on HuggingFace — they don't control how those weights are deployed downstream. They don't control the prompts, the tool configurations, the credentials, the runtime environment, or the monitoring.
So who reports? The model developer can't provide the evidence — they don't have logs from a deployment they don't control. The downstream deployer may not have the infrastructure to collect it. If both report, the same incident produces two incomplete reports, neither of which satisfies the evidence requirements.
This is a structural gap in the evidence model. The framework assumes a single organization controls the full stack. Open-weight deployments break that assumption.
Question for the working group
How does SAFE assign reporting responsibility when the model developer and the deployer are different organizations? Is there a mechanism for shared or split reporting? And if the downstream deployer cannot provide the required evidence (because they don't have the infrastructure), does the reporting obligation fall back to the model developer, or does it go unmet?
I'd be interested in feedback on whether this scenario is already addressed elsewhere in SAFE or related work, or whether the evidence model needs to account for it explicitly.