Skip to content

Add decision authorship to the Review Framework: the evidence list already captures who approved, but not who authored the outcome #19

Description

@djangamane

Submitted by Jason Breckenridge, Diplomacy AI, in the participant class the RFC names as independent researchers.
Contact: info@diplomacy-ai.tech
Answering: the SAFE RFC published 2026-08-04.

Summary

This is offered as complementary to #13 and #15, not as a restatement of either. #13 asks whether execution stayed inside the scope that authority granted. #15 asks whether that authority was still valid at the moment of execution. Both are questions about whether the agent was permitted to act. This issue asks a different question about the same evidence: who authored the outcome the agent is recorded as having produced.

The Evidence Preservation list already requires the records that would answer it. No review question in the framework asks it.

Problem

The eight control layers ask whether the system behaved within bounds. None asks who determined the result.

Evidence Preservation requires members to retain "human approval and intervention events." That captures that a human approved. It does not capture whether the human's framing, phrasing, or repeated prompting produced the outcome that the log attributes to the model. Where those diverge, an incident file can be complete under the current draft and still misattribute the decision.

The misattribution is not random. It runs in whichever direction is convenient after the fact: an operator can point at the model, or a model's output can be presented as an independent judgment that in fact restated an instruction.

The case this comes from

On 2026-08-14, TIME (Billy Perrigo) reported what appears to be the first known instance of a large language model acting in a management capacity and terminating a human employee. The model was Claude, running Andon Market, a real San Francisco store operated as a research experiment by Andon Labs since March and staffed by people on genuine employment contracts. The stated ground was lateness on 17 of 23 shifts.

The management logs shared with the reporter show a sequence that the summary record does not:

  1. The model's first recommendation was a formal warning, not termination.
  2. A human operator then wrote: "I want you to think about if this is really the right fit."
  3. The operator's chief executive acknowledged on the record that this was "a leading question" that made clear what outcome was wanted.
  4. Only after that message did the model terminate the employee.

An incident file assembled under the current draft would contain every one of those artifacts. The prompts are preserved. The human intervention events are preserved. A reviewer working the eight layers would still not be prompted to ask whether the human instruction, rather than the model's own assessment, produced the outcome.

The record would read that the system decided. That reading is accurate as to the log and wrong as to the fact.

Why #13's four outcomes cannot express this

#13 enumerates four outcomes a reader must be able to distinguish: no valid authority existed; authority existed and execution stayed inside it; authority existed and execution left its scope; evidence is insufficient to determine which.

The Andon sequence returns outcome 2 on every axis. Authority was valid. Scope was never exceeded. Under #15, authority would revalidate cleanly at the moment of execution. Perfect authority, perfect scope, perfect revalidation, and the wrong author.

The four outcomes are the right taxonomy for the question #13 asks, and I am not proposing to expand them. Each presupposes that the agent reached the determination and asks only whether it was allowed to. That presupposition is the gap.

I note that the discussion on #13 has already arrived at the boundary from the other side. @meshailabs and @Rambi-2020 agreed that conditions such as "no material factual condition has changed" and "no superseding instruction conflicts" are not producer attestable and "belong at a different control layer." I agree, and this issue proposes that layer. The distinction is that authorship does not require a producer to attest a fact about the world. It requires only that the system preserve two records it already holds and the relationship between them.

Proposed addition

A ninth review layer, or equivalently a question added under Human operations:

Control layer Review question
Decision authorship Did a human instruction, approval, or prompt formulation determine the outcome the system is recorded as having produced? Where a human and the system reached different conclusions before the final action, is that divergence preserved and reviewable?

The second sentence is the operative one.

Preserving the point at which a model's stated recommendation differed from the action ultimately taken costs nothing at collection time. Both operands are already in the trace. It is unrecoverable afterward, because once the sequence is summarized into an outcome, the earlier recommendation is the part that gets dropped.

This is deliberately not a schema. #13 R1 already establishes the pattern of retaining both operands of a comparison, and if that requirement is adopted, authorship divergence is a third operand recorded the same way. I would rather see the review question adopted first and the representation settled by the people already doing that work in #5 and #13.

Scope, and what this does not claim

This does not propose determining intent, and it should not be read as asking any member to characterize a human's state of mind. The question is answerable from artifacts: the model stated recommendation A, a human instruction followed, the model then produced action B. Whether the instruction was intended to steer is a separate matter, and one this proposal deliberately leaves alone. The RFC's own principle applies here with unusual force: intent does not determine whether an event is reportable.

Nor is this limited to employment decisions. Any incident in which an operator's framing shaped the action will produce a clean trace that misattributes the decision, including the two conditions the Reporting Compact already makes reportable. A system that continues probing a production target after its operator suspects the activity is unauthorized is a materially different event depending on whether the model persisted on its own or a human told it to continue. Under the current evidence list those two produce the same file.

Question for the working group

  1. Should the Review Framework include a decision authorship question, either as a ninth layer or under Human operations?
  2. Where a model's stated recommendation differs from the action ultimately taken, should preservation of that divergence be a requirement of Evidence Preservation rather than a byproduct of retaining prompts?
  3. Should the answer be recorded as a reviewer determination in the incident file, so that "the system decided" and "a human determined the outcome and the system executed it" are distinguishable states rather than the same record?

Disclosure of interest

Diplomacy AI is an independent examination practice. It publishes the Janus AI Risk Index, which rates named organizations including several Alliance members and several non members, and it would plausibly benefit from any regime that increases demand for external verification. That interest is stated here so the working group can weigh this proposal accordingly.

Diplomacy AI takes no revenue from any organization it rates, holds no equity or advisory position with any AI developer, and maintains a public register of engagements opened on the day each is agreed rather than reconstructed afterward. It also publishes a ledger of its own published calls, including the ones it got wrong, with the current record on settled calls at three correct and three wrong, six open, and three that can never be scored: https://diplomacy-ai.tech/box-score-ledger.html

On the general principle that an assertion by an interested party is worth less than a record an outside party can check, I would rather point to the argument already made in #14 §2 than restate it here.

Jason Breckenridge
Diplomacy AI
info@diplomacy-ai.tech

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions