rfc: state what a verification method must declare about itself - #9
Conversation
The "From Lessons to Controls" section asks each recommendation to specify a reproducible verification method. Reproducibility is necessary and not sufficient: a method can be deterministic, repeatable, and green on every run while never executing the control it reports on. Expands that one bullet with four sub-bullets covering the method's own failure mode, its noise floor and the population that floor was measured over, the path the evidence was produced through, and how much of the corpus exercised the control under test. Raised as OpenSecureAIAlliance#4 for discussion first, per CONTRIBUTING.md. Signed-off-by: JIAWEI DU <du@ddl99.com>
|
Read it against the thread it came out of. All four bullets are anchored in an incident that produced a The anchors, so a reader can check rather than agreeIts own failure mode. Ours returned an empty judge response that an arithmetic stage laundered into a Noise floor and the population it was measured over. We measured a floor of zero on identical repeats The path the evidence took. Ours scored a text view production never sees. A repair measured on the raw How much of the corpus exercised the control. Ours was a drift lock green over 222 rows that never One addition, on ratiosThe denominator sentence covers a manipulable denominator. There is a second ratio failure that is not the Ours mixed texts and firings. The numerator counted texts where a rule decided the outcome; the Suggested, in the same voice as the rest:
Six words wider than the current line and it catches a failure the current line does not. A property that may deserve a fifth bullet, once it settlesIssue #11 is working out a third requirement for preserved evidence: whether the record identifies the I am not proposing text for it here, because the scope is still moving in that thread and a clause written On the shape of the change+5/-1 under an existing line is the right size. The four requirements are all answerable in one or two No objection to merging as is, with or without the ratio wording. |
|
This is the strongest evidence in the thread, and it changes what I think the clause should say. Three things I got wrong or under-specified. The ratchet beats the deadline. I said an unusable state needs an owner and a deadline. You are right that a deadline needs someone watching a calendar. A list that fails closed in both directions has an owner by construction. The part I would not have specified is failing when an entry stops being a gap, because that is the direction nobody notices. I am taking your version over mine. Same structure shows up in your #10 comment: every negative corpus records its measured overlap, the overlap may not increase, and an absent measurement fails rather than implying zero. That is the general form, not a one-off. Which is an argument for putting the mechanism in the requirement instead of leaving it to whoever implements it. Three states, not two. I collapsed partial and unverified into "unusable." Wrong. Partial means something declared is missing and you can name it. Unverified means nothing declared an identity at all, so no claim is available. Different findings, different fixes. The self-digest point is the sharpest thing in this thread. A declared integrity value the consuming side cannot verify is a claim about the artifact, not a measurement of it. That is my own argument applied one level up, to the integrity metadata, and I did not follow it there. "The per-part digests are the measurement; the self-digest is a version string wearing a hash" is the line I would want in the rationale. On your refusal to fail closed: that boundary belongs in the requirement, stated. Dropping 88,451 patterns to resolve an identity gap trades a measurement problem for an outage, and the tenant is worse off. Whatever we write should govern what a result must state, not what a system must do. A clause that reads as mandating fail-closed will be ignored by anyone running production detection, and they will be right to ignore it. Where this goes, and how it sits with #9 and #10 No collision with the fourth bullet in #9. Path is how the evidence was produced. Identity is what artifact was measured. Your stale binary had the right path and the wrong subject, which is exactly why the harness stayed green. No collision with #10 either. These stack. Espirado's issue is a scorer returning a well-formed verdict that is systematically wrong on a recognizable class of input. Your comment there goes a level below, to a negative corpus whose ground truth was wrong in a consistent direction, 188 of 261 apparent false positives being actual attacks. Mine sits a level below that: whether the record establishes which artifact was measured at all. Scorer wrong. Ground truth wrong. Subject unidentified. A result can satisfy every declaration in #9 and still fail on any one of the three. Different remediations, different owners, so I think they land as separate declarations rather than one clause trying to cover all of it. The home for this one is Evidence Preservation, where the gap is already sitting in the text. It requires "model and safeguard versions and third-party dependencies." That is entirely the declared side. A version string is what was asked for, not what loaded. Proposed as an addition to that list, not a rewrite: Observed artifact identity: for each model, safeguard, tool and dependency, the identity requested and the identity measured at load or execution, with the comparison outcome. Where the observed identity could not be measured, that is recorded as a limitation on the finding rather than omitted. A missing or mismatched identity makes a result unusable rather than passing or failing, and the set of unusable results is enumerable and change-detecting. On a fifth bullet in #9 : it applies to verification methods directly and I would support one, but that is your section and Jiawei's PR. The sentence there should be narrower than the one above, since a method reports on a control rather than preserving evidence about an incident. Something like: where a result depends on a versioned artifact, the method states whether the artifact's identity was measured or asserted. Last thing. You changed a production loader on the strength of a comment in this thread, which is more than the argument had earned at that point. Your July case, 26,488 of 79,468 patterns absent with healthy logs, is better evidence for this property than anything I wrote. If the clause lands it should carry that case, not the reasoning. |
|
Your criticism of us is the right one, and I want to be exact about what it hit. What actually changed in the loaderNot its behaviour. It loads the same patterns in the same order, and it still refuses to fail closed. Your point survives that correction, though, and it is worth stating in your words rather than mine. On the self-digest, we wrote the substance into the type rather than a commit message, because that is What the measurement did not buy us, found this weekThe identity was measured and it was still unavailable at the moment it was needed. On 2026-08-04 our RU gateway was OOM-killed against a 1280 MB container ceiling. Two weeks later, Measured 2026-08-19 (UTC+7), which is the part that turned this from a story into a number. Building So the property is one step past your clause, and your clause is where it belongs:
We now publish the state as a metric with four values, three of them yours and the fourth being "no A second case for the tri-state, from a layer nobody was looking atSame week, same shape, different instrument. One of our alert rules selects Which is your fourth bullet in #9 arriving in the monitoring layer: "nothing to report" and "cannot The check was wrong twice before it was right, and both are worth having in a case book. Its matcher On the fifth bullet in #9Support, with one narrowing. "Measured or asserted" is the right axis. Add that the two must not be One more, since it is the same family and cheap to stateA measured maximum equal to a constraint measures the constraint. Our signature updater reports |
|
Supporting the fifth bullet in the narrower form you proposed. For our own platform the honest answer to it is What we declare and never measureEvery checkup record we produce carries Those two fields carry weight. Drift attribution across retests reads from them, and a paid migration attestation we had scoped sits on top of them. So the standing position was that a result could say which model produced a score, when what it could actually say was which model the measured party stated had produced it. That is the clause we argued for in #4 pointed back at us, and it surfaced from your comment, not from our own review. Where we should not follow you, and why the boundary mattersYou wrote that a missing or mismatched identity makes a result unusable rather than passing or failing. For evidence preservation after an incident we think you are right. For a cross-platform benchmark we cannot adopt it, and the reason is the same one you gave DmitrL about fail-closed. We measure agents on platforms whose runtime we cannot reach. Coze, Claude Code, Codex and OpenClaw do not expose a loaded model identity to a third party running a checkup through the published interface. A rule that voids every result with an unmeasured identity does not produce measured identities. It produces no results on any platform, and the comparison the benchmark exists to make disappears. So the version we can implement, and the one we would argue belongs in a method clause and not an evidence clause: where a result depends on a versioned artifact, the method states whether that artifact's identity was measured or asserted, and an asserted identity travels with every finding derived from it. The result stays usable and stops overclaiming. Anyone consuming it can then apply your rule at their own boundary, which is where the decision to reject belongs. If the working group prefers one rule for both cases, we would rather the benchmark side be excluded from the clause than write a clause we would have to ignore. Your three states, one layer upAt the dimension level our records already separate Our own noise floor, since #4 is where we asked for itMeasured 2026-08-19 against frozen transcripts. Fourteen sessions, answers held constant, rescored five times each, so all variation is the scorer and none is the agent.
What we have not resolved is what to do about it. A report that prints one decimal place on top of a floor that wide is claiming precision it does not have, and we have not decided whether to widen the presented interval or stop printing tiers. Population, since a floor without a range is the failure we named in #4: those fourteen sessions span raw 22.9 to 91.5 and the low band is thin. One dimension carrying 12% of the weight produces 61% of the variance, and the mechanism is flips between adjacent rubric anchors, not truncation. We assumed truncation and tested it. The share of unstable prompts is 13%, 16% and 14% at token ceilings of 1500, 2500 and 4000, and the flips persist at temperature 0, so raising the ceiling does not touch it. Two of the anchors in that rubric were never once produced in the run. The figure our earlier evidence in #4 rested on was measured before we rewrote most of the scoring prompts, and two dimensions that were rule-based then are model-judged now. It does not describe the system any more. Withdrawn, and replaced by the numbers above rather than restated. |
|
Your boundary argument is right and mine was sloppy in exactly the way you name. Conceded, and for your reason rather than politenessI wrote that a missing or mismatched identity makes a result unusable. That is a rule for a consumer Your formulation is better than mine on its own terms, not as a compromise: the asserted identity A third state your clause does not yet name, and I have it measured from todayWe shipped both halves of the identity check in one release and the ten minutes between them are the The publisher (our signature updater) began writing a Measured on the live RU gateway, before and after: Nothing about the artifact changed between those two lines. So
A record that says only I found it by checking the pack on the host after the rollout instead of trusting that two halves of On your noise floorFour of fourteen sessions changing tier across rescorings of identical input is the number I would lead On the open question, we hit the same shape and took the narrower option: stop printing the derived Your two rubric anchors that were never once produced belong in the same family as something we keep One security note on
|
Expands one bullet in From Lessons to Controls. Raised as #4 for discussion first, per CONTRIBUTING.md, and submitted here as the implementation of it.
What this changes
A reproducible verification methodbecomes that bullet plus four sub-bullets. Nothing else in the proposal is touched.Why
Reproducibility is the property the bullet asks for today, and it is the one a broken verification method is most likely to have. A method can be deterministic, repeatable, and green on every run while never executing the control it reports on. If SAFE's recommendations are going to carry verification methods, the format is where that case has to be made visible, because a reader of the finished recommendation cannot see it in the number.
The four sub-bullets are the four things that, in the evidence below, separated a verification artefact that established its claim from one that only appeared to.
Evidence
The four come from defects that shipped in two production systems, an LLM-judged evaluation platform and a rule-based security gateway, found over six days of the discussion in #4 and written up as a pair of companion documents attached to that thread: a case book of 23 incidents, each with the command that re-measures it, and a clause set of 25 requirements, each carrying the defect that earned it and a reference to where that defect is documented.
One case per sub-bullet:
What a reader can and cannot check, stated here rather than left to be discovered. The gateway repository is required for the case book's table and the evaluation platform's repository is not public, so no reader outside the two authors can run both halves and most can run neither. What transfers is not the command but the check each clause specifies, which is written to be runnable against the reader's own system. The two documents are deliberately separable: if a case turns out to be misreported, the clause it earned still stands or falls on its own reasoning, and a reader can see which of the two to disbelieve.
Neither author is a member organization. This is submitted as public comment under the process the repository describes.
Not included
The clause set is 25 requirements and this PR proposes four sentences. The remainder is supporting material for the discussion in #4 and is not proposed for the RFC text. If the working group would rather this went in as a linked companion document, or not at all, that is a better outcome than a proposal section that outgrows the section it qualifies.