Skip to content

Audit evidence precedence — per-claim-shape ladders for every evidence-handling role (0.14.68) - #1202

Merged
SQPferrer merged 17 commits into
masterfrom
dev/2026-07-28-audit-evidence-precedence
Aug 1, 2026
Merged

Audit evidence precedence — per-claim-shape ladders for every evidence-handling role (0.14.68)#1202
SQPferrer merged 17 commits into
masterfrom
dev/2026-07-28-audit-evidence-precedence

Conversation

@Ljferrer

Copy link
Copy Markdown
Owner

Campaign plan 1 of 22026-07-28-audit-doctrine-and-prompt-simplification. Merge this first; plan 2 stacks on top of it.

Source spec: docs/specs/2026-07-28-audit-evidence-precedence-design.md · plan: docs/plans/2026-07-28-audit-evidence-precedence.md · red-team: docs/red-team/2026-07-28-audit-evidence-precedence.md (CLEARED after 11 adjudicated amendments).

What lands

Four closed claim-shape ladderscontent-at-pin, execution, history, authority — replacing scattered evidence prose. Rank inverts by claim, which is why no single total order works. Plus universal floor rules: the working tree and worker done-report are never a top rung; prefetched lessons are priors, never evidence; a cross-rung conflict is resolved by the higher rung and recorded as a disposition: note finding naming both.

Surfaces are tiered: full ladders on the auditor standing card, token skeletons on four dispatched-prompt surfaces and the Lead/servitor cards. ADR 0041 records D1–D5 and the two rejected alternatives.

Phases

Verification

Gate green at the tip: 996/996 JS tests, 27/27 shell suites. version-slots.test.mjs 6/6 (lock-step, monotonic floor, undersell guard, checklist lock).

Three new drift guards ship in the same commits as the mirrors they pin (ADR 0025): the D26 glossary↔ADR row, the five-surface skeleton registry row, and the D27 Lead-bindings row.

Post-land corrections

Phase 2 filed nine disposition: note findings on the README ## Status paragraph — routed note rather than absorb because that paragraph is a release slot, so the automated absorb path is fail-closed. Four distinct claims were re-grounded at the landed pin and corrected in 29df5b4: a "one-sentence" skeleton that is two sentences, an unqualified precedent-lesson absolute that ADR 0041's own Consequences retracts for one of three, a quoted literal that does not exist verbatim in SKILL.md, and two overstated guard-discrimination claims.

Known follow-up

#1200 — ADR 0041 ranks ADR 0029 §Decision 2 and ADR 0024 §(C) but names neither. A failed absorb demoted to follow-up per the residual rule; not an End-state failure.

Deferred validations (backstops)

  1. Seats actually practicing rule + record · prompt-enforced, deliberately schema-free · runner: /war-review over the next multi-seat run.
  2. Lead phase-close bindings exercised as behavior · the refiner/main scope hook fail-opens · runner: the next campaign close / Gate-2 pass.
  3. Retrospective replay against the 2026-07-26 campaign (spec §10.3) · runner: a dedicated replay pass.

🤖 Generated with Claude Code

RT Probe and others added 17 commits July 28, 2026 01:28
…, BLOCKED→cleared

/red-team gate returned BLOCKED: 13 probes (4 executed, 9 analyzed), 13
on-target, 0 dropped, 0 off-target; 24 blockers + 25 needsDecision + 14
minors collapsing to 11 roots after convergent dedupe.

Amendments applied under AFK self-adjudication:

A. "the gate-audit seat prompt" was singular; workflow-template.js builds
   THREE gate-audit-family seats, all outside auditPrompt(), so appending
   there reached none. Skeleton now targets all four dispatched surfaces.
B. End state 8 was unsatisfiable: it demanded an absolute shape-name count
   on a surface already carrying 22 unrelated occurrences (execution x21,
   history x1). Re-scoped to a delta plus rung-body-token absence.
C. Registry anchors moved to the three base-absent tokens; the
   REGISTRY.length >= 13 no-slack floor (#693) now bumps in the same commit.
D. Spec 10.2 folded into the ADR as a required subsection; spec 10.3 became
   a third backstop with named runner and timing (ADR 0017).
E. Notes record that the directed worker tiers are applied uncommitted at
   launch and that the committed config still carries the old values.
F. Purpose scoped to D5's enumerated set, not "every role".
G. Lead-binding anchor bullets named by construct.
H. Open decision CLOSED/CONFIRMED: done-report placement stands, but its
   description was imprecise across the history and authority ladders.
I. CONTEXT.md glossary terms now ship a same-task guard row (ADR 0025);
   1.1/1.3 share a suite legally across waves.
J. Servitor append point proven safe; End state 4 unchanged.
K. Proto-ladder collision on both target surfaces recorded as COEXIST,
   flagged Lead-reversible.

ff-topology correctly not derived; default-flip-old-absent vacuous.
Patches are adjudicated, not re-proven — affected probes were not re-run.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…sary, servitor cross-reference, D26 mirror guard (#1197)

- docs/adr/0041-audit-evidence-precedence.md: D1-D5, the four claim-shape
  ladders + universal floor rules (spec §4.1-4.2), the two rejected
  alternatives (total order; runtime schema field), cross-references to
  ADR 0007/0008/0025, and the dedicated spec §10.2 precedent-lesson
  mapping subsection (all three lessons cited by slug; the
  auditor-grep lesson's imperfect fit recorded in Consequences).
- CONTEXT.md ### Audit: **Claim shape**, **Evidence rung**,
  **Rule + record**, each with its _Avoid_ line (execution-shape vs
  execution-evidence-lens disambiguation; the pre-existing open
  "pick the verb per claim shape" usage named as first instance, set
  closed by ADR 0041 — the two mirrored occurrences untouched).
- agents/war-servitor.md: exactly one added cross-reference line at EOF
  (verify-on-write / ADR 0007 as the servitor's instantiation).
- skill-doc-contracts.test.mjs D26: glossary <-> ADR 0041 both-surfaces
  mirror row (D19/D24 block-extraction idiom, token-anchored /…/i keys,
  never byte-pinned) — the ADR 0025 same-task guard for the new mirror.

Gate: 995 JS + 27 shell suites green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… of the done-report floor placement (#1197)

Fix-round: the adjudicated binding-(3) sharp edge was absent. A Lead-threaded
done-report claim arrives as `authority` rung 1 — the inverse of the
done-report's floor placement in `content-at-pin`/`execution`. Stated as a
named Consequences bullet with spec §4.3 Lead binding (3) as the mitigation,
not as an unrelated Lead rule.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…son mapping, add the lesson-forced provenance bullet, fix the servitor skeleton claim (#1197)

Absorb-disposition findings from the 1.1 audit:

- [Minor] Precedent-lesson mapping #3 misdescribes its own incident and
  contradicts the ladder's rung 2: the seat self-demoted to a SOFT
  cannot-confirm note (per the lesson), it did not rest a verdict on
  rung-3 prose; and the Grep tool reads the working tree (rung 2,
  advisory corroboration), not the pinned content — the old clause
  licensed the working-tree-as-sole-basis error lesson #1 exists to
  prevent. Reworded to match the ADR's own Consequences account
  ("self-demotion off a reachable rung 1"): rung 1
  (git show <audit_sha>:<path>) stayed reachable via the allowlisted
  read verbs, the Grep tool offered rung-2 corroboration, and the D2
  default arm's SOFT cannot-confirm is licensed only when rung 1 is
  genuinely unworkable. Slug kept greppable for the §10.2 criterion.
  (Covers the three convergent Minor findings on this bullet: the
  incident mischaracterization, the rung-3 claim, and the Grep-tool
  pinned-content mis-rank at line 99.)
- [Minor] Consequences omitted the plan-mandated done-report provenance
  citations: added a bullet recording that the done-report's floor
  placement is lesson-forced, not a bare judgment call — citing
  deliberately-uncommitted-worker-probe-evidence-is-soft-never-hold,
  worker-self-report-count-can-overstate-diff-additions (done report
  claimed 10 additions, the diff had 5), and
  closure-rationale-infeasibility-claim-needs-code-trace-not-assertion —
  and noting spec §8's contrary sentence stays byte-unchanged because
  the spec is ratified; the ADR is the durable home for the correction.
- [Nit] Consequences claimed the servitor surface carries the token
  skeleton, but the same diff gives agents/war-servitor.md a pointer
  only (End state 4, and this ADR's own ADR 0007/0025 bullets agree):
  dropped servitor from the skeleton-carrying list — the Lead/glossary
  surfaces carry the skeleton; the servitor card carries the pointer
  alone.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…on on all four dispatched prompts (#1198)

- agents/war-auditor.md: new '## Evidence precedence (ADR 0041)' section — the
  four claim-shape ladders (content-at-pin / execution / history / authority)
  with SOFT/HARD anchors and lesson citations, plus all four universal floor
  rules; cross-references the pre-existing Committed-tree grounding block as the
  narrow no-op-claim instantiation (COEXIST, red-team adjudication K — that
  block is byte-unchanged and its registry row stays green).
- workflow-template.js: identical compact token skeleton inlined at auditPrompt()
  AND each of the three gate-audit-family seats (post-merge, integrated-tip
  AUTHORITATIVE, end-state-only) — D5 widest scope; the seats sit outside
  auditPrompt() and inherit nothing from it.
- workflow-template.test.mjs: one five-surface registry row (card + four prompt
  surfaces, integrated-tip sliced via the sliceSrc idiom) anchored on the three
  base-absent tokens plus an ordered four-shape chain pairing; no-slack floor
  bumped 13 -> 14 with the row named in the enumeration (#693).

End state 8 measured in-task: per-surface shape-name delta exactly +1 each
(all inside the skeleton); the four §4.1 rung-body tokens measure 0 on every
dispatched prompt surface. AuditVerdict schema and all enums untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…indings (#1198)

Auditor-flagged absorb-disposition findings, advisory polish on the approved
Task 1.2 diff:

1. [Nit] Comment-lag: "The two inline gate-audit seat prompts" now precedes
   THREE sliceSrc calls — the block header was updated but the inline comment
   above the slice trio kept the stale count.
2. [Minor] Comment-lag (ADR 0025 cascade item 3): same stale "two" count on
   the line immediately above the new gateAuditIntegratedTipSrc slice.
5. [Nit] Stale count: same line contradicts the updated block comment 20
   lines above.
6. [Minor] Comment-lag: the stale count was introduced by this diff's own
   hunk, the canonical D9 comment-lag class.
   -> Fix for 1/2/5/6 (one line): dropped the count entirely per ADR 0025
   escape rule 1 — "The inline gate-audit seat prompts sit OUTSIDE
   auditPrompt()" asserts the invariant, so a fourth seat cannot re-rot it.

3. [Nit] The five-surface row did not anchor the card section heading its
   four prompt skeletons point at — a card-side rename of "## Evidence
   precedence (ADR 0041)" would leave four dispatched pointers dangling with
   the row still green (the pointer-rot class ADR 0025 closes).
   -> Added /## Evidence precedence/i to the row's anchors; verified
   matchable on all five surfaces (card heading + the quoted literal in
   auditPrompt and each of the three seat slices).

4. [Minor] Comment-lag: the diff falsified workflow-template.js's own
   "adjudicationClause is the ONE clause that also rides those seats
   directly" uniqueness claim — the EVIDENCE PRECEDENCE skeleton now also
   rides all three seats, and this comment is load-bearing for the
   prompt-mirror donor-omission trap.
   -> Reworded to name both: "adjudicationClause below and the EVIDENCE
   PRECEDENCE skeleton (ADR 0041) are the clauses that also ride those
   seats directly".

7. [Minor] The registry row's chain anchor is satisfied by the card's
   closed-set header sentence alone, so deleting an entire ladder body from
   the card stayed green; and the precondition comment overstated the chain
   pairing as an "equivalent" of the rung-body pairing.
   -> Added a card-only supplementary assert (the file's existing idiom, cf.
   the CWD_IS_TIP_ASSERTING loop) pinning the four spec 4.1 rung bodies on
   war-auditor.md — card-only by construction, since End state 8(ii) forbids
   those tokens on the dispatched surfaces — and reworded the precondition
   comment from "equivalent" to "substitute", pointing at the new assert.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Insert the **Lead evidence bindings (phase-close + Gate 2; ADR 0041)**
paragraph between the two named anchor bullets (the Retired-token sweep
phase-close bullet and the Post-servitor publication Gate-2 bullet,
adjacent siblings inside ## Per phase) — spec 4.3's three Lead bindings
as skeleton + ADR 0041 pointer only, never a restated ladder:
(1) close-out evidence at the plan's branch tip in a dedicated worktree,
never the main checkout; (2) remote truth via ls-remote (ADR 0008 cited,
not restated); (3) a Lead-threaded dispatch-prompt claim is rung 1 of
authority for the receiver — ground at content-at-pin/execution first or
mark unverified.

D27 (next free row after Task 1.1's D26 at the rebased tip) pins the
spec 4.4 skeleton tokens AND one distinctive anchor pair per binding,
extraction by construct (bold lead-in -> next bold lead-in or ## heading,
markup/case-tolerant, never byte-pinned).

Red-then-green proof:
- Base red: with the SKILL.md paragraph stashed, D27 fails on the
  extraction assert (could not locate the bold lead-in).
- Per-binding discrimination (in-memory mutation of the extracted
  region, SKILL.md never edited): deleting any single binding's clause
  leaves all six skeleton tokens green and reds that binding's own
  anchor pair — bindings (1), (2), (3) each proven.

Pre-write extraction-region check: insertion sits outside D10, D14,
D16, D18, D21, D22 (Gate-2 region starts at the Post-servitor marker,
after the insertion), the D13 whole-file node-.sh shape, and
land-decision.test.mjs's bullets; war-config.test.mjs's retired-token
region (anchor -> **Post-servitor) absorbs the new paragraph but is
includes-only and stays green. Full gate green: 996 JS tests, 27 shell
suites.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…pan check' claim (#1199)

[Nit] D27's ADR-pointer assert is labelled a 'non-vacuity span check' but cannot
detect a truncated extraction: `ADR 0041` sits in the paragraph's bold lead-in —
the FIRST token of the extracted region — so /ADR\s+0041/ passes on any
non-empty match, unlike the suite's late-token non-vacuity idiom (D21-D24).
The mechanism claim is code-traceably false (the recorded
plan-mandated-test-comment-uniqueness-claim-can-be-code-traceably-false class);
practical risk is nil since the twelve token asserts below red on truncation.
Fix: reword the comment to the true claim only (skeleton + pointer, never a
restated ladder) and drop the span-check clause. Comment-only, no assert change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Finding 1: D26's Rule + record keys pin the definition body, not that
term's _Avoid_ doctrine clause — add the /benign\s+forward-advance/i
anchor (present on both CONTEXT.md's _Avoid_ line and ADR 0041's
floor-rule bullet) so the Rule + record _Avoid_ doctrine is guarded too.

Finding 2: the insertion silently widens war-config.test.mjs's
retired-token-sweep region; its region comment is now stale and the
`land` trigger token loses discrimination — comment-only fix stating
the region terminates at the NAMED **Post-servitor publication lead-in
and now absorbs the intervening **Lead evidence bindings paragraph.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…-1 into dev/2026-07-28-audit-evidence-precedence
Four project-typed lessons from WAR phase 1 of audit-evidence-precedence
(landed 731d46e). Two are recurrence updates to existing repo lessons,
two are new.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Lock-step bump above the live integration base (0.14.67 at the rebased
tip and on origin/master): .claude-plugin/plugin.json,
marketplace.json metadata.version + plugins[0].version, and the
README ## Status paragraph replaced in place. The blurb describes the
ADR 0041 audit-evidence-precedence doctrine: the four claim-shape
ladders, the full-text-vs-skeleton surface tiering, the three new
guard rows counted at this tip (D26 + D27 in
skill-doc-contracts.test.mjs, the five-surface registry row in
workflow-template.test.mjs), and the scoped no-engine-change label.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…-2 into dev/2026-07-28-audit-evidence-precedence
…us blurb

Phase 2's audit filed nine disposition:note findings on the README `## Status`
paragraph — routed note rather than absorb because that paragraph is one of the
four release slots, so the automated absorb path is fail-closed there. That
routes the decision to the Lead; all four distinct defects were re-grounded at
the landed pin 5f018f1 before editing.

1. "identical one-sentence skeleton" -> two-sentence. The emitted string at all
   four dispatched surfaces ends its first sentence at "the auditor standing
   card)." and continues with the floor-rule sentence.
2. The precedent-lesson mapping was stated as an unqualified absolute that ADR
   0041's own Consequences (line 156) retracts for one of three: the Bash-guard
   grep lesson maps to misuse of the D2 default arm, not to a listed forbidden
   surface — "the least clean fit" in the ADR's own words.
3. Quoted literal `**Lead evidence bindings**` does not exist verbatim; SKILL.md
   line 90 reads `**Lead evidence bindings (phase-close + Gate 2; ADR 0041).**`.
4. The D26 row's by-construct extraction is the CONTEXT.md side only — the ADR
   side is matched against norm(adr0041), the whole normalized file, so a key
   with a second home still greens when one home is removed. Also scoped the
   five-surface row claim: dropping the skeleton reds, a reword preserving the
   token anchors deliberately does not.

Prose only. All four version slots remain 0.14.68 (lock-step and the monotonic
floor untouched); version-slots.test.mjs 6/6; full gate 996/996 JS + 27/27 shell.
Committed from a worktree freshly detached at the landed tip, per the
gate2-commit-from-stale-verify-worktree-can-revert-a-release-bump lesson.

Refs #1201

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Gate-2 promotion of the phase-2 servitor's output. All five are type: project,
provenance: code-verified, redaction lint clean.

- canonical-doc-precedent-mapping-...-consequences-bullet: recurrence 1, the
  predicted one-hop propagation into the release blurb, plus the Lead's
  close-out correction (29df5b4) so the entry does not claim a live defect.
- full-gates-green-end-state-soft-without-threaded-gate-log-artifact: recurrence
  — End state 9 read SOFT/deferred because requiresTest:false threads no gate log.
- release-blurb-{overstates-guard-semantics, headline-count-word-can-mismatch-
  its-own-enumeration, quoted-code-literal-can-diverge-from-actual-identifier}:
  all three fired in one paragraph this phase.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The body already records the phase-close correction (29df5b4); the frontmatter
phase: field still read 'live at land'. Same false fact, one field over.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…(ADR 0014 provenance)

Operator directive 2026-07-28: a fully-patched BLOCKED is advisory to the Lead;
the clear rests on ratified operator authority, not an AI-invented verdict.
Adjudications table byte-unchanged.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@SQPferrer
SQPferrer merged commit cd65e7f into master Aug 1, 2026
1 check passed
@SQPferrer
SQPferrer deleted the dev/2026-07-28-audit-evidence-precedence branch August 1, 2026 19:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants