Runbook and standing-record coherence — ledger adjudications contract, two-path land-failed prose, doctrine anchor + locks, manifest envelope, guard deny truth, T2.9 census (0.14.59) - #1109
Merged
Ljferrer merged 19 commits intoJul 26, 2026
Conversation
…tanding-record-coherence Pre-red-team plan edit, operator-ratified at the plan-1 Phase-1 checkpoint. Plan 1 (2026-07-24-land-advance-exit-contract-truth) Task 1.2 landed a T2.9 census comment asserting the push-error branch is "the only SILENT exit-3 route", with a universal "every one of which dies LOUDLY" lead-in. Two auditor seats independently code-traced both claims false — cmd_land_advance has a second silent bare-exit-3 arm (the post-push origin-readback mismatch on the push-success path) — and dispositioned them `note` rather than `absorb` solely because the wording was plan-mandated. Adds Task 1.6 (comment-only, deps: [], file-disjoint — owns provision-worktrees.test.sh, which no other task in this plan touches) and End state 11. Purpose count seven → eight; Decomposition note five → six Phase-1 tasks. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…-coherence Checkpointing in-progress /red-team patches so two rounds of work do not ride a compaction uncommitted. Verdict is still BLOCKED; round 3 (wf_4856642e-cba) is verifying these patches. NOT yet CLEARED. Round 1 (12 root defects, 16 blockers): - End state 9 was unsatisfiable on the plan's own happy path (a byte-untouched Files member can never appear in git diff --name-only) -> subset predicate with one named permitted absentee. - Task 1.2(e) mandated "mines one from arbitrary prose", not the censused literal -> End state 5's expected home would have returned zero hits. - End state 5 expected-homes incomplete (plan + spec carry the phrase); "one occurrence wrapped" was wrong -> census now enumerates, never counts. - End state 4's §3 row shape did not exist; envelope aggregates had no in-repo provenance -> row split mandated + attestation from three observed runs. - End state 7 absence floor too narrow; End state 2 had NO committed guard. - Task 1.6 (campaign carry-over) broke the stacking note and the backstop sweep. - #1085 spans three surfaces; Task 1.1 understated land-decision.test.mjs reads. Round 2 (9 defects, all in the round-1 patches themselves): - The new End-state-7 absence key was itself wrap-blind (zero hits at HEAD). - Lock (c) as an absence key would have RED-ed on the CORRECT two-path text -> respecified as a both-arms presence key. - The held:land-failed anchor literal has zero occurrences in SKILL.md. - Task 1.1's slice still contradicted its own End state 4. - #1085 does not close on plan 4's PR either -> closes on NEITHER. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…herence
Round-3 gate returned BLOCKED (7 Major). Adjudicated each against primary
evidence; 3 of 7 did not reproduce.
REFUTED (probe claimed the absence floors match zero and are vacuous):
an enumerated read flag with =-attached values -> 1 hit (tr-only AND sed+tr)
only silent / every one of which -> 1 hit each, all forms
The floors are non-vacuous and discriminating at HEAD. The true, narrower
defect: `tr '\n' ' '` leaves the `# ` continuation marker mid-phrase, so the
recipe does not defend a FUTURE re-wrap the way the plan's prose claimed.
Downgraded to one Minor and fixed by stripping the marker before joining
(proved 0/0/1 on a re-wrap fixture: raw miss, tr-only miss, sed+tr catch).
CONFIRMED and patched:
B1 provision-worktrees.test.sh contention is 1,2,5 -- not "exactly one file".
Plan 5 declares the same file; the campaign roadmap still records 1,5 and
predates the Task 1.6 fold. Recorded as a known stale record with a
Lead-owned correction route (roadmap is a campaign artifact, not a task
footprint).
B2 `T2\.` matches 90 lines file-wide and 3 in the census block -- one the
byte-pinned (b) line. End state 11 + floor (iii) were jointly
unsatisfiable. Re-scoped to the replacement sentence, as plan 1 scoped it.
B3 Task 1.5's "corrected" guard text named 3 of 6 value flags and 5 of 9 bare
tokens, then claimed "enumerated exactly" -- re-committing the exact
misdescription defect the task exists to remove. Now mandates transcribing
all fifteen tokens from the case arms; presence floor requires 15/15; the
completeness adjective is banned unless the list is complete.
B7 Task 1.3 slice said "two new tests"/"both locks"; it carries three.
Also swept: the census rationale claimed the lesson carries "one occurrence
wrapped" -- it carries four, two of them wrapped (round-1 blocker #4 recurring
in prose the enumerate-don't-count fix left behind).
Round-4 gate: BLOCKED, 7 blockers / 4 root defects. All four were introduced by my own round-3 patches (R4 by round-1). Reproduced every one before patching. R1 (3 probes, needsDecision) -- the round-3 region anchor was the DEFECTIVE sentence's own opener. Plan line 550 suggests the replacement open "Exit 3 is reached by several routes" while lines 198/543 anchored the scanned region on "Exit 3 is shared by": take the suggested wording and the locator evaporates; keep the locator and the reword is forbidden. Resolved toward the structural option -- the region is now the one the WORKER DECLARES (replacement sentence quoted verbatim in the done report), with no opening literal pinned at any of the three sites (End state 11, floor (i), backstop key (5)). R2 (2 probes) -- the 15/15 presence floor was a per-token SUBSTRING grep, and -a is a substring of --all, -r of --remotes, -v of --verbose/-vv. Measured: a message naming only the long forms scores 15/15 while omitting three tokens -- the partial-enumeration defect sailing through the floor built to stop it (worse than the probe's 14/15 estimate). Replaced with token-set EQUALITY derived mechanically from the two case arms' |-split patterns; the hardcoded fifteen is gone from both homes. R3 -- "three floors" with four enumerated (round-3 added the third). Floors are now numbered (i)-(iv); the count word is gone. R4 -- the /war-review Notes rationale still argued from "the spec's design tree deliberately chose exactly two new locks, a third would add contention" while this plan ships three (lock (c), added round 1). Reworded: three is the count, a FOURTH is where the line is drawn. Every fix REMOVES a rot-prone literal (opening anchor, hardcoded token count, floor count word, "exactly two") rather than adding a more specific one -- that literal-interlock is the class that produced new defects in all three prior rounds. Also added the step missing from rounds 2-4: a mechanical self-consistency sweep over every count word, anchor literal, and cross-reference BEFORE handing to verification. Sweep clean.
… fixes + minor sweep Round-5 gate: BLOCKED, 3 blockers / 1 needsDecision, 26 minors. The entire blocker surface was the two round-4 constructs; the whole-plan sweep, spine probes, and ff-topology returned Minors only. Operator present — the three decisions were grilled interactively (the skill's native Step-5 mode) instead of afk self-adjudication: ADJUDICATION 1 (B1, t2-region-declaration): Task 1.6's scanned region is STRUCTURAL and worker-independent — the census comment lines strictly between the block's last bare `#` paragraph-break line and its `# (b) ` detail line, located from the PAIR9= fixture anchor. Never an opening literal (evaporates on reword), never the done-report quote alone (a partial quote hid a T2. in the unquoted remainder — proven). Applied at End state 11, floor (i), backstop key (5). Fourth and final definition of this region: file-wide -> opening-literal -> worker-declared -> structural. ADJUDICATION 2 (B2+B3, token-set-equality-floor): floor (iii) equality is WHOLE-MESSAGE with the derivation stated at both homes: arm side |-split with trailing glob `*` stripped; message side boundary-anchored extraction (=-attached yields no token) with =<rev> collapsed to =. Consequence stated: the deny string names no denied token — teaching lives in the header comment. ADJUDICATION 3: narrow round 6 (the two construct probes only) closes the loop. Minor sweep (15 deduped, all probe-evidenced): ES7 key (2) + backstop key (4) switched to the #-comment normalized form (Task 1.5 rewrites that wrapped comment; line-local goes green on a re-wrap); ES4 never-carries grep now -i; census example now grep -oi|wc -l (grep -c caps at 1 on a joined line); `already spent` pre-change count corrected to 2 with "retry provably spent" reclassified as a hand-scan keep (ES2 + Task 1.2 + backstop key (1)); Task 1.2(b) re-anchored on the token-only 2-space prefix (compact wrap has zero occurrences in SKILL.md); status:"landed" D9 example replaced with the label form that actually trips (status:landed) + carve-out cited; "four sub-bullets" count dropped (five live; counts rot); "End state 3-style" cross-ref fixed; backstop why-deferred "six parallel tasks" corrected to six tasks across three waves; line-resident parenthetical scoped to keys (1)(2)(4)(5) with key (3)'s wrap-tolerance load-bearing at base; ES9 per-task diff base pinned (own dispatch base, not frozen phase base for deps tasks); ES9 "integrated tip" pinned (integration-branch tip at post-merge gate-audit, pre-Land, pre-docs(learnings)); ES9 straggler hedge scoped to the one in-footprint home; Task 1.3 classification list gains plan+spec homes and the fixed-in-flight ruling becomes present-at-base; latitude enumerations gain Task 1.6's sentence (Method + Open decisions). Self-consistency sweep run pre-commit: clean (the one flagged gap was my own line-local grep missing the wrapped structural phrase — the wrap-tolerant form finds all three sites).
…pletion, boundary hardening
Round-6 gate: BLOCKED, 2 blockers (1 new Major + 1 A2-completion Major/nd) +
9 minors. First round where claims-vs-reality and coverage-vs-source passed
with ZERO findings; A1 (structural region) held with one Minor edge.
NEW Major (dependency-feasibility, reproduced against live SKILL.md:68):
Task 1.2(d) added the envelope aggregates to schemas.md's MUST-carry list
and the On-phase-return stamp list but NOT to SKILL.md's own binding
"Field names follow spec 4.A ... MUST-carry set is binding" mirror sentence
-- landing that ships a fresh two-record divergence from a plan whose
purpose is killing record divergence. Fixed: 1.2(d) extends the mirror
sentence in the same edit; End state 4 checks the pair moves in lock-step.
A2 completion (token-set-equality-final, both probe-validated readings taken
under --afk as forced choices -- the rejected readings RED correct work):
- token TERMINATOR stated at both homes: boundary to first char outside
[A-Za-z0-9=<>_-] (trailing , ; ) . never part of the token);
- collapse rule widened: everything after a token's = discarded, placeholder
or concrete value alike (--sort=refname -> --sort=).
Region hardening (M1/M7): "last bare #" -> "FIRST bare #" at all three sites
(identical at base -- exactly one bare # -- immune to breaks a reword adds);
floor (iii) additionally pins "no new '# ('-prefixed detail-label line".
Minors: ES4's never-carries key restated wrap-tolerant (file hard-wraps at
~100 cols) and folded into backstop key (2); ES5 + Task 1.3 stop expecting
the new test lock as a census home (lock (a) is \s+-spelled by mandate --
invisible to the census by construction); backstop preamble taxonomy split
three ways ((4)/(5) absence, (3) census, (1)/(2) adjudicated sweeps).
…rence + r7 minor sweep Micro-round 7 (wf_02123020-0ef): 0 blockers, 0 needsDecision, 3 Minors. final-spec-execution PASSED with zero findings -- all three round-6 texts executed clean (equality floor GREEN/RED/RED, MUST-carry lock-step discriminates, first-bare-# region catches the two-paragraph evasion). Report at docs/red-team/2026-07-24-runbook-and-standing-record-coherence.md with the full operator adjudication table (structural region, whole-message equality + Lead-completed terminator/collapse, contention set 1-2-5, #1085 closes on neither PR). r7 Minors auto-fixed in the plan: schemas.md heading anchors noted as PREFIX matches (live headings continue past the short forms); the vacuous-key rationale corrected to rejected-as-line-local (the phrase matches once under the plan's own normalized form); backstop taxonomy's key-(2) count clause scoped to the manifest census half (its never-carries sub-key pins 1).
…view sourcing truth (#1101) skills/war/references/schemas.md: - `## ledger.json — run state` jsonc block gains a top-level `adjudications?` key (sibling of phases/pr_url?), documented as carrying rows verbatim as threaded in either args-contract row shape (string or { adjudicated|value, supersedes }); absent means none recorded. Closes the gap the doc-promises-ledger-field lesson flagged: the args-contract paragraph promised a "re-thread ... from the run-ledger record" that the ledger schema never actually defined a home for. - The `Optional adjudications (array|null)` args-contract paragraph now cites that key by name. - `## Run manifest` per-phase record gains `envelope: { totalTokens, totalToolCalls, agentCount } | null` (the Workflow task-completion envelope, stamped at phase return; unsourceable ⇒ null, never transcript-derived), added to the MUST-carry list with the same binding-to-attempt/null-tolerated posture as `workflowRunId`. The fail-open / never-resume-input charter sentences are untouched. skills/war-review/SKILL.md (same commit — spec §8 coherence pair): - §2: replaced the "the manifest never carries them" parenthetical with the new sourcing rule — prefer the manifest's per-phase `envelope` for totals when present; the input/output/cache split stays transcript-mined, n/a when unsourceable, and is now labeled best-effort/possibly undercounting (~20x calibration) so it's never read as cross-summable against an envelope total. Also extended the adjacent "manifest supplies" enumeration to include the envelope totals when present (a survey-derived reconciliation of the same claim, same footprint). - §3: split the one combined "total tokens — input/output/cache" row into `total tokens` (envelope-else-mined) and `token split — input/output/cache` (mined-or-n/a); set `total tool calls`' Source to envelope-else-mined too. - §4: added the "unfinalized phase record" friction signal (endedAt/tasks/ land missing although the run ended or a later phase started), with the killed-run discriminator carve-out so a genuinely dead phase stays silent. Token sweep: grep `manifest` over war-review/SKILL.md (32 hits post-edit) and hand-scanned §3 + ## Scavenge; wrap-tolerant `never carries` census is 0 post-edit (was 1). No other "never carries"-style paraphrase found; ## Scavenge's transcript-only tokens/tool-calls claim is correct as-is (those runs predate the manifest entirely, so the widened contract doesn't apply). requiresTest: false (docs-tier contract + consumer prose per the plan; the doc-contract lock guarding the new ledger key arrives with Task 1.3). node --test 'skills/**/*.test.mjs' (926/926) and the full anchored *.test.sh sweep are green, including land-decision.test.mjs's D9 enum-leak scan and doc-parity checks — new prose introduces no equality/label-shaped status/landDecision examples. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…pe-else-transcripts (#1101) Advisory polish (absorb-disposition findings, task 1.1): 1. [Minor] war-review §honesty-invariant still sources token/tool-call counts solely from transcripts — stale against this same commit's envelope-else-mined rule (skills/war-review/SKILL.md:23). Rationale: as of this commit, total tokens and total tool calls prefer the manifest phases[].envelope (§2, §3) and only fall back to transcript mining; the invariant's sourcing sentence misattributed where the totals come from — the exact record-vs-behavior drift this task removes. The headline claim holds (the envelope is itself a harness-surfaced read, not billing truth). Survey-derived correction: the sentence carries no `manifest` token so the plan's grep sweep cannot see it, and it sits outside the declared hand-scan scope (§3 tally rows + ## Scavenge). The frontmatter description stays untouched — it remains literally true (transcripts are still mined for the split). 2. [Nit] §2 preamble honesty-invariant still attributes token/tool-call counts solely to transcript files (skills/war-review/SKILL.md:23). Rationale: same sentence, same drift — after this diff a review following the old wording would attribute envelope-derived numbers to transcripts. Not a plan-faithfulness violation or regression; preamble is outside the swept scope, so per the grep-sweep-floor calibration it is a survey-derived correction. Normative content stays true; only the named mechanism under-described. One edit satisfies both: the invariant's first sentence now reads "come from the harness — the manifest `envelope` aggregates where present, otherwise Claude Code's transcript files — whose formats are harness-internal and may change". No other line moves; the never-fabricate invariant above is untouched. No test reads this file (plan Notes decline a /war-review prose lock); skills/war-review/SKILL.md is in this task's Files, in-footprint for End state 9. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
) docs/plans/2026-07-22-war-memory-hardening.md's own red-team round 1 added CONTEXT.md to that plan's Task 1.2 Files (reversing an original keep-no-edit call on two CONTEXT.md glossary entries); round 4 then caught that the 2026-07-22 roadmap was never updated to match — row 6's Files-owned cell omitted CONTEXT.md and the CONTEXT.md contention row's plan list ("2, 3, 4, 5, 7") omitted plan 6, per that plan's own Notes. Bookkeeping only, per this plan's Task 1.4 / End state 8 — the roadmap is a record, not the live queue, and the source issue itself records there is no landing hazard. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The READ-FORM branch arm admits two token shapes, but both the deny message
and the header comment described only one ("takes only =-attached read flags"
/ "EVERY token must be an enumerated read flag with =-attached values"). An
auditor seat hitting the deny read a blanket that is false for nine of the
fifteen accepted tokens, and the enumeration itself was partial — 3 of 6
value flags, 5 of 9 bare tokens.
Deny message now names the accepted set in full, split by shape:
value-carrying flags =-attached (six) and bare read flags (nine). Header
comment corrected to the same mixed-shape truth, with no completeness
adjective over a partial list.
Floor (iii) token-set equality, derived mechanically at this base (arm side:
the two case patterns |-split, trailing * stripped; message side:
boundary-anchored flag tokens, everything after = discarded) — both sets are
the same fifteen:
--all --contains= --list --merged= --no-contains= --no-merged=
--points-at= --remotes --show-current --sort= --verbose -a -r -v -vv
The same derivation on the pre-change message yields 8, missing seven.
Case arms, allow/deny behavior, and every other arm are byte-untouched; no
flag added to or removed from the accepted set. hooks/validate-auditor-git.test.sh
is byte-untouched — 87/87 green, J16's =-attached micro-teach pin still hits.
Partially addresses #1085 (the agents/war-auditor.md and workflow-template.js
mirrors are out of footprint here — issue stays open).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…im (#1106) Campaign carry-over from plan 1 (2026-07-24-land-advance-exit-contract-truth, Task 1.2): the mandated T2.9 census wording claimed exit 3 "is shared by multiple routes, every one of which dies LOUDLY with route-naming text" and that the push-error branch is "the only SILENT exit-3 route". Two auditor seats code-traced both claims false and dispositioned them `note` (not `absorb`) solely because the wording was plan-mandated; the operator ratified routing the correction here. Code trace of cmd_land_advance in provision-worktrees.sh AT THIS BASE (floor (iv) — re-read, not taken from the slice summary). die() is `printf … >&2; exit "${2:-1}"` and EX_FOREIGN=3, so: LOUD exit-3 routes (die + $EX_FOREIGN, all print route-naming text): - ls-remote rc-guard ("could not read the origin tip") - phantom land ("refusing to report a land that did not advance") - unresolvable HEAD ("could not resolve HEAD to a commit") - unresolvable <new-sha> ("does not resolve to a commit") SILENT exit-3 routes (bare `exit 3`, print nothing) — TWO, not one: - post-push origin readback: `[ "$actual" = "$new_sha" ] || exit 3`, on the push-SUCCESS path (push_rc -eq 0) - the trailing push-error branch, on the push-FAILURE path (push_rc != 0, no [rejected] token) — the route T2.9 itself exercises Both "every … dies LOUDLY" and "the only SILENT" are therefore false. Comment text only: no assertion, fixture, or case byte moved; the (b)/(c)/(d) detail lines are byte-unchanged and the block gains no new `# (`-prefixed detail-label line. Diff is 4 comment lines for 4; zero non-comment lines. Floor evidence (non-vacuity — every key matched at base, so each discriminates): (i) replacement sentence, structurally delimited from the PAIR9="$(setup_origin_pair)" fixture anchor (strictly between the block's first bare `#` and its `# (b) ` line), greps 0 for `T2\.` — sentence-scoped, as plan 1's End state 4 scoped it (file-wide is 90 lines and block-wide 3, either of which false-REDs) (ii) normalized (sed strip of the `# ` continuation marker BEFORE joining, then tr join + squeeze), case-insensitive: `only silent` 1 -> 0 and `every one of which` 1 -> 0 (v) count-free: no numeric route count ("two"/"three") — "several" Hand-scan of the rest of the T2.9 case found no other FALSE universality claim: "rejects every push" (the pre-receive hook is an unconditional `exit 1`), "the ONLY rejection reason is the hook" (fixture-scoped — new-sha is a ff descendant, so no non-ff rejection is reachable), and "the rejected push never advanced it" are each true and asserted; "never widen to bare `rejected`" is a directive on the byte-pinned (c) line. Verify: bash skills/war/assets/provision-worktrees.test.sh -> 380/380 PASS (no assertion moved). Full gate green: 926/926 JS tests, all shell suites. End state 11. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…, adjudication continuity, manifest stamps, doctrine framing Five construct-anchored edits to skills/war/SKILL.md (#1016, #1039, #1084, #1078 part 2, #1087): (a) Step 5: "record each row in the run ledger" now cites the ledger.json top-level `adjudications` key (references/schemas.md), closing the ledger-contract citation loop Task 1.1 opened. (b) The `held:land-failed` Outcome-handling bullet's environment-retry sentence no longer states an unconditional "already spent" claim. Verified against workflow-template.js's actual land-time routing (lines ~1900-1960): a **primary-land arm** (the land gate failure was itself classified `environment`, one environment-proceed re-land ran, and it failed `environment`-classified a second time) has spent its retry — the manual re-run is the second line of defense. A **baseline-proceed arm** (the land gate failure was classified `baseline`, a baseline-proceed re-land ran instead, and *that* re-land's own failure came back `environment`-classified) reaches the same hold with no environment-proceed retry ever dispatched for it — the two `*-proceed` flavors deliberately never chain — so the manual re-run there is genuinely the first attempt. The sibling D18-pinned `environment` bullet under "gate_failed routing by class" is byte-untouched; root-cause (a)/(b)/(c) prose and the resumeFromRunId warning inside the held:land-failed bullet are untouched (verified via a script mirroring land-decision.test.mjs's extraction: the region still reaches "dead land agent" and the resumeFromRunId negation, and now also carries both "primary-land arm" and "baseline-proceed arm" markers for Task 1.3's forthcoming both-arms-presence lock). (c) Recovery relaunch > Shared mechanics gains a fourth bullet, "Adjudication continuity" — re-threads the full accumulated args.adjudications set from the ledger record alongside args.recovery, the same duty the held-partial-phase runbook's step 4 already performs, so neither relaunch entry point is adjudication-blind. (d) Run manifest: the "On phase return" stamp list and the "Field names follow spec §4.A" MUST-carry mirror sentence both gain the envelope aggregates (totalTokens/totalToolCalls/agentCount), sourced from the Workflow task-completion notification's envelope, moving in lock-step with schemas.md's MUST-carry list (Task 1.1) per the r6 red-team finding. Checkpoint gains a new first bullet, "Manifest stamp (telemetry, fail-open)", immediately before the issue-lifecycle floor bullet. (e) Step 5's Provenance-discipline sentence gains, verbatim, "and a row is **never mined from arbitrary prose** — rows come only from the two producers above" — restoring the doctrine anchor's literal phrasing (never the "mines one from" paraphrase, per the r1 red-team finding). Token sweep: `already spent` over skills/war/ + agents/ — 2 pre-change hits (the target sentence, now conditional per-arm; workflow-template.js's ace-demotion string, a different retry budget, out of footprint, left). Hand-scanned the Outcome-handling list, Recovery relaunch, and held-partial-phase runbook sections plus agents/war-refiner.md for paraphrase echoes of the unconditional claim — none found; every existing mention of a second-environment-classification "retry spent" is already correctly scoped to that one case. Gate: node --test 'skills/**/*.test.mjs' (926/926, incl. skill-doc-contracts.test.mjs's D10/D18 rows byte-untouched) and the full *.test.sh sweep across hooks/ + skills/ all green. Refs #1102 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…three doc-contract locks (#1103) The "rows come only from the two named producers" doctrine had ZERO operative anchors. The 2026-07-22 audit-adjudication-threading spec justified overwriting the CONTEXT.md **Adjudication** _Avoid_ clause by citing a duplicate home at skills/red-team/references/lenses.md that never carried it, so the plan-faithful rewrite removed the doctrine's only live copy ([[spec-non-goal-citation-of-a-doctrines-home-file-can-be-wrong]]). - CONTEXT.md **Adjudication** term: the _Avoid_ line regains the clause. It composes with, rather than duplicates, the term body's two named producers and SKILL.md step 5's provenance-discipline sentence (hand-scanned both). - docs/specs/2026-07-22-audit-adjudication-threading-design.md §9: a dated bracketed correction note appended at the citing non-goal sentence. ANNOTATIVE — the ratified sentence is byte-unchanged. It records the false citation, the orphaning it caused, the restored anchor + its new lock, and the lesson. Re-confirmed no test reads that spec's bytes (grep of its filename over skills/ + hooks/ test files: zero hits), so the note cannot red anything. - skill-doc-contracts.test.mjs: three construct-anchored rows, house style. D19 (lock a) — the CONTEXT.md **Adjudication** block, extracted bolded-term -> next-bolded-glossary-term, keeps the clause. The key spells its inner spaces \s+, never literal spaces: CONTEXT.md wraps near 100 columns so the restored line may legitimately wrap mid-phrase, and the \s+ form also keeps the contiguous literal out of this file, which the plan's wrap-tolerant repo-wide census expects to return zero hits here (a guard must never trip its own floor). D20 (lock b) — the `## ledger.json — run state` jsonc block (heading matched by PREFIX, the live line continues `at .claude/teams/<run-id>/`; region ends at the closing fence) declares the top-level `adjudications` key. D21 (lock c) — the held:land-failed bullet, extracted EXACTLY as land-decision.test.mjs does it (2-space token-only prefix, SAME-INDENT `- **` terminator, never a top-level one), names BOTH environment arms. A PRESENCE key on the two-path shape, deliberately NOT an /already\s+spent/i absence key: the sanctioned replacement keeps that token inside the CONDITIONAL primary-land arm, so an absence key would red the correct text and green the wrong one. Red proofs — run in a throwaway copy of the read set; the worktree was never broken. Baseline 11/11 green; each break reds EXACTLY its own row (10 pass, 1 fail): (a) CONTEXT.md _Avoid_ line reverted to its pre-restore text -> only D19 fails. (b) the `adjudications?,` line deleted from the ledger jsonc block -> only D20 fails. One `adjudications` hit REMAINS in the file (the args-contract paragraph), so a whole-file key would have passed here — the construct scoping is load-bearing, not decorative. (c) the held:land-failed bullet collapsed to its pre-Task-1.2 single unconditional sentence ("...has already spent its in-workflow environment-proceed retry") -> only D21 fails, on the primary-land arm assert; "baseline-proceed arm" occurs 0 times in the collapsed bullet. D10, D12-D18 pass with their extraction constructs untouched. Full gate green: node --test 'skills/**/*.test.mjs' 929/929, 26 shell suites, exit 0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ase-1 into dev/2026-07-24-runbook-and-standing-record-coherence
Bump all four release slots in lock-step to the next free patch above the live integration base (0.14.58 -> 0.14.59), resolved from the slots as they stand at this base, never from a plan literal: plugin.json `version`, marketplace.json `metadata.version` AND `plugins[0].version`, and the README `## Status` line (replace-in-place, never emptied, no badge). Stack context: master 0.14.57, campaign plan 1 landed 0.14.58, so 0.14.59 is the next free patch. Release blurb is additive and claims no behavior change. Phase 1 corrected standing records and one deny string only: verified against the plan-1 tip f9fc4a4, `skills/war/assets/workflow-template.js`, `land-decision.mjs`, every merge-path floor, and every guard `case` arm are byte-untouched (the auditor guard diff is exactly one header comment + one deny string, so every allow/deny outcome is unchanged). The blurb deliberately does NOT quote the restored doctrine phrase literally. End state 5's census is repo-wide against a set of expected homes, and a README release note is not one of them; quoting it would plant a fresh census hit -- the recorded trap where a release blurb re-fires its own change's guard. The blurb describes the restore instead. Re-ran the census after the edit: the hit set is byte-identical to base (8 files, README absent), and README trips none of the plan's absence keys (`takes only =-attached`, `never carries`, `already spent` all 0). Gate green at HEAD: node --test 'skills/**/*.test.mjs' 929/929 pass 0 fail; 26 shell suites, incl. validate-auditor-git.test.sh 87/87 and provision-worktrees.test.sh 380/380. version-slots.test.mjs (the lock-step arbiter) green on all four slots. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ase-2 into dev/2026-07-24-runbook-and-standing-record-coherence
This was referenced Jul 25, 2026
Base automatically changed from
dev/2026-07-24-land-advance-exit-contract-truth
to
master
July 26, 2026 18:26
Ljferrer
deleted the
dev/2026-07-24-runbook-and-standing-record-coherence
branch
July 26, 2026 18:26
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Campaign 2026-07-24-standing-record-and-guard-hardening, plan 2/6 (stack-and-plow, ADR 0011 — merges bottom-up at campaign end; do not merge ahead of plan 1's #1097).
Phases landed (run
2026-07-24-runbook-and-standing-record-coherence-2026-07-24):3f136c0(6/6 tasks, unanimous approvals, 0 fix rounds, 0 escalations; gate-audit approve @992c0c5): ledgeradjudicationskey + manifestenvelopecontract +/war-reviewsourcing truth; five SKILL.md runbook edits incl. the two-pathheld:land-failedbullet and the verbatim doctrine clause; CONTEXT.md doctrine anchor restored + THREE new doc-contract locks (red-proofs in commit bodies); roadmap record cells; auditorgit branchdeny-string/header truth (token-set equality floor, operator-adjudicated derivation); T2.9 census uniqueness overclaim dropped (structural worker-independent region).3444016: all four version slots 0.14.59 in lock-step (version-slots.test.mjsgreen); additive blurb closing "No behavior change".c2fe413(phase 1, 3 lessons),53e15d9(phase 2, 3 lessons — tip).Engine untouched by design:
workflow-template.js,land-decision.mjs, every floor script, and every guard case arm byte-identical (spec constraint 1) — the phase-1 diff was a strict subset of the declared Files union with the one expected absentee (hooks/validate-auditor-git.test.sh, byte-untouched).Escalations: none. Follow-up filed: #1107 (war-review §3 run totals can silently mix envelope and mined sources — demoted from an unabsorbed Minor per the residual rule).
Unexecuted backstops: (1) Integrated-tip five-sweep re-check (End states 2, 4, 5, 7, 11) · whole-repo property spanning six tasks/three waves · runner: the Lead at Phase-1 land — EXECUTED at the Phase-1 checkpoint, PASS @ 3f136c0 (doctrine-census extra home = the red-team report, adjudicated leave). (2) "Unfinalized phase record" friction signal fires in practice · trigger is a future run's manifest, unreachable from this tree · runner: the next
/war-reviewover a post-change run — deferred (source: plan). Neither entry AI-declared.Red-team: CLEARED-WITH-NOTES after 7 rounds —
docs/red-team/2026-07-24-runbook-and-standing-record-coherence.md(4 machine-readable Adjudications rows, threaded into every auditor seat).Closes #1016. Closes #1039. Closes #1053. Closes #1078. Closes #1084. Closes #1087.
#1085 is partially addressed here (the hook's deny string + header comment) and deliberately carries no closing keyword — the
agents/war-auditor.md+ dispatched-prompt mirrors close via plan 4's Lead-filedwar-followup.🤖 Generated with Claude Code