Skip to content

fix(ci): let the free-threaded canary report a problem — it was structurally unable to - #27

Merged
wshallwshall merged 2 commits into
mainfrom
claude/freethread-liveness
Jul 28, 2026
Merged

fix(ci): let the free-threaded canary report a problem — it was structurally unable to#27
wshallwshall merged 2 commits into
mainfrom
claude/freethread-liveness

Conversation

@wshallwshall

Copy link
Copy Markdown
Collaborator

freethread-smoke.yml could not fail in any path. Every step was continue-on-error, every step
after setup was gated on steps.setup.outcome == 'success' (so a failed setup skipped the rest), and
the job itself was continue-on-error on top. A runner that couldn't provision 3.14t reported
success; a failing GIL assertion or smoke test was swallowed and reported success.

Five runs, five successes — and nothing in that record could distinguish a canary that flew from one
that never left the ground. The header even claimed "a red canary is an informational red check"; no
red canary was reachable.

It is currently flying. Run 30256316219 shows sys._is_gil_enabled() = False on a genuine
free-threading build with 73 tests passing. The defect was latent, not active — which is exactly
why it needed fixing: nothing would have told us when it stopped. An early-warning tripwire whose
silence reads as good news is worse than no tripwire.

What changed

The job-level continue-on-error is gone. It bought nothing — this workflow runs only on a weekly
cron and manual dispatch, never on pull_request, so it produces no PR check and structurally cannot
gate a merge — and it cost the entire signal. The steps stay continue-on-error so they all run and
every outcome stays collectable; a terminal verdict step rules on those outcomes and is allowed
to fail.

Verdict Meaning
did-not-fly 3.14t couldn't be provisioned — a dead tripwire, not a clean run
regressed install / GIL assertion / smoke tests failed under 3.14t
vacuous pytest exited 0 having collected nothing — success that measured nothing

That last one is #25's rubric §4.0 liveness rule
applied here.

Verified by execution, not inspection

The verdict shell was extracted from the YAML and run against every scenario:

exit=0  HEALTHY
exit=1  SETUP DEAD      :: canary did not fly
exit=1  GIL RE-ENABLED  :: canary regression
exit=1  SMOKE FAILED    :: canary regression
exit=1  INSTALL FAILED  :: canary regression
exit=1  VACUOUS PASS    :: canary measured nothing

Also removed a | tee that would have masked pytest's exit code behind tee's, and the smoke step now
records how many tests actually ran so the verdict can say what flew.

tests/test_freethread_smoke_liveness.py (7 tests) pins the two properties that make the signal real
— the verdict step exists and can fail — and the one that makes failing safe: the workflow never runs
on a pull request.

zizmor real exit 0 · YAML valid · new tests pass.

…turally unable to

`freethread-smoke.yml` could not fail in any path. Every step was `continue-on-error`, every step
after setup was gated on `steps.setup.outcome == 'success'`, AND the job itself was
`continue-on-error`. So a runner that could not provision 3.14t skipped everything and reported
success; a failing GIL assertion or smoke test was swallowed and reported success. Five runs, five
successes, and nothing in that record could distinguish a canary that flew from one that never left
the ground.

It IS currently flying — run 30256316219 shows `sys._is_gil_enabled() = False` on a free-threading
build with 73 tests passing. The defect was latent, not active, which is exactly why it needed
fixing: nothing would have told us when it stopped. An early-warning tripwire that cannot warn is
worse than no tripwire, because its silence is mistaken for good news.

The header also claimed "a red canary is an informational red check". No red canary was reachable.

WHAT CHANGED. The job-level `continue-on-error` is gone. It bought nothing — this workflow runs only
on a weekly cron and manual dispatch, never on `pull_request`, so it produces no PR check and cannot
gate a merge regardless — and it cost the entire signal. The steps stay `continue-on-error` so they
all run and every outcome stays collectable; a terminal verdict step then rules on those outcomes and
IS allowed to fail. Green now means the canary flew.

Three verdicts, each executed locally against the extracted shell rather than reasoned about:
  * did-not-fly    — 3.14t could not be provisioned. A dead tripwire, not a clean run.
  * regressed      — install, the GIL assertion, or the smoke tests failed under 3.14t.
  * vacuous        — pytest exited 0 having collected nothing. Same shape as the gates fixed in #25:
                     success that measured nothing.
All six scenarios verified: healthy exits 0; setup-dead, GIL-re-enabled, smoke-failed, install-failed
and vacuous-pass each exit 1 with the right diagnosis and a step summary naming it.

The smoke step also stops hiding pytest behind a pipe (`| tee` yields tee's status, not pytest's) and
now records how many tests actually ran, so the verdict can say what flew rather than merely that
nothing errored.

tests/test_freethread_smoke_liveness.py pins the two properties that make the signal real — the
verdict step exists and can fail — and the one that makes failing safe: the workflow never runs on a
pull request.
@wshallwshall
wshallwshall enabled auto-merge (squash) July 28, 2026 20:46
@wshallwshall
wshallwshall merged commit 1400e9a into main Jul 28, 2026
33 checks passed
@wshallwshall
wshallwshall deleted the claude/freethread-liveness branch July 28, 2026 22:57
wshallwshall added a commit that referenced this pull request Aug 7, 2026
…e, which is where it rots

Caught by the b5-backlog-archival session. The marker exists precisely to outlive
#1073, and it cited `docs/BACKLOG.md #1073` -- but #1073 is closed, and closed items
LEAVE that file for docs/archive/backlog/BACKLOG-CLOSED.md.

The failure is invisible to every check: the markdown link to BACKLOG.md still
resolves, so nothing 404s. Only the human-readable "#1073" silently stops being
findable there. That is exactly how section 12's EXISTING #26/#27 pointers already
went stale -- verified against origin/main: neither heading is in docs/BACKLOG.md,
both are in the archive.

Now names both locations rather than guessing an archive anchor that may not match
the slug the archival pass generates. The repo's deep-link-to-archive convention
(44 occurrences in docs/BACKLOG.md) is the precedent.

NOT FIXED HERE, deliberately: the pre-existing #26/#27 staleness in the same list.
It is the same defect class but not this PR's scope, and the session that found it
is taking it.

I had noticed the #26/#27 staleness earlier in this session and judged it "not worth
chasing", then reproduced the identical defect in a marker whose whole purpose is
durability. Recording that, because the lesson is not "fix the pointer" -- it is that
a rot I declined to fix is one I had already stopped seeing.
wshallwshall added a commit that referenced this pull request Aug 7, 2026
…ose #1073, file #1089-#1093 (#269)

* backlog: ship the ASCQM catalogue pass, file #1089-#1093, mark the 5055 decline in CLAUDE.md (closes #1073)

The BACKLOG #1073 pass ran over all 74 live ASCQM 1.1 elements with an adversarial
refutation stage on every non-not-applicable verdict. The measure stays declined; the
catalogue earned its keep on two narrow grounds and this commit records both.

FINDINGS FILED
  #1089  HL7 parse_path accepts component 0, so PID-5.0 reads AND OVERWRITES the last
         component. Reproduced by execution. The X12 twin validates and is tested; the
         HL7 side -- the default content type, the one carrying PHI -- has neither.
  #1090  write_reference_snapshot's json.dumps has no default hook; the FILE reference
         source does not coerce where its DATABASE sibling does. Every existing test
         uses CSV, so the suite cannot reach it.
  #1091  Credential detection reads neither .ps1 nor .yaml, and the entropy gate floors
         out on low-entropy secrets. Written in the conditional: this is a control that
         cannot see the class, NOT a live exposure.
  #1092  Eight verdicts flipped under measurement and ALL EIGHT flipped covered->gap.
         Three falsify a specific written claim, including the quality record's
         "import/layer rules are machine-checked in CI".
  #1093  Inventory of the rest, with the ~19 metric findings explicitly NOT to be filed
         (section 4.1: a count is not an item) and two unverified premises flagged.

THE DECLINE MARKER IS THE PART THAT OUTLIVES THE ITEM. CLAUDE.md section 12 now carries
it. A decline recorded only in a backlog item disappears when that item archives --
verified: #26 and #27 are both in BACKLOG-CLOSED.md today and survive as binding
decisions only because their markers were lifted into section 12.

COUNTS RESOLVED: the CISQ-vs-ASCQM conflict was a UNITS problem. CISQ counts CWEs
including children; ASCQM counts elements. Performance Efficiency = 15 is now confirmed
and its unverified mark is lifted; the "139 total" stays unconfirmed and that mark stays.

RECORDED AGAINST MYSELF: the first run silently judged 62 of 74 elements after one
triage batch died, and returned a confident report that gave no sign a sixth of the
catalogue was unread. Caught by arithmetic, not by the run. One of the missed elements
became a filed finding, so it was not harmless. It is this repo's own section 4.0
failure mode reproduced inside the tool built to hunt for it.

VERIFIED: backlog_status_check.py exit 0, 350 items each declaring one status; five new
headings, one banner each; zero prose glyphs in the new section.
NOT RUN: pytest, mypy -- no .venv in this worktree. Docs-only diff.

* docs(CLAUDE): the 5055 decline marker cited only the LIVE backlog file, which is where it rots

Caught by the b5-backlog-archival session. The marker exists precisely to outlive
#1073, and it cited `docs/BACKLOG.md #1073` -- but #1073 is closed, and closed items
LEAVE that file for docs/archive/backlog/BACKLOG-CLOSED.md.

The failure is invisible to every check: the markdown link to BACKLOG.md still
resolves, so nothing 404s. Only the human-readable "#1073" silently stops being
findable there. That is exactly how section 12's EXISTING #26/#27 pointers already
went stale -- verified against origin/main: neither heading is in docs/BACKLOG.md,
both are in the archive.

Now names both locations rather than guessing an archive anchor that may not match
the slug the archival pass generates. The repo's deep-link-to-archive convention
(44 occurrences in docs/BACKLOG.md) is the precedent.

NOT FIXED HERE, deliberately: the pre-existing #26/#27 staleness in the same list.
It is the same defect class but not this PR's scope, and the session that found it
is taking it.

I had noticed the #26/#27 staleness earlier in this session and judged it "not worth
chasing", then reproduced the identical defect in a marker whose whole purpose is
durability. Recording that, because the lesson is not "fix the pointer" -- it is that
a rot I declined to fix is one I had already stopped seeing.
wshallwshall added a commit that referenced this pull request Aug 7, 2026
Resolves the CLAUDE.md section 12 conflict predicted before #269 landed. #269
merged as f28359e while this PR was being opened.

KEPT BOTH SIDES, as the two changes are semantically independent:
  ours   -- the #26 / #27 / #222 pointers re-derived onto BACKLOG-CLOSED.md
  theirs -- #269's new ISO/IEC 5055 / ASCQM decline bullet

DROPPED, and this is the only deletion: main's stale
`([docs/BACKLOG.md](docs/BACKLOG.md) #27, ...)` line. It is the pre-fix text of
the very bullet this branch rewrites, carried in as context by #269's insertion
directly beneath it -- not content #269 authored. Keeping it would have restored
the rot.

Verified after resolution rather than assumed:
  - no conflict markers remain
  - all four spans present: #26, #27, #222 rewrites AND the ISO 5055 bullet
  - the stale #27 line is gone
  - every link target in the merged section 12 resolves (8 unique, all OK),
    with a known-bad path run through the same checker to prove it reports a miss
  - re-resolved every cited number against the post-merge tree: #26/#27/#222
    archived, #232 live, #1073 live-but-closed -- so #269's "once archived"
    wording is correct and its dual citation is satisfied

Merged rather than rebased so the PR's auto-merge arming survives.
wshallwshall added a commit that referenced this pull request Aug 7, 2026
…ed (BACKLOG #1073 R1/R2) (#271)

* docs(CLAUDE.md): re-derive every section 12 pointer; 3 of 12 had rotted

Section 12's decline markers exist so a decision stays binding after its backlog
item closes and archives. The mechanism works -- #26 and #27 survive only because
they were lifted here. Their own pointers were the ones that had decayed.

Scanned all 12 pointers in section 12 against origin/main (the primary checkout
runs behind and returns confident false negatives). Re-derived each target rather
than trusting the text.

ROTTED, now fixed -- all three named the live ledger for an archived item:
  #26  visual/template authoring -> docs/archive/backlog/BACKLOG-CLOSED.md
  #27  serial / ASTM             -> docs/archive/backlog/BACKLOG-CLOSED.md
  #222 typed action vocabulary   -> docs/archive/backlog/BACKLOG-CLOSED.md
       (#222 was a bare "BACKLOG #222" with no path, which reads as the live
       ledger; it is closed. Not named in the brief -- found by the sweep.)

LEFT AS-IS, verified to resolve:
  ADR 0037 / 0063 / 0039 paths     -- all three files exist
  ADR 0007 (bare, "see section 1") -- section 1 carries the path, file exists
  ADR 0076 Amendment D             -- exists, 0076-typed-action-...md:658
  BACKLOG #232                     -- genuinely still open in docs/BACKLOG.md
  docs/CONNECTIONS.md              -- still carries the serial decline, :2436
  parse_items, section 11, section 1 -- resolve

Added the ADR 0076 path inline, since Amendment D was cited by bare number only.

Verification: extracted every markdown link target in section 12 and resolved it
(7 unique, all OK), with a deliberately-broken path run through the same checker
to prove it can report a miss.

No gate added. A backlog number resolving to an archived item is not mechanically
distinguishable from one resolving to nothing without encoding the archive's
shape, and a gate that fails on a legitimate archive is one people delete.

NOT included: #1073's marker. It is not on origin/main -- it is in PR #269, still
open, and it already cites the archive correctly. Basing on an unmerged PR head is
the stacking trap, so this branch is cut from origin/main. PR #269 inserts a new
bullet immediately after the #27 bullet this commit edits; the two are
semantically independent but adjacent, so expect a textual conflict and keep both.

* docs: DECIDED -- Code_Quality_Standards.md section 4.1 gets NO back-pointer to #1073

R2 of the #1073 leftovers. The B5 brief left this "optionally" and nobody had
chosen. Deciding it NO, and recording the decision, because an unresolved
"optionally" is indistinguishable from a deliberate omission six months later --
which is exactly how #1073's own unpinned status came about.

Deliberately no file change. The decision is not to add an artifact, so an empty
commit is the whole record.

The case for was real: section 4.1 is the anti-metric rule the ISO 5055 decline
turns on, and a reader wondering whether 4.1 has ever been applied to a live
proposal gets no answer from 4.1 itself.

Rejected because:

1. R1, in this same branch, is the evidence. Three of section 12's twelve
   pointers had rotted, and the two rotted ones named in the brief were #26 and
   #27 -- the very entries cited as proof that lifting a decline into section 12
   makes it outlive its item. The decision survived; the trail back to it did
   not. This estate's demonstrated failure mode is pointers decaying, not
   decisions being unfindable.

2. It would be a fourth copy of the same pointer (the backlog item, the section
   12 marker, PR #269's prose, and 4.1), in a repo already bitten by the install
   procedure in three copies and the reference table in two.

3. It would be born rotted. #1073 is closed by PR #269 and archives on merge, so
   a clause written today naming docs/BACKLOG.md acquires the exact defect R1
   just cleaned up.

4. Section 4.1 is a four-sentence hard rule carrying no worked examples for any
   of the metrics it bans. A #1073 example would be the only one, which reads as
   though 5055 were the rule's primary case rather than one application of it.

Section 12 holds the binding decision, which is what that mechanism is for.
Reopening needs a reason that outweighs the maintenance cost, not just the
observation that the cross-reference is absent.
wshallwshall added a commit that referenced this pull request Aug 7, 2026
…rchives (#272)

* backlog: file #1094, section 12 decline markers rot when their item archives (BACKLOG #1094)

Two changes, both small, both about the same load-bearing property: a decline
lifted into CLAUDE.md section 12 exists to OUTLIVE the backlog item that recorded
it, so the pointer back to that item has to survive archiving too.

#1094 records that it does not. Section 12 cites docs/BACKLOG.md #26 and #27;
retiring an item moves it verbatim into docs/archive/backlog/BACKLOG-CLOSED.md.
Measured on origin/main: both headings are absent from the live file and present
in the archive. Both pointers are already dead, and nothing here can catch the
class -- the markdown link still resolves, because it targets the file, not the
item; only the human-readable number stops being findable, and no tool reads it.
That is the argument for whatever check gets proposed, and it is why two
instances sat unnoticed alongside 44 correctly-formed archive citations in
docs/BACKLOG.md.

Code_Quality_Standards.md section 4.1 gains a POINTER, deliberately not a
restatement. Section 4.1 is one of the three reasons the ISO 5055 decline rests
on, so naming the relationship closes a loop; the reasons themselves stay in
section 12, stated once, per CLAUDE.md section 11. A second copy would be a
second thing to keep in sync, and section 4.1's subject is the anti-metric rule,
not the standard.

Verified rather than assumed:
  - backlog_status_check.py exits 0 on this tree (346 items, 151 live).
  - The green is evidence: planting a second status banner on #1094 makes it
    exit 1 naming line 5846 and item #1094, so the checker does reach the item
    this commit adds.
  - The number came from scripts/coord/alloc.ps1, not from grepping, and no
    other ref carries a #1094 heading.
  - ledger_check.py exits 0.
No pytest or mypy run: this worktree has no .venv, so the local quartet is not
available here and CI is the real check. Docs-only change, no code touched.

One negative result worth recording, because it was nearly filed as a defect.
Running the checker BARE, with docs/BACKLOG.md emptied, prints OK and exits 0
(scanned: 0) -- which looks like the section 4.0 liveness class, a gate that
cannot fail. It is not. ci.yml:148 runs it as `--min-items 300`, and under that
invocation the same empty file gives exit 1: "found 195 backlog items, below the
required floor of 300", naming both scanned paths. The tool's docstring already
says --min-items is the anti-narrowing floor and is not optional in CI, and CI
honours it. The bare CLI is not the deployed gate, and measuring it answered a
narrower question than the one that mattered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs(quality): address the 5055 decline by name, not only by section number

The 4.1 clause added in the previous commit cited "../CLAUDE.md section 12" and
nothing else. That is the same rot #1094 is about, one file over: a section
number cannot be re-resolved once it is wrong, and nothing checks one, so a
renumbering leaves the pointer aimed confidently at whatever now occupies the
slot. Pointing nowhere is recoverable; pointing convincingly at the wrong thing
is not.

The clause now names the target three ways -- the "ISO/IEC 5055" string, the
section title "Do / Don't Quick Reference", and the number. Only the number dies
in a renumbering, and either survivor re-resolves it by grep. This mirrors the
form already shipped in the other direction: CLAUDE.md section 12's marker cites
"the anti-metric rule in docs/Code_Quality_Standards.md 4.1", where the NAME is
what makes it survivable.

Measured rather than assumed, because the guidance came with a count taken
before this clause existed:
  - "anti-metric rule": 9 occurrences across the two files (CLAUDE.md 1,
    Code_Quality_Standards.md 8) -- a good label, too common to be an address.
  - "ISO/IEC 5055": 1 in CLAUDE.md on the branch carrying the section 12 marker,
    plus 1 here. So it is TWO post-merge, not the one reported. That is fine and
    is the intended shape -- one occurrence is the pointer, one is the target --
    but the premise "unique" stops being true the moment this clause lands, so it
    is recorded here rather than carried forward as a stale measurement.
  - Section title verified verbatim: "## 12. Do / Don't Quick Reference".

Docs-only, one line changed. No pytest or mypy: no .venv in this worktree.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
wshallwshall added a commit that referenced this pull request Aug 7, 2026
…napshot, file #1095 (#276)

Three corrections to the ledger's own accuracy, in one commit because they
cross-reference: #1094 and the ranking note both point at #1095, so splitting
them leaves an intermediate commit citing an item that does not exist yet.

1. #1094 CLOSED as already satisfied when filed; no work performed.

   Its premise is false on origin/main. The repoint it asks for merged as
   befe997 (PR #271) ONE COMMIT BEFORE the item itself landed (7ecff8a, PR
   #272) -- a filing race, not a wrong finding. Re-verified after both: CLAUDE.md
   section 12 now reads "BACKLOG #26 -- closed, so it lives in
   docs/archive/backlog/BACKLOG-CLOSED.md, not in the live ledger", same for #27.

   Banner flipped from the OPEN glyph to a CLOSED one -- replaced, not added, so
   the item still declares exactly one status. The analysis is kept: its point
   that no gate in this repo can catch the class is the argument any future
   check has to answer, and it is now attached to #1095 at true scale.

2. The "Connector & feature-breadth gaps vs. Mirth Connect" section marked a
   historical snapshot.

   All TEN backlog numbers it cites -- #7, #20-#27, #35 -- have closed and moved
   to the archive; none is in this file. So "#7 above" and "#35 below" are false
   directions out of the document, and "P1 -- close first" names work that
   shipped: #20 (FHIR, ADR 0022) and #21 (observability, PR #407). The section
   marks #24 and #35 SHIPPED inline, which makes the unmarked #20/#21 read as
   still open. A reader planning from this picks up finished work.

   Deliberately NOT repointed per-number. Every cited item is archived, so
   attaching an archive path to only the two decline-by-design lines would assert
   by contrast that the other eight are live. Uniform staleness is at least
   detectable; differentiated staleness is not.

3. #1095 filed for the systemic class. Number allocated via
   scripts/coord/alloc.ps1, never grepped.

   Measured on origin/main with parse_items (imported, not re-derived): of 129
   path-bearing BACKLOG.md citations, AT LEAST 69 distinct sites across AT LEAST
   35 files name the live ledger for an archived item. Plus 13 hrefs that do not
   resolve at all, 12 line anchors past EOF (file is 6318 lines; one cites 8429),
   and 31 in-range anchors that drifted onto unrelated text.

   The item's central point is DETECTABILITY, because getting this wrong means
   someone closes it with a linter having fixed a third of it: the 13 broken
   hrefs and 12 past-EOF anchors are catchable, but the 69 wrong-file citations
   and the 31 drifted anchors are NOT -- those links resolve perfectly, and what
   rots is the number or the line beside them.

   It also records that the test is "does the cited FILE contain the item", not
   "is the item CLOSED". Those differ: #1073 is closed and still legitimately in
   the live ledger, so a sweep keyed on closure would corrupt correct citations.

   Prior art found and named rather than duplicated: MIG-35 is already "the
   BACKLOG-reference classifier" folding into MIG-74 in the master test plan
   (:128). The item notes MIG-74 as worded -- "every doc path resolves" -- would
   pass the largest class untouched, since those paths do resolve.

Verification:
  - parse_items diffed before and after: exactly two items changed state, #1094
    (open -> closed) and #1095 (new). No unintended banner churn.
  - backlog_status_check.py: OK, 365 items, each declaring exactly one status.
  - All 7 link targets introduced were resolved from docs/, with a known-missing
    path run through the same checker to prove it can report a miss.
  - The MIG-74 quote was confirmed verbatim in the source file, not paraphrased
    from an agent's summary.
  - Line endings normalized to CRLF to match the file; diff stayed at 50/2
    rather than whole-file churn.

Not included: the ~69-site sweep itself and any gate. Those are #1095's scope,
and a partial repoint is worse than none for the reason given in item 2.
wshallwshall added a commit that referenced this pull request Aug 7, 2026
…251 #252 #257 #260 #273) (#274)

* feat(asvs): prove an absence claim bites, not just that its pattern matches (BACKLOG #1006)

`check_absences` admits an ASVS absence claim on `re.search(a.pattern, a.mutation)`
-- one TOML field matched against another. That proves the mutation is well-formed;
it never proves the mutation BITES. A reintroduction raised into a swallowing
handler, written to a field nobody reads, or behind a flag nobody branches on
satisfies every failure mode `check_absences` has and changes nothing observable.
A green gate that is not evidence.

Add an opt-in `--prove-absences` mode (`scripts/asvs/scorecard.py`) that executes
the claim rather than grepping it:

- Two optional `Absence` fields, `mutation_path` and `observable` (a pytest node
  id). When both are set the mode copies the tree to a scratch dir, runs the
  observable (baseline must be green), appends the mutation, and requires the
  observable to go RED -- and to fail as a test failure (exit 1). It fails closed
  on every other code: an already-red baseline, an uncollectable node, or a
  mutation that only breaks import is a PROVE-ERROR, never a proof. A claim that
  reddens nothing is UNPROVEN and fails the mode.
- A coarse same-file static backstop screens claims carrying `mutation_path` but no
  `observable`: a `raise` landing in a file whose every handler swallows. It is a
  screen, not a proof (it cannot see a swallow in a caller), documented as such.
- The whole pass runs in a TemporaryDirectory scratch copy, so it never mutates the
  tracked tree and never trips the committed-tree scan on itself.

Both fields default empty and load without being refused: the vault's ~81 existing
absence claims carry neither and must stay loadable (ADR 0156 §7). Absent means
"not yet proven by execution", surfaced by the mode, never "proven vacuous".

Review hardening carried in this change (the mode's own helpers):

- `_scratch_ignore` refuses `.env*`, `*.db` (+ WAL sidecars) and `docs/security`
  when copying the tree. The vault runs this module against the REAL tree (ADR 0156
  §7); a scratch copy carrying those would spill secrets / the local store / vault
  posture data into a world-default temp dir, which CLAUDE.md §9 forbids. The
  public-repo path never sees them; this is defence for the eventual vault run.
- `_is_within_tree` refuses a `mutation_path` that is absolute or contains `..`
  before anything is applied, so an authored path cannot escape the scratch copy.

Tests (tests/test_asvs_scorecard.py): eight fixture tests drive `prove_absences`
directly (proved / UNPROVEN / already-red baseline / collection-error / two static
backstop arms including a re-raise reach control / root-untouched / load
round-trip), plus three that drive the CLI contract CI depends on -- `main([...,
"--prove-absences"])` exit 0 on a biting fixture and 1 on a non-biting one, and
`main([...])` without `--corpus` exit 2 -- plus the secrets-exclusion and
path-traversal guards. Every new test was falsified (broken on purpose, watched
red, restored).

MessageFoundry is a not-deployed beta: the mode is opt-in, the default `verify`
path is byte-unchanged, and no authored claim carries an `observable` yet, so
nothing new is blocked by this alone today. Wiring the mode over the vault claims
and backfilling their observables is the owner's follow-up.

* docs(backlog): flip #1006 to shipped -- the absence gate can now prove behaviour (BACKLOG #1006)

Flip the #1006 banner from filed to shipped. It is written as a capability claim,
not a closure claim: the `--prove-absences` mode CAN catch a well-formed-but-vacuous
reintroduction once a claim carries an `observable`, but the default `verify` path
is byte-unchanged and no authored claim carries one yet, so nothing new is blocked
by this alone today -- the honest present-tense state for a not-deployed beta.

Banner lines of #1006 ONLY. The ranked table, the four census distribution lines,
and every other item's banner are untouched. The status census was NOT recomputed.

* test(connscale): dynamic contiguous inbound-port allocation, drop the flaky marker (BACKLOG #1014)

The connscale SQLite smoke test hard-coded base_port=41000 and needs 24
contiguous inbound ports, so two worktrees running the suite at once
contended for the same fixed block; a @pytest.mark.flaky(reruns=2) marker
retried past the collision, relabelling a determinate resource conflict as
CI noise. On the first parallel run it would keep masking exactly this class.

Replace the fixed block with _free_contiguous_ports(), which anchors an
n-wide block at a RANDOM base inside a bounded window, probes each port with
a no-REUSEADDR bind, and returns the range only when all n bind. The random
anchor over a wide window de-correlates concurrent worktrees; a genuine
future collision now surfaces as a red, not a masked retry. Contiguity is
asserted at the acquisition site and the allocator fails loudly -- never a
silent fixed fallback -- via two branches: an up-front width guard when the
block cannot fit the window, and a post-loop raise when no free block is
found after `tries` attempts.

The window is [20000,30000): the lower bound sits ABOVE the sibling MLLP
fixed-port band (other tests bind fixed inbound ports in the 11xxx-19xxx
range, e.g. 15099/19601), and the upper bound stays BELOW the OS ephemeral
floors (Linux 32768+, Windows/macOS 49152+) so a kernel-assigned ephemeral
port -- the sink/API ports, or any unrelated connection -- can never land in
the block after it is probed.

Drop the @pytest.mark.flaky marker: the collision was the cause, so keeping
it would re-hide the class this removes. Add three helper tests -- contiguity
and in-window, post-loop exhaustion (tries=0), and the width guard -- each
pinned to its branch (match=) and falsified by mutation.

Test-only change; no product code is touched.

* backlog: close #1014 -- dynamic connscale port allocation ships, flaky marker dropped

Flip #1014's status banner from open (filed) to shipped: the dynamic
contiguous inbound-port allocation and the flaky-marker removal land in the
same branch (commit 3450c3f).

This edits the #1014 banner line ONLY. The ranked table and the four census
distribution lines are untouched, and the census was NOT recomputed.

* feat(anon): structural PHI-shape detectors + coverage report + token-floor signal on the leak-check (BACKLOG #331)

The fail-closed leak-check verified only that MAPPED fields were pseudonymized;
PHI sitting in a field no rule mapped would pass the check clean on first
deployment (a real MRN is not a denylisted string). Scoped to the fields
anonymize did NOT rewrite, add:

- high-precision structural detectors over unmapped fields: dashed SSN,
  punctuated NANP phone, and CX MR/MRN-typed identifier. Deliberately narrow, to
  avoid the mass false-positives a broad digit-run search produces on HL7 bodies
  dense with dates/order-numbers/set-ids (ADR 0030 section 5).
- LeakReport: an unmapped-field coverage report (addresses only, never a value)
  plus token_tables_live / token_floor_reason. Reasons name the shape + field
  ADDRESS only, so a raised LeakError / log line never carries PHI.
- token_floor_failure() folded into the fail-closed decision under the
  require_live_denylist opt-IN lever (default off, so token-less CI/OSS/fork runs
  stay green with the structural detectors as the live backstop).

The whole structural block is mirrored byte-identical into tee/anon/leak.py; a
new engine/tee leak_report parity test pins it. Each detector was falsified
(removed it, watched the unmapped-PHI dataset slip through, restored); a
false-positive guard proves a benign unmapped field (14-digit EVN timestamp,
order number, coded observation id) does not trip.

Docs are written to the shipped DEFAULT behaviour, not an overclaim: the
coverage report and token-floor reason are RECORDED and surfaced on a refusal or
via the on_report hook (not an unconditional clean-path catch-all), and the
strict refusal is opt-IN. anon/__init__.py (both copies) necessarily changed to
export LeakReport/leak_report/coverage_clause and wire the lever. ADR 0030
section 5 / section 7 / Consequences amended (the "deferred" phrasing was stale
against the shipped code).

NOT-DEPLOYED beta: worded as "would let unmapped PHI through on first
deployment", no present-tense exposure claim.

* docs(backlog): flip #331 banner to SHIPPED, worded to the default behaviour (BACKLOG #331)

Banner-only flip of #331 to SHIPPED. Worded to the shipped DEFAULT behaviour, not
an overclaim: the coverage report and token_floor_reason are RECORDED and surfaced
on a refusal or via the on_report hook (not an unconditional clean-path catch-all),
and require_live_denylist is the strict opt-IN lever (default off).

Census NOT recomputed: only the #331 banner line changed. The ranked table, the
four census distribution lines, and every other item's banner are untouched.

* test(sandbox): a static ast guard pins the codec+worker import boundary (BACKLOG #346)

The sandbox's import boundary (DEFAULT_FORBIDDEN_MODULES in pipeline/sandbox.py --
socket/ssl/asyncio, the I/O- and secret-bearing messagefoundry.* subpackages, cryptography)
is enforced only at RUNTIME and only inside the off-by-default [sandbox].mode=subprocess
child. Nothing statically pins that the two modules which run inside that boundary --
_sandbox_codec.py and _sandbox_worker.py -- do not themselves import a forbidden module. Both
are clean today; a future edit reintroducing a forbidden import would make mode=subprocess DOA
on first deployment while the default-mode suite stayed green -- the failure inverts, hitting
the most security-conscious installs hardest and quietest.

This is defence-in-depth test coverage, not a code change: neither sandbox.py nor the codec is
touched. tests/test_sandbox_import_boundary.py walks the two files' own ast import nodes
(ast.Import/ast.ImportFrom, including nested/function-level and relative imports resolved to
absolute) and asserts none resolves under a DEFAULT_FORBIDDEN_MODULES prefix. The forbidden set
is imported from the runtime constant, never copied, so the guard tracks whatever the sandbox
forbids. It ships with a positive control (each static import form the walker handles is seen,
including the load-bearing from-parent alias-append) and a negative control (benign
messagefoundry.* imports raise zero flags).

Scope is the two files' DIRECT imports, deliberately not a transitive walk: importing the codec
pulls asyncio/cryptography/store/transports/auth into sys.modules, so a transitive walker would
red on clean shipped code and prove nothing. sandbox.py is out of scope per BACKLOG #346 even
though the worker child imports it; the docstring records that residual for the owner.

Falsified: planting `import socket` into the real _sandbox_codec.py reddens the live guard
naming it; removing the walker's alias-append reddens only the alias-append positive-control
case; an over-broad matcher reddens the negative control. All plants restored before commit.

* backlog: flip #346 to SHIPPED -- the static ast import-boundary guard landed (BACKLOG #346)

The #346 banner alone: OPEN -> SHIPPED, pointing at tests/test_sandbox_import_boundary.py
(the static ast guard added in the preceding commit). The completeness wording is softened from
"every forbidden import form is seen" to "each static import form the walker handles" -- a
static walker cannot see dynamic importlib/__import__ forms, and CLAUDE.md section 11 prefers a
bounded claim to an enumeration.

Only the #346 banner line changed; the ranked table, the four census distribution lines, and
every other item's banner are untouched. The census was NOT recomputed.

* fix(tls): route four insecure-TLS escape cells through the ADR-0092 clamp (BACKLOG #329)

LDAPS (auth/ldap.py), SFTP host-key (transports/remotefile.py), the webhook sink
(pipeline/alert_sinks.py) and the AI-broker (transports/ai_broker.py) read the raw
MEFOR_ALLOW_INSECURE_TLS escape directly; on an enforcing-PHI instance each would
otherwise honour the env var on first deployment. Each now routes through the ADR-0092
weakened_tls_escape helper: SFTP is built in-gate so it uses _here(); the other three
are built outside the hop scope, so the instance posture is threaded explicitly through
AuthService / notifier_from_settings / ai_broker_from_settings (additive, default None =
byte-identical for existing callers). The fifth cell the item names (direct.py) was
already clamped in #323, so this converts the remaining four. Docs (CONNECTIONS/DEPLOYMENT/
PHI) corrected from 'not clamped'/'unclamped' to clamped.

* docs(backlog): flip #329 to shipped -- four insecure-TLS cells clamped (BACKLOG #329)

Banner line only (leaves the 2026-08-03 amendment note); census not recomputed.

* fix(dev): anchor setup-leak-gate.ps1 to its own checkout, not the cwd (BACKLOG #1063)

`$repo` came from `git rev-parse --show-toplevel`, which resolves against the CURRENT
directory rather than the path the script was handed. Invoked by absolute `-File` path
from another worktree -- the ordinary shape on a clone carrying dozens of them -- it armed
the CALLER's checkout and printed CONFIGURED about that one, while the checkout the
operator named kept no token source and went on failing closed. An absolute `-File`
invocation is naming the checkout to act on; it must not then consult a different one.

Now `Split-Path -Parent (Split-Path -Parent $PSScriptRoot)`, the form postgres.ps1:37 and
sqlserver.ps1:56 in the same directory already use, plus an assert that the derived root
actually carries scripts/security/ -- a wrong root should say so where it is derived
rather than surface later as a confusing scanner failure.

Tested by the DIVERGENCE, which is the only shape that can fail: two temp checkouts that
both carry scripts/security/, the script invoked by absolute path while the shell stands
in the other one. A test run from inside the target passes with the bug still in, because
cwd and script root are then the same directory. Reverted to the old line, the same test
reports "the named checkout was not armed" -- the negative control was run, not assumed.

The fixture copies only the three files the script reaches for, never the whole of
scripts/security/: a maintainer running this suite has the real token list sitting in that
directory, and a copytree would sweep it into a temp dir.

Also corrects this item's own prose. It called alloc.ps1:51 "byte-equivalent"; it is not
-- alloc.ps1 carries --path-format=absolute and this script does not. The defect is
identical, the bytes are not, and "byte-equivalent" is the kind of claim a later reader
greps for and then trusts.

* fix(coord): anchor alloc.ps1 and claim.ps1 to their own checkout (BACKLOG #1060)

Both took `$repo` from an unanchored `git rev-parse --show-toplevel`, which resolves
against the CURRENT directory rather than the path the script was handed. Invoked by
absolute `-File` path from worktree A while intending to commit from worktree B, the claim
was recorded to A; the ledger gate then refused B's commit -- correctly, it fails closed --
but far from the cause and with a message about the wrong thing, and it cost a number.

`git -C $PSScriptRoot`, not `Split-Path`. The recorded `worktree` value has THREE readers:
ledger_check.py:227, this script's own `-List`, and prune-merged.ps1's orphan-claim release,
whose comment at :787 names the producing command -- "records `worktree = $repo` from
`git rev-parse --path-format=absolute`" -- and matches on the full normalised path because
a false positive there hands a live session's key to someone else. All three fold
separators, so `Split-Path` would not have broken anything; it would have silently
falsified that comment, in a destructive tool, for no gain.

THE FILING NAMED ONE OF FOUR CWD-DERIVED READS, and the other three are measured in the
negative control below:

  * the `branch` recorded with the claim was the CALLER's branch;
  * the floor's boundary was parsed from the CALLER's scripts/hooks/ledger_check.py;
  * the floor's WORKING-TREE term read the CALLER's docs/BACKLOG.md.

The third is not friction and the item's severity paragraph is corrected in the same
commit. That term exists to catch a number written but committed NOWHERE. Reading the
caller's tree makes a number drafted in the target worktree invisible, so the allocator
hands it out as free and two items share it -- both owned by that worktree, so owns()
passes and the ledger gate never fires. The silent collision the docstring says this script
exists to prevent, reached through the script. Narrow, since anything committed on any ref
is still caught by the all-refs term, but a correctness hole rather than friction.

claim.ps1:54 carried the same construct and was never filed -- found by inspection here,
fixed in the same commit. Its enforcing hook, claim_check.py, reads the repo from cwd and
is RIGHT to: a commit hook's cwd IS the committing worktree. Hook right, tool wrong, and
only the tool can be invoked from somewhere else.

Both scripts now print a NOTE when the shell is standing somewhere else. Anchoring is
correct but surprising, and the item's other half -- showing the recorded worktree -- was
already built (`claimed by:` / `by :`); what was missing is saying so when it diverges,
instead of leaving it to surface as a refused commit later. Silent on the ordinary
same-tree invocation, so it stays worth reading.

THE FIX TURNED TWO SANDBOXED TEST FILES INTO WRITERS ON THE LIVE REGISTRY, which is worse
than the red suite it also caused, and is the reason those fixtures changed here.
test_coord_claim_{refresh,liveness}.py ran the REAL scripts/coord/claim.ps1 with cwd set to
a temp repo -- scoped to a throwaway registry purely by ambient cwd, and one of them said so
("it scopes itself to the cwd's repo"). Once the script stopped consulting cwd, the passing
half of the run wrote real claims into this clone's shared registry: two strays, `k` and a
date-shaped key, were created and removed by hand. Both fixtures now stage and COMMIT a copy
of the script inside the temp repo, so the sandbox is structural rather than ambient, and a
linked worktree of the fixture carries its own copy -- which is how the peer-holds-the-key
tests still produce a claim recorded against the peer. test_ledger_check.py already did
exactly this for alloc.ps1, which is why it was the one that did not break.

Tested by the DIVERGENCE, with -ShowFloor so no numbers are burned: allocation is a one-way
door and a test that allocated would leave permanent holes in the shared registry for every
worktree of this clone. Two temp checkouts draft different numbers and carry different
PUBLIC_BACKLOG_FLOOR stubs; the caller's number is deliberately HIGHER, because the floor is
a maximum and an equal or lower one would pass with the bug in. Reverted to the old lines,
the same test reports floor 7777, boundary 1900 and a watermark under Caller/.git -- three
independent signals, all pointing at the wrong tree.

* docs(backlog): record the test-isolation trap #1060's fix walked into (BACKLOG #1060)

A cwd-dependence that reads as a defect in the tool can be load-bearing ISOLATION in its
tests. Both claim test files were scoped to a throwaway registry purely by ambient cwd --
one said so in a docstring -- so anchoring the script turned the passing half of the run
into a writer on this clone's shared registry before the rest of it went red.

Recorded in the item rather than only in the commit message, because #1057 and #1059 are
the remaining instances of the same class and will hit the same trap: check what a test is
isolated BY before changing what the code reads.

* docs(CONNECTIONS): repoint the serial/ASTM decline at the archive; #27 is closed

The connector-parity row for Serial (RS-232) / ASTM E1381/E1394/E1318 cited the
decline as ([BACKLOG.md](BACKLOG.md) #27). Item 27 is closed and lives at
docs/archive/backlog/BACKLOG-CLOSED.md:994; it is not in the live ledger. The
pointer sent a reader to the wrong file.

This is the second half of a designated two-marker pair. The archived item's own
banner names both markers -- "marker landed in PR #411 (CLAUDE.md section 12 +
docs/CONNECTIONS.md Serial row)" -- and commit 8a14602 repointed the CLAUDE.md
half while logging this one as still carrying the decline, because that commit
was scoped to section 12. The two halves disagreed about where #27 lives until
now.

Form: an anchored link, matching the sibling convention already used in this same
directory for this same target (docs/AOAG-DEPLOYMENT.md:389 and :476). The cell
already opens with "declined-by-design (v0.2+)", so the citation's only job is to
resolve; restating "closed" in the cell would duplicate a fact the cell asserts
two clauses earlier.

Relative path: (archive/backlog/BACKLOG-CLOSED.md), NOT (docs/archive/...). The
link is repo-relative from inside docs/. CLAUDE.md is at the repo root and
correctly uses the docs/-prefixed form; copying that form here would resolve to
docs/docs/archive/... and 404.

Verified, not assumed:
  - The anchor slug was derived by a rule first replayed against three anchors
    already committed in the repo (#100, #101, #52) -- 3 of 3 exact -- then
    applied to #27's heading, then confirmed to match exactly one real "## "
    heading in the target file. A bogus anchor was run through the same check
    and found nothing, so the check can report a miss.
  - Item locations come from parse_items imported from
    scripts/docs/backlog_status_check.py, per CLAUDE.md section 11 -- not a
    hand-rolled scan of the banner alphabet.
  - backlog_status_check.py still reports 363 items, unchanged.
  - Read from origin/main throughout; the primary checkout runs behind.

NOT changed, deliberately, with the reason:
  - docs/BACKLOG.md:574 -- bare number inside the section headed "Value &
    priority analysis (recorded 2026-06-19) - superseded". A superseded snapshot
    is a historical record; it has no path to rot.
  - docs/testing/FEATURE-COVERAGE-PLAN.md:41 -- names the features in prose and
    carries no number or path at all. Nothing to rot; adding a pointer would be
    new scope, not a repair.
  - docs/testing/master-test-plan/00-strategy-and-governance.md:699 -- bare
    #26/#27 that resolve to nothing rather than to wrong content. Repairing one
    link here would leave a single correct relative link among 28 broken
    root-relative ones in the same file; it belongs in the doc-set-wide sweep
    that class needs.
  - docs/BACKLOG.md:906 -- a real defect, but larger than a pointer repair and
    in a file several sessions are editing. Reported separately for a decision.

A wider scan (127 path-bearing BACKLOG citations) found roughly 90 more naming
the live ledger for an archived item. Not touched here: the staleness is
currently uniform, and repointing a subset would assert by contrast that the
untouched siblings are live. That class needs one pass, not a trickle.
wshallwshall added a commit that referenced this pull request Aug 7, 2026
Retiring a backlog item moves it verbatim from docs/BACKLOG.md into
docs/archive/backlog/BACKLOG-CLOSED.md. Every citation that named the live file
keeps pointing at a file the item is no longer in. The link still resolves, so
nothing in CI can see it. #1094 fixed two such markers in CLAUDE.md section 12;
this is the same defect at repo scale.

73 citations across 34 files, href-only. No prose was rewritten. Visible labels
changed ONLY where leaving them would contradict the target -- a label reading
`BACKLOG.md` pointing at the archive -- and then only to `BACKLOG-CLOSED.md`.

THE TEST IS "DOES THE CITED FILE CONTAIN THE ITEM", NOT "IS THE ITEM CLOSED".
Those differ, and keying on closure would corrupt correct citations: #1073 is
closed and still legitimately in the live ledger. Item locations came from
parse_items imported from scripts/docs/backlog_status_check.py, per CLAUDE.md
section 11 -- never a hand-rolled scan of the banner alphabet.

DELIBERATELY NOT TOUCHED, each for a stated reason:

  Both ledger files -- ZERO edits to docs/BACKLOG.md and BACKLOG-CLOSED.md.
    Only two sites named them and both are excluded, so this change costs no
    conflict against the merge trains or the pending #1096 filing. The one real
    site (#322 at BACKLOG.md:2720) is left because the file is the most
    contended in the repo and the item number is visible in plain text a search
    away.

  docs/CONNECTIONS.md:2436 -- the #27 serial/ASTM row. Already fixed on a branch
    inside merge train #274. Sweeping it from origin/main would re-fix stale
    text and collide.

  QUOTATIONS OF THE DEFECT. docs/BACKLOG.md:6319, inside #1094, reads "Two of
    its markers CITED [`docs/BACKLOG.md`](BACKLOG.md) #26 and #27" -- past
    tense, describing rot that is already fixed. Repointing it would corrupt a
    historical record. A regex cannot tell this from a live pointer, which is
    the reason this was not done with sed.

  THE WRONG-NUMBER CLASS, which is a different defect and must not be swept into
    this one. ADR 0068:10 cites #11 and ADR 0113:9 cites #239; both numbers are
    absent from the live ledger, but the ARCHIVE's #11 ("`check` dry-run
    cross-products") and #239 ("Re-measure Steps view estate coverage") are
    unrelated to WebAuthn passkeys and to a Windows tray manager respectively.
    Repointing would convert a vague reference into a confidently wrong one that
    lands the reader on the wrong item. Left, and reported.

  MIXED-LOCATION LINKS, where one link covers items in both files so no single
    target is correct: docs/AI-OFF-MATRIX.md:50 (six items), docs/adr/0001:13
    (#1 archived, #3 live), THROUGHPUT-IMPROVEMENTS.md:215 (#62 live, so its
    link is already correct).

VERIFICATION
  - Plan applied by literal replacement on the named line only, requiring the
    quoted string to occur EXACTLY ONCE there; a mismatch aborts rather than
    fuzzy-matching. Dry run: 73/73 clean, 0 problems, before anything was
    written.
  - All 35 distinct anchor fragments introduced match exactly one real "## N."
    heading in the archive, checked after applying, with a known-bad fragment
    run through the same check to prove it can report a miss. Fragments were
    derived with a slugger that does NOT collapse consecutive spaces -- the
    doubled hyphens are correct, not typos.
  - All 74 archive hrefs in the changed files resolve to the archive from their
    own directory depth; the relative prefix differs by depth and was computed
    per file, not pattern-matched.
  - Coverage confirmed with a DELIBERATELY LOOSER regex than the one that built
    the work list: it finds exactly one wrong-file site outside this change set,
    docs/CONNECTIONS.md:2436, which is the intended exclusion.
  - backlog_status_check.py: OK, 365 items. No mixed line endings introduced.
    All 34 changed files are markdown; diff is 67 insertions / 67 deletions,
    line-for-line.
  - Staged by explicit path from the plan, cross-checked against git's modified
    set, so nothing another session is editing was swept in.

Not included: the broken-href class (13 sites, mostly (docs/BACKLOG.md) written
from inside docs/testing/master-test-plan/), the 12 line anchors past EOF, and
the 31 in-range anchors that drifted onto unrelated text. Those are separate
classes under #1095 and are catchable by a link checker, which this repo still
does not run.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant