From 2d633652912a89867112411d1add55c6f9b9e30b Mon Sep 17 00:00:00 2001 From: wshallwshall Date: Fri, 31 Jul 2026 09:55:20 -0500 Subject: [PATCH] docs: state the number-resolution rule generally, and the stale-branch blind spot The erratum recorded that the two ledger sequences diverged at #231. A peer's fair criticism: that makes the record defensible without telling a reader what to DO, and a reader resolving a cited number against this file still lands on the wrong item. So the rule is now stated generally rather than scoped to the ASVS numbers: for ANY cited #N above #231, resolve against the internal ledger unless the citation is demonstrably about this file, and ask rather than guess if you cannot see it. A number appearing in both places is not evidence they are the same item. Records that this has already shipped once. An item was filed here at a number the internal ledger had used for unrelated work and reached main before anyone noticed, cited in a gate comment, an operator-facing refusal message, a test docstring, two docs and the CHANGELOG. Described without naming the internal item, because naming it would republish exactly the material SECURITY-DOCS-POLICY withholds. The point that matters is that backlog_status_check.py CANNOT catch this class: it reads one file, where the number appears exactly once. Replaces the "suspect range" framing with the trigger that is actually correct -- the allocation TIMESTAMP. Any number issued before the floor fix landed is suspect regardless of value. The bounded range I gave peers excluded the very number that turned out to be colliding, and one session checked it only because it was theirs. Adds the all-refs check command, because a single-ref check reported a number as free that was in use on 28 refs. That was my error and it is worth stating as a method rule. WORKTREES.md: a watcher checking merged?/failing? is blind to the outcome that happens most -- main moved and the branch went BEHIND again, which produces neither signal. Adds the third arm, and the sibling case of polling before a new run's legs exist, where "nothing pending" reads as "all settled". Also sharpens the auto-merge note: it wins the race against checks finishing, not against main moving. Co-Authored-By: Claude Opus 4.8 --- docs/BACKLOG.md | 40 ++++++++++++++++++++++++++++++++++++---- docs/WORKTREES.md | 17 ++++++++++++++++- 2 files changed, 52 insertions(+), 5 deletions(-) diff --git a/docs/BACKLOG.md b/docs/BACKLOG.md index 38738b25..ee59a74b 100644 --- a/docs/BACKLOG.md +++ b/docs/BACKLOG.md @@ -20,19 +20,51 @@ already say. They were never published here, and they are not back-filled: their defined only in the `docs/security/` remediation plan, which [`SECURITY-DOCS-POLICY.md`](SECURITY-DOCS-POLICY.md) withholds. Read those citations the way the `docs/reviews/` and `docs/security/` paths above are read — **provenance into the internal ledger, not a -pointer into this file.** The same applies to every ASVS-programme number above #231; that programme -continued well past #246. Whether any of it is republished here is an owner decision. +pointer into this file.** Whether any of it is republished here is an owner decision. + +**The resolution rule, stated generally — it is not only the ASVS numbers.** For **any** cited `#N` +above **#231** (in an ADR, a plan, a commit message, a code comment, or an operator-facing string), +resolve it against the **internal** ledger unless the citation is demonstrably about this file. The two +sequences diverged at #231 and have been allocated independently since, so a number appearing both here +and in a citation is *not* evidence they are the same item. If you cannot see the internal ledger, +**ask rather than resolving it here** — landing on a same-numbered but unrelated item is the failure +mode, and it looks like success. **Consequently the numbers in this file above #231 are a second, independent sequence**, and items #232–#239 and #248–#251 do not correspond to the internal items sharing those numbers. This is recorded, not repaired: renumbering would rewrite ratified ADRs, and republishing would cross the policy above. +⚠️ **This has already shipped once, and nothing in CI can catch it.** On 2026-07-30 an item was filed +here at a number the internal ledger had already used for unrelated work, and it reached `main` before +anyone noticed; the number was cited in a gate comment, an operator-facing refusal message, a test +docstring, two docs and the CHANGELOG. It was corrected on 2026-07-31 by re-allocating against the fixed +floor. **`backlog_status_check.py`'s duplicate detection cannot see this class of collision** — it reads +this one file, where the number appears exactly once. The published baseline is not a safe place to +check a number against; only the allocator is. + **#240–#247 are permanent holes — do not file there.** They were allocated on 2026-07-30 by repeated runs for the same four titles; only the last run's numbers (#248–#251) were filed. #240–#243 are held by a worktree that no longer exists, and `alloc.ps1` has no release verb by design ("holes are free, collisions are not"), so those claims stand permanently and the ledger gate will refuse a commit that -files there. **#315** is a deliberate probe allocation used to verify the floor fix below; it is also a -hole. Always allocate with `scripts/coord/alloc.ps1`; never pick a number by reading this file. +files there. **#315** and **#317** are deliberate probe allocations used to verify the floor fix below; +they are holes too. Always allocate with `scripts/coord/alloc.ps1`; never pick a number by reading this +file. + +**If you allocated a backlog number before 2026-07-31T00:31Z, re-check it — the trigger is the +timestamp, not the value.** That is when the floor fix landed. Any number issued before it came from a +floor that could not see most of the namespace, so it is suspect **regardless of how low or high it +looks**; a bounded "suspect range" is the wrong instrument and gave at least one session false comfort. +To check one, search **every** ref, not the published file and not a single branch: + +```bash +for r in $(git for-each-ref --format='%(refname)' refs/heads refs/remotes); do + git show "$r:docs/BACKLOG.md" 2>/dev/null | grep -qE '^## \.' && echo "$r" +done +``` + +Checking one ref is not checking the namespace — the numbers are scattered across hundreds of refs, +which is exactly why the floor sweeps them all. A single-ref check reported one number as free that was +in fact in use on 28 refs. The root cause is fixed: the backlog floor in `alloc.ps1` now sweeps **every** local and remote ref, as its own header comment always promised and as the ADR path already did. Before the fix it read only diff --git a/docs/WORKTREES.md b/docs/WORKTREES.md index 95e77941..849a6c90 100644 --- a/docs/WORKTREES.md +++ b/docs/WORKTREES.md @@ -78,12 +78,27 @@ Three things that cost real time here: - **`BLOCKED` is the one that looks actionable and usually isn't.** Right after a push, every required check is pending and the PR reads `BLOCKED` — identical to genuinely failing. Count failures in `statusCheckRollup` before diagnosing; zero failures plus pending checks means wait. -- **Armed auto-merge does *not* update a `BEHIND` branch here.** Landing PR A puts PR B `BEHIND`, and B +- **Armed auto-merge wins the race against checks finishing, not against `main` moving.** It does *not* + update a `BEHIND` branch. Landing PR A puts PR B `BEHIND`, and B sits armed and stalled indefinitely. Someone has to rebase it. If you queue two PRs, expect to rebase the second after the first lands. - **`BEHIND` and `DIRTY` are easy to confuse and the wrong fix is destructive.** Treating `DIRTY` as `BEHIND` means resolving conflicts in a hurry to make a force-push succeed. +### If you poll for "is it merged yet", watch for three outcomes, not two + +A watcher that checks *merged?* and *failing?* is blind to the outcome that actually happens most: +**`main` moved and the branch went `BEHIND` again.** That state produces no failure and no merge, so a +two-armed watcher reports "still running" right up to its timeout while nothing is progressing. Add the +third arm — merged / failing / **went stale** — and act on the third by re-syncing. + +The same blindness has a second form: polling immediately after a push, when the new run's check legs +do not exist yet. "Nothing pending" then reads as "all checks settled" when it means "no checks have +started". Assert on the count of legs you *expect*, not on the absence of pending ones. + +Both are the same failure as taking `--ours` on a conflict: **the instrument was accurate about what it +looked at and silent about what it did not.** `main` moved seven times during one pair of PRs. + ### Resolving a conflict: never take a side wholesale `docs/BACKLOG.md` and `CHANGELOG.md` are single large files every session appends to, so they conflict