From c27e5539750fad8fdb89d1ffd0b94c7480bbc390 Mon Sep 17 00:00:00 2001 From: mecattaf Date: Thu, 17 Sep 2026 07:17:37 +0200 Subject: [PATCH 1/3] tally-b: serve --evaluator-lock from a packaged fixed-argv evaluator (DF-U-D13-2 discharged) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit U-D13 (dotfiles#316), integration gap Int-G6. Without `--evaluator-lock` the served kernel derives NO verdict at all (tally docs/socket.md §4), so the admission -> lease -> execute -> witness -> verdict -> receipt loop stopped one rung short of a receipt. DEFERRED.md DF-U-D13-2 deferred passing it on one condition — `apps/evaluator` existing in the lake to be locked. It exists, in the pinned `tally-lake` input's store path, so the row is discharged and the flag is served. pkgs/tally-evaluator/default.nix is the FIXED ARGV: a writeShellApplication (runtimeInputs jq/coreutils/git/gnused) whose text is `exec /bin/sh ${tallyLake}/tools/e2e-evaluator.sh "$@"`. The argv the kernel hashes is therefore ONE store word, not the two-word checkout form (`/bin/sh /tools/e2e-evaluator.sh`) the audit's probe C ran: a checkout path in a locked argv is a lock over a file `git pull` can move. Everything per-evaluation still travels on stdin (T7-3), which is the only reason an argv can be pinned at all. modules/tally-b.nix builds the lock rather than transcribing one: a runCommand over the KERNEL's own ${inputs.tally-b}/tools/make-evaluator-lock.sh with --root ${inputs.tally-lake}, --argv ${tallyEvaluator}/bin/tally-evaluator and the seven --file rows (tools/e2e-evaluator.sh plus apps/evaluator's bin/evaluate.mjs and src/{evaluate,env,normalize,usage,receipt-ingestion}.mjs). No digest is typed into this repository, so the lock cannot drift from either pin without the derivation changing. `evaluatorLock`'s default becomes that derivation; null is still accepted and still means "derive no verdict". Guard G3 of tests/tally-b/probe-u-d13-guards.sh flips with it: it builds the lock offline and runs the kernel's own scripts/verify-evaluator-lock.sh --lock/--root/--argv (rc 0 — every file row recomputed from the pinned lake, the argv row recomputed from the wrapper), then takes the same verifier RED (rc 4) under the pre-wrapper two-word argv so part C is known to compare. The null branch is kept as the other half of the wire. test-tally-b-input.sh gains clause E: the flag reaches ExecStart and names a STORE path. docs/local-ai/tally-b-input.md: a new section on the lock, and the rev is now b3a040e (flake.nix is untouched — it already pinned it). DEFERRED.md: DF-U-D13-2 gone, with the discharge recorded; new DF-U-D13-4 [OTHER-REPO] — the CUBS and `eval(build:LOCAL-SMOKE)` kit entries stay `/bin/sh -c true` no-ops, so serving the lock changes nothing for the campaign (the 58 hand-driven runs were graded under them). What a mechanical verdict for CUBS IS is Tom's line (integration G7), and the per-item stdin an `eval(...)` entry needs is the uplink render tool's second step. MEASURED: verify-evaluator-lock.sh rc 0 over the built lock; foreign argv rc 4; probe-u-d13-guards.sh all [P]; tally-b-topology and util-sampler-topology build offline; `nix flake check --offline --no-build` is rc 1 for the pre-existing `nas-personal-tailnet` ENV red alone — byte-identical on a `main` worktree. Co-Authored-By: Claude Opus 5 --- DECISIONS.md | 35 ++++++++++ DEFERRED.md | 14 +++- docs/local-ai/tally-b-input.md | 99 ++++++++++++++++++++++++++--- modules/tally-b.nix | 74 +++++++++++++++++++-- pkgs/tally-evaluator/default.nix | 68 ++++++++++++++++++++ tests/tally-b/probe-u-d13-guards.sh | 85 ++++++++++++++++++++++--- tests/tally-b/test-tally-b-input.sh | 18 +++++- 7 files changed, 370 insertions(+), 23 deletions(-) create mode 100644 pkgs/tally-evaluator/default.nix diff --git a/DECISIONS.md b/DECISIONS.md index c759e6c59..591e7158d 100644 --- a/DECISIONS.md +++ b/DECISIONS.md @@ -1,5 +1,40 @@ # DECISIONS +2026-09-17 the served kernel derives verdicts. + +**The evaluator lock is BUILT, not pinned by hand (default, unruled).** No +ruling says where the served `--evaluator-lock` comes from, and there were two +ways to serve one: transcribe a generated `EVALUATOR.sha256` into this +repository as data (the way tally carries its own `ORACLE.sha256` locks), or +generate it in a derivation from the two inputs this repo already pins. Taken: +the derivation. `modules/tally-b.nix` runs the KERNEL's own +`${inputs.tally-b}/tools/make-evaluator-lock.sh` with `--root +${inputs.tally-lake}` at build time, so not one digest is typed here and the +lock cannot disagree with either pin without the derivation changing. The cost +is that reading the lock means building it (guard G3 now builds one small +derivation, offline); the gain is that a transcribed lock can rot silently +against a bumped input and a generated one cannot. The consequence to know: a +`tally-lake` bump CHANGES the lock, which changes the `evaluator.lock` cell of +every verdict derived after it — that is the point, a verdict names the bytes +that judged it, and it is why the lake is not in `rollingInputOverrides`. + +**The locked argv is ONE store word.** `pkgs/tally-evaluator` wraps +`tools/e2e-evaluator.sh` so the argv the kernel hashes is +`/bin/tally-evaluator` rather than the two words the audit's probe C +ran (`/bin/sh /tools/e2e-evaluator.sh`). A checkout path in a locked +argv is a lock over a file `git pull` can move; a store path is not. Everything +per-evaluation still travels on stdin — T7-3, and the only reason an argv can +be pinned at all. `node` is deliberately absent from the wrapper's +`runtimeInputs`: the stdin item names the interpreter by absolute path, because +the node the lake's `scripts/node-env.sh` records is the one its packages were +resolved against. + +**Serving the lock changes nothing for the campaign.** The kit's `eval(...)` +entries stay `/bin/sh -c true` no-ops. `DEFERRED.md` `DF-U-D13-4` carries why: +what a mechanical verdict for CUBS IS is Tom's line (integration G7), and the +per-item stdin an `eval(...)` entry would need is the uplink render tool's +second step, not this module's constant `stdin` cell. + 2026-09-13 flake checkouts no longer ride into host closures, chrome-stream is installed, and a switch refuses a stale raw-dotfiles checkout. diff --git a/DEFERRED.md b/DEFERRED.md index f6d23b177..cdfa0dba3 100644 --- a/DEFERRED.md +++ b/DEFERRED.md @@ -10,7 +10,7 @@ about. | DF-U-D12-3 | Refreshing the expired `cc3` OAuth token | D-B5 records this as the TOM LINE "run `claude` once on cc3". This feeder may read the credential through `stamp-receipt.py window`; it may not perform an interactive login | Tom runs `claude` once with the cc3 seat configuration. Until then the enabled feeder writes a fresh UNKNOWN row with the reader's reason | | DF-U-D12-4 | Releasing work onto the `codex` row | D-B6 says the login is third-party and the row is never proposed onto. Reading `rate_limits` from its rollout is not a spend and does not change ownership | A later explicit ownership ruling (TL-6); this unit preserves `owner: third-party` in every write | | DF-UTIL-TOK-1 | Counting a Halogen token window (a request with >= 128 completion tokens in the journal) as an evidence event for a box-hour | `/home/tom/research-methods/cards/UTIL-01.md:81-82`: "A metrics delta is NOT an evidence event … only Tom widens it". Since 2026-09-13 (#312) the worker's sampler reads Halogen's per-request `serve_api:` journal line and util-row sums `tokens_in` / `tokens_out` from it, but a token window carries no fingerprint and no receipt, so `receipt_is_evidence()` is deliberately unchanged and box-hours are still graded on receipts and drain-ledger rows alone | Tom, by editing the card's evidence definition. Until then the windows feed `tokens_in` / `tokens_out` / `tokens_grade` only, never a disposition | -| DF-U-D13-2 | Passing `--evaluator-lock` to the served kernel, so a verdict is derived at all | `services.tally-kernel.evaluatorLock` defaults to null, and null is the honest state: the lock pins `apps/evaluator`'s bytes, which do not exist until U-A17 delivers them in mecattaf/tally-ts-sdk (tally docs/socket.md §4: "Without --evaluator-lock, no verdict is ever derived at all"). Locking a path that does not exist would be a stub standing in for a deliverable | U-A17's delivery, then a follow-up pin of the lock file's path in this module — an edit reviewed like any other, which the topology check will see | +| DF-U-D13-4 | `[OTHER-REPO]` Giving the kit's `eval(...)` entries a real mechanical verdict — the CUBS entries and `eval(build:LOCAL-SMOKE)` stay `/bin/sh -c true` no-ops | Serving the lock changes what the KERNEL can derive, not what the CAMPAIGN asks for, and the two must not move in one act: the 58 hand-driven runs of `~/mecattaf/cubs-campaign/receipts/REVIEW-2026-09-17.md` were graded under the no-op entries, so flipping them in the same PR would change the campaign's semantics mid-flight with no ruling behind it. WHAT a mechanical verdict for CUBS is (integration G7: the validation exit plus revert-and-rerun as AC#4's verdict) is Tom's line and is asked, not assumed. The other half is mechanical and belongs elsewhere: an `eval(...)` entry must hand `tools/e2e-evaluator.sh` a whole per-item JSON object on stdin (card, deliverable repo, commit file, usage source, receipt out, node, repo), and rendering that per item is the SECOND step of the uplink render tool that D-A owns — this module's `stdin` cell is a constant string | Tom rules G7; then the kit entry is a reviewed edit to `home/tally-uplink.nix` (D-A's render tool supplying the per-item stdin), and `tally-uplink-topology` sees it. Discharged when a `verdict` on the chain names a CUBS item and the lock this module serves | | DF-U-D14-3 | Setting `services.tally-uplink.plan` — the acceptor's plan body | Null is the honest state, not an oversight. The plan body is the ACCEPTOR's, re-POSTed to arm and re-arm; authoring one in this module would be the lake proposing from the wrong side of the seam, and arming is Tom's act. (The kit half of this row is discharged: `home/tally-uplink.nix` sets `kit` to the reviewed store kit file, and `tally-uplink-topology` asserts it reaches the rendered argv.) | Tom arms the acceptor; then a follow-up edit of this option, reviewed like any other change, which the `tally-uplink-topology` check sees because it asserts `cfg.plan == null` and no `--plan` in the argv today | | DF-U-D18-2 | Setting an anti-starvation number for the filler lane, and implementing D-B10's age-based promotion (a filler item passed over by 10 non-filler releases proposed at level 1 for one release) | Spec §4.4.3 names TL-10 as the anti-starvation number and says that "until it exists the lane runs whenever the row is free". The promotion is a RELEASE-STATION decision over a backlog — the acceptor's `levels` document and the lake's proposals — and a clock has no backlog to age. Putting a promotion rule in a timer would be this unit deciding admission, which is the one thing its own card forbids ("the timer only wakes the uplink's filler pass, the kernel admits") | Tom for TL-10's number; the release station (§2.2c) and the lake's selector for the promotion itself. Nothing in this repository shortens the wait, and the five-minute cadence stands meanwhile: it is the drain's own, so neither filler can crowd the other out | | DF-U-D18-3 | Moving the academic drain onto a LEASE over the kernel's socket, so the two fillers alternate UNDER the kernel's admission rather than beside it | Spec §2.4 lists the drain as an *unleased* tenant of the GPU row "until TL-15 moves the drain onto a lease over the socket (U-D11 wires the call)". Until that happens, D-B10's round-robin is carried here by equal cadence plus the lane's own `/running`-empty gate — two mechanisms outside the kernel — and this unit cannot lease on another unit's behalf. MEASURED 2026-09-07: `tally-drain.service` runs `tally --socket /run/user/1000/tally/tally.sock daemon drain`, the LIVE daemon's socket, not the rewrite kernel's `~/.local/state/tally-rewrite/kernel.sock` | U-D11 wires the call; TL-15 is the line. Discharged when the drain's GPU turn is a lease on the chain and the `tally-filler-topology` cadence equality can be replaced by an admission-side assertion | @@ -29,6 +29,18 @@ about. | DF-CLIENT-11 | Setting sshd `ClientAliveInterval` / `ClientAliveCountMax`, so the coordinator reaps a dead projector's `remote-client-bridge` in seconds instead of at TCP keepalive | #385 is a client-seat change and deliberately changes no sshd behaviour: a `ClientAlive*` line in modules/common.nix applies to every host's sshd and every live ssh session (the orchestration running inside herdr included). The laptop side is already bounded — herdr's managed ssh config sends ServerAlive 15 s × 4, and `home/dot_local/bin/herdr-projector` exits its own master and removes its sockets — so what lingers is one idle bridge process on the coordinator. MEASURED 2026-09-13: `grep -rn ClientAlive modules hosts` in this tree is empty | Tom, if lingering bridges ever pile up (`pgrep -af remote-client-bridge` on the coordinator); then a reviewed edit to modules/common.nix | | DF-SCREEN-OCR-1 | Putting `screen-ocr` (Mod+Ctrl+S) under a prioritized GPU lease instead of calling Halogen directly | Tom, 2026-09-15: a busy Halogen must not break the flow, and hierarchical prioritization of GPU work belongs to `tally`, which is still being built. Until then the script has no idle gate: it waits behind in-flight requests (curl `--max-time 600`) and fails without touching the clipboard | Tom, once `tally` offers a lease a short interactive request can take; then a reviewed edit to `home/dot_local/bin/screen-ocr` | +**DF-U-D13-2 is gone (2026-09-17).** It deferred PASSING `--evaluator-lock` to +the served kernel, on one condition: `apps/evaluator` existing in the lake to be +locked. It exists, in the pinned `tally-lake` input's own store path, so the +condition occurred and the row left. `modules/tally-b.nix` now BUILDS the lock — +the kernel's own `tools/make-evaluator-lock.sh` over `pkgs/tally-evaluator`'s +fixed argv and `apps/evaluator`'s bytes — and the coordinator's ExecStart carries +it. Nothing about it is transcribed by hand, and guard G3 of +`tests/tally-b/probe-u-d13-guards.sh` recomputes every row with the kernel's own +`scripts/verify-evaluator-lock.sh` (and takes it red under a foreign argv). What +the lock does NOT do is change any kit entry: `DF-U-D13-4` above carries that +half. + **The U-D19 switch rows are gone (2026-09-13, #380).** DF-U-D12-1, DF-U-D13-1, DF-U-D14-1, DF-U-D14-2, DF-U-D14-4, DF-U-D15-1, DF-U-D16-1, DF-U-D16-2, DF-U-D17-1, DF-U-D18-1, DF-CAP-1-1 and DF-MEM-2-1 all waited on one act, the diff --git a/docs/local-ai/tally-b-input.md b/docs/local-ai/tally-b-input.md index 608de8372..423af111c 100644 --- a/docs/local-ai/tally-b-input.md +++ b/docs/local-ai/tally-b-input.md @@ -17,7 +17,7 @@ pins: the spec history that grew the rewrite inside it. ```nix tally-b = { - url = "git+https://github.com/mecattaf/tally?rev=26d758049bf0e89126157b3ea743085bb1b918f0"; + url = "git+https://github.com/mecattaf/tally?rev=b3a040e423926c542d794736d3976ec514bad02f"; flake = false; }; ``` @@ -46,10 +46,18 @@ printed to establish any of this. The consequence is stated plainly in on a host whose git can authenticate to github.com. After that act the git cache and the store path make every gate `--offline`-clean anywhere. -**The rev** `26d7580` is `origin/main` of mecattaf/tally at the pin: the merged -head past U-B13 (PR #45, `k/socket`) plus the evaluator's own probe commit. It -is the commit U-B13's deliverable sits on, and the card's oracle requires -exactly that — "the input pinned to a pushed commit of mecattaf/tally". +**The rev** `b3a040e` is `origin/main` of mecattaf/tally at the pin. U-D13 was +delivered against the merged head past U-B13 (PR #45, `k/socket`, plus the +evaluator's own probe commit); the pin has since moved 27 commits along the same +`main` — FOLD-1/FOLD-2K, the kernel answering an unowned row as `row_unknown` +(#50) instead of crashing the box uplink, FIX-E03, and the two `docs/rows.md` +publications the uplink reads (capacity per row, D-E2E-5; the envelope, +D-W03-5). It is a pushed commit, and the card's oracle requires exactly that — +"the input pinned to a pushed commit of mecattaf/tally". The ONE place the rev +is authoritative is `flake.nix`; this document names it so a reader can see +which `main` the prose below was measured against, and clause D of +`tests/tally-b/test-tally-b-input.sh` re-reads the pin from `flake.nix` and +`flake.lock` rather than from here. `tests/tally-b/test-tally-b-input.sh` clause D checks the pin is an ancestor of the clone's `origin/main` without touching the network. Bump by editing the rev in `flake.nix` and running `nix flake lock --update-input tally-b`, @@ -99,9 +107,84 @@ pin must be a NO-OP, and clause A0 asserts it. legible failure, never a silent no-op. They coexist with seat-feeder's user-bus rules over the same two paths — `d` lines are idempotent and both say 0700 tom. -- **`evaluatorLock`**, default null: without `--evaluator-lock` no verdict is - ever derived (tally `docs/socket.md` §4), which is the right state until - U-A17's `apps/evaluator` exists in the lake to be locked. +- **`evaluatorLock`**, default the BUILT lock below: with it the served kernel + derives a `verdict`; without it none is ever derived (tally + `docs/socket.md` §4). Null is still accepted and still means "derive no + verdict" — the honest state for a host that serves a kernel with no evaluator + to lock. + +## The evaluator lock this module serves + +`DEFERRED.md` `DF-U-D13-2` deferred one thing: PASSING `--evaluator-lock`, and +it waited on one condition — `apps/evaluator` existing in the lake to be +locked. It exists, in the pinned `tally-lake` input's own store path, so the +row is discharged here and the flag is served. + +**What a lock is.** The kernel derives a `verdict` only from the attestation of +an `exec.run` whose `argv_sha256` matches the evaluator's pinned lock (spec +§2.1). The lock file is `sha256sum`'s format with one reserved row: + +``` +<64 hex> argv the digest of the evaluator's ARGV +<64 hex> tools/e2e-evaluator.sh one row per pinned file +<64 hex> apps/evaluator/bin/evaluate.mjs +… +``` + +The `argv` row names no file: it is the digest of the executor's own prefix-free +argv preimage (`crates/tally-kernel/src/exec.rs`, `argv_preimage`; transcribed +in `tools/make-evaluator-lock.sh`). An argv that carried the card would hash +differently every run and could not be pinned at all, which is why everything +per-evaluation travels on STDIN (T7-3, a RULED-NEVER row of the lake's +`DEFERRED.md`). + +**The fixed argv, as a package.** `pkgs/tally-evaluator/default.nix` is a +`writeShellApplication` whose whole text is `exec /bin/sh +${inputs.tally-lake}/tools/e2e-evaluator.sh "$@"`, with `runtimeInputs` +`[ jq coreutils git gnused ]`. Two things follow. The argv the kernel hashes is +ONE WORD — `/bin/tally-evaluator` — so it moves only when the +`tally-lake` pin moves, never when a checkout is `git pull`ed under nobody's +review. And the evaluator's PATH is the wrapper's closure rather than whatever +environment started the kernel. `node` is deliberately NOT in that closure: the +item on stdin names the interpreter by absolute path, because the node the +lake's own `scripts/node-env.sh` records is the one its packages were resolved +against. + +**The lock, as a derivation.** `tallyEvaluatorLock` is a `runCommand` that runs +the KERNEL's own `${inputs.tally-b}/tools/make-evaluator-lock.sh` with +`--root ${inputs.tally-lake}`, `--argv ${tallyEvaluator}/bin/tally-evaluator` +and the seven `--file` rows (`tools/e2e-evaluator.sh` plus `apps/evaluator`'s +`bin/evaluate.mjs` and `src/{evaluate,env,normalize,usage,receipt-ingestion}.mjs`). +No digest is transcribed into this repository: the argv row is computed from +the store path of the wrapper this repo builds, and every file row is computed +from the pinned input, so the lock cannot drift from either pin without the +derivation itself changing. The rendered unit therefore reads + +``` +tally-kernel serve --state … --rows … --socket … --evaluator-lock /nix/store/…-tally-evaluator-lock +``` + +**How it is checked.** Guard G3 of `tests/tally-b/probe-u-d13-guards.sh` builds +the lock offline and runs the kernel's own +`${inputs.tally-b}/scripts/verify-evaluator-lock.sh --lock +--root ${inputs.tally-lake} --argv /bin/tally-evaluator`: part C +recomputes the argv digest and part D recomputes every pinned file from the lake +input. rc 0 is the green; the guard also runs the same verifier with the +two-word argv that preceded the wrapper (`/bin/sh /tools/e2e-evaluator.sh`) +and requires rc 4, so part C is known to be comparing something. Clause E of +`tests/tally-b/test-tally-b-input.sh` asserts the flag reaches ExecStart and +names a STORE path — a lock under `$HOME` would be a file anyone could edit +between two evaluations, which is the one thing a lock exists to prevent. + +**A file added to `apps/evaluator` runs UNPINNED** unless it is added to that +`--file` list. The list lives beside the argv it belongs to, in +`modules/tally-b.nix`, rather than in a generated manifest, precisely so that +adding one is a reviewed edit. + +**MEASURED**, 2026-09-17 (the integration audit's probe C): this same tool pair +generated a lock the served kernel loaded, an `exec.run` of the wrapped launcher +produced an attestation whose `argv_sha256` matched it, and the chain carried a +`verdict`. ## Coexistence with the live daemon, as bytes diff --git a/modules/tally-b.nix b/modules/tally-b.nix index 0f6cc4f9f..5ca1df818 100644 --- a/modules/tally-b.nix +++ b/modules/tally-b.nix @@ -96,6 +96,58 @@ let rowsFile = pkgs.writeText "tally-b-rows.json" (builtins.toJSON cfg.rows); + # THE EVALUATOR, AND THE LOCK OVER IT (DEFERRED.md DF-U-D13-2, discharged). + # + # Without `--evaluator-lock` the served kernel derives NO verdict at all + # (tally docs/socket.md §4), so the admission -> lease -> execute -> witness + # -> VERDICT -> receipt loop stops one rung short of a receipt. The row that + # deferred it waited on one thing — `apps/evaluator` existing in the lake to + # be locked — and it exists: `${inputs.tally-lake}/apps/evaluator` is in the + # pinned input's store path, behind the fixed-argv launcher + # `tools/e2e-evaluator.sh` beside it. + # + # Two derivations, in the order the kernel reads them: + # + # tallyEvaluator the FIXED ARGV: one store path, `pkgs/tally-evaluator`, + # exec-ing that launcher. Everything per-evaluation + # (card, deliverable, usage source, receipt path, node) + # travels on STDIN, never in argv — T7-3, and the only + # reason an argv can be pinned at all. + # tallyEvaluatorLock the lock itself, generated at BUILD TIME by the + # KERNEL's own tools/make-evaluator-lock.sh over the + # LAKE's own bytes. Neither digest is transcribed by + # hand into this repository: the argv row is computed + # from the store path above, and every file row is + # computed from the pinned input, so the lock cannot + # drift from either pin without the derivation changing. + # + # MEASURED (2026-09-17, audit probe C): this exact tool pair generated a lock + # the served kernel loaded, and an `exec.run` of the wrapped launcher produced + # an attestation whose `argv_sha256` matched it and a `verdict` on the chain. + # + # The seven pinned files are apps/evaluator's whole reachable source plus the + # launcher — the bytes that decide, and nothing else. A file added to the + # evaluator that is not named here would run UNPINNED, which is why this list + # lives beside the argv it belongs to rather than in a generated manifest. + tallyEvaluator = pkgs.callPackage ../pkgs/tally-evaluator { + tallyLake = inputs.tally-lake; + }; + + tallyEvaluatorLock = pkgs.runCommand "tally-evaluator-lock" { } '' + sh ${inputs.tally-b}/tools/make-evaluator-lock.sh \ + --root ${inputs.tally-lake} \ + --title "the served coordinator kernel's evaluator, pinned by modules/tally-b.nix" \ + --argv ${tallyEvaluator}/bin/tally-evaluator \ + --file tools/e2e-evaluator.sh \ + --file apps/evaluator/bin/evaluate.mjs \ + --file apps/evaluator/src/evaluate.mjs \ + --file apps/evaluator/src/env.mjs \ + --file apps/evaluator/src/normalize.mjs \ + --file apps/evaluator/src/usage.mjs \ + --file apps/evaluator/src/receipt-ingestion.mjs \ + --out $out + ''; + package = pkgs.rustPlatform.buildRustPackage { pname = "tally-b-kernel"; version = "0.0.1"; # the workspace's own [workspace.package] version @@ -176,11 +228,25 @@ in evaluatorLock = lib.mkOption { type = lib.types.nullOr lib.types.path; - default = null; + default = tallyEvaluatorLock; + defaultText = lib.literalExpression "a runCommand over inputs.tally-b's tools/make-evaluator-lock.sh, pinning pkgs/tally-evaluator's argv and apps/evaluator's bytes out of inputs.tally-lake"; description = '' - Path to apps/evaluator's pinned lock (EVALUATOR.sha256). Null means no - verdict is ever derived — the right state until U-A17's evaluator - exists to be locked (tally docs/socket.md §4). + Path to apps/evaluator's pinned lock (the EVALUATOR.sha256 shape: one + reserved `argv` row plus one row per pinned file). Without it the kernel + derives no verdict at all (tally docs/socket.md §4), so the loop stops + one rung short of a receipt. + + The default is BUILT, never transcribed: the kernel's own + `tools/make-evaluator-lock.sh` runs over the pinned `tally-lake` input + with `pkgs/tally-evaluator`'s store path as the single argv word, so the + lock, the argv it pins and the bytes it pins all move together with the + two input pins and with nothing else. + `scripts/verify-evaluator-lock.sh --lock --root --argv /bin/tally-evaluator` recomputes every row and + is guard G3 of tests/tally-b/probe-u-d13-guards.sh. + + Null is still accepted and still means "derive no verdict" — the honest + state for a host that serves a kernel with no evaluator to lock. ''; }; }; diff --git a/pkgs/tally-evaluator/default.nix b/pkgs/tally-evaluator/default.nix new file mode 100644 index 000000000..21ed2e569 --- /dev/null +++ b/pkgs/tally-evaluator/default.nix @@ -0,0 +1,68 @@ +{ + coreutils, + git, + gnused, + jq, + lib, + tallyLake, + writeShellApplication, +}: +# tally-evaluator — the FIXED ARGV the kernel derives a verdict under. +# +# UNIT: the served `--evaluator-lock` (dotfiles DEFERRED.md DF-U-D13-2, the row +# this package discharges). SPEC: tally `docs/socket.md` §4 and the rewrite's +# spec §2.1 — "`verdict` payloads are derived by the kernel from the attestation +# of an `exec.run` whose `argv_sha256` matches apps/evaluator's pinned lock". +# +# WHY A WRAPPER AND NOT A PATH. The lock pins a DIGEST OF THE ARGV, computed +# over the executor's own prefix-free preimage (tally +# `crates/tally-kernel/src/exec.rs` `argv_preimage`, transcribed in +# `tools/make-evaluator-lock.sh`). An argv that carries the card, the +# deliverable or the usage source hashes differently on every evaluation and +# therefore cannot be pinned at all — which is exactly why T7-3 is a +# RULED-NEVER row of the lake's DEFERRED.md and why `tools/e2e-evaluator.sh` +# takes its whole item as ONE JSON OBJECT ON STDIN. This package is the last +# step of that discipline: it turns the two-word argv probe-C measured +# (`/bin/sh /tools/e2e-evaluator.sh`) into ONE word that is a store +# path, so the argv the kernel hashes moves only when the pinned `tally-lake` +# input moves — never when a checkout is `git pull`ed under nobody's review. +# +# WHAT IT EXECS. `${tallyLake}/tools/e2e-evaluator.sh`, out of the flake input's +# store path, through `/bin/sh` — the same interpreter probe-C ran it under +# (MEASURED 2026-09-17: rc 0, `argv_sha256` matching the generated lock, a +# `verdict` on the chain). The script's own `#!` line is not used, so the +# interpreter is part of the wrapper's text and thus part of what the lock's +# FILE rows cannot silently change. +# +# WHAT IT DOES NOT DO. It passes no judgement, adds no flag, and reads no +# credential: the evaluator's exit code IS the verdict (spec §2.2d) and the +# kernel only records it. `"$@"` is forwarded so a hand invocation can still be +# debugged, but the LOCKED argv is the bare program name with no words after +# it — an argv with extra words hashes to something the lock does not carry and +# the kernel derives no verdict from it, which is the refusal working. +# +# THE RUNTIME CLOSURE IS THE ORACLE'S PATH. A `writeShellApplication` sets PATH +# from `runtimeInputs` alone, so what the evaluator can reach is declared here +# rather than inherited from whatever started the kernel: `jq` (the script +# parses its stdin item with it), `coreutils` (`cat`, `date`, `mkdir`, `wc`), +# `git` (apps/evaluator's deliverable checkout) and `gnused`. `node` is NOT in +# this list on purpose — the item on stdin names the interpreter by absolute +# path (`.node`), because the node the lake's own `scripts/node-env.sh` records +# is the one its packages were resolved against. +writeShellApplication { + name = "tally-evaluator"; + runtimeInputs = [ + jq + coreutils + git + gnused + ]; + text = '' + exec /bin/sh ${tallyLake}/tools/e2e-evaluator.sh "$@" + ''; + meta = { + description = "the fixed-argv mechanical evaluator the tally kernel derives a verdict under"; + mainProgram = "tally-evaluator"; + platforms = lib.platforms.linux; + }; +} diff --git a/tests/tally-b/probe-u-d13-guards.sh b/tests/tally-b/probe-u-d13-guards.sh index 8277e24af..265bd5a2a 100755 --- a/tests/tally-b/probe-u-d13-guards.sh +++ b/tests/tally-b/probe-u-d13-guards.sh @@ -8,8 +8,10 @@ # tmpfiles rule sets this unit puts over the SAME two paths from two different # buses do not fight at switch time. # -# Everything here is eval-time and --offline: nothing is built, nothing is -# switched, no state is written, no credential is read. Each guard is shown +# Everything here is --offline, nothing is switched, no state is written and no +# credential is read. G1/G2/G4 are eval-time only; G3 additionally BUILDS one +# derivation — the evaluator lock the module now serves — because a lock whose +# rows were never recomputed is a claim and not a guard. Each guard is shown # GREEN at the delivered tree and RED under a one-line override, because a # guard nobody has seen red is not a guard. set -uo pipefail @@ -59,15 +61,80 @@ else bad "G2 an empty row set evaluated: $out" fi -# --- G3: evaluatorLock is a wire, not a stub ------------------------------- -# DF-U-D13-2 defers PASSING the lock (U-A17 has not delivered the bytes to -# lock). It does not excuse an option that is never read. Null must omit the -# flag; a set value must appear verbatim in ExecStart. +# --- G3: the SERVED evaluator lock recomputes, row for row ----------------- +# DF-U-D13-2 is discharged: the module no longer defers PASSING the lock, it +# builds one. So this guard flipped with it. The old clause asserted that the +# default (null) omitted the flag — the honest state while `apps/evaluator` had +# not been delivered to the lake. The bytes exist now +# (`${inputs.tally-lake}/apps/evaluator`), the module generates a lock over them +# with the KERNEL's own tools/make-evaluator-lock.sh, and what must be guarded +# is no longer "is the option read" but "does the served lock still mean what it +# says": its argv row must be the digest of the wrapper this repository builds, +# and every file row must recompute byte-for-byte from the PINNED lake input. +# +# That is exactly `scripts/verify-evaluator-lock.sh --lock … --root … --argv …` +# from the pinned kernel input — the kernel's own refusal, run ahead of the +# kernel. Nothing here is transcribed: the lock, the wrapper, the lake root and +# the verifier all come out of this tree's two pins. +lake=$(nix eval --offline --raw --impure --expr \ + "(builtins.getFlake \"git+file://$repo\").inputs.tally-lake.outPath" 2>/dev/null) +kern=$(nix eval --offline --raw --impure --expr \ + "(builtins.getFlake \"git+file://$repo\").inputs.tally-b.outPath" 2>/dev/null) +# The wrapper is re-derived here from pkgs/ rather than read off the lock's own +# `# argv:` comment: reading the argv out of the artifact under test would make +# part C of the verifier tautological. +wrapper=$(nix eval --offline --raw --impure --expr \ + "let f = builtins.getFlake \"git+file://$repo\"; + in (f.nixosConfigurations.coordinator.pkgs.callPackage $repo/pkgs/tally-evaluator { + tallyLake = f.inputs.tally-lake; + }).outPath" 2>/dev/null) +lock=$(nix build --offline --no-link --print-out-paths \ + "$repo#nixosConfigurations.coordinator.config.services.tally-kernel.evaluatorLock" 2>/dev/null) +case "$lock" in + /nix/store/*) pass "G3 the default evaluatorLock is a built store path: $lock" ;; + *) bad "G3 the default evaluatorLock did not build to a store path: ${lock:-}" ;; +esac +if [ -n "$lock" ] && [ -n "$lake" ] && [ -n "$kern" ] && [ -n "$wrapper" ] \ + && sh "$kern/scripts/verify-evaluator-lock.sh" \ + --lock "$lock" --root "$lake" --argv "$wrapper/bin/tally-evaluator" >/dev/null 2>&1; then + pass "G3 verify-evaluator-lock.sh rc 0: every file row recomputes from the pinned lake, and the argv row is the digest of $wrapper/bin/tally-evaluator" +else + bad "G3 verify-evaluator-lock.sh refused the served lock (lock=${lock:-} root=${lake:-} argv=${wrapper:-}/bin/tally-evaluator)" +fi +# THE RED SIDE. The same verifier, over the same lock, with the argv probe-C ran +# BEFORE the wrapper existed (`/bin/sh /tools/e2e-evaluator.sh`, two +# words). It must DIVERGE (rc 4) — otherwise part C is not comparing anything +# and the green above would mean only "the file parses". +if [ -n "$lock" ] && [ -n "$lake" ] && [ -n "$kern" ]; then + sh "$kern/scripts/verify-evaluator-lock.sh" \ + --lock "$lock" --root "$lake" \ + --argv /bin/sh --argv "$lake/tools/e2e-evaluator.sh" >/dev/null 2>&1 + rc=$? + if [ "$rc" -eq 4 ]; then + pass "G3 the pre-wrapper two-word argv DIVERGES against the same lock (rc 4) — part C really compares" + else + bad "G3 a foreign argv did not diverge against the served lock (rc $rc, wanted 4)" + fi +else + bad "G3 could not run the negative control: lock/root/kernel input unresolved" +fi +# The flag reaches the rendered unit, naming that same store lock. base=$(nix eval --offline --raw \ '.#nixosConfigurations.coordinator.config.systemd.services.tally-kernel.serviceConfig.ExecStart' 2>/dev/null) case "$base" in - *--evaluator-lock*) bad "G3 the default (null) lock still emitted --evaluator-lock: $base" ;; - *) pass "G3 evaluatorLock=null omits the flag (the deferred, honest state)" ;; + *"--evaluator-lock $lock"*) + pass "G3 ExecStart serves it: …${base##*--evaluator-lock}" ;; + *) + bad "G3 ExecStart does not carry --evaluator-lock $lock: ${base:-}" ;; +esac +# And null is STILL a wire, not a hole: a host with no evaluator to lock can set +# it back and the flag disappears rather than pointing at a stale store path. +out=$(ext 'services.tally-kernel.evaluatorLock = null;' \ + 'c.config.systemd.services.tally-kernel.serviceConfig.ExecStart') +case "$out" in + *--evaluator-lock*) bad "G3 evaluatorLock=null still emitted --evaluator-lock: $out" ;; + *tally-kernel*) pass "G3 evaluatorLock=null omits the flag (the option is still a wire in both directions)" ;; + *) bad "G3 could not render ExecStart with evaluatorLock=null: $out" ;; esac # G0: the delivered tree's own toplevel evaluates — the GREEN side of G1/G2, # so a guard that fired on everything would be caught here. @@ -80,7 +147,7 @@ fi out=$(ext 'services.tally-kernel.evaluatorLock = /etc/hostname;' \ 'c.config.systemd.services.tally-kernel.serviceConfig.ExecStart') if [ $? -eq 0 ] && printf '%s' "$out" | grep -q -- '--evaluator-lock /nix/store/.*'; then - pass "G3 a set evaluatorLock reaches ExecStart: ...${out##*--evaluator-lock}" + pass "G3 a SET evaluatorLock still overrides the default and reaches ExecStart: ...${out##*--evaluator-lock}" else bad "G3 evaluatorLock is not wired to ExecStart: $out" fi diff --git a/tests/tally-b/test-tally-b-input.sh b/tests/tally-b/test-tally-b-input.sh index 5c4330d03..14f3191c6 100755 --- a/tests/tally-b/test-tally-b-input.sh +++ b/tests/tally-b/test-tally-b-input.sh @@ -18,7 +18,8 @@ # network, no credential — `git ls-remote` is NOT called here) # E the unit's shape: ExecStart is the store-built tally-kernel binary with # `serve`, the state root carries the tally-rewrite component, the socket -# is kernel.sock BESIDE that root, the rows file names exactly the +# is kernel.sock BESIDE that root, `--evaluator-lock` names a STORE lock +# (so a verdict is derived at all), the rows file names exactly the # three kernel-owned rows of the rewrite's docs/rows.md, and every row's # `running` is `{kind: none}` — no row probes an endpoint, because no # device on this estate publishes one @@ -147,6 +148,21 @@ case "$exec_start" in *) bad "E --socket is not /kernel.sock: $exec_start" ;; esac +# The served EVALUATOR LOCK (DEFERRED.md DF-U-D13-2, discharged). Without the +# flag the kernel derives no verdict at all (tally docs/socket.md §4) and the +# loop stops one rung short of a receipt, so its presence is part of the unit's +# shape and not an option's detail. It must be a STORE path: a lock under $HOME +# is a file anyone can edit between two evaluations, and the whole point of the +# lock is that the bytes which judged a deliverable are named and immutable. +# Which rows it carries, and that they recompute from the pinned lake input, is +# guard G3 of tests/tally-b/probe-u-d13-guards.sh — the clause here is the +# unit's, not the lock's. +case "$exec_start" in + *"--evaluator-lock /nix/store/"*) + pass "E --evaluator-lock names a store lock: ${exec_start##*--evaluator-lock }" ;; + *) + bad "E ExecStart carries no --evaluator-lock : $exec_start" ;; +esac # The state root must NEVER be the live estate's: this is the eval-time twin of # the kernel's own Ledger::open refusal (ledger.rs:31-35). case "$exec_start" in From af56e72f1668ca55cb69b7c1187ff54bcb4c0a69 Mon Sep 17 00:00:00 2001 From: mecattaf Date: Thu, 17 Sep 2026 07:18:01 +0200 Subject: [PATCH 2/3] seat-feeder: bound the Claude reader at 16 s and retain from the reader's own cache MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit U-D12 (dotfiles#315) / CAP-1 (dotfiles#337), integration gap Int-G11. MEASURED over ~/.local/state/tally-rewrite/uplink/events.jsonl for 2026-09-16T04Z -> 09-17T04Z: the `cc` row read STALE-MEASURED on 135 of 286 wakes. Two causes, both here. (1) CLAUDE_READER_TIMEOUT_SECONDS 12 -> 16. D-B54's bound is unchanged (30 + 1 + 20 = 51 < 60); the three readers are CONCURRENT, so three 16-second reads fit in ONE window inside the 20-second TimeoutStartSec and leave four seconds for shaping and one os.replace per row. A read that lands in thirteen seconds was a MEASURED reading this box threw away for a retained one; the cap is there so a HUNG reader cannot push the service past its deadline, not to budget the shaping. (2) On a failed read the retained reading now comes from the READER'S OWN cache `.window-cache-.json` in preference to this feeder's last published row whenever the two name the same measured instant (`last_measured`'s tuple order IS the tie-break, and it is now documented as such). The cache is the measurement; the published row is a projection that drops cells (the five-hour `resets_at` in the sentinel-window branch, a `model_split` the endpoint answered null for), so re-publishing the row degraded the retained reading a little further on every failed read. `reading_age_seconds` is now the cache's TRUE age — which is what `reading_source` already claimed by naming that file by path. `stale_reason` is unchanged. tests/tally-b/probe-seat-feeder-timeout.sh is the fixture: a fake reader and a scratch meters dir, no credential opened, the live meters dir neither read nor written, no unit started. A reader that sleeps 14 s over a 30-second-old cache -> grade MEASURED, reading_age_seconds 30, no stale_reason; a reader that sleeps 20 s -> cut off at 16 s, STALE-MEASURED, stale_reason naming TimeoutExpired, reading_source the .window-cache-cc.json path, age 300. Run against a `main` worktree it goes RED in exactly the four places this change moves (recorded in its header). home/seat-feeder.nix carries the new delivered sha256 (59fcc77063f19767d0fdfcf94effb48cafb7f2ef26c4441c92301b147efc8813) and the corrected arithmetic comment; docs/local-ai/seat-feeder.md says both halves. Neither change can manufacture a reading from an expired token: re-logging in on cc and cc3 stays Tom's (DEFERRED.md DF-U-D12-3). Co-Authored-By: Claude Opus 5 --- DECISIONS.md | 22 +- docs/local-ai/seat-feeder.md | 32 ++- home/dot_local/bin/tally-seat-feeder | 48 +++- home/seat-feeder.nix | 16 +- tests/tally-b/probe-seat-feeder-timeout.sh | 261 +++++++++++++++++++++ 5 files changed, 355 insertions(+), 24 deletions(-) create mode 100755 tests/tally-b/probe-seat-feeder-timeout.sh diff --git a/DECISIONS.md b/DECISIONS.md index 591e7158d..bbb725fc0 100644 --- a/DECISIONS.md +++ b/DECISIONS.md @@ -1,6 +1,7 @@ # DECISIONS -2026-09-17 the served kernel derives verdicts. +2026-09-17 the served kernel derives verdicts, and a slow seat read is not a +stale one. **The evaluator lock is BUILT, not pinned by hand (default, unruled).** No ruling says where the served `--evaluator-lock` comes from, and there were two @@ -35,6 +36,25 @@ what a mechanical verdict for CUBS IS is Tom's line (integration G7), and the per-item stdin an `eval(...)` entry would need is the uplink render tool's second step, not this module's constant `stdin` cell. +**The Claude seat reader is bounded at 16 s, and a failed read re-publishes the +READER'S cache (default, unruled).** D-B54 fixes the arithmetic (30 + 1 + 20 = +51 < 60) and leaves the reader's own bound free inside the 20-second +`TimeoutStartSec`; the reads are concurrent, so three of them fit in one window. +12 seconds was cutting off reads that were going to land — MEASURED over the +uplink's `events.jsonl` for 2026-09-16T04Z → 09-17T04Z, `cc` read +`STALE-MEASURED` on 135 of 286 wakes. And when a read genuinely cannot land, the +retained reading is now taken from the reader's own `.window-cache-.json` +in preference to this feeder's own last row on a tie: the cache is the +measurement, the row is a projection of it that drops cells (the five-hour +`resets_at`, the null `model_split`), so re-publishing the row degrades the +retained reading a little further on every failed read while re-publishing the +cache does not — and `reading_age_seconds` becomes the age of the measurement, +which is what `reading_source` already claimed by naming that file. +`tests/tally-b/probe-seat-feeder-timeout.sh` holds both halves down and was +MEASURED red against the pre-change feeder in exactly those places. Neither +change can manufacture a reading from an expired token: re-logging in on `cc` +and `cc3` stays Tom's (`DEFERRED.md` DF-U-D12-3). + 2026-09-13 flake checkouts no longer ride into host closures, chrome-stream is installed, and a switch refuses a stale raw-dotfiles checkout. diff --git a/docs/local-ai/seat-feeder.md b/docs/local-ai/seat-feeder.md index cd4e69c2e..c16392acc 100644 --- a/docs/local-ai/seat-feeder.md +++ b/docs/local-ai/seat-feeder.md @@ -33,10 +33,21 @@ staleness bound. A service still active at 20 seconds is terminated and its unit fails; it cannot silently publish outside that envelope. The worker configuration evaluates with no feeder unit. -The Claude service runs its three readers concurrently. Each reader retains -its 12-second timeout, leaving eight seconds inside the service cap for row -shaping and publication. As each read returns, that seat is shaped and written -immediately; a slow seat cannot hold completed seats in a batch. +The Claude service runs its three readers concurrently. Each reader is bounded +at 16 seconds, leaving four seconds inside the service cap for row shaping and +publication. Because the reads are concurrent, three of them fit inside ONE +16-second window and not inside three. As each read returns, that seat is shaped +and written immediately; a slow seat cannot hold completed seats in a batch. + +The bound was 12 seconds until 2026-09-17. MEASURED over the uplink's +`events.jsonl` for 2026-09-16T04Z → 09-17T04Z, `cc` read `STALE-MEASURED` on +**135 of 286** wakes: a read that lands in thirteen seconds was a MEASURED +reading this box threw away for a retained one, and a read that cannot land in +sixteen is a reader that is not answering this tick. The four seconds left over +are ample — shaping is arithmetic and publication is one `os.replace` per row; +the cap exists so a HUNG reader cannot push the service past +`TimeoutStartSec`, not to budget the shaping. `tests/tally-b/probe-seat-feeder-timeout.sh` +is the fixture that holds both halves down. ## Source boundaries @@ -116,10 +127,15 @@ endpoint answers in whole percents. The feeder follows that through: `STALE-MEASURED` with `reading_age_seconds`, `reading_observed_at`, `reading_source` and the failure's own `stale_reason`. -Two retained sources, and the newer wins: the row this feeder last published -(already in the contract's shape) and the reader's own -`.window-cache-.json`, whose directory the reader hard-codes and which -`TALLY_WINDOW_CACHE_DIR` names so a fixture can redirect it. Ages chain off +Two retained sources, and the newer wins — **and on a tie the reader's own cache +wins**: `.window-cache-.json`, whose directory the reader hard-codes and +which `TALLY_WINDOW_CACHE_DIR` names so a fixture can redirect it, is preferred +over the row this feeder last published. The cache is the reading; the published +row is a projection of it, and going back through the projection loses cells the +row never promised to carry (the five-hour `resets_at` the sentinel-window branch +drops, a `model_split` the endpoint answered null for). So the cache is what +`reading_source` names and `reading_age_seconds` is the cache's true age, not the +age of whatever the last row happened to say. Ages chain off `reading_observed_at`, the instant the numbers were MEASURED, never off publication, so re-publishing a re-publication cannot make a reading look younger than it is. Only when nothing at all was retained does the row fall diff --git a/home/dot_local/bin/tally-seat-feeder b/home/dot_local/bin/tally-seat-feeder index be72b4e59..2bcec6d8b 100755 --- a/home/dot_local/bin/tally-seat-feeder +++ b/home/dot_local/bin/tally-seat-feeder @@ -79,7 +79,10 @@ AND A FAILED READ KEEPS THE LAST MEASURED ONE. When the sanctioned reader times out, hands back an expired token or answers UNKNOWN, the honest row is not a blank one: it is the last MEASURED reading this box retained, re-published with `reading_age_seconds`, `reading_observed_at`, -`reading_source` and the grade STALE-MEASURED. D-B92 already rules that a +`reading_source` and the grade STALE-MEASURED. It is retained from the +READER'S OWN cache (`.window-cache-.json`) whenever that cache is at +least as new as the last row this feeder published, so the age published is +the age of the measurement and not of a projection of it. D-B92 already rules that a reading under 45 s old is the same measurement, so under that age the grade stays MEASURED and only the age cell moves; over it the row says plainly how old its numbers are. An old reading that says its age is @@ -115,10 +118,21 @@ KILL_GRACE_SECONDS = 10 # section 1.1 measured. pi-qwencloud has none it can prove: absent, not zero. CONTEXT_WINDOW = {"claude": 200_000, "codex": 258_400} -# The three readers run concurrently. Each is individually bounded at 12 -# seconds, leaving eight seconds inside the systemd service's 20-second hard -# cap for row shaping and atomic publication (D-B54). -CLAUDE_READER_TIMEOUT_SECONDS = 12 +# The three readers run concurrently. Each is individually bounded at 16 +# seconds, leaving four seconds inside the systemd service's 20-second hard cap +# for row shaping and atomic publication (D-B54: 30 + 1 + 20 = 51 < 60, and the +# 20-second term is what this bound sits inside). The reads are CONCURRENT, so +# three of them fit in one 16-second window and not in three. +# +# WHY IT MOVED FROM 12. MEASURED over ~/.local/state/tally-rewrite/uplink/ +# events.jsonl for 2026-09-16T04Z -> 09-17T04Z: `cc` read STALE-MEASURED on 135 +# of 286 wakes. A read that lands in thirteen seconds is a MEASURED reading this +# box threw away for a retained one; a read that cannot land in sixteen is a +# reader that is not going to answer this tick. The four seconds left over are +# for shaping and one os.replace per row, which is microseconds of work — the +# cap exists so a hung reader cannot take the service past TimeoutStartSec, not +# to budget the shaping. +CLAUDE_READER_TIMEOUT_SECONDS = 16 # The one sentinel this program writes into a cell it cannot fill. The kernel's # reader reads it as ABSENT, never as a number and never as a refusal @@ -477,15 +491,28 @@ def reading_from_published_row(directory, seat): def last_measured(directory, seat): """The most recent MEASURED reading this box retained for the seat. - Two sources, written by two clocks and neither reliably ahead of the other: - the row this feeder last published (already in the contract's shape) and - the sanctioned reader's own D-B92 cache. The newer wins. + Two sources, written by two clocks: the sanctioned reader's own D-B92 cache + and the row this feeder last published. The newer wins — and on a TIE the + CACHE wins, which is the whole of this function's opinion and is deliberate. + + The cache is the reading; the published row is a projection of it. Going + back through the row loses cells the row never promised to carry — a + five-hour `resets_at` the sentinel-window branch drops, a `model_split` the + endpoint answered null for and the row wrote as UNKNOWN — so re-publishing + from a row degrades the retained reading a little more on every failed read, + while re-publishing from the cache re-publishes the measurement itself. It + also makes `reading_age_seconds` the CACHE's true age rather than the age of + whatever the last row happened to say, which is what `reading_source` + already claims when it names the cache file by path. + + `max` returns the FIRST maximal element, so the order of this tuple IS the + tie-break; it is not incidental. """ kept = [ candidate for candidate in ( - reading_from_published_row(directory, seat), reading_from_cache(seat), + reading_from_published_row(directory, seat), ) if candidate is not None ] @@ -559,7 +586,8 @@ def claude_row(directory, seat, reading, reason): STALE-MEASURED over it, with the age in the row either way; * the read did not land at all (a timeout, an expired token, an endpoint error) — the last MEASURED reading this box retained is re-published - with its age, its origin and the failure's own reason. A row that says + with its age, its origin and the failure's own reason, taken from the + reader's own cache in preference to this feeder's own last row. A row that says "44% of five hours, measured 2,400 s ago" is evidence; a row with a null utilization is a hole, and CAP-1 exists because holes were being written. * the reading carries the five-hour utilization and NO reset instant (which diff --git a/home/seat-feeder.nix b/home/seat-feeder.nix index 850a36139..e686cd4d2 100644 --- a/home/seat-feeder.nix +++ b/home/seat-feeder.nix @@ -69,9 +69,10 @@ let # home/tally.nix's capacity oracle and home/harness-records.nix's recorders: # a threshold or a row field is retuned by editing a file, not by a rebuild. # Its delivered sha256 is - # 4d3dff21a6ee7d61ebe70564c090d7944c06a7a0d8f3e51f08bc4b5c1f7db4d1 - # (CAP-1; the D-B54 repair commit recorded 76fb8cb9… before every row stated - # its window), following UTIL-01's motion. + # 59fcc77063f19767d0fdfcf94effb48cafb7f2ef26c4441c92301b147efc8813 + # (the reader timeout raised to 16 s and the retained reading taken from the + # reader's own cache; CAP-1 recorded 4d3dff21…, and the D-B54 repair commit + # 76fb8cb9… before every row stated its window), following UTIL-01's motion. # A systemd user unit inherits no interactive PATH, so each unit supplies its # own. feeder = "%h/.local/bin/tally-seat-feeder"; @@ -91,8 +92,13 @@ let # D-B48: the period, timer accuracy, AND whole service duration belong in the # bound. 30 + 1 + 20 = 51 seconds, strictly inside the 60-second refusal # boundary. TimeoutStartSec below makes the 20-second term an enforced cap, - # not a timing hope; the feeder's concurrent 12-second readers leave eight - # seconds for shaping and atomic publication. + # not a timing hope; the feeder's readers are CONCURRENT, so three of them fit + # in ONE reader window and not in three, and at 16 seconds each that window + # leaves four seconds for shaping and atomic publication (one os.replace per + # row). 12 seconds was cutting off reads that were going to land: MEASURED + # over the uplink's events.jsonl for 2026-09-16T04Z -> 09-17T04Z, `cc` read + # STALE-MEASURED on 135 of 286 wakes. The cap is there so a HUNG reader cannot + # push the service past TimeoutStartSec, not to budget the shaping. policyTickSeconds = 60; feederPeriodSeconds = 30; timerAccuracySeconds = 1; diff --git a/tests/tally-b/probe-seat-feeder-timeout.sh b/tests/tally-b/probe-seat-feeder-timeout.sh new file mode 100755 index 000000000..927c4c016 --- /dev/null +++ b/tests/tally-b/probe-seat-feeder-timeout.sh @@ -0,0 +1,261 @@ +#!/usr/bin/env bash +# THE SEAT-FEEDER READER BOUND, AS A FIXTURE (Int-G11). +# +# Two facts about `home/dot_local/bin/tally-seat-feeder` that only a running +# clock can establish, and that the flake's eval-time checks cannot see: +# +# 1. A read that takes FOURTEEN seconds now LANDS. The bound was 12 s, and +# MEASURED over ~/.local/state/tally-rewrite/uplink/events.jsonl for +# 2026-09-16T04Z -> 09-17T04Z the `cc` row read STALE-MEASURED on 135 of +# 286 wakes. A reading thrown away at 12 s for a retained one is a row +# that says it is older than the measurement it could have had. So: a +# reader that sleeps 14 s and then serves its own 30-second-old D-B92 +# cache must produce grade MEASURED (30 <= the 45 s D-B92 boundary), +# `reading_age_seconds` 30, and NO `stale_reason`. +# +# 2. A read that CANNOT land is still bounded, and what it re-publishes is +# the READER'S OWN CACHE. A reader that sleeps 20 s is cut off at 16 s +# (`subprocess.TimeoutExpired`), and the row must be STALE-MEASURED with +# `stale_reason` naming TimeoutExpired, `reading_source` naming +# `.window-cache-cc.json` by path, and `reading_age_seconds` equal to that +# cache's TRUE age — not the age of the row the feeder last published. +# The second run of clause 2 is deliberately run after a successful one, +# so BOTH retained sources exist at the SAME observed instant: the tie is +# the whole point, and the cache is what must win it. +# +# WHAT THIS TOUCHES. A scratch directory and nothing else. The reader is a fake +# written here — no credential file is opened, no token exists, nothing leaves +# the box, and `TALLY_REWRITE_METERS` / `TALLY_WINDOW_CACHE_DIR` both point into +# the scratch tree, so the live ~/.local/state/tally-rewrite/meters is neither +# read nor written. No unit is started, restarted or enabled. +# +# THE CLOCK. `TALLY_FEEDER_NOW` pins the feeder's own clock so the published +# ages are exact integers rather than a race with the sleeps; the SLEEPS are +# real wall time, because the bound under test is a wall-clock bound. +# +# A GUARD NOBODY HAS SEEN RED IS NOT A GUARD. This one was run against the +# PRE-CHANGE feeder (a worktree at `main`) by passing that tree as `repo`, and +# it went red in exactly the four places the change moves and nowhere else: +# +# [F] T0 CLAUDE_READER_TIMEOUT_SECONDS is '12', wanted 16 +# [F] T1 wall time 12s is not a landed 14-second read +# [F] T1 the row carries stale_reason 'the reader could not be run: +# TimeoutExpired' after a successful read +# [F] T2 wall time 12s: the read was not bounded at 16 s +# [F] T2 reading_source is 'the last row this feeder published at …/cc.json', +# wanted the .window-cache-cc.json path +# +# Re-run it that way — `bash tests/tally-b/probe-seat-feeder-timeout.sh +# ` — whenever either half is +# touched. +# +# Usage: bash tests/tally-b/probe-seat-feeder-timeout.sh [repo] [scratch-dir] +set -uo pipefail + +repo="${1:-$(cd "$(dirname "${BASH_SOURCE[0]}")/../.." && pwd)}" +scratch="${2:-${TMPDIR:-/tmp}/tally-seat-feeder-timeout.$$}" +cd "$repo" || { echo "FAIL: cannot cd $repo"; exit 2; } + +fail=0 +pass() { printf '[P] %s\n' "$*"; } +bad() { printf '[F] %s\n' "$*"; fail=1; } + +feeder="$repo/home/dot_local/bin/tally-seat-feeder" +[ -f "$feeder" ] || { echo "FAIL: no feeder at $feeder"; exit 2; } + +# The python the coordinator's own unit runs the feeder under (home/seat-feeder.nix +# puts pkgs.python3 on the unit's PATH and nothing else). Evaluated, never a +# hardcoded store path. +python_store=$(nix eval --offline --raw "$repo#nixosConfigurations.coordinator.pkgs.python3.outPath" 2>/dev/null) +python="$python_store/bin/python3" +if [ ! -x "$python" ]; then + printf '[S] the evaluated python3 is unavailable (%s); this is an ENV limit, not a divergence\n' "${python_store:-}" + exit 3 +fi + +work="$scratch" +rm -rf -- "$work" +mkdir -p "$work/meters" || { echo "FAIL: cannot create $work/meters"; exit 2; } +trap 'rm -rf -- "$work"' EXIT + +NOW="2026-09-17T12:00:00Z" + +# --- the fake reader -------------------------------------------------------- +# stamp-receipt.py's interface only: `window --seat `, one JSON object on +# stdout, exit 0 on MEASURED. It sleeps for TALLY_FIXTURE_SLEEP seconds and then +# serves the cache file the REAL reader would have served (D-B92), stamping +# `from_cache_seconds` the way the real one does. It opens no credential. +cat > "$work/fake-reader.py" <<'PY' +import json, os, sys, time +from datetime import datetime, timezone + +def stamp(text): + return datetime.fromisoformat(text.strip().replace("Z", "+00:00")).astimezone(timezone.utc) + +def main(argv): + if len(argv) < 4 or argv[1] != "window" or argv[2] != "--seat": + sys.stderr.write("fake-reader: usage: window --seat \n") + return 64 + seat = argv[3] + time.sleep(float(os.environ.get("TALLY_FIXTURE_SLEEP", "0"))) + path = os.path.join(os.environ["TALLY_WINDOW_CACHE_DIR"], f".window-cache-{seat}.json") + with open(path, encoding="utf-8") as fh: + cached = json.load(fh) + observed = stamp(cached["observed_at"]) + now = stamp(os.environ["TALLY_FEEDER_NOW"]) + out = {"grade": "MEASURED", "seat": seat, "observed_at": cached["observed_at"], + "from_cache_seconds": int((now - observed).total_seconds())} + out.update(cached["usage"]) + print(json.dumps(out)) + return 0 + +if __name__ == "__main__": + sys.exit(main(sys.argv)) +PY + +write_cache() { # write_cache + "$python" - "$work/meters" "$NOW" "$1" <<'PY' +import json, os, sys +from datetime import datetime, timedelta, timezone +meters, now_text, age = sys.argv[1], sys.argv[2], int(sys.argv[3]) +now = datetime.fromisoformat(now_text.replace("Z", "+00:00")).astimezone(timezone.utc) +observed = (now - timedelta(seconds=age)).strftime("%Y-%m-%dT%H:%M:%SZ") +doc = { + "observed_at": observed, + "usage": { + "five_hour": {"utilization": 43, "resets_at": "2026-09-17T16:39:59Z"}, + "seven_day": {"utilization": 79, "resets_at": "2026-09-23T09:59:00Z"}, + "seven_day_opus": {"utilization": None, "resets_at": None}, + "seven_day_sonnet": {"utilization": None, "resets_at": None}, + }, +} +with open(os.path.join(meters, ".window-cache-cc.json"), "w", encoding="utf-8") as fh: + json.dump(doc, fh) +print(observed) +PY +} + +run_feeder() { # run_feeder + TALLY_REWRITE_METERS="$work/meters" \ + TALLY_WINDOW_CACHE_DIR="$work/meters" \ + TALLY_STAMP_RECEIPT="$work/fake-reader.py" \ + TALLY_CLAUDE_SEATS=cc \ + TALLY_FEEDER_NOW="$NOW" \ + TALLY_FIXTURE_SLEEP="$1" \ + "$python" "$feeder" claude >"$work/feeder.out" 2>"$work/feeder.err" +} + +cell() { # cell + "$python" -c 'import json,sys; row=json.load(open(sys.argv[1])); v=row.get(sys.argv[2]); print("" if v is None else v)' \ + "$work/meters/cc.json" "$1" 2>/dev/null +} + +# --- the declared bound, as bytes ------------------------------------------ +declared=$(sed -n 's/^CLAUDE_READER_TIMEOUT_SECONDS = \([0-9]*\)$/\1/p' "$feeder") +if [ "$declared" = "16" ]; then + pass "T0 CLAUDE_READER_TIMEOUT_SECONDS is 16 (30 + 1 + 20 = 51 < 60 keeps D-B54; three CONCURRENT 16 s reads fit the 20 s TimeoutStartSec)" +else + bad "T0 CLAUDE_READER_TIMEOUT_SECONDS is '${declared:-}', wanted 16" +fi + +# --- T1: a slow-but-successful read is MEASURED with the cache's own age ---- +observed=$(write_cache 30) +started=$(date +%s) +run_feeder 14 +rc=$? +elapsed=$(( $(date +%s) - started )) +if [ "$rc" -ne 0 ]; then + bad "T1 the feeder exited $rc: $(head -3 "$work/feeder.err")" +fi +if [ "$elapsed" -ge 14 ] && [ "$elapsed" -lt 30 ]; then + pass "T1 the 14-second read LANDED (${elapsed}s wall) — under the old 12-second bound it could not have" +else + bad "T1 wall time ${elapsed}s is not a landed 14-second read" +fi +grade=$(cell grade); age=$(cell reading_age_seconds); stale=$(cell stale_reason) +if [ "$grade" = "MEASURED" ]; then + pass "T1 grade MEASURED (the cache was $age s old, inside D-B92's 45 s)" +else + bad "T1 grade is '${grade:-}', wanted MEASURED" +fi +if [ "$age" = "30" ]; then + pass "T1 reading_age_seconds is the cache's true age: $age (cache observed $observed)" +else + bad "T1 reading_age_seconds is '${age:-}', wanted 30" +fi +if [ -z "$stale" ]; then + pass "T1 no stale_reason — a read that landed is not a stale row" +else + bad "T1 the row carries stale_reason '$stale' after a successful read" +fi +util=$(cell utilization_pct) +if [ "$util" = "43.0" ]; then + pass "T1 the measurement itself is published: utilization_pct $util" +else + bad "T1 utilization_pct is '${util:-}', wanted 43.0" +fi + +# --- T2: a read that cannot land -> STALE-MEASURED, from the CACHE ---------- +# First a successful pass over a 300-second-old cache, so the feeder publishes a +# row whose reading_observed_at EQUALS the cache's. Both retained sources then +# sit at the same instant and the tie-break is what is under test. +observed=$(write_cache 300) +run_feeder 0 +seeded=$(cell reading_observed_at) +if [ "$seeded" = "$observed" ]; then + pass "T2 seeded: the published row and the cache now name the same measured instant ($observed)" +else + bad "T2 could not seed the tie: row says '${seeded:-}', cache says '$observed'" +fi +started=$(date +%s) +run_feeder 20 +rc=$? +elapsed=$(( $(date +%s) - started )) +if [ "$rc" -ne 0 ]; then + bad "T2 the feeder exited $rc: $(head -3 "$work/feeder.err")" +fi +if [ "$elapsed" -ge 16 ] && [ "$elapsed" -lt 20 ]; then + pass "T2 the 20-second read was CUT OFF at the bound (${elapsed}s wall < the reader's own 20 s)" +else + bad "T2 wall time ${elapsed}s: the read was not bounded at 16 s" +fi +grade=$(cell grade); age=$(cell reading_age_seconds) +stale=$(cell stale_reason); origin=$(cell reading_source) +if [ "$grade" = "STALE-MEASURED" ]; then + pass "T2 grade STALE-MEASURED — the numbers are old and the row says so" +else + bad "T2 grade is '${grade:-}', wanted STALE-MEASURED" +fi +case "$stale" in + *TimeoutExpired*) pass "T2 stale_reason names the failure: $stale" ;; + *) bad "T2 stale_reason is '${stale:-}', wanted one naming TimeoutExpired" ;; +esac +case "$origin" in + *"/.window-cache-cc.json") + pass "T2 reading_source is the READER'S OWN cache, not this feeder's last row: $origin" ;; + *) + bad "T2 reading_source is '${origin:-}', wanted the .window-cache-cc.json path" ;; +esac +if [ "$age" = "300" ]; then + pass "T2 reading_age_seconds is the cache's true age: $age" +else + bad "T2 reading_age_seconds is '${age:-}', wanted 300" +fi +if [ "$(cell utilization_pct)" = "43.0" ]; then + pass "T2 the retained measurement is re-published, not blanked: utilization_pct 43.0" +else + bad "T2 the retained measurement was lost: utilization_pct '$(cell utilization_pct)'" +fi + +# --- T3: nothing outside the scratch tree was written ---------------------- +# The fixture's whole state is $work. The live meters dir is named here only to +# say that it was never opened for writing: the two environment variables above +# point elsewhere, and this clause is the assertion that they were honoured. +live="$HOME/.local/state/tally-rewrite/meters" +if [ ! -e "$work/../.window-cache-cc.json" ] && [ -f "$work/meters/cc.json" ]; then + pass "T3 every row and cache this fixture wrote is under $work (the live $live was not a target)" +else + bad "T3 the fixture wrote outside its scratch tree" +fi + +exit "$fail" From 415459f1ffe135da1f2dd9d5b5926a1c89dba1cf Mon Sep 17 00:00:00 2001 From: mecattaf Date: Thu, 17 Sep 2026 08:08:47 +0200 Subject: [PATCH 3/3] docs: name the inherited flake-check red correctly (nas-topology 8731) and defer it MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The PR body, commit c27e553 and the item report all called the one red `checks.x86_64-linux.nas-personal-tailnet` ("path '…-source' is not valid") and typed it as an offline-input `[ENV]` problem. MEASURED 2026-09-17 on this branch, `nix flake check --offline --no-build` instead dies at `checks.x86_64-linux.nas-topology`: error: assertion '(! ((builtins).elem 8731 (coordinator).networking.firewall.interfaces.wlp192s0.allowedTCPPorts))' failed at flake.nix:1694 That is a real eval assertion about a coordinator firewall port, not a missing store path. It is still inherited — this branch edits no flake.nix and no firewall or NAS configuration, and a detached worktree at main (202d9c31) run through the same command gives the byte-identical tail — but a reader would otherwise chase a tail that does not reproduce, and read a real assertion as an environment excuse. - DEFERRED.md gains DF-FLAKE-1 `[ENV]`, carrying BOTH tails verbatim, saying which surfaces first depends on evaluation order, and naming who discharges it (whoever owns the wlp192s0 line that opened 8731; branch fix/fleet-connectivity-artifact-count is already on it). - DECISIONS.md's 2026-09-17 entry gains one paragraph correcting the characterisation and pointing at that row. - DF-U-D13-4 is retyped `[OTHER-REPO]` -> `[SCOPE]`: only the G7 ruling is outside this repository, while the actionable half is a reviewed edit to home/tally-uplink.nix right here, in D-A's lane. No code, no test logic and no nix is touched, and no `[ENV]` fence is added to any probe: clause A of tests/tally-b/test-tally-b-input.sh stays red and the composite oracle stays rc 1, which is correct and expected. Co-Authored-By: Claude Opus 5 --- DECISIONS.md | 19 +++++++++++++++++++ DEFERRED.md | 3 ++- 2 files changed, 21 insertions(+), 1 deletion(-) diff --git a/DECISIONS.md b/DECISIONS.md index bbb725fc0..d24200c68 100644 --- a/DECISIONS.md +++ b/DECISIONS.md @@ -55,6 +55,25 @@ MEASURED red against the pre-change feeder in exactly those places. Neither change can manufacture a reading from an expired token: re-logging in on `cc` and `cc3` stays Tom's (`DEFERRED.md` DF-U-D12-3). +**The one red is INHERITED, and it is the `nas-topology` assertion — not an +offline-input problem.** `nix flake check --offline --no-build` is rc 1 in this +repository, which is why clause A of `tests/tally-b/test-tally-b-input.sh` is +red and this item's composite oracle exits 1. The tail that reproduces now — +MEASURED 2026-09-17 on this branch and byte-identical on a detached worktree at +`main` (202d9c31) — is `checks.x86_64-linux.nas-topology` → `error: assertion +'(! ((builtins).elem 8731 +(coordinator).networking.firewall.interfaces.wlp192s0.allowedTCPPorts))' failed` +at `flake.nix:1694`: a real eval assertion about a coordinator firewall port, +belonging to whoever opened 8731, and NOT the "an input is missing offline" +`[ENV]` reading an earlier pass of this item gave it. That earlier pass reached +`checks.x86_64-linux.nas-personal-tailnet` → `error: path +'pl6rmijq3dkwqw9cf16wzpycsc7gb9m2-86byf0qaz7f0vgj7x4km4zc9q446skd0-source' is +not valid` first instead; which of the two surfaces first depends on evaluation +order, so a rerun may show either. Either way nothing here introduced it — this +branch edits no `flake.nix` and no firewall or NAS configuration — and nothing +here fences it out of a probe to make an rc green. `DEFERRED.md` `DF-FLAKE-1` +carries both tails and who discharges them. + 2026-09-13 flake checkouts no longer ride into host closures, chrome-stream is installed, and a switch refuses a stale raw-dotfiles checkout. diff --git a/DEFERRED.md b/DEFERRED.md index cdfa0dba3..8f33ca479 100644 --- a/DEFERRED.md +++ b/DEFERRED.md @@ -10,7 +10,7 @@ about. | DF-U-D12-3 | Refreshing the expired `cc3` OAuth token | D-B5 records this as the TOM LINE "run `claude` once on cc3". This feeder may read the credential through `stamp-receipt.py window`; it may not perform an interactive login | Tom runs `claude` once with the cc3 seat configuration. Until then the enabled feeder writes a fresh UNKNOWN row with the reader's reason | | DF-U-D12-4 | Releasing work onto the `codex` row | D-B6 says the login is third-party and the row is never proposed onto. Reading `rate_limits` from its rollout is not a spend and does not change ownership | A later explicit ownership ruling (TL-6); this unit preserves `owner: third-party` in every write | | DF-UTIL-TOK-1 | Counting a Halogen token window (a request with >= 128 completion tokens in the journal) as an evidence event for a box-hour | `/home/tom/research-methods/cards/UTIL-01.md:81-82`: "A metrics delta is NOT an evidence event … only Tom widens it". Since 2026-09-13 (#312) the worker's sampler reads Halogen's per-request `serve_api:` journal line and util-row sums `tokens_in` / `tokens_out` from it, but a token window carries no fingerprint and no receipt, so `receipt_is_evidence()` is deliberately unchanged and box-hours are still graded on receipts and drain-ledger rows alone | Tom, by editing the card's evidence definition. Until then the windows feed `tokens_in` / `tokens_out` / `tokens_grade` only, never a disposition | -| DF-U-D13-4 | `[OTHER-REPO]` Giving the kit's `eval(...)` entries a real mechanical verdict — the CUBS entries and `eval(build:LOCAL-SMOKE)` stay `/bin/sh -c true` no-ops | Serving the lock changes what the KERNEL can derive, not what the CAMPAIGN asks for, and the two must not move in one act: the 58 hand-driven runs of `~/mecattaf/cubs-campaign/receipts/REVIEW-2026-09-17.md` were graded under the no-op entries, so flipping them in the same PR would change the campaign's semantics mid-flight with no ruling behind it. WHAT a mechanical verdict for CUBS is (integration G7: the validation exit plus revert-and-rerun as AC#4's verdict) is Tom's line and is asked, not assumed. The other half is mechanical and belongs elsewhere: an `eval(...)` entry must hand `tools/e2e-evaluator.sh` a whole per-item JSON object on stdin (card, deliverable repo, commit file, usage source, receipt out, node, repo), and rendering that per item is the SECOND step of the uplink render tool that D-A owns — this module's `stdin` cell is a constant string | Tom rules G7; then the kit entry is a reviewed edit to `home/tally-uplink.nix` (D-A's render tool supplying the per-item stdin), and `tally-uplink-topology` sees it. Discharged when a `verdict` on the chain names a CUBS item and the lock this module serves | +| DF-U-D13-4 | `[SCOPE]` Giving the kit's `eval(...)` entries a real mechanical verdict — the CUBS entries and `eval(build:LOCAL-SMOKE)` stay `/bin/sh -c true` no-ops | Serving the lock changes what the KERNEL can derive, not what the CAMPAIGN asks for, and the two must not move in one act: the 58 hand-driven runs of `~/mecattaf/cubs-campaign/receipts/REVIEW-2026-09-17.md` were graded under the no-op entries, so flipping them in the same PR would change the campaign's semantics mid-flight with no ruling behind it. WHAT a mechanical verdict for CUBS is (integration G7: the validation exit plus revert-and-rerun as AC#4's verdict) is Tom's line and is asked, not assumed. The other half is mechanical and belongs elsewhere: an `eval(...)` entry must hand `tools/e2e-evaluator.sh` a whole per-item JSON object on stdin (card, deliverable repo, commit file, usage source, receipt out, node, repo), and rendering that per item is the SECOND step of the uplink render tool that D-A owns — this module's `stdin` cell is a constant string. (Typed `[SCOPE]`, not `[OTHER-REPO]`: only the G7 ruling is outside this repository — the actionable half is a reviewed edit to `home/tally-uplink.nix` right here, in D-A's lane rather than this one) | Tom rules G7; then the kit entry is a reviewed edit to `home/tally-uplink.nix` (D-A's render tool supplying the per-item stdin), and `tally-uplink-topology` sees it. Discharged when a `verdict` on the chain names a CUBS item and the lock this module serves | | DF-U-D14-3 | Setting `services.tally-uplink.plan` — the acceptor's plan body | Null is the honest state, not an oversight. The plan body is the ACCEPTOR's, re-POSTed to arm and re-arm; authoring one in this module would be the lake proposing from the wrong side of the seam, and arming is Tom's act. (The kit half of this row is discharged: `home/tally-uplink.nix` sets `kit` to the reviewed store kit file, and `tally-uplink-topology` asserts it reaches the rendered argv.) | Tom arms the acceptor; then a follow-up edit of this option, reviewed like any other change, which the `tally-uplink-topology` check sees because it asserts `cfg.plan == null` and no `--plan` in the argv today | | DF-U-D18-2 | Setting an anti-starvation number for the filler lane, and implementing D-B10's age-based promotion (a filler item passed over by 10 non-filler releases proposed at level 1 for one release) | Spec §4.4.3 names TL-10 as the anti-starvation number and says that "until it exists the lane runs whenever the row is free". The promotion is a RELEASE-STATION decision over a backlog — the acceptor's `levels` document and the lake's proposals — and a clock has no backlog to age. Putting a promotion rule in a timer would be this unit deciding admission, which is the one thing its own card forbids ("the timer only wakes the uplink's filler pass, the kernel admits") | Tom for TL-10's number; the release station (§2.2c) and the lake's selector for the promotion itself. Nothing in this repository shortens the wait, and the five-minute cadence stands meanwhile: it is the drain's own, so neither filler can crowd the other out | | DF-U-D18-3 | Moving the academic drain onto a LEASE over the kernel's socket, so the two fillers alternate UNDER the kernel's admission rather than beside it | Spec §2.4 lists the drain as an *unleased* tenant of the GPU row "until TL-15 moves the drain onto a lease over the socket (U-D11 wires the call)". Until that happens, D-B10's round-robin is carried here by equal cadence plus the lane's own `/running`-empty gate — two mechanisms outside the kernel — and this unit cannot lease on another unit's behalf. MEASURED 2026-09-07: `tally-drain.service` runs `tally --socket /run/user/1000/tally/tally.sock daemon drain`, the LIVE daemon's socket, not the rewrite kernel's `~/.local/state/tally-rewrite/kernel.sock` | U-D11 wires the call; TL-15 is the line. Discharged when the drain's GPU turn is a lease on the chain and the `tally-filler-topology` cadence equality can be replaced by an admission-side assertion | @@ -28,6 +28,7 @@ about. | DF-136-2 | Turning on the nightly bulk-admission timer (`myNas.paperless.bulk.enable = true`) and writing the measured canary throughput, idle/import RSS and projected completion into DECISIONS.md | The 2026-09-13 flip shipped the gate off by design. A timer that runs unattended OCR on the house router needs a measurement first, and no measurement can exist before the deploy (DECISIONS.md 2026-09-13) | The orchestrator or Tom, after the post-deploy canary run in hosts/nas/paperless.nix steps 3 to 5; one manual `systemctl start paperless-bridge-bulk` run may precede it | | DF-CLIENT-11 | Setting sshd `ClientAliveInterval` / `ClientAliveCountMax`, so the coordinator reaps a dead projector's `remote-client-bridge` in seconds instead of at TCP keepalive | #385 is a client-seat change and deliberately changes no sshd behaviour: a `ClientAlive*` line in modules/common.nix applies to every host's sshd and every live ssh session (the orchestration running inside herdr included). The laptop side is already bounded — herdr's managed ssh config sends ServerAlive 15 s × 4, and `home/dot_local/bin/herdr-projector` exits its own master and removes its sockets — so what lingers is one idle bridge process on the coordinator. MEASURED 2026-09-13: `grep -rn ClientAlive modules hosts` in this tree is empty | Tom, if lingering bridges ever pile up (`pgrep -af remote-client-bridge` on the coordinator); then a reviewed edit to modules/common.nix | | DF-SCREEN-OCR-1 | Putting `screen-ocr` (Mod+Ctrl+S) under a prioritized GPU lease instead of calling Halogen directly | Tom, 2026-09-15: a busy Halogen must not break the flow, and hierarchical prioritization of GPU work belongs to `tally`, which is still being built. Until then the script has no idle gate: it waits behind in-flight requests (curl `--max-time 600`) and fails without touching the clipboard | Tom, once `tally` offers a lease a short interactive request can take; then a reviewed edit to `home/dot_local/bin/screen-ocr` | +| DF-FLAKE-1 | `[ENV]` Making `nix flake check --offline --no-build` rc 0 in this repository. It is rc 1 today, which is what makes clause A of `tests/tally-b/test-tally-b-input.sh` red and this item's composite oracle rc 1 | Inherited, never introduced here: this branch edits no `flake.nix` and no firewall or NAS configuration (`git diff --name-only main...HEAD` names 11 files, none of them `flake.nix`), and BOTH tails are `main`'s. (a) The one reproducing now — MEASURED 2026-09-17 on this branch AND, byte-identical, on a detached worktree at `main` (202d9c31): `checks.x86_64-linux.nas-topology` → `error: assertion '(! ((builtins).elem 8731 (coordinator).networking.firewall.interfaces.wlp192s0.allowedTCPPorts))' failed` at `flake.nix:1694`. That is a real eval assertion about a coordinator firewall port, NOT an offline-input problem — the `[ENV]` tag on this row means inherited from the tree this lane was handed, not that the store is short of a path. (b) The one an earlier pass of this item hit first: `checks.x86_64-linux.nas-personal-tailnet` → `error: path 'pl6rmijq3dkwqw9cf16wzpycsc7gb9m2-86byf0qaz7f0vgj7x4km4zc9q446skd0-source' is not valid`. Which of the two surfaces first depends on evaluation order, so a rerun may show either tail. Neither is fenced out of any probe or oracle here: the red stays visible and clause A is reported red rather than laundered | Whoever owns the coordinator `wlp192s0` line that opened 8731 — or the `nas-topology` assertion itself, if the port is intended — and, for (b), whoever can supply that input. Another lane is already on it (branch `fix/fleet-connectivity-artifact-count`). Discharged when `nix flake check --offline --no-build` is rc 0 in this repository and clause A of `test-tally-b-input.sh` passes on its own | **DF-U-D13-2 is gone (2026-09-17).** It deferred PASSING `--evaluator-lock` to the served kernel, on one condition: `apps/evaluator` existing in the lake to be