Repository navigation
Conversation
DTALEX66
added a commit
that referenced
this pull request
Oct 6, 2026
… a stale exe CI is green at 6c63218 (both workflows, both events; PR #162 head, observer and aggregate SUCCESS), so the "pull_request run still in flight" note is withdrawn. The remaining blocker for the real-desktop-surface proof is now named: u19_webview_e2e.py:446 hardcodes src-tauri/target/release/app.exe, but the documented local recipe sets CARGO_TARGET_DIR, so a local U19 verdict keys on a superseded binary. CI is unaffected because its build uses the default target dir - which is why the defect never showed up there.
DTALEX66
added a commit
that referenced
this pull request
Oct 6, 2026
…snapshot The owner's executable prompt asked for two things in this project: correct the descriptions, and land the full project description and future blueprint with a repeatable audit. Documentation, normal commits, branch push, PR update and the GitHub About are authorized by it; merge, release, install, global config and other projects are not, and none were touched. Input verified: the blueprint DOCX hashes to the value the prompt declares (f3784a99...bea3440). The prompt's own digest is recorded for traceability; it had no declared value to compare. docs/future/WORK-LAB-BLUEPRINT-20261006.md carries all 19 chapters plus the appendix onto repository anchors. Where the source is silent the document says SOURCE_GAP instead of filling the hole: the conditional enhancements in 15.6 have no scenario, acceptance, phase slot or exit condition, and Lite, tray HUD and Quick Entry have no field or permission spec anywhere. .project/governance/blueprint-coverage.json is the single mutable source; verify_blueprint_coverage.py re-extracts AG-01..AG-20 from the atlas card and all 22 U rows from the single open register, so no row can be dropped or invented, and renders the 87-row projection. Status reuses the ledger's existing vocabulary; disposition reuses the prompt's own four decisions. Twelve injected violations were required to fail before the gate was trusted, and the first round exposed two defects in the gate itself: a hand-edited projection passed because nothing compared it against the regeneration, and the "no self-referential digest" rule could never fail since a file cannot contain its own hash — a check that reports protection it cannot provide is worse than none. Both fixed; the digest rule now forbids generated artifacts from hashing each other, with two new injections proving it bites. README corrected: it described the project as a global-configuration layer, which under-defines it against Authority section 3. Now the mother definition, the owns/does-not-own boundary, the evidence vocabulary, the current/future/candidate zoning at Ongoing, the mandatory reading order and the audit entry points are all on the first screen. Corrections recorded rather than rewritten: the blueprint's PR #162 head 6f323a3 has advanced to f3b5dec live, while its main cd4daa8 still matches; the 2026-10-05 local full-gate FAIL stays on the record alongside the affected-group passes, because a group pass is not the aggregate.
DTALEX66
added a commit
that referenced
this pull request
Oct 6, 2026
ERR-109. Adding tests/test_artifact_freshness.py to observer-python-skeleton broke the integration job's negative-control step, because tests/ci/test_failfast_group.py pins each observer group's command count at 8 and I never ran it: I verified the two groups whose contents changed and pushed, treating that as sufficient. The guard is fixed by strengthening it rather than by editing a number. It now pins the full command list of all four groups, so dropping a required check or swapping it for a cheaper one both fail. Verified in both directions: the baseline passes, and a swap that keeps the count at 9 — replace test_artifact_freshness.py with a duplicate of the cheap skeleton test — now reports missing=[[python, tests/test_artifact_freshness.py]], which the old count-only guard could never have seen. Afterwards all 25 CI-invoked governance and integration steps were run locally and judged by exit code. Two traps worth recording: a negative-control suite prints its expected FAIL lines, so reading its verdict from the last line of output invents failures in test_aggregate_gate.py and test_error_ledger.py that never happened; and test_exact_tree_review.py genuinely fails on a feature branch because it asserts HEAD equals origin/main and every task COMPLETED, so CI does not invoke it — recorded so nobody chases it. Also lands the double-end readback the snapshot generator now observes itself: branch pushed to a018fbe, PR #162 OPEN/MERGEABLE at that head, GitHub About before -> target -> live readback matching, homepage and topics deliberately untouched, and the exact-SHA CI state recorded honestly at 16 pass / 2 fail / 4 pending rather than as green.
DTALEX66
added a commit
that referenced
this pull request
Oct 6, 2026
…a foreign SHA PR #162 head is 44b778d; both runs for it completed/success with 24 checks COMPLETED. The chain stays visible rather than smoothed over: a018fbe failed twice (ERR-109, the manifest guard I skipped), 4115d63 onward is green. Tooling trap recorded: gh run list --branch placed an unrelated 2026-10-02 run (2bc9168, not an ancestor of the current head) at the top of the output, which nearly made me report a foreign SHA as this round's result. Pin the head with gh pr view --json headRefOid first, then filter runs by SHA or read statusCheckRollup.
…e only in the legacy suites Sharpening the previous commit's conclusion with an assertion-by-assertion inventory, and correcting my own first read twice along the way (ledger ERR-118). The parity matrix now carries a measured 2026-10-07 section instead of an opinion. Per-suite counts are taken by scanning the suites themselves: test_projection_contract 13, test_read_only_surface 17, test_render_v3 19, test_responsive_contract 11, test_visual_assets_r2 5 — 65 assertions in total — and the production side is described by its test titles (84 titled cases across 14 files) plus the 11 desktop contract cases, with a per-suite verdict: no production equivalent for the projection-truth, read-only-surface and v3-render groups; partial for responsive (only the top clearance and grid ownership are pinned) and for visual assets (theme tokens and the legacy-purple removal are pinned). Method note kept because both of my shortcuts would have produced a false green: - a keyword census said the concepts exist (UNKNOWN appears 36 times in frontend/src). Presence of a word is not an assertion guarding it, so the matrix only accepts evidence from test titles. - my first count of React-side cases came back 0 because a grep bracket expression containing a backtick did not survive argument passing. A wrong count written into a tracked document is exactly the failure this session keeps hitting, so it was recounted in-process and the line patched, with the stale zero verified gone. What this changes about the plan: the browser entry was never the blocker — that is now pinned to the same Vite dist by 8 sidecar tests. The blocker is that retiring apps/observer/web would delete the only home of the read-only and no-fabrication contracts, since those suites read web/index.html, web/styles/*.css and web/scripts/*.js. Deleting the project's truth assertions to make a cleanup pass is the wrong direction, so U03 stays PARTIAL and the matrix records the four ordered steps: re-anchor the read-only and projection contracts to the production surfaces, re-anchor responsive and brand assets, then delete web/ together with the required_groups glob and the pinned fail-fast list (ERR-109), the guarded-path tuple, the README/architecture text and the retirement stub, with a hash manifest plus a full copy of web/ as the retention point and a dated amendment to WORK-LAB-AUTHORITY.md §7. No deletions this round, no rebuild, nothing published; the eight default gates plus the JS suites (83 passed / 0 failed) remain green.
… surfaces (U03 step 1, half) apps/observer/web cannot be retired while it is the only home of the read-only and no-fabrication assertions, so the first porting step lands the file-level half of those guarantees on the production tree (ledger ERR-118; the parity matrix carries the assertion-by-assertion inventory). test_production_surface_static_contract.js — 11 static contracts, wired into run_all_tests.js so the existing observer-web-contracts group runs it with no group-manifest change: no remote runtime/CDN in the entry html; no url(http or @import http in the skins; no write method anywhere in the UI layer (POST/PUT/PATCH/DELETE and any fetch with an options object); no credential, env, bearer or cookie read; theme and layout state never touch web storage; horizontal overflow suppressed by construction; tabular numerals; long CJK/SHA wrapping; prefers-reduced-motion honoured; no rule drops the base text under 12px; and every sub-12px declaration must belong to a named micro role (.winctl-zoom, .load-strip, .brand small, .tag/.badge, .kpi small) so a future 10px paragraph fails even when the base rule still passes. All four of the security- and legibility-bearing assertions were falsified by injecting the violation into the real files — a CDN script tag, a PATCH in a component, a 10px body rule, a localStorage write in lib/a11y.ts — each turning exactly its own assertion red with the file restored byte-for-byte. Two assertions were wrong on my side first, and both were fixed by correcting the test rather than loosening it: paths arrive with the platform separator, so the editor-lane exclusion regex matched nothing and reported the editor's four sanctioned storage uses as violations; and an early version applied the body-size rule to every px declaration, which flagged the B10 brand lockup's letter-spaced 10px caption as a defect. The replacement states the rule the legacy suite actually enforced and names its exceptions. Half of step 1 only: the projection-truth group (13 assertions — fixture numbers, estimate versus bill, strict RFC3339, coverage, forward compatibility) and the v3 render group (19 assertions) are still un-ported, and those are the assertions that would be lost by deleting web/, so the directory stays and U03 stays PARTIAL. The remaining order is recorded in the matrix. run_all_tests.js 94 passed 0 failed (was 83), the ten default gates green, and ERROR_LEDGER_PASS entries=117 counts_consistent=true. No deletions, no rebuild, nothing published.
…cts, snapshot re-recorded Coverage is credited only against a named production test title, never against a keyword census: 5 of the 13 legacy projection assertions have one, 8 are OPEN, and the shared themes of the open ones (fixture numeric agreement, strict RFC3339, the coverage triple, forward compatibility with unknown keys, graceful degradation with missing keys, the 10 required schema keys, not-metered rendering as 订阅未计量 rather than 0) are all pure-function properties portable from lib/api.ts without any real material. Why the legacy fixtures are not simply reused: apps/observer/tests/fixtures/*.json are the v1 shape (usage/summary/quality, no coverage, no tokenSummary, no revision), so porting assertions about them into the SnapshotV3 typed layer would mean inventing production behaviour that does not exist. The map is the work order instead, and apps/observer/web stays until both halves are ported. The audit snapshot was re-recorded from live state: 8 checks, 0 failing, and its live-claims block still reads the U19 register status, the ledger's status_after values, candidate-pool coverage and the observed artifact's receipt drift straight from the files that own them.
… the 8 ported truth contracts The seven date render sites were `value ? new Date(value).toLocaleString() : 'UNKNOWN'`. That leaks the literal "Invalid Date" for an unparseable string and, worse, presents a corrupt timestamp as a plausible one: Date.parse does not fail on 2026-02-30T10:00:00Z, it rolls over to March 2. `isRfc3339` now checks the shape AND the calendar (month, day-in-month with the leap rule, hour/minute/ second, offset), and `fmtTimestamp` renders UNKNOWN otherwise. The producers were read first (collector_scheduler._now_iso and durable_worker both emit `...Z`) so strictness cannot hide real data. The projections also degrade a missing transport / coverage / projects / executions / token object to UNKNOWN/null instead of throwing, because the snapshot crosses a process boundary. That is the U03 step-1 close: all 8 OPEN projection-truth assertions from tests/test_projection_contract.js are re-anchored to the production tree, 8 named behavior cases in frontend/src/lib/projectionTruthContract.test.ts plus 3 file-level cases in test_production_surface_static_contract.js (11 to 14). The file-level half cannot live in vitest: frontend/tsconfig.json has no @types/node, so node:fs is TS2307 and CI runs `npm run typecheck`. Falsified 11/11 by mutation, each file restored byte-for-byte with SHA-256 verified (.project-local/runs/convergence-20261007-c/falsify_u03_ports.py). Two first-version injections failed to hit their case because my expectation was wrong, not the gate (`?? 'LIVE'` and zero-padding tokenTruth cannot reach a case that passes transportState='UNKNOWN' explicitly); and a bare-$ needle flagged api.ts's own legitimate 'http://$1' URL replacement. All three were corrected against the facts rather than by loosening an assertion. Recorded as ERR-119. vitest 15 files / 95 cases green, run_all_tests.js 97 passed 0 failed, tsc --noEmit clean, vite build green, gate battery BATTERY_FAILURES=0, AUDIT_SNAPSHOT checks=8 failing=0. apps/observer/web still is NOT deleted: the per-assertion attribution for test_render_v3.js (19), test_read_only_surface.js (17), test_responsive_contract.js (11) and test_visual_assets_r2.js (5) is still open, so U03 stays PARTIAL.
…ad stops claiming LIVE
Porting the render_v3 contracts surfaced two ways the production surface reported
an unknown as a positive.
ProjectPanel.tsx and the cross-project grid in Views.tsx branched on
`dirty ? 脏 N : 干净`, so a project whose git.dirtyCount is null (never observed)
rendered 干净. That contradicts the repository's own wording in RulesPolicyView
("缺失即 UNKNOWN, 不伪造『全部干净』"). Both sites now branch on `== null` first and
render 脏 UNKNOWN; 干净 requires an observed 0.
useLiveSnapshot only called setError() on a failed fetch, so the compact HUD kept
replaying whatever transportState the retained snapshot was read with - a LIVE
badge on a connection the browser had lost. A failed or throwing read now stops the
LIVE claim and returns to the non-live source label (the projection itself is kept:
no wipe, no fabricated zero), and the HUD word comes from the new
frontTransportState(snap, live, error).
The port itself: 8 mounted cases in
apps/observer/frontend/src/views/renderTruthContract.test.tsx cover the six
properties the 10 OPEN render_v3 assertions collapse to - registry identity,
platform, execution count, branch@sha and dirty count; real token figures with no
invented cache-hit rate (v3 has no cache field); no progress bar without a source;
the overview KPI strip from the snapshot; 覆盖 UNKNOWN rather than 覆盖 0/0 for an
absent triple while a measured 0/0 still reads as a number; no metric card for
fields v3 does not carry; and last-good flipping to OFFLINE when the read fails.
Falsified 8/8 by mutating the production components, each restored byte-for-byte
with SHA-256 verified (falsify_render_ports.py).
Falsification caught the gate being blind, not the tree being wrong: the
no-fabricated-metric needle used `\bCPU\b`, but DOM textContent concatenates the
label with its value into `CPU11`, so an injected CPU card stayed green. Recorded
as ERR-120 together with the two defects.
The parity matrix is corrected and completed: the read-only suite measures 24
assertions (17 t() + 7 asyncTest()), not 17, so the legacy total is 72 not 65;
render_v3 (19), read-only (24), responsive (11) and visual (5) now have a
per-assertion verdict each, with every cited production title, the Rust
`observer_api_is_loopback_get_only` test, index.html:5 and index.css:220 checked
back against the tree. 18 OPEN items remain and are ordered as the step-3 work
order; the first six are defects that can produce a false impression (revision
monotonicity, payload validation, SSE failure staying OFFLINE, cache:'no-store',
?api= trust on the browser path, static-preview freshness). apps/observer/web is
still not deleted; U03 stays PARTIAL.
vitest 16 files / 103 cases green, run_all_tests.js 97/0, tsc --noEmit clean,
gate battery BATTERY_FAILURES=0, OBSERVER_READONLY_PASS, AUDIT_SNAPSHOT 8/0.
…ted endpoint
The read-only transport contracts survived only in the retired static suite
(test_read_only_surface.js, measured at 24 assertions: 17 t() + 7 asyncTest()),
so nothing in the production tree asserted them and no gate could notice the
drift. Six of them were defects that could put a false impression on screen:
- the snapshot GET had no cache directive, so a cached projection could read as
current. It now declares method:'GET' and cache:'no-store'.
- the body was cast straight to SnapshotV3. A legacy/v2 or truncated 200 response
rendered as truth. parseSnapshotPayload now checks schemaVersion, revision,
arrays and the nested objects and fails closed; SSE frames go through it too.
- apply() overwrote unconditionally, so a delayed response with a LOWER revision
replaced newer facts already on screen. The last accepted revision is now
tracked and a lower one is refused without touching the display.
- a browser-supplied ?api= was accepted as authoritative with no validation, so
anything reachable from the webview (external host, https, credentials in the
URL, an added ?write=1, a non-/api/v1 path) became the data source.
isTrustedObserverEndpoint mirrors the Rust gate observer_api_is_loopback_get_only
on the client; an untrusted value is refused outright, never used, never
marked authoritative.
- a payload read through the non-authoritative static-preview endpoint could set
live=true and announce LIVE. live now requires both an authoritative descriptor
and a LIVE verdict in the payload.
- an EventSource error was an empty handler, so the surface kept its LIVE claim
after the stream dropped, and a snapshot frame that failed to parse was
swallowed. A stream error or a corrupt frame now stops the LIVE claim at once;
we never close the source, so the browser's own reconnect still recovers, and
the next good read clears the failure without wiping the projection.
Burst behaviour follows the legacy contract as well: onOpen and resync_required
schedule exactly one re-read, heartbeats cost no read, and concurrent requests
collapse to one in flight plus one follow-up.
Ports: 11 mounted cases in
apps/observer/frontend/src/lib/transportTruthContract.test.ts (real useLiveSnapshot
against a controllable EventSource stand-in) and 2 file-level cases in
test_production_surface_static_contract.js (14 to 16). Falsified 15/15 by mutating
production files with byte-for-byte restore and SHA-256 verification.
Three gate needles were wrong before the code was, and each was fixed by matching
the shape of the violation instead of by loosening: the old clause treated any
fetch with an options object as suspicious, which is the opposite of the explicit
GET the legacy contract required; a bare-text scan flagged the validator's own
schema constant as an embedded snapshot; and the endpoint scan matched
@tauri-apps/api/window (a module path) and a regex replacement ('$1') as URLs. All
three now fail when they match nothing, so an empty scan can not read as a PASS.
Recorded as ERR-121.
Remaining retirement blockers for apps/observer/web are down to five: the brand
SVGs are still only in web/assets/brand, src/theme/tokens.ts declares no read-only
constraint block, the compact surface's dense project list is undecided, the
read-only label wording has no named owner, and the legacy 800/640 reflow
breakpoints have no production analogue. U03 stays PARTIAL; web/ is not deleted.
vitest 17 files / 114 cases green, run_all_tests.js 99/0, node static contracts
16/0, tsc --noEmit clean, vite build green, gate battery BATTERY_FAILURES=0,
OBSERVER_READONLY_PASS, AUDIT_SNAPSHOT 8/0.
…ish the parity attribution
Copies the five brand files out of the tree that is scheduled for retirement and
pins them where the gate can see them:
web/assets/brand/{design-tokens.json,observer-icons.svg,
work-lab-observer-symbol.svg,work-lab-observer-tray.svg,app-icon-512.png}
-> frontend/src/assets/brand/
Each was compared by SHA-256 before and after the copy (5/5 identical). The new
cases assert presence, that each SVG starts with an <svg> root and reaches out to
no remote href/src, and that design-tokens.json still declares themes
dark/light, views full/compact and the machine-readable read-only constraints
(readOnly true, externalMutation false, modelSummary false). Nothing is mounted
this round: the production header is a text lockup, and inserting a graphic is a
visual change that belongs in its own task.
Two more attribution gaps closed. The compact surface now has a named guard that
it keeps exactly the four contracted KPI cells, and the shell's narrow reflow is
asserted by behaviour rather than by the legacy numbers - production collapses
.two-col/.three-col/.split inside @container page (max-width: 760px), and the
case reads the declaration value, so 1fr 1fr fails just like a missing block. The
read-only and no-fabrication wording in Views.tsx and App.tsx now has an owner.
The dense project list from the legacy compact hierarchy is recorded as an
intentional drop, not an oversight: the HUD is a single-column status strip, the
project truth already renders in the projects lane with mounted assertions, and a
second project surface inside the HUD is exactly what the legacy suite's own CC2
grid case argued against.
One legacy assertion is deliberately NOT ported: having the front recompute the
LIVE gate (coverage equality, timestamps, loopback eventsUrl) would make it a
second Update Authority. The client-side job - refuse an untrusted endpoint, never
claim LIVE from a non-authoritative source, and never let an out-of-order or
corrupt read replace good facts - is already pinned by step 3.
Falsified 5/5 against the production files (falsify_step4_node.py), each restored
byte-for-byte with SHA-256 verified: a remote href in a brand SVG,
constraints.readOnly flipped to false, the second-authority sentence removed, a
fifth KPI cell added, and 1fr widened to 1fr 1fr. My first version of the wording
case scanned the whole source blob, which would have counted even the test's own
quotation as a hit; it now reads the two named files.
With this, all 72 assertions in the five legacy suites carry a verdict. What is
left for U03 is the switch itself: deleting web/ with its five JS suites and
helpers.js, the four manifest bindings, the docs and the authority amendment.
node static contracts 16 to 21 green, run_all_tests.js green, tsc --noEmit clean.
…ace, before any deletion 24 files / 242007 bytes under apps/observer/web and 6 suites / 63363 bytes (five JS contract suites plus the helpers.js they read) are listed with per-file SHA-256, re-verified against the working tree after writing. The recovery point is the already-pushed commit named in the manifest, so git checkout 6de25fe -- apps/observer/web apps/observer/tests restores the whole surface; a second copy sits under .project-local/artifacts/observer-web-retired-20261007. No deletion happens in this commit. The parity evidence that makes the retirement lawful is apps/observer/parity-matrix-u03.md: all 72 assertions across the five suites now carry a verdict, with named production owners (31 behavior cases across three vitest files, 21 file-level cases, 8 browser-entry cases).
…e only UI U03 asked for static-web -> React parity, migration of valid capability, then retirement of the static surface as production. Parity is now evidence, not a claim: all 72 assertions across the five legacy suites carry a verdict in apps/observer/parity-matrix-u03.md, credited only by named production titles (31 behavior cases in three vitest files, 21 file-level cases, 8 browser-entry cases), and each port was falsified by injecting the violation. Deleted: apps/observer/web (24 files, 242007 bytes) plus the five suites and the helpers.js that read it (63363 bytes). The audit list and the recovery point exist before the deletion, in the parent commit d6976dc: docs/audits/OBSERVER_WEB_RETIREMENT_MANIFEST_2026-10-07.json records every file with its SHA-256 (re-read and re-verified against the working tree after writing), names the pushed pre-deletion commit 6de25fe, and gives the restore command git checkout 6de25fe -- apps/observer/web apps/observer/tests. A working-tree copy also sits under .project-local/artifacts/ observer-web-retired-20261007/. No new tag is pushed - recoverability comes from already-published history plus the manifest. Bindings and text updated in the same change, because the deletion would otherwise leave four manifests pointing at paths that no longer exist: run_all_tests.js drops the five legacy suites; required_groups.json drops the web/scripts/*.js glob from observer-web-contracts and tests/ci/test_failfast_group.py keeps its pinned command list in sync (the ERR-109 discipline); the observer-no-business-write trigger path becomes apps/observer/frontend/src; the observer_live_server.py stub text, observer-source-architecture.md (with a dated normative amendment, and the run instructions now the sidecar --frontend-root instead of http.server --directory apps/observer/web), the Observer skill reference and WORK-LAB-AUTHORITY.md section 7 all state the switch. Also fixed a pointer I broke myself: the browser-entry suite imported the sidecar module without the sys.path preamble its sibling tests carry, so the very command recorded in ERR-118 and in the register row as its regression check failed with ModuleNotFoundError when run standalone. It now sets its own path and passes 8/8 on three consecutive standalone runs. One thing stays unknown and is recorded rather than smoothed over: on the first standalone run after that fix, 1 of 8 tests errored and the three runs after it were clean; the socket binds port 0 so this was not a port collision, and nothing yet names the cause (ERR-122). Verified after deletion: failfast_group observer-web-contracts commands=1 PASS, node suites green, the pinned failfast list green, CURRENT_STATE_FRESHNESS_PASS, blueprint coverage green, gate battery BATTERY_FAILURES=0, error ledger PASS at 121 entries. Breaking by design (a whole surface is removed), recoverable by manifest and history.
…iewport captures The parity work proved what the surface says; it never proved what it looks like. Reading CSS is not a render, so this round drives the declared playwright-bundled headless Chromium against the live sidecar serving frontend/dist (real v3 snapshot: one project, zero executions, all-null tokens, transport OFFLINE) and captures 26 PNGs, one per lane plus theme/layout/width variants, with no window raised. Machine-checked, no eyes needed: every render produced a plausible PNG (thin=[]), and no two different lanes produced a byte-identical PNG - which is what would catch a route that silently renders the same surface twice. Dark differs from light, full from compact, and 320/820/1440 each differ, so the switches are live. Judged by eye, two commercial-grade shortfalls are recorded and NOT fixed this round: at 320px the top bar stacks hamburger, search and the four action buttons into one row each and eats about 40% of the first screen (navigation itself still works through the drawer, and the action-cluster contract still holds), and a lane heading plus its meta line both wrap in a narrow column. Honesty held up screen by screen: Token and coverage read UNKNOWN rather than 0 or 0/0, the trend panel says there is no series, the empty list says registry has no active execution, and revision 0 is a real value. Also recorded because it cost a debugging cycle: --dump-dom with --virtual-time-budget hangs against this app because the page holds an SSE connection; screenshot-only mode renders and exits. Status correction inside the register: frontend/dist has been rebuilt since the first round (vite build green); the release binary is still deliberately not rebuilt, and the migrated brand SVGs are still not mounted in the header.
…ow 841px Measured with a dependency-free CDP layout probe (Node 22 global WebSocket + Emulation.setDeviceMetricsOverride), not inferred from a screenshot: at a 320px viewport .app computed '210px 110px' and .main was 210px wide, with a 263px top bar whose four action buttons each took their own row. Same 210px at 560 and 840. What looked like a top-bar wrapping bug was the whole application column being squeezed into the rail's grid track. Cause: a display:none grid child is not a grid item, so hiding the rail below 841px auto-placed main.main into the FIRST track and left the 1fr column empty. B10 does collapse .app to one track at <=840px, but the shell declares the two-track clamp(210px,17vw,280px) minmax(0,1fr) form unconditionally and later in the cascade, so it wins everywhere. That unconditional form was deliberate when written - it made the content track shrinkable - and nobody had measured the sub-841px case. The two-track form now lives inside @media (min-width: 841px), the default is a single shrinkable track, and a <=560px tier lets the action cluster take a full row and the search box shrink to what is left. Re-measured: .app is one track equal to the viewport at 320/560/840, .main equals the viewport, the four buttons share one row (320px: x=14/56/110/152, top bar 263px to ~104px), and the 1440px desktop layout is unchanged at 244.8px + 1195.2px. Locked by a new static contract case, 'the shell keeps a single app track below the rail breakpoint', falsified two ways: restoring the unconditional two-track declaration, and moving a clamp() declaration outside its media gate. My first version of the lock counted clamp() declarations instead of checking that each sits inside the gate - the shell legitimately declares it twice - and that version failed against the fixed CSS, which is how I found it. Two self-corrections recorded in the audit: the earlier round attributed this to the top bar itself, and an closest() probe that omitted .top-actions from its selector list misattributed the buttons. A probe's selector set decides what you can see. docs/audits/OBSERVER_UI_RENDER_AUDIT_2026-10-07.md gains the measured before/after table and the probe's reproduction command. ERR-123. vitest 17 files / 114 cases, node contracts 33 (22 static + 11 desktop), gate battery BATTERY_FAILURES=0, vite build green; the audit sidecar and probe profiles were stopped and removed.
…beneath it
Reading rendered screens catches contradictions that reading source does not. Two came out of the headless captures:
1. The overview status stack derived its alert chip from alerts.length, and alerts is built by iterating snap?.ci and the execution rows. With no data source both are empty, so the chip painted a green 无告警信号 while the panel directly beneath it said 数据源未接入 - 无法判断告警(保持 UNKNOWN,不伪造「全部正常」). Counting an absent collection is not measuring zero - the same confusion as dirtyCount in ERR-120, one layer up. The chip is now three-state: no snapshot -> 告警状态 UNKNOWN (muted), alerts -> N 条告警信号 (warning), snapshot with none -> 无告警信号 (success).
2. The permanently-disabled 导出状态 / 新建执行 buttons looked exactly like live ones: b10.css ships no :disabled rule and keeps button{cursor:pointer} plus a hover lift, and D-11 pins b10 verbatim, so the shell now carries the affordance (not-allowed cursor, dimming, no lift).
Locked by a vitest case asserting all three chip states and a static contract case asserting the disabled affordance shape. Falsified 3/3: two-state chip restored, cursor:not-allowed deleted, hover override deleted - each reds its case, each file restored byte-for-byte with SHA-256 verified (falsify_ui_round4.py).
Two process notes. Appending to a CRLF working copy with a shell heredoc produced mixed line endings, and two of my own falsification needles then reported NEEDLE-ERROR until the file was normalised - the needle was right, the file shape was not. And the first wording of ERR-124's repeat_prevention was rejected by the ledger gate as 'not enforceable' because it read as description rather than rule; the gate caught me, which is the point of it. ERR-124.
vitest 17 files / 115 cases, node contracts 34 (23 static + 11 desktop), tsc clean, vite build green, battery BATTERY_FAILURES=0. CI on the retirement commit 891c871 is green across both workflows, so deleting web/ and rebinding the four manifests did not break the pipeline.
…sion U08 was still marked PARTIAL because its Windows Tauri real-WebView layer was deferred to U19. Both of the things U08 actually owns are done and CI-gated: the observer-frontend-typecheck group runs npm ci -> typecheck -> build -> test, where vitest is now 17 files / 115 named assertions (this session added the projection-truth 8, render-truth 9 and transport-truth 11), and the observer-desktop-crate group runs cargo test --locked plus cargo check --locked. The layer it handed over has landed too - U19 is PASS with an owner-observed desktop surface and a CI real-WebView E2E readback step. The genuinely open piece, rebuilding and publishing the release binary, is now recorded where it belongs: on U19 and the release work order, not on U08. SESSION-HANDOFF-CONVERGENCE-20261007-C.md records the structural change for the next session: apps/observer/web is retired (891c871, CI green) with an audited per-file hash manifest and a recovery commit; all 72 legacy assertions carry a verdict; six classes of production truth defect fixed (ERR-119..124); the CDP layout probe and headless render pipeline that turned 'looks wrong' into measured numbers; the next-round priorities (U02 remainder, manifest-first greening with its do-not-touch list, remaining UI items, release chain); and the traps hit - --dump-dom --virtual-time-budget hangs on the app SSE connection, the in-app browser reports a 0x0 viewport so it cannot measure layout, Git Bash has no node/python on PATH and Windows PYTHONPATH needs ';', heredoc appends create mixed line endings, and the ledger gate rejects repeat_prevention written as description rather than rule. docs/future/WORK-LAB-BLUEPRINT-COVERAGE.md regenerated after the register edit, as the coverage gate instructed; it now verifies PASS at 87 rows.
…ite never covered U02 listed the .hermes/ fallback in hermes-project-data.py as a leftover to remove. Reading it first says otherwise: the docstring states it as a deliberate cross-project compatibility contract (WL-010/020/030) - projects whose .gitignore lacks .project-local keep working, and the guard fails closed when neither root is ignored. Deleting it would break other projects and cross the owner's do-not-write-outside-this-project boundary, so the item is adjudicated as audited-and-kept rather than fixed. The real risk was never that the fallback exists, but that it could win inside this repo. Measured: .gitignore ignores both .hermes/ (line 2) and .project-local/ (line 42), and candidate order prefers .project-local, so behaviour is correct - but nothing pinned that order. Every case in test_project_data_boundary.py builds a fixture that ignores only .hermes/, i.e. the entire suite exercises the fallback path and none of it exercises the preference. Adds test_project_local_is_preferred_when_both_runtime_roots_are_ignored: with both roots ignored, tmp must land under .project-local/runs and .hermes must not appear in the path. Falsified by swapping the candidate order - the case goes red; the script is restored byte-for-byte with SHA-256 verified. 17 tests OK. U02 now has one named remainder: publishing the corrected managed skill into Hermes Home via sync_hermes_workflow_assets.py --apply, a global change that rewrites live agent guidance and needs its own staged execution.
…actually check AG-20 asked for source-registry increments - path, time, hash, original location, coverage relation. The registry those increments were appended to is the atlas one, under .project-local: ERR-114 says in its own remaining_boundary that it is machine-local, that nothing there is enforced by CI, and that a rebuilt atlas reverts the pins to BLOCKED_NOT_VISIBLE. So the missing piece was not more rows in an unversioned file but a tracked mirror with a gate. .project/governance/recovered-source-registry.json carries 12 entries in five honest statuses: machine-local recovered copies (the recovered WORK-LAB-SUMMARY, the atlas registry itself, the scan manifest), absent (the 15,558,839 B / 432,344-line timeline, with its negative proof kept as text and no digest claim), unpinned (the two 2026-09-28 items that have a name but no expected hash, so no claim is made in either direction), tracked (the web retirement manifest and the five migrated brand assets, each recording apps/observer/web/assets/brand as its original location and 6de25fe as the recovery point). Every digest is computed at write time by generate_source_registry.py and re-read from the written bytes; nothing is transcribed. tests/workflow-assistance/test_recovered_source_registry.py (7 cases) is discovered dynamically by run_quality_gate.py governance - it is in the 182-file set, so CI runs it with no manifest edit. It re-hashes every tracked entry, refuses a versioned file labelled machine-local, refuses a hash claim on an absent or unpinned original, and fails if the timeline row disappears. Falsified 5/5 with the registry restored byte-for-byte and SHA-256 verified. AG-19 stays PARTIAL: the timeline is unrecoverable from here and history_complete is still false - what changed is that the verifiable half is now pinned by a gate instead of by a prose claim.
…aller gone, three cited trees kept The four remaining cleanup candidates were measured and citation-checked rather than deleted by size. p0c-oracle-20261006 (1,661 files / 158,786,208 B), topbar-measure-20261006 (1,615 / 151,633,914 B) and ui-suite (344 / 73,450,325 B) are all cited by tracked files - the error ledger for the first two, five UI manifests and the visual QA report for the third - so the citation guard refused them and they stay. Their per-file listings were still captured so the next round can compare without re-walking. dsh-reconfig-20260904 had zero citations, but deleting the whole directory would have thrown the run's own evidence away with the bulk. Splitting it: 134,265,015 B of its 141 MB was a single DSH-Desktop-2.0.5-x64-Setup.exe, a community build AGENTS.md already marks superseded by the official DeepSeek Harness 0.2.0-rc.2, and uncited anywhere in the repository. Only that file was removed - digest 777cb50c86b0... recorded first, manifest written and re-read, deletion confirmed by the file no longer existing - while the logs, JSON and python-test-env stay in place so any path that does resolve still does. A spill-ledger line with all four verbs (trace/locate/clean/migrate) was appended to .project-local/artifacts/spill-ledger.jsonl, and PROJECT_DATA_BOUNDARY_PASS. The lesson recorded in the register: greening is not 'delete the big ones', it is 'only touch what nothing cites, measure and record before touching, and leave evidence and cited paths where they are'. The three retained trees need their citations migrated before they can go, which is the next round's task, not a reason to bypass the guard.
…nnot disagree Work-order item 14 was the last open row in the U03 parity list. The brand JSON already carried constraints/readOnly/externalMutation/modelSummary after the migration, but the code that the UI actually imports declared nothing - two places writing a rule down is not a single source, so the fix is a cross-check, not a copy. theme/tokens.ts now exports VIEW_CONSTRAINTS, and a new static-contract case compares it value by value against assets/brand/design-tokens.json and pins the law itself (readOnly true, the other two false). The case fails on a disagreement in either direction, so neither file can drift alone. Node static contracts 23 to 24, tsc clean, vitest unaffected. Falsified with three injections: externalMutation flipped to true in the JSON, the VIEW_CONSTRAINTS export deleted from tokens.ts, and readOnly flipped to false in tokens.ts only. The third initially reported SURVIVED and that was my expectation string being wrong, not the gate: it trips the disagreement assertion first and prints 'readOnly disagrees: tokens.ts=false design-tokens.json=true'. Re-checked against the real message; the file was restored byte-for-byte each time. The register row and the parity matrix now record the closure of all 18 work-order items, and two things that are NOT part of the old list: the migrated brand SVGs are still not mounted in the header (a visual change needing its own review), and the CDP layout probe measures but is not wired into CI, so narrow-viewport defects are still caught by CSS-shape contracts rather than by real box metrics.
The recovered-source registry recorded sha256 over working-tree bytes. With `* text=auto` the same commit is CRLF here and LF on the runner, so three of six tracked entries could only ever match on this machine: CI failed at 1cffd9f/e7ede1f/ab33537 while the identical 45-command step list passed 45/45 locally. Tracked entries now record the blob at HEAD, and a new test runs each recorded recovery command so `byte-identical` becomes a computed relation rather than an adjective.
21 rows said DEFERRED and the gate only checked their shape, so three rows contradicted the code (OpenHands already an honest fleet adapter, Hindsight already a POC on the nine-operation memory contract, n8n already barred as a fourth task core) and two rows carried TBD triggers. Every row now carries a GitHub API discovery readback, an absorption level, native_evidence paths that must exist and a decision: 15 retired with reversal conditions, 7 kept with blockers named, 0 promoted because no candidate has pilot evidence and a status word is not an experiment. The gate refuses a licence the readback did not return, an absent evidence path, a PILOT row with no code and placeholder text; falsified 10/10, and the committed verifier reproduced as 0/10 (ERR-126).
Nine models carried 64-hex sha256 values and health words reading DOWNLOADED_HASH_VERIFIED while nothing recorded what those digests were digests of or when anyone recomputed them, and the integrity gate passed on shape. This re-hashes all 48,531,408,516 B: nine match, and the directory-backed zipformer entry gets twelve per-file digests instead of one truncated prefix. The two pending-retirement weights turn out to be present (5.2 GB and 18.6 GB), so the open decision now carries its cost; the reranker leftover is the same length as the registered weight but a different digest, so deleting it as a duplicate is not justified; the ollama-partials row was written as prose and is marked unresolvable rather than absent; and a 5.97 GB weight layer belonging to no entry is registered and left in another runtime's store. ERR-127, gate rules falsified 14/14 and 0/14 against the committed verifier.
…stale URL The AG-05g falsifications first lived in git-ignored scratch, where CI would never have re-run them. They are now thirteen named fixture cases inside nf22_registry_and_acp_honesty_gates.py (49 tests OK), so the governance group guards the rules permanently. The retired-model cost check moved outside the early return so a model with no digest claim still has to state its bytes, and 0-when-released is accepted as its own honest answer. Separately, source-ledger's agent-skills row pointed at https://agent-skills.org/, which resolves to nothing; the readback names agentskills/agentskills at Apache-2.0 and the old value is kept inside identityReadback rather than quietly replaced.
The integration job went red 17 seconds after the previous push: the agent-skills row had gained an identityReadback object, verify_source_ledger_v4.py let it through and the CI-invoked verify_source_ledger.py validates a schema that closes its property set. Third occurrence of one shape this session (ERR-109, ERR-125, now ERR-128): validating a slice of the CI command set and calling it CI. The provenance moved into integrationNote, where the contract allows it, keeping the previous URL and the readback. scripts/ci/reproduce_ci_commands.py now parses both workflow files, groups each run block the way a shell does, honours working-directory, and names what it cannot reproduce (CI_CONTEXT, STEP_ENV, TOOL_NOT_ON_PATH) instead of faking or hiding it — 107 interpreter commands, including the 12 integration-job commands my earlier harness never ran.
…leased with zero cited paths lost, and the three obligations that remain
`command` says what ran then; `lifecycle.regressionCommand` promises what can be run now. Nothing in the ledger distinguished them, so 17 of 43 promises pointed into the git-ignored runtime root the greening rounds delete, and the two that looked broken (`cd apps/observer && node tests/run_all_tests.js`) turned out to be fine once the scan resolved them against the cd prefix they carry — my first scanner also inflated the absent count to 49 by matching `.js` inside `.json` and treating prose as commands. The 17 real cases were promoted byte-for-byte into scripts/audit and scripts/maintenance with digests recorded, the entries keep their original string in regressionCommandPrior plus a dated correction, and a new CI-discovered module (183 modules, no manifest edit) refuses an untracked promise, a .project-local promise, a dropped promise and an undated correction. Falsified 4/4 against the live ledger; no exemptions needed.
…s, and a gate that re-derives them AG-11 had been NOT_RUN since the atlas gap was opened: no inference path was ever bound, so no OCR/ASR claim could be made. LM Studio serves qwen2.5-vl-7b-instruct at localhost:1234, so the matrix could finally be run against a real model instead of asserted. Measured 2026-10-07 over nine declared cells: probe recovery is 11/11 on the mixed zh/en page, 11/11 scanned-degraded, 4/4 zh-prompt, 4/4 en-prompt, 8/8 pdf-page, 15/15 table, 3/3 page-order; ASR on the fixture reaches CER 0.0 at RTF 1.23. The vendor wav set is recorded as MEASURED_NO_TRUTH, not a pass — it ran, and nothing in it matched a declared probe, which is the honest state rather than a zero scored as 100%. Evidence: docs/audits/AG11_OCR_ASR_MATRIX_2026-10-07.json (fixture identities hashed, package versions and served models named, per-cell hits/probes/rates). Instrument: scripts/audit/ag11_ocr_asr_matrix.py — cells are declared up front, so a cell that cannot be measured is reported as not_run rather than quietly dropped. Gate: tests/workflow-assistance/test_ag11_matrix_evidence.py re-derives every rate from hits/probes, refuses a score on a MEASURED_NO_TRUTH cell, requires RTF to follow its own seconds, and rejects a summary that disagrees with its cells. 10 tests, 8 negative controls. The register row keeps its original NOT_RUN sentence and records the closure after it. Real-user audio and multi-page PDF OCR stay UNKNOWN: no such material exists in the repo.
…ontract breach closed The head carries the config-compiler conformance fix, the governance vacuity guard and the authority-index reference gate, so a verdict there covers those records rather than an ancestor.
…s currency They sat under docs/current and two live documents called them "本轮", so a reader at the front door was pointed at some past round's record as current procedure; one of those links had not resolved since the 2026-09-17 convergence. The reference gate now judges four surfaces and distinguishes a path from a backticked term, and the frozen files keep their pre-convergence paths on purpose -- ten of them -- because correcting a dated record to today's names would make it describe a moment that never happened.
… currency Measured per file: ten pre-convergence paths remain inside the frozen records and are deliberately neither rewritten nor guarded, which is written down instead of left implicit.
… carrying both That head is the first one to contain the register-shape gate and the widened reference checker alongside the records they were measured against.
…s them, and gate it Containment by a ref is the test, not whether the object parses: three of these hashes return HTTP 422 from GitHub and exist only on this machine, yet every existence check called them resolvable. 34 other pins sit on the open PR branch and are legitimate today, which is why an earlier "not an ancestor of main" count of 40 would have failed a healthy register.
…uld have used The earlier 13-tokens/24-rows figure came from judging reachability as "ancestor of origin/main", which also swept up 34 pins that sit on this open branch and are perfectly checkable today.
The watcher decided "changed" from row counts, status tallies and aggregates, and an acquire_lease moves none of them: measured, writing holder/token/checkpoint onto a live task left the canonical fingerprint byte-identical. The witness is now content over the mutable table plus a PRAGMA data_version fast path, and acquire/heartbeat/release finally stamp updated_at, which they silently did not before.
The first version of this fix was caught by its own gate: acquire_lease did not stamp updated_at either, so the witness stayed put until it did.
…ck guards A watcher that reads only when another connection commits cannot prove its own readback works, cannot notice a publish failure to retry, and never enters the read whose blocking shutdown the guards in test_sidecar_v3_snapshot.py hold. The content was never measured as a cost worth trading those for; the witness and the updated_at stamps stay, the loop goes back to reading every tick.
…the retraction The ledger entry names my own regression: gating the watcher's content read behind PRAGMA data_version made a watcher that never reads unable to fail, so the shutdown guard could not see it enter the read, a raising readback could not push it to STALE, and a failed publish was never retried because the cursor advanced first. ERR-199's remaining_boundary gets a dated correction rather than a silent rewrite -- the two halves that still hold (newest_changes as the witness, updated_at stamped by acquire/heartbeat/release) stay asserted, the fast path does not. The two republished audits that carry ERR-200's binding are committed; the citation and tool-inventory regenerations were timestamp-only and are discarded, not smuggled into a record commit.
…ale from my own pin rewrite The integration job has been red at d76deef and cdfd27d with BLUEPRINT_COVERAGE_FAIL while the canonical local gate printed PASS on the same tree: verify_blueprint_coverage.py ran in no local gate at all. The projection restates register cells verbatim, so my 12 dangling-pin corrections made it stale, and the only reader that could see that was CI -- a freshness rule a writer cannot run before pushing is a rule that honest edits will break. Regenerated with --write (rows=87, anchors_resolve=true) and wired into VERIFY_ORDER as blueprint-projection. Falsified rather than assumed: the new gate passes on the regenerated projection, and editing one word of a U19 cell makes it exit 1.
…e caught it on my own run The governance batch went red at 951dbcc in exactly the way this branch is supposed to fail: test_quality_gate_runner_is_canonical_and_just_is_optional pins VERIFY_ORDER as a tuple AND retypes a long prefix of the runner's `verify: Run ...` line, so adding blueprint-projection had to be written in two places and the second one was missed. The tuple stays a hand-written pin -- that is the point of it, a gate can only join by someone accepting the order. The second assertion now derives from VERIFY_ORDER instead of copying it, which pins more than before (the whole line, not a truncated prefix) and stops disagreeing with itself on every addition.
…current pages went unjudged Widened to every tracked markdown under docs/current (38 surfaces, 340 references, broken=0) and fixed what that found. Sixty references did not resolve: eleven pointed at scripts a reader would run and watch fail (python scripts/workflow/sync_codex_global_assets.py and friends, dead since the 2026-09 convergence, invisible to the gate because they sit inside fenced blocks), sixteen had lost a root prefix, thirteen named a client home or an installed app in bare form, twenty-two were a dated page writing today's tree, and three named things that never existed in any ref. Client-home references now use the notation this repo already standardises ($HERMES_HOME/, ~/, <dshInstallRoot>/, <Cognitive-Loop-OS>/) instead of growing a global allowlist, because the string-keyed list hid a stale README entry once already; DECLARED_NON_PATHS is still exactly one entry and a test says so. The DSH adapter page gained a dated SUPERSEDED banner: its whole identity is the community build AGENTS.md retired, and five dead pointers were the symptom, not the disease. Two rule changes rather than two excuses: a fenced-map entry is judged as a path only in file or directory shape, which is what stopped origin/main and /interrupt reading as missing files, and the run refuses to pass below a measured reference floor so a collapsed extractor cannot claim the cleanest tree yet. Verified the loosening excused nothing: no originally-broken literal survives as a bare backticked path. Also: apps/token-monitor/src-tauri/target/ was not ignored by anything (the handoff asserted it was, and the only /target/ rule in the repo is observer-scoped), so the build output is now ignored and the sentence says what is true. The live-command audit's own floor was pinned to its debt count, so fixing one reference made the guard red for doing its job; the floor is now a not-blind floor, with the measured debt noted beside it.
…out to be false Two new rows. DOC-REFERENCE-WIDENING-20261008 closes the widened scan with its measured before and after. REGISTER-CELL-TRUTH-20261008 registers what a full literal-by-literal audit of the register found and does not quietly rewrite: 21 observer paths missing their apps/observer root, and two claims that are not a path problem at all -- AG-20 credits a generator named generate_source_registry.py that appears in no added or deleted path in any ref, and U03-PARITY-20261007 says seven observer JS consumers are still to be retired while the very commit the same cell cites deleted all seven. Those two need the owner's intent, not my guess. The tool-inventory regeneration was timestamp-only and is not smuggled into this commit.
…id now have a local route Measured: CI invokes 64 python operands across two workflows; 11 of them were reachable from no local route at all -- eight scripts/ci verifiers, a generator whose committed output CI regenerates, and the WebView2 E2E. The blueprint projection was this same class two commits ago: red in CI, PASS locally, and the writer had no way to know because the rule they broke was a rule they could not run. verify_ci_check_reachability.py decides reachability from repository text (named gate, a test module that names it, or the module being a discovered test itself) and refuses a run whose declared exemptions have quietly become reachable -- which is how I found that three of my own draft exemptions were wrong, and why the list is now two entries with measured reasons rather than five with guesses. Steps carrying working-directory: are resolved before judging, so apps/observer/scripts/ u19_webview_e2e.py is not mistaken for a root-level path. tests/ci/test_ci_invoked_verifiers_have_a_local_route.py then runs the eight verifiers here and checks the committed contracts.ts against a regeneration whose bytes are restored afterwards. All five tests pass; the reachability run goes green by being satisfied, not by being declared away.
…ed itself before shipping The REGISTER-CELL-TRUTH row's first draft said "seven JS consumers deleted" and "STALE_PREFIX 21". Per literal, the measured numbers are six deleted (run_all_tests.js was modified by that commit and is still tracked) and 22 literals across 11 lines, and AG-20 already carries a dated correction saying its generator script does not exist -- so the remaining fact there is that no script produces the registry, which is an owner decision about a producer, not an unmarked lie. The row says what the commands showed and records that its own opening claim was rewritten after verification.
…state freshness mode is locally reachable Two reds that were mine, both found by running things rather than by reading my own notes. The witness rolled up count, fencing token, checkpoint length, holder length and MAX(updated_at), so the only column a heartbeat necessarily moves -- lease_expires_at -- was not witnessed at all. My renewal test passed by accident: it relied on updated_at, and on this host _now() returned a single value across 2000 back-to-back calls (~1 ms granularity), so the same statement pair can write byte-identical rows. Red at 9fb20f2 in the full batch, green standalone. The expiry is now part of tasks_state, and the test freezes the store clock before the lease is taken so the renewal can change nothing else -- which is how the blindness shows rather than being asserted away. CURRENT_STATE_FRESHNESS_FAIL source-digest-mismatch was red in CI at 6e6d4fb and red locally, and the canonical gate said PASS: the CI step runs generate_current_state.py --check-current, while the local route for that script is a test module that imports it and never runs that mode. A route naming a file without its mode is not a route. The mode now runs in the local batch, the projection is regenerated, and the failure names the moved sources via git history instead of printing a hash pair -- measured: it pointed straight at config/skill-provenance.yaml and the skill I edited today.
Both are my own defects from this round, recorded with the command that showed them: the witness that never looked at lease_expires_at behind a test that only passed because the host clock was coarse (~0.95 ms, two distinct values in 2000 calls), and the freshness mode that CI ran and no local route executed -- which my own new reachability rule counted as covered, because it judged the file rather than the flag. ERR-204's phase was first written as CI_VERIFICATION and the ledger verifier refused it; the enum is LOCAL_VERIFICATION, EVIDENCE_CLOSEOUT, LOCAL_IMPLEMENTATION, AUDIT_ONLY, and the check caught my invention rather than the record standing.
… so CI's verdict at one SHA was luck run_forever writes the collector rows inside run_once() and only stamps worker_loop after that tick returns, which makes a health table holding just the collector a legal intermediate state. The test polled until the table was non-empty and then asserted worker_loop was in it -- so the outcome depended on where the poll landed in that window. Measured: the push-triggered work-lab-gate run at 048e616 (tree bd3a678) was success while the pull-request-triggered run at the identical SHA and tree failed exactly there ("'worker_loop' not found in {'healthy_collector'}"). I had already stamped eight verifiedCommit fields from the green run; they are withdrawn, because one red run at a head means the head is not verified. The wait is now on the set the sidecar itself declares as expected, with the deadline raised to 5s and the failure naming which collectors stayed silent. Falsified rather than assumed: suppressing the worker_loop write makes the test fail in 5.3s and name the gap; the module is green 5 runs in a row and 41 tests pass across the four sidecar/worker modules.
…two anchors pointed at the wrong line A literal-by-literal audit of the register's path claims (686 distinct literals, resolution against git ls-files only) turned up assertions that are simply false today: - AG-06o calls references/frontend-baseline-contract.md "confirmed absent from the repository source". It is tracked, at packages/client-neutral-core/skills/software-development/windows-development-environment/ references/, and the very skill whose deployment the row audits cites it at SKILL.md:130-131. - AG-14 records b10.css at 21,002 B and l10b-shell.css at 2,907 B "verified against main@cd4daa83"; git cat-file -s at that commit says 20,469 B and 2,832 B, and l10b-shell.css is 22,012 B today. - U03-PARITY-20261007 asserts the observer JS suites are "退役仍未做" while the commit the same cell cites, 891c871, deleted all six of them (run_all_tests.js was modified and survives, and its own header comment records the removal). Two anchors in that row are wrong too: required_groups.json has no web/scripts/*.js glob anywhere and its :22 group runs node tests/run_all_tests.js, and the observer-no-business-write watch list lives at run_quality_gate.py:1381 and protects frontend/src, not apps/observer/web/. The dated corrections keep the original judgement as evidence instead of erasing it. 22 root-prefixed path literals across 11 rows were rewritten, each replacement re-checked with git ls-files --error-unmatch (20/20 tracked), and the table shape (170 rows, wrong_width=0) plus pin reachability (67 pins) still pass. AG-20's claim that a script named generate_source_registry.py generated a tracked registry stands as an owner ask: no script in the repo produces that file.
…uced them ERR-205 keeps the flake honest: one trigger green and one red at the same SHA and tree, and the verifiedCommit stamps I took from the green run are recorded as withdrawn rather than quietly re-taken. ERR-206 names the four git commands behind the register corrections and states plainly that no machine guard yet covers the register's factual claims -- applying the docs rule to it yields 61 unresolved out of 321, and nearly all of those are legitimate client-home, build-output, cross-project or deleted-file mentions, so gating it needs the per-row exemption design, which is the next open step rather than a claim of closure.
…p-alive connection, in process ERR-187 recorded that this test could not be written without a subprocess boot because the handler class is nested inside the serving function. That was never checked and it is false: create_server is module level and binds port 0, so the test speaks raw bytes to it in the same process. The boundary stayed open because the excuse for not writing it was never measured. One connection, two requests: a POST that must be refused with 405 and must not echo its own body, then a GET that must begin with a clean status line -- which is the symptom a non-draining refusal actually produces, invisible to any client that only parses statuses. The control case serves the sloppiness on purpose: a refusal that never reads its body makes the second response something other than HTTP/1.1 200 OK, so the assertion above is not green because nothing can fail it. Three consecutive runs stable; ERR-187's boundary now carries the dated closure instead of the false obstacle.
…is repo's convention and my omission The repoint tool rewrote 330 lines of the ledger and the diff looked alarming until it was compared semantically: records head=204 working=204, semantic changes in pre-existing records=0, added=[] and removed=[]. Every one of those lines was \uXXXX escapes becoming readable CJK. The escape form was mine -- my append scripts used json.dumps' default while the ten ledger writers already in scripts/audit all pass ensure_ascii=False -- so this commit restores the house convention rather than introducing one, and the next append must keep it or the same phantom diff returns.
… for everything, now pinned Measured on this checkout: `git check-ignore -v --no-index services/` exits 0 and names `.gitignore:44`, which is an empty line, and it does the same for docs/current/ and for a directory that does not exist. With a file inside the directory the answer is honest -- services/orchestration/sidecar.py exits 1 while .project-local/runs/x.json exits 0. So "is this path ignored?" is only a question you may ask about a file, and a claim built on a trailing-slash query is not evidence. The guard pins the trap, asserts the file-shaped probe still discriminates (including the Cargo target rule added today and apps/observer/web, which is deleted rather than ignored), and scans tracked call sites for directory-shaped operands. The first scanner used a regex over a text window starting at the word check-ignore -- which in every real call site is mid-list, so the quotes paired off by one and it reported the wrong literals; the planted case now proves the ast-based reader catches build/ and ignores a runtime concatenation it honestly cannot judge.
…ster guard designed with numbers Every Actions run at 970fef4 -- both triggers of work-lab-gate plus the production gates -- completed success, so ten records finally carry a verifiedCommit that means what it says. 970fef4 contains all ten records, and the ancestry and ledger-present checks are the stamper's own, not my reading. ERR-207 records the check-ignore probe trap and the off-by-one in the first version of my own guard for it. Three register rows: the sidecar 405 byte-level proof, the probe-shape rule, and REGISTER-PATH-GUARD-20261008 carrying the measured 61-unresolved breakdown and the design decision (declarations live inside the row they excuse, because a code-side map goes stale before the prose does and hides the exemption from the reader). That row also states a gate hole I found while designing it: the refs>0 requirement and the reference floor only apply to the default widened scan, so a --index run can still pass on an extractor that matched nothing -- the register is bootstrap step 6 and belongs in the navigation set.
…ich is what a broken extractor looks like The capability floor I added this morning only applied to the default widened scan, so `--index <one prose page>` printed refs=0 broken=0 and exited 0. A page legitimately making no tree claim stays fine inside the widened scan, where the other surfaces carry the floor; a whole run that examined nothing is not a pass. Found while designing the register's own guard, which is exactly the kind of hole that design work is for.
The ledger requires a non-zero exit_code, and the vacuity I am recording is precisely a run that exited 0 after judging nothing. The record now says which number is which: the defect's code was 0, the 1 on file is the refusal measured after the fix on the same command. The REGISTER-PATH-GUARD row is closed halfway the same way -- the --index hole it named is fixed, the in-row declarations, the nine genuine corrections, the 280 floor and pulling the register into the navigation set remain unimplemented and stay written as such.
…he head is green Three runs exist at that SHA -- work-lab-gate on push and on pull_request, plus the production gates -- and all three concluded success, which is the condition ERR-205 itself set after one trigger lied about the head last time. Measured remaining binding debt, stated rather than implied by the stamping: 185 PASS records, 94 with no fixedCommit and 91 with no verifiedCommit, and the triage reports bindableByBirth=0 for that population. Those are pre-lifecycle records, so closing them is an owner/evidence decision, not a command I can run.
…g its own dead names in the row Widening the reference gate to docs/current stopped one short of the surface the authority chain actually tells a reader to navigate first: the live open-task register sat outside the scanned root with 65 unresolved references out of 330, while the default run exited 0. It joins as EXTRA_SURFACES under the strict refs>0 rule and its own measured floor (280 against 333 observed). Exemptions are declared in the row that needs them -- `[no-tree-claim <CODE> ref=<name>]` over nine reason codes -- rather than in a checker-side map or a sidecar, because a map keyed by row ID goes stale before the prose does (the register gained ~13 dated rows today) and hides the excuse from the reader; the string-keyed global allowlist already cost this repo a hidden stale README entry. Every declaration is re-falsified each run: a declared name that resolves in the tracked tree fails, a name the surface never asked about fails, an unknown code fails, and DELETED / NEVER_EXISTED are settled against git history. That last check needed three corrections while being written: --diff-filter=D misses the R100/R081 renames, a row names the tail of a path so the suffix form must be tried, and a directory is proven by its contents -- `-- '*web/'` finds nothing while `-- '*web/*'` finds five commits. Reconciled the 65: 14 literals rewritten to unique tracked successors (each proved with git ls-files before use, including the two archived-AGENTS.md and the sync_hermes_workflow_assets.py real home), 40 declarations covering 50 occurrences, plus two shape rules (an elided `…` and any backslash-containing literal are no longer read as paths, which retires the regex `/\bCPU\b/` false positive). Measured: targets=39 refs=690 broken=0 declared_in_row=50, 28 tests in the ci module pass including four planted-failure cases. active-authority-index.md's claim that the gate covers "four current surfaces" is corrected in the same commit, and the design row's own quoted syntax no longer matches the extractor it describes.
…uard that measures it The row now states what was built rather than what was designed, including two self-inflicted traps found on the way: the design prose quoted the declaration syntax in its matching form, so the guard flagged its own documentation until the example moved to placeholder spelling, and my apply script's "[no-tree-claim" substring check reported that prose as an existing declaration. Reconciled totals: register refs 350, broken 0, declared in-row 50 across 40 tokens; the whole scan reads targets=39 refs=691 broken=0.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
摘要
对账一份独立收敛审计(
WORK-LAB_审计裁决_2026-10-01.json,anchord43d07a,overall_verdict = NOT_READY_FOR_REINSTALL_OR_DEPLOYMENT_ACCEPTANCE,23 项发现 / 2 blocker / 9 high,其中 18 项needs_local),把每一项拿到当前树上逐项核实,而不是采信任何一方的叙述。23 项全部有结论,无遗漏、无多余(用脚本对 F01–F23 全集校验,计数平衡于 23):
完整摘要与错误台账见
reports/SESSION-CONSOLIDATION-20261001.md。主要修复
假断言与不可支撑的声明(证据包自身)
global-workflow-coverage.json断言has_worklab_managed_marker: true,而 live 与归档两份CODEX_HOME/AGENTS.md都是 16,805 B / 691 行且不含渲染器标记。改为false+ 可复核更正块;MANIFEST 摘要同步。external_roots_touched: []只能证明记录内容;SECRET-SCAN-REPORT.json的扫描器源码与规则集未保留,第三方无法复现。新增behavior_declarations分级。陈旧登记与静默漂移
check_skill_provenance的 live 检查被if live_root is not None包住,而 canonical 门不传--live-root→ 门报live_checked=False,live_sha256从未被比对;model-switch因此长期带陈旧值。修正 + 新增离线自洽规则。chrome-profiles的revision 5b9c3257是准确的,但工作树被本地改过:__init__.py已装 24,433 B vs 该 revision 的 23,764 B。只记 revision 等于给错的字节发认证。model-switch(根级枚举缺口)。skills-inventory.json14 条中 6 条哈希陈旧,用权威生成器重新生成。用户状态与危险形状
disposition,两者标REJECT_USER_DATA。CC-1(.tmp-*)与CC-3(*.bak*)非互斥:实测 15 / 11 / 3 个两者都命中。加双向overlaps_with+ 优先级。MEMORY_BACKENDseam 写着 "no new callers allowed" 却无任何校验、扫描显示调用者为 0(既无人执行、又碰巧满足)。加 seam-caller 基线,注入调用者即失败。ArcheAxis-Knowledge-OS/AGENTS.md被工具链自动当作治理指引注入会话(本会话实测发生),且其00-governance引用已不存在。三份归档指令文件加解除指令横幅。commit=null+provenance_status: UNKNOWN+ 理由,并明写"未证明"清单。部署内容保全(F04)
windows-development-environment实机比仓库多:经验 30(.cmd不能经node <path>.cmd执行)、经验 31(工具链缺失必须如实报 BLOCKED,不得编造 PASS)、以及references/frontend-baseline-contract.md。已 land 进仓库源码(144 行,1.3.0 → 1.4.0),并把 live 中指向不存在文件的悬空引用换成真实路径。新增机器门(41 → 47)
evidence-tiering(46)plugin-inventory-honesty(47)负向对照:nf27 43 / nf28 13 / nf29 22 / nf30 11,全部验证过"能失败"。
验证
run_quality_gate.py verify→ 47 门 PASS(本地)EVIDENCE_TIER_PASSbundles=7;PLUGIN_INVENTORY_PASS;THREE_PROJECT_BOUNDARY_PASSsplits=6 markers=5 seams=3%USERPROFILE%类路径与origin/main同为 86 个文件 → 不引入新暴露证据等级声明:以上均为本地 canonical。exact-SHA CI / 原生运行验收 / 用户验收 / Release 全部
NOT_RUN。明确未做
apply(只跑了只读plan)auth.json、.env一律未碰)E:\/F:\UNKNOWN)本会话自身的错误(已撤销并留痕)
摘要第五节列出 14 项,其中值得注意的三项:
security-guidance/web-ddgs存在"声明 vs 观测不一致",是假发现:两者type/upstream/spdx均写明bundled,实为随 Hermes 应用发布的组件,不在用户安装目录是正确状态。已撤销并加门规则禁止再犯。reports/audit-archive/20260930完全一致,重复部分已删除。需要人类决定(本 PR 不包含)
ArcheAxis-Knowledge-OS/DESIGN-LAB的授权DSH_HOME数据根迁移