Skip to content

audit: reconcile all 23 independent-audit findings + land recovered skill content - #162

Open
DTALEX66 wants to merge 374 commits into
mainfrom
task-decomposition/atlas-gap-archive-20261001
Open

DTALEX66 wants to merge 374 commits into
mainfrom
task-decomposition/atlas-gap-archive-20261001

Conversation

@DTALEX66

@DTALEX66 DTALEX66 commented Oct 1, 2026

Copy link
Copy Markdown
Owner

摘要

对账一份独立收敛审计(WORK-LAB_审计裁决_2026-10-01.json,anchor d43d07a,overall_verdict = NOT_READY_FOR_REINSTALL_OR_DEPLOYMENT_ACCEPTANCE,23 项发现 / 2 blocker / 9 high,其中 18 项 needs_local),把每一项拿到当前树上逐项核实,而不是采信任何一方的叙述。

23 项全部有结论,无遗漏、无多余(用脚本对 F01–F23 全集校验,计数平衡于 23):

结论 数量
已修复 16(F01、F02、F04、F05、F06、F07、F08/F09、F12、F13、F14、F15、F16、F18、F19、F21)
已确认、仅报告 7(F03、F10、F11、F17、F20、F22、F23)
另 +裁决 Q5(FAIL,指出共同根因)
未对账 0

完整摘要与错误台账见 reports/SESSION-CONSOLIDATION-20261001.md。


主要修复

假断言与不可支撑的声明(证据包自身)

  • F21 — global-workflow-coverage.json 断言 has_worklab_managed_marker: true,而 live 与归档两份 CODEX_HOME/AGENTS.md 都是 16,805 B / 691 行且不含渲染器标记。改为 false + 可复核更正块;MANIFEST 摘要同步。
  • F12 — 归档 README 称 "hermes 363 文件中含 186 SKILL.md";实测 hermes 178 + codex 8 = 186,归属写错。
  • F15/F16 — external_roots_touched: [] 只能证明记录内容;SECRET-SCAN-REPORT.json 的扫描器源码与规则集未保留,第三方无法复现。新增 behavior_declarations 分级。

陈旧登记与静默漂移

  • F08/F09 — check_skill_provenance 的 live 检查被 if live_root is not None 包住,而 canonical 门不传 --live-root → 门报 live_checked=False,live_sha256 从未被比对;model-switch 因此长期带陈旧值。修正 + 新增离线自洽规则。
  • F07 — chrome-profiles 的 revision 5b9c3257 是准确的,但工作树被本地改过:__init__.py 已装 24,433 B vs 该 revision 的 23,764 B。只记 revision 等于给错的字节发认证。
  • F05 — 归档声称"所有技能主文件均已完整复制",实际5 个根级技能从未归档,含受管技能 model-switch(根级枚举缺口)。
  • F06 — skills-inventory.json 14 条中 6 条哈希陈旧,用权威生成器重新生成。

用户状态与危险形状

  • F14(blocker) — 内容本非破坏性(CC-4/HH-1 原本就写"官方 prune、禁止删除活动库"/"只 OBSERVE");缺陷在容器 —— 安全的话放进"清理候选"清单。改为机器可读 disposition,两者标 REJECT_USER_DATA。
  • F13 — CC-1(.tmp-*)与 CC-3(*.bak*)非互斥:实测 15 / 11 / 3 个两者都命中。加双向 overlaps_with + 优先级。
  • F19 — 边界验证器从不扫任何调用者;且 SSOT 的 MEMORY_BACKEND seam 写着 "no new callers allowed" 却无任何校验、扫描显示调用者为 0(既无人执行、又碰巧满足)。加 seam-caller 基线,注入调用者即失败。
  • F18 — 归档进本仓库的 ArcheAxis-Knowledge-OS/AGENTS.md 被工具链自动当作治理指引注入会话(本会话实测发生),且其 00-governance 引用已不存在。三份归档指令文件加解除指令横幅。
  • F01 — 归档含两个外项目的文件却只绑自己的冻结提交,无任何源 revision 绑定。记 commit=null + provenance_status: UNKNOWN + 理由,并明写"未证明"清单。

部署内容保全(F04)

windows-development-environment 实机比仓库多:经验 30(.cmd 不能经 node <path>.cmd 执行)、经验 31(工具链缺失必须如实报 BLOCKED,不得编造 PASS)、以及 references/frontend-baseline-contract.md。已 land 进仓库源码(144 行,1.3.0 → 1.4.0),并把 live 中指向不存在文件的悬空引用换成真实路径。


新增机器门(41 → 47)

门 作用
evidence-tiering(46) 行为声明必须标等级(DECLARED/OBSERVED/VERIFIED);低于 VERIFIED 却不写"不能证明什么"→ 拒绝
plugin-inventory-honesty(47) 插件记录不得超出证据;revision 字符串不能证明字节匹配;bundled 组件不得按"安装"语义评估
seam-caller 基线 声明"禁止新增调用者"的 seam,增长即失败
归档深度覆盖 只枚举深层、浅层为空 → 拒绝,除非显式记录缺口
归档指令文件 归档中可自动加载的治理文件名必须带解除指令标记

负向对照:nf27 43 / nf28 13 / nf29 22 / nf30 11,全部验证过"能失败"。


验证

  • run_quality_gate.py verify → 47 门 PASS(本地)
  • EVIDENCE_TIER_PASS bundles=7;PLUGIN_INVENTORY_PASS;THREE_PROJECT_BOUNDARY_PASS splits=6 markers=5 seams=3
  • register 47 行 AG、无重复 id、无畸形行、无 U+FFFD;权威校验器 PASS
  • push 前安全扫描:凭证模式 0 命中;私密会话正文 0 命中;%USERPROFILE% 类路径与 origin/main 同为 86 个文件 → 不引入新暴露
  • fast-forward,无历史改写(ahead 81 / behind 0)

证据等级声明:以上均为本地 canonical。exact-SHA CI / 原生运行验收 / 用户验收 / Release 全部 NOT_RUN。


明确未做

  • 未运行 sync 通道的 apply(只跑了只读 plan)
  • 未修改任何 live 客户端文件(只读)
  • 未读取或复制任何凭证(auth.json、.env 一律未碰)
  • 未访问 E:\ / F:\
  • 未安装、未联网抓取依赖
  • 未解决:AG-15、AG-17、DSH_HOME、AG-12/13/18、Rust/Tauri U19(缺 MSVC 工具链)
  • 未达成:F01 两个外项目的真实 commit/tree 绑定(需访问那两个仓库的授权,现为显式 UNKNOWN)

本会话自身的错误(已撤销并留痕)

摘要第五节列出 14 项,其中值得注意的三项:

  • E1 — 我报告 security-guidance/web-ddgs 存在"声明 vs 观测不一致",是假发现:两者 type/upstream/spdx 均写明 bundled,实为随 Hermes 应用发布的组件,不在用户安装目录是正确状态。已撤销并加门规则禁止再犯。
  • E3 — 我把 13 个受管技能重复归档一遍;核实后 11/13 哈希与既有 reports/audit-archive/20260930 完全一致,重复部分已删除。
  • E4 — tally 计数三次算错;第 25 轮起改为由列表长度计算并打印校验。

需要人类决定(本 PR 不包含)

  1. 两个外项目的 commit/tree 绑定 —— 需访问 ArcheAxis-Knowledge-OS / DESIGN-LAB 的授权
  2. AG-15 可写控制面 —— 授权 loopback 端点,或暂缓
  3. AG-17 web-GPT 传输方式
  4. DSH_HOME 数据根迁移
  5. Rust/Tauri U19 —— 需 MSVC 工具链,可能不是"授权抓依赖"能解决

DTALEX66 added a commit that referenced this pull request Oct 6, 2026
… a stale exe

CI is green at 6c63218 (both workflows, both events; PR #162 head, observer and
aggregate SUCCESS), so the "pull_request run still in flight" note is withdrawn.

The remaining blocker for the real-desktop-surface proof is now named:
u19_webview_e2e.py:446 hardcodes src-tauri/target/release/app.exe, but the documented
local recipe sets CARGO_TARGET_DIR, so a local U19 verdict keys on a superseded
binary. CI is unaffected because its build uses the default target dir - which is why
the defect never showed up there.
DTALEX66 added a commit that referenced this pull request Oct 6, 2026
…snapshot

The owner's executable prompt asked for two things in this project: correct
the descriptions, and land the full project description and future blueprint
with a repeatable audit. Documentation, normal commits, branch push, PR
update and the GitHub About are authorized by it; merge, release, install,
global config and other projects are not, and none were touched.

Input verified: the blueprint DOCX hashes to the value the prompt declares
(f3784a99...bea3440). The prompt's own digest is recorded for traceability;
it had no declared value to compare.

docs/future/WORK-LAB-BLUEPRINT-20261006.md carries all 19 chapters plus the
appendix onto repository anchors. Where the source is silent the document
says SOURCE_GAP instead of filling the hole: the conditional enhancements in
15.6 have no scenario, acceptance, phase slot or exit condition, and Lite,
tray HUD and Quick Entry have no field or permission spec anywhere.

.project/governance/blueprint-coverage.json is the single mutable source;
verify_blueprint_coverage.py re-extracts AG-01..AG-20 from the atlas card and
all 22 U rows from the single open register, so no row can be dropped or
invented, and renders the 87-row projection. Status reuses the ledger's
existing vocabulary; disposition reuses the prompt's own four decisions.

Twelve injected violations were required to fail before the gate was
trusted, and the first round exposed two defects in the gate itself: a
hand-edited projection passed because nothing compared it against the
regeneration, and the "no self-referential digest" rule could never fail
since a file cannot contain its own hash — a check that reports protection
it cannot provide is worse than none. Both fixed; the digest rule now
forbids generated artifacts from hashing each other, with two new
injections proving it bites.

README corrected: it described the project as a global-configuration layer,
which under-defines it against Authority section 3. Now the mother
definition, the owns/does-not-own boundary, the evidence vocabulary, the
current/future/candidate zoning at Ongoing, the mandatory reading order and
the audit entry points are all on the first screen.

Corrections recorded rather than rewritten: the blueprint's PR #162 head
6f323a3 has advanced to f3b5dec live, while its main cd4daa8 still matches;
the 2026-10-05 local full-gate FAIL stays on the record alongside the
affected-group passes, because a group pass is not the aggregate.
DTALEX66 added a commit that referenced this pull request Oct 6, 2026
ERR-109. Adding tests/test_artifact_freshness.py to observer-python-skeleton
broke the integration job's negative-control step, because
tests/ci/test_failfast_group.py pins each observer group's command count at 8
and I never ran it: I verified the two groups whose contents changed and
pushed, treating that as sufficient.

The guard is fixed by strengthening it rather than by editing a number. It now
pins the full command list of all four groups, so dropping a required check or
swapping it for a cheaper one both fail. Verified in both directions: the
baseline passes, and a swap that keeps the count at 9 — replace
test_artifact_freshness.py with a duplicate of the cheap skeleton test — now
reports missing=[[python, tests/test_artifact_freshness.py]], which the old
count-only guard could never have seen.

Afterwards all 25 CI-invoked governance and integration steps were run locally
and judged by exit code. Two traps worth recording: a negative-control suite
prints its expected FAIL lines, so reading its verdict from the last line of
output invents failures in test_aggregate_gate.py and test_error_ledger.py that
never happened; and test_exact_tree_review.py genuinely fails on a feature
branch because it asserts HEAD equals origin/main and every task COMPLETED, so
CI does not invoke it — recorded so nobody chases it.

Also lands the double-end readback the snapshot generator now observes itself:
branch pushed to a018fbe, PR #162 OPEN/MERGEABLE at that head, GitHub About
before -> target -> live readback matching, homepage and topics deliberately
untouched, and the exact-SHA CI state recorded honestly at 16 pass / 2 fail /
4 pending rather than as green.
DTALEX66 added a commit that referenced this pull request Oct 6, 2026
…a foreign SHA

PR #162 head is 44b778d; both runs for it completed/success with 24 checks COMPLETED. The chain stays visible rather than smoothed over: a018fbe failed twice (ERR-109, the manifest guard I skipped), 4115d63 onward is green.

Tooling trap recorded: gh run list --branch placed an unrelated 2026-10-02 run (2bc9168, not an ancestor of the current head) at the top of the output, which nearly made me report a foreign SHA as this round's result. Pin the head with gh pr view --json headRefOid first, then filter runs by SHA or read statusCheckRollup.
…e only in the legacy suites

Sharpening the previous commit's conclusion with an assertion-by-assertion inventory, and
correcting my own first read twice along the way (ledger ERR-118).

The parity matrix now carries a measured 2026-10-07 section instead of an opinion. Per-suite
counts are taken by scanning the suites themselves: test_projection_contract 13,
test_read_only_surface 17, test_render_v3 19, test_responsive_contract 11,
test_visual_assets_r2 5 — 65 assertions in total — and the production side is described by
its test titles (84 titled cases across 14 files) plus the 11 desktop contract cases, with a
per-suite verdict: no production equivalent for the projection-truth, read-only-surface and
v3-render groups; partial for responsive (only the top clearance and grid ownership are
pinned) and for visual assets (theme tokens and the legacy-purple removal are pinned).

Method note kept because both of my shortcuts would have produced a false green:
- a keyword census said the concepts exist (UNKNOWN appears 36 times in frontend/src).
  Presence of a word is not an assertion guarding it, so the matrix only accepts evidence
  from test titles.
- my first count of React-side cases came back 0 because a grep bracket expression containing
  a backtick did not survive argument passing. A wrong count written into a tracked document
  is exactly the failure this session keeps hitting, so it was recounted in-process and the
  line patched, with the stale zero verified gone.

What this changes about the plan: the browser entry was never the blocker — that is now pinned
to the same Vite dist by 8 sidecar tests. The blocker is that retiring apps/observer/web would
delete the only home of the read-only and no-fabrication contracts, since those suites read
web/index.html, web/styles/*.css and web/scripts/*.js. Deleting the project's truth
assertions to make a cleanup pass is the wrong direction, so U03 stays PARTIAL and the matrix
records the four ordered steps: re-anchor the read-only and projection contracts to the
production surfaces, re-anchor responsive and brand assets, then delete web/ together with the
required_groups glob and the pinned fail-fast list (ERR-109), the guarded-path tuple, the
README/architecture text and the retirement stub, with a hash manifest plus a full copy of
web/ as the retention point and a dated amendment to WORK-LAB-AUTHORITY.md §7.

No deletions this round, no rebuild, nothing published; the eight default gates plus the JS
suites (83 passed / 0 failed) remain green.
… surfaces (U03 step 1, half)

apps/observer/web cannot be retired while it is the only home of the read-only and
no-fabrication assertions, so the first porting step lands the file-level half of those
guarantees on the production tree (ledger ERR-118; the parity matrix carries the
assertion-by-assertion inventory).

test_production_surface_static_contract.js — 11 static contracts, wired into
run_all_tests.js so the existing observer-web-contracts group runs it with no group-manifest
change: no remote runtime/CDN in the entry html; no url(http or @import http in the skins; no
write method anywhere in the UI layer (POST/PUT/PATCH/DELETE and any fetch with an options
object); no credential, env, bearer or cookie read; theme and layout state never touch web
storage; horizontal overflow suppressed by construction; tabular numerals; long CJK/SHA
wrapping; prefers-reduced-motion honoured; no rule drops the base text under 12px; and every
sub-12px declaration must belong to a named micro role (.winctl-zoom, .load-strip,
.brand small, .tag/.badge, .kpi small) so a future 10px paragraph fails even when the base
rule still passes.

All four of the security- and legibility-bearing assertions were falsified by injecting the
violation into the real files — a CDN script tag, a PATCH in a component, a 10px body rule, a
localStorage write in lib/a11y.ts — each turning exactly its own assertion red with the file
restored byte-for-byte.

Two assertions were wrong on my side first, and both were fixed by correcting the test rather
than loosening it: paths arrive with the platform separator, so the editor-lane exclusion
regex matched nothing and reported the editor's four sanctioned storage uses as violations;
and an early version applied the body-size rule to every px declaration, which flagged the
B10 brand lockup's letter-spaced 10px caption as a defect. The replacement states the rule
the legacy suite actually enforced and names its exceptions.

Half of step 1 only: the projection-truth group (13 assertions — fixture numbers, estimate
versus bill, strict RFC3339, coverage, forward compatibility) and the v3 render group (19
assertions) are still un-ported, and those are the assertions that would be lost by deleting
web/, so the directory stays and U03 stays PARTIAL. The remaining order is recorded in the
matrix.

run_all_tests.js 94 passed 0 failed (was 83), the ten default gates green, and
ERROR_LEDGER_PASS entries=117 counts_consistent=true. No deletions, no rebuild, nothing
published.
…cts, snapshot re-recorded

Coverage is credited only against a named production test title, never against a keyword
census: 5 of the 13 legacy projection assertions have one, 8 are OPEN, and the shared
themes of the open ones (fixture numeric agreement, strict RFC3339, the coverage triple,
forward compatibility with unknown keys, graceful degradation with missing keys, the
10 required schema keys, not-metered rendering as 订阅未计量 rather than 0) are all
pure-function properties portable from lib/api.ts without any real material.

Why the legacy fixtures are not simply reused: apps/observer/tests/fixtures/*.json are the
v1 shape (usage/summary/quality, no coverage, no tokenSummary, no revision), so porting
assertions about them into the SnapshotV3 typed layer would mean inventing production
behaviour that does not exist. The map is the work order instead, and apps/observer/web
stays until both halves are ported.

The audit snapshot was re-recorded from live state: 8 checks, 0 failing, and its live-claims
block still reads the U19 register status, the ledger's status_after values, candidate-pool
coverage and the observed artifact's receipt drift straight from the files that own them.
… the 8 ported truth contracts

The seven date render sites were `value ? new Date(value).toLocaleString() :
'UNKNOWN'`. That leaks the literal "Invalid Date" for an unparseable string and,
worse, presents a corrupt timestamp as a plausible one: Date.parse does not fail
on 2026-02-30T10:00:00Z, it rolls over to March 2. `isRfc3339` now checks the
shape AND the calendar (month, day-in-month with the leap rule, hour/minute/
second, offset), and `fmtTimestamp` renders UNKNOWN otherwise. The producers were
read first (collector_scheduler._now_iso and durable_worker both emit `...Z`) so
strictness cannot hide real data. The projections also degrade a missing
transport / coverage / projects / executions / token object to UNKNOWN/null
instead of throwing, because the snapshot crosses a process boundary.

That is the U03 step-1 close: all 8 OPEN projection-truth assertions from
tests/test_projection_contract.js are re-anchored to the production tree, 8 named
behavior cases in frontend/src/lib/projectionTruthContract.test.ts plus 3
file-level cases in test_production_surface_static_contract.js (11 to 14). The
file-level half cannot live in vitest: frontend/tsconfig.json has no @types/node,
so node:fs is TS2307 and CI runs `npm run typecheck`.

Falsified 11/11 by mutation, each file restored byte-for-byte with SHA-256
verified (.project-local/runs/convergence-20261007-c/falsify_u03_ports.py). Two
first-version injections failed to hit their case because my expectation was
wrong, not the gate (`?? 'LIVE'` and zero-padding tokenTruth cannot reach a case
that passes transportState='UNKNOWN' explicitly); and a bare-$ needle flagged
api.ts's own legitimate 'http://$1' URL replacement. All three were corrected
against the facts rather than by loosening an assertion. Recorded as ERR-119.

vitest 15 files / 95 cases green, run_all_tests.js 97 passed 0 failed,
tsc --noEmit clean, vite build green, gate battery BATTERY_FAILURES=0,
AUDIT_SNAPSHOT checks=8 failing=0. apps/observer/web still is NOT deleted: the
per-assertion attribution for test_render_v3.js (19), test_read_only_surface.js
(17), test_responsive_contract.js (11) and test_visual_assets_r2.js (5) is still
open, so U03 stays PARTIAL.
…ad stops claiming LIVE

Porting the render_v3 contracts surfaced two ways the production surface reported
an unknown as a positive.

ProjectPanel.tsx and the cross-project grid in Views.tsx branched on
`dirty ? 脏 N : 干净`, so a project whose git.dirtyCount is null (never observed)
rendered 干净. That contradicts the repository's own wording in RulesPolicyView
("缺失即 UNKNOWN, 不伪造『全部干净』"). Both sites now branch on `== null` first and
render 脏 UNKNOWN; 干净 requires an observed 0.

useLiveSnapshot only called setError() on a failed fetch, so the compact HUD kept
replaying whatever transportState the retained snapshot was read with - a LIVE
badge on a connection the browser had lost. A failed or throwing read now stops the
LIVE claim and returns to the non-live source label (the projection itself is kept:
no wipe, no fabricated zero), and the HUD word comes from the new
frontTransportState(snap, live, error).

The port itself: 8 mounted cases in
apps/observer/frontend/src/views/renderTruthContract.test.tsx cover the six
properties the 10 OPEN render_v3 assertions collapse to - registry identity,
platform, execution count, branch@sha and dirty count; real token figures with no
invented cache-hit rate (v3 has no cache field); no progress bar without a source;
the overview KPI strip from the snapshot; 覆盖 UNKNOWN rather than 覆盖 0/0 for an
absent triple while a measured 0/0 still reads as a number; no metric card for
fields v3 does not carry; and last-good flipping to OFFLINE when the read fails.
Falsified 8/8 by mutating the production components, each restored byte-for-byte
with SHA-256 verified (falsify_render_ports.py).

Falsification caught the gate being blind, not the tree being wrong: the
no-fabricated-metric needle used `\bCPU\b`, but DOM textContent concatenates the
label with its value into `CPU11`, so an injected CPU card stayed green. Recorded
as ERR-120 together with the two defects.

The parity matrix is corrected and completed: the read-only suite measures 24
assertions (17 t() + 7 asyncTest()), not 17, so the legacy total is 72 not 65;
render_v3 (19), read-only (24), responsive (11) and visual (5) now have a
per-assertion verdict each, with every cited production title, the Rust
`observer_api_is_loopback_get_only` test, index.html:5 and index.css:220 checked
back against the tree. 18 OPEN items remain and are ordered as the step-3 work
order; the first six are defects that can produce a false impression (revision
monotonicity, payload validation, SSE failure staying OFFLINE, cache:'no-store',
?api= trust on the browser path, static-preview freshness). apps/observer/web is
still not deleted; U03 stays PARTIAL.

vitest 16 files / 103 cases green, run_all_tests.js 97/0, tsc --noEmit clean,
gate battery BATTERY_FAILURES=0, OBSERVER_READONLY_PASS, AUDIT_SNAPSHOT 8/0.
…ted endpoint

The read-only transport contracts survived only in the retired static suite
(test_read_only_surface.js, measured at 24 assertions: 17 t() + 7 asyncTest()),
so nothing in the production tree asserted them and no gate could notice the
drift. Six of them were defects that could put a false impression on screen:

- the snapshot GET had no cache directive, so a cached projection could read as
  current. It now declares method:'GET' and cache:'no-store'.
- the body was cast straight to SnapshotV3. A legacy/v2 or truncated 200 response
  rendered as truth. parseSnapshotPayload now checks schemaVersion, revision,
  arrays and the nested objects and fails closed; SSE frames go through it too.
- apply() overwrote unconditionally, so a delayed response with a LOWER revision
  replaced newer facts already on screen. The last accepted revision is now
  tracked and a lower one is refused without touching the display.
- a browser-supplied ?api= was accepted as authoritative with no validation, so
  anything reachable from the webview (external host, https, credentials in the
  URL, an added ?write=1, a non-/api/v1 path) became the data source.
  isTrustedObserverEndpoint mirrors the Rust gate observer_api_is_loopback_get_only
  on the client; an untrusted value is refused outright, never used, never
  marked authoritative.
- a payload read through the non-authoritative static-preview endpoint could set
  live=true and announce LIVE. live now requires both an authoritative descriptor
  and a LIVE verdict in the payload.
- an EventSource error was an empty handler, so the surface kept its LIVE claim
  after the stream dropped, and a snapshot frame that failed to parse was
  swallowed. A stream error or a corrupt frame now stops the LIVE claim at once;
  we never close the source, so the browser's own reconnect still recovers, and
  the next good read clears the failure without wiping the projection.

Burst behaviour follows the legacy contract as well: onOpen and resync_required
schedule exactly one re-read, heartbeats cost no read, and concurrent requests
collapse to one in flight plus one follow-up.

Ports: 11 mounted cases in
apps/observer/frontend/src/lib/transportTruthContract.test.ts (real useLiveSnapshot
against a controllable EventSource stand-in) and 2 file-level cases in
test_production_surface_static_contract.js (14 to 16). Falsified 15/15 by mutating
production files with byte-for-byte restore and SHA-256 verification.

Three gate needles were wrong before the code was, and each was fixed by matching
the shape of the violation instead of by loosening: the old clause treated any
fetch with an options object as suspicious, which is the opposite of the explicit
GET the legacy contract required; a bare-text scan flagged the validator's own
schema constant as an embedded snapshot; and the endpoint scan matched
@tauri-apps/api/window (a module path) and a regex replacement ('$1') as URLs. All
three now fail when they match nothing, so an empty scan can not read as a PASS.
Recorded as ERR-121.

Remaining retirement blockers for apps/observer/web are down to five: the brand
SVGs are still only in web/assets/brand, src/theme/tokens.ts declares no read-only
constraint block, the compact surface's dense project list is undecided, the
read-only label wording has no named owner, and the legacy 800/640 reflow
breakpoints have no production analogue. U03 stays PARTIAL; web/ is not deleted.

vitest 17 files / 114 cases green, run_all_tests.js 99/0, node static contracts
16/0, tsc --noEmit clean, vite build green, gate battery BATTERY_FAILURES=0,
OBSERVER_READONLY_PASS, AUDIT_SNAPSHOT 8/0.
…ish the parity attribution

Copies the five brand files out of the tree that is scheduled for retirement and
pins them where the gate can see them:

  web/assets/brand/{design-tokens.json,observer-icons.svg,
  work-lab-observer-symbol.svg,work-lab-observer-tray.svg,app-icon-512.png}
  -> frontend/src/assets/brand/

Each was compared by SHA-256 before and after the copy (5/5 identical). The new
cases assert presence, that each SVG starts with an <svg> root and reaches out to
no remote href/src, and that design-tokens.json still declares themes
dark/light, views full/compact and the machine-readable read-only constraints
(readOnly true, externalMutation false, modelSummary false). Nothing is mounted
this round: the production header is a text lockup, and inserting a graphic is a
visual change that belongs in its own task.

Two more attribution gaps closed. The compact surface now has a named guard that
it keeps exactly the four contracted KPI cells, and the shell's narrow reflow is
asserted by behaviour rather than by the legacy numbers - production collapses
.two-col/.three-col/.split inside @container page (max-width: 760px), and the
case reads the declaration value, so 1fr 1fr fails just like a missing block. The
read-only and no-fabrication wording in Views.tsx and App.tsx now has an owner.

The dense project list from the legacy compact hierarchy is recorded as an
intentional drop, not an oversight: the HUD is a single-column status strip, the
project truth already renders in the projects lane with mounted assertions, and a
second project surface inside the HUD is exactly what the legacy suite's own CC2
grid case argued against.

One legacy assertion is deliberately NOT ported: having the front recompute the
LIVE gate (coverage equality, timestamps, loopback eventsUrl) would make it a
second Update Authority. The client-side job - refuse an untrusted endpoint, never
claim LIVE from a non-authoritative source, and never let an out-of-order or
corrupt read replace good facts - is already pinned by step 3.

Falsified 5/5 against the production files (falsify_step4_node.py), each restored
byte-for-byte with SHA-256 verified: a remote href in a brand SVG,
constraints.readOnly flipped to false, the second-authority sentence removed, a
fifth KPI cell added, and 1fr widened to 1fr 1fr. My first version of the wording
case scanned the whole source blob, which would have counted even the test's own
quotation as a hit; it now reads the two named files.

With this, all 72 assertions in the five legacy suites carry a verdict. What is
left for U03 is the switch itself: deleting web/ with its five JS suites and
helpers.js, the four manifest bindings, the docs and the authority amendment.

node static contracts 16 to 21 green, run_all_tests.js green, tsc --noEmit clean.
…ace, before any deletion

24 files / 242007 bytes under apps/observer/web and 6 suites / 63363 bytes (five JS contract suites plus the helpers.js they read) are listed with per-file SHA-256, re-verified against the working tree after writing. The recovery point is the already-pushed commit named in the manifest, so git checkout 6de25fe -- apps/observer/web apps/observer/tests restores the whole surface; a second copy sits under .project-local/artifacts/observer-web-retired-20261007.

No deletion happens in this commit. The parity evidence that makes the retirement lawful is apps/observer/parity-matrix-u03.md: all 72 assertions across the five suites now carry a verdict, with named production owners (31 behavior cases across three vitest files, 21 file-level cases, 8 browser-entry cases).
…e only UI

U03 asked for static-web -> React parity, migration of valid capability, then
retirement of the static surface as production. Parity is now evidence, not a
claim: all 72 assertions across the five legacy suites carry a verdict in
apps/observer/parity-matrix-u03.md, credited only by named production titles
(31 behavior cases in three vitest files, 21 file-level cases, 8 browser-entry
cases), and each port was falsified by injecting the violation.

Deleted: apps/observer/web (24 files, 242007 bytes) plus the five suites and the
helpers.js that read it (63363 bytes).

The audit list and the recovery point exist before the deletion, in the parent
commit d6976dc: docs/audits/OBSERVER_WEB_RETIREMENT_MANIFEST_2026-10-07.json
records every file with its SHA-256 (re-read and re-verified against the working
tree after writing), names the pushed pre-deletion commit 6de25fe, and gives the
restore command git checkout 6de25fe -- apps/observer/web apps/observer/tests.
A working-tree copy also sits under .project-local/artifacts/
observer-web-retired-20261007/. No new tag is pushed - recoverability comes from
already-published history plus the manifest.

Bindings and text updated in the same change, because the deletion would
otherwise leave four manifests pointing at paths that no longer exist:
run_all_tests.js drops the five legacy suites; required_groups.json drops the
web/scripts/*.js glob from observer-web-contracts and tests/ci/test_failfast_group.py
keeps its pinned command list in sync (the ERR-109 discipline); the
observer-no-business-write trigger path becomes apps/observer/frontend/src; the
observer_live_server.py stub text, observer-source-architecture.md (with a dated
normative amendment, and the run instructions now the sidecar --frontend-root
instead of http.server --directory apps/observer/web), the Observer skill
reference and WORK-LAB-AUTHORITY.md section 7 all state the switch.

Also fixed a pointer I broke myself: the browser-entry suite imported the sidecar
module without the sys.path preamble its sibling tests carry, so the very command
recorded in ERR-118 and in the register row as its regression check failed with
ModuleNotFoundError when run standalone. It now sets its own path and passes 8/8
on three consecutive standalone runs. One thing stays unknown and is recorded
rather than smoothed over: on the first standalone run after that fix, 1 of 8
tests errored and the three runs after it were clean; the socket binds port 0 so
this was not a port collision, and nothing yet names the cause (ERR-122).

Verified after deletion: failfast_group observer-web-contracts commands=1 PASS,
node suites green, the pinned failfast list green,
CURRENT_STATE_FRESHNESS_PASS, blueprint coverage green, gate battery
BATTERY_FAILURES=0, error ledger PASS at 121 entries.

Breaking by design (a whole surface is removed), recoverable by manifest and
history.
…iewport captures

The parity work proved what the surface says; it never proved what it looks like. Reading CSS is not a render, so this round drives the declared playwright-bundled headless Chromium against the live sidecar serving frontend/dist (real v3 snapshot: one project, zero executions, all-null tokens, transport OFFLINE) and captures 26 PNGs, one per lane plus theme/layout/width variants, with no window raised.

Machine-checked, no eyes needed: every render produced a plausible PNG (thin=[]), and no two different lanes produced a byte-identical PNG - which is what would catch a route that silently renders the same surface twice. Dark differs from light, full from compact, and 320/820/1440 each differ, so the switches are live.

Judged by eye, two commercial-grade shortfalls are recorded and NOT fixed this round: at 320px the top bar stacks hamburger, search and the four action buttons into one row each and eats about 40% of the first screen (navigation itself still works through the drawer, and the action-cluster contract still holds), and a lane heading plus its meta line both wrap in a narrow column. Honesty held up screen by screen: Token and coverage read UNKNOWN rather than 0 or 0/0, the trend panel says there is no series, the empty list says registry has no active execution, and revision 0 is a real value.

Also recorded because it cost a debugging cycle: --dump-dom with --virtual-time-budget hangs against this app because the page holds an SSE connection; screenshot-only mode renders and exits.

Status correction inside the register: frontend/dist has been rebuilt since the first round (vite build green); the release binary is still deliberately not rebuilt, and the migrated brand SVGs are still not mounted in the header.
…ow 841px

Measured with a dependency-free CDP layout probe (Node 22 global WebSocket + Emulation.setDeviceMetricsOverride), not inferred from a screenshot: at a 320px viewport .app computed '210px 110px' and .main was 210px wide, with a 263px top bar whose four action buttons each took their own row. Same 210px at 560 and 840. What looked like a top-bar wrapping bug was the whole application column being squeezed into the rail's grid track.

Cause: a display:none grid child is not a grid item, so hiding the rail below 841px auto-placed main.main into the FIRST track and left the 1fr column empty. B10 does collapse .app to one track at <=840px, but the shell declares the two-track clamp(210px,17vw,280px) minmax(0,1fr) form unconditionally and later in the cascade, so it wins everywhere. That unconditional form was deliberate when written - it made the content track shrinkable - and nobody had measured the sub-841px case.

The two-track form now lives inside @media (min-width: 841px), the default is a single shrinkable track, and a <=560px tier lets the action cluster take a full row and the search box shrink to what is left. Re-measured: .app is one track equal to the viewport at 320/560/840, .main equals the viewport, the four buttons share one row (320px: x=14/56/110/152, top bar 263px to ~104px), and the 1440px desktop layout is unchanged at 244.8px + 1195.2px.

Locked by a new static contract case, 'the shell keeps a single app track below the rail breakpoint', falsified two ways: restoring the unconditional two-track declaration, and moving a clamp() declaration outside its media gate. My first version of the lock counted clamp() declarations instead of checking that each sits inside the gate - the shell legitimately declares it twice - and that version failed against the fixed CSS, which is how I found it.

Two self-corrections recorded in the audit: the earlier round attributed this to the top bar itself, and an closest() probe that omitted .top-actions from its selector list misattributed the buttons. A probe's selector set decides what you can see.

docs/audits/OBSERVER_UI_RENDER_AUDIT_2026-10-07.md gains the measured before/after table and the probe's reproduction command. ERR-123. vitest 17 files / 114 cases, node contracts 33 (22 static + 11 desktop), gate battery BATTERY_FAILURES=0, vite build green; the audit sidecar and probe profiles were stopped and removed.
…beneath it

Reading rendered screens catches contradictions that reading source does not. Two came out of the headless captures:

1. The overview status stack derived its alert chip from alerts.length, and alerts is built by iterating snap?.ci and the execution rows. With no data source both are empty, so the chip painted a green 无告警信号 while the panel directly beneath it said 数据源未接入 - 无法判断告警(保持 UNKNOWN,不伪造「全部正常」). Counting an absent collection is not measuring zero - the same confusion as dirtyCount in ERR-120, one layer up. The chip is now three-state: no snapshot -> 告警状态 UNKNOWN (muted), alerts -> N 条告警信号 (warning), snapshot with none -> 无告警信号 (success).

2. The permanently-disabled 导出状态 / 新建执行 buttons looked exactly like live ones: b10.css ships no :disabled rule and keeps button{cursor:pointer} plus a hover lift, and D-11 pins b10 verbatim, so the shell now carries the affordance (not-allowed cursor, dimming, no lift).

Locked by a vitest case asserting all three chip states and a static contract case asserting the disabled affordance shape. Falsified 3/3: two-state chip restored, cursor:not-allowed deleted, hover override deleted - each reds its case, each file restored byte-for-byte with SHA-256 verified (falsify_ui_round4.py).

Two process notes. Appending to a CRLF working copy with a shell heredoc produced mixed line endings, and two of my own falsification needles then reported NEEDLE-ERROR until the file was normalised - the needle was right, the file shape was not. And the first wording of ERR-124's repeat_prevention was rejected by the ledger gate as 'not enforceable' because it read as description rather than rule; the gate caught me, which is the point of it. ERR-124.

vitest 17 files / 115 cases, node contracts 34 (23 static + 11 desktop), tsc clean, vite build green, battery BATTERY_FAILURES=0. CI on the retirement commit 891c871 is green across both workflows, so deleting web/ and rebinding the four manifests did not break the pipeline.
…sion

U08 was still marked PARTIAL because its Windows Tauri real-WebView layer was deferred to U19. Both of the things U08 actually owns are done and CI-gated: the observer-frontend-typecheck group runs npm ci -> typecheck -> build -> test, where vitest is now 17 files / 115 named assertions (this session added the projection-truth 8, render-truth 9 and transport-truth 11), and the observer-desktop-crate group runs cargo test --locked plus cargo check --locked. The layer it handed over has landed too - U19 is PASS with an owner-observed desktop surface and a CI real-WebView E2E readback step. The genuinely open piece, rebuilding and publishing the release binary, is now recorded where it belongs: on U19 and the release work order, not on U08.

SESSION-HANDOFF-CONVERGENCE-20261007-C.md records the structural change for the next session: apps/observer/web is retired (891c871, CI green) with an audited per-file hash manifest and a recovery commit; all 72 legacy assertions carry a verdict; six classes of production truth defect fixed (ERR-119..124); the CDP layout probe and headless render pipeline that turned 'looks wrong' into measured numbers; the next-round priorities (U02 remainder, manifest-first greening with its do-not-touch list, remaining UI items, release chain); and the traps hit - --dump-dom --virtual-time-budget hangs on the app SSE connection, the in-app browser reports a 0x0 viewport so it cannot measure layout, Git Bash has no node/python on PATH and Windows PYTHONPATH needs ';', heredoc appends create mixed line endings, and the ledger gate rejects repeat_prevention written as description rather than rule.

docs/future/WORK-LAB-BLUEPRINT-COVERAGE.md regenerated after the register edit, as the coverage gate instructed; it now verifies PASS at 87 rows.
…ite never covered

U02 listed the .hermes/ fallback in hermes-project-data.py as a leftover to remove. Reading it first says otherwise: the docstring states it as a deliberate cross-project compatibility contract (WL-010/020/030) - projects whose .gitignore lacks .project-local keep working, and the guard fails closed when neither root is ignored. Deleting it would break other projects and cross the owner's do-not-write-outside-this-project boundary, so the item is adjudicated as audited-and-kept rather than fixed.

The real risk was never that the fallback exists, but that it could win inside this repo. Measured: .gitignore ignores both .hermes/ (line 2) and .project-local/ (line 42), and candidate order prefers .project-local, so behaviour is correct - but nothing pinned that order. Every case in test_project_data_boundary.py builds a fixture that ignores only .hermes/, i.e. the entire suite exercises the fallback path and none of it exercises the preference.

Adds test_project_local_is_preferred_when_both_runtime_roots_are_ignored: with both roots ignored, tmp must land under .project-local/runs and .hermes must not appear in the path. Falsified by swapping the candidate order - the case goes red; the script is restored byte-for-byte with SHA-256 verified. 17 tests OK.

U02 now has one named remainder: publishing the corrected managed skill into Hermes Home via sync_hermes_workflow_assets.py --apply, a global change that rewrites live agent guidance and needs its own staged execution.
…actually check

AG-20 asked for source-registry increments - path, time, hash, original location, coverage relation. The registry those increments were appended to is the atlas one, under .project-local: ERR-114 says in its own remaining_boundary that it is machine-local, that nothing there is enforced by CI, and that a rebuilt atlas reverts the pins to BLOCKED_NOT_VISIBLE. So the missing piece was not more rows in an unversioned file but a tracked mirror with a gate.

.project/governance/recovered-source-registry.json carries 12 entries in five honest statuses: machine-local recovered copies (the recovered WORK-LAB-SUMMARY, the atlas registry itself, the scan manifest), absent (the 15,558,839 B / 432,344-line timeline, with its negative proof kept as text and no digest claim), unpinned (the two 2026-09-28 items that have a name but no expected hash, so no claim is made in either direction), tracked (the web retirement manifest and the five migrated brand assets, each recording apps/observer/web/assets/brand as its original location and 6de25fe as the recovery point). Every digest is computed at write time by generate_source_registry.py and re-read from the written bytes; nothing is transcribed.

tests/workflow-assistance/test_recovered_source_registry.py (7 cases) is discovered dynamically by run_quality_gate.py governance - it is in the 182-file set, so CI runs it with no manifest edit. It re-hashes every tracked entry, refuses a versioned file labelled machine-local, refuses a hash claim on an absent or unpinned original, and fails if the timeline row disappears. Falsified 5/5 with the registry restored byte-for-byte and SHA-256 verified.

AG-19 stays PARTIAL: the timeline is unrecoverable from here and history_complete is still false - what changed is that the verifiable half is now pinned by a gate instead of by a prose claim.
…aller gone, three cited trees kept

The four remaining cleanup candidates were measured and citation-checked rather than deleted by size. p0c-oracle-20261006 (1,661 files / 158,786,208 B), topbar-measure-20261006 (1,615 / 151,633,914 B) and ui-suite (344 / 73,450,325 B) are all cited by tracked files - the error ledger for the first two, five UI manifests and the visual QA report for the third - so the citation guard refused them and they stay. Their per-file listings were still captured so the next round can compare without re-walking.

dsh-reconfig-20260904 had zero citations, but deleting the whole directory would have thrown the run's own evidence away with the bulk. Splitting it: 134,265,015 B of its 141 MB was a single DSH-Desktop-2.0.5-x64-Setup.exe, a community build AGENTS.md already marks superseded by the official DeepSeek Harness 0.2.0-rc.2, and uncited anywhere in the repository. Only that file was removed - digest 777cb50c86b0... recorded first, manifest written and re-read, deletion confirmed by the file no longer existing - while the logs, JSON and python-test-env stay in place so any path that does resolve still does.

A spill-ledger line with all four verbs (trace/locate/clean/migrate) was appended to .project-local/artifacts/spill-ledger.jsonl, and PROJECT_DATA_BOUNDARY_PASS. The lesson recorded in the register: greening is not 'delete the big ones', it is 'only touch what nothing cites, measure and record before touching, and leave evidence and cited paths where they are'. The three retained trees need their citations migrated before they can go, which is the next round's task, not a reason to bypass the guard.
…nnot disagree

Work-order item 14 was the last open row in the U03 parity list. The brand JSON already carried constraints/readOnly/externalMutation/modelSummary after the migration, but the code that the UI actually imports declared nothing - two places writing a rule down is not a single source, so the fix is a cross-check, not a copy.

theme/tokens.ts now exports VIEW_CONSTRAINTS, and a new static-contract case compares it value by value against assets/brand/design-tokens.json and pins the law itself (readOnly true, the other two false). The case fails on a disagreement in either direction, so neither file can drift alone. Node static contracts 23 to 24, tsc clean, vitest unaffected.

Falsified with three injections: externalMutation flipped to true in the JSON, the VIEW_CONSTRAINTS export deleted from tokens.ts, and readOnly flipped to false in tokens.ts only. The third initially reported SURVIVED and that was my expectation string being wrong, not the gate: it trips the disagreement assertion first and prints 'readOnly disagrees: tokens.ts=false design-tokens.json=true'. Re-checked against the real message; the file was restored byte-for-byte each time.

The register row and the parity matrix now record the closure of all 18 work-order items, and two things that are NOT part of the old list: the migrated brand SVGs are still not mounted in the header (a visual change needing its own review), and the CDP layout probe measures but is not wired into CI, so narrow-viewport defects are still caught by CSS-shape contracts rather than by real box metrics.
The recovered-source registry recorded sha256 over working-tree bytes. With
`* text=auto` the same commit is CRLF here and LF on the runner, so three of
six tracked entries could only ever match on this machine: CI failed at
1cffd9f/e7ede1f/ab33537 while the identical 45-command step list passed 45/45
locally. Tracked entries now record the blob at HEAD, and a new test runs each
recorded recovery command so `byte-identical` becomes a computed relation
rather than an adjective.
21 rows said DEFERRED and the gate only checked their shape, so three rows
contradicted the code (OpenHands already an honest fleet adapter, Hindsight
already a POC on the nine-operation memory contract, n8n already barred as a
fourth task core) and two rows carried TBD triggers. Every row now carries a
GitHub API discovery readback, an absorption level, native_evidence paths that
must exist and a decision: 15 retired with reversal conditions, 7 kept with
blockers named, 0 promoted because no candidate has pilot evidence and a status
word is not an experiment. The gate refuses a licence the readback did not
return, an absent evidence path, a PILOT row with no code and placeholder text;
falsified 10/10, and the committed verifier reproduced as 0/10 (ERR-126).
Nine models carried 64-hex sha256 values and health words reading
DOWNLOADED_HASH_VERIFIED while nothing recorded what those digests were digests
of or when anyone recomputed them, and the integrity gate passed on shape. This
re-hashes all 48,531,408,516 B: nine match, and the directory-backed zipformer
entry gets twelve per-file digests instead of one truncated prefix. The two
pending-retirement weights turn out to be present (5.2 GB and 18.6 GB), so the
open decision now carries its cost; the reranker leftover is the same length as
the registered weight but a different digest, so deleting it as a duplicate is
not justified; the ollama-partials row was written as prose and is marked
unresolvable rather than absent; and a 5.97 GB weight layer belonging to no
entry is registered and left in another runtime's store. ERR-127, gate rules
falsified 14/14 and 0/14 against the committed verifier.
…stale URL

The AG-05g falsifications first lived in git-ignored scratch, where CI would
never have re-run them. They are now thirteen named fixture cases inside
nf22_registry_and_acp_honesty_gates.py (49 tests OK), so the governance group
guards the rules permanently. The retired-model cost check moved outside the
early return so a model with no digest claim still has to state its bytes, and
0-when-released is accepted as its own honest answer. Separately, source-ledger's
agent-skills row pointed at https://agent-skills.org/, which resolves to nothing;
the readback names agentskills/agentskills at Apache-2.0 and the old value is
kept inside identityReadback rather than quietly replaced.
The integration job went red 17 seconds after the previous push: the agent-skills
row had gained an identityReadback object, verify_source_ledger_v4.py let it
through and the CI-invoked verify_source_ledger.py validates a schema that closes
its property set. Third occurrence of one shape this session (ERR-109, ERR-125,
now ERR-128): validating a slice of the CI command set and calling it CI. The
provenance moved into integrationNote, where the contract allows it, keeping the
previous URL and the readback. scripts/ci/reproduce_ci_commands.py now parses both
workflow files, groups each run block the way a shell does, honours
working-directory, and names what it cannot reproduce (CI_CONTEXT, STEP_ENV,
TOOL_NOT_ON_PATH) instead of faking or hiding it — 107 interpreter commands,
including the 12 integration-job commands my earlier harness never ran.
…leased with zero cited paths lost, and the three obligations that remain
`command` says what ran then; `lifecycle.regressionCommand` promises what can be
run now. Nothing in the ledger distinguished them, so 17 of 43 promises pointed
into the git-ignored runtime root the greening rounds delete, and the two that
looked broken (`cd apps/observer && node tests/run_all_tests.js`) turned out to
be fine once the scan resolved them against the cd prefix they carry — my first
scanner also inflated the absent count to 49 by matching `.js` inside `.json`
and treating prose as commands. The 17 real cases were promoted byte-for-byte
into scripts/audit and scripts/maintenance with digests recorded, the entries
keep their original string in regressionCommandPrior plus a dated correction,
and a new CI-discovered module (183 modules, no manifest edit) refuses an
untracked promise, a .project-local promise, a dropped promise and an undated
correction. Falsified 4/4 against the live ledger; no exemptions needed.
…s, and a gate that re-derives them

AG-11 had been NOT_RUN since the atlas gap was opened: no inference path was ever bound, so no
OCR/ASR claim could be made. LM Studio serves qwen2.5-vl-7b-instruct at localhost:1234, so the
matrix could finally be run against a real model instead of asserted.

Measured 2026-10-07 over nine declared cells: probe recovery is 11/11 on the mixed zh/en page,
11/11 scanned-degraded, 4/4 zh-prompt, 4/4 en-prompt, 8/8 pdf-page, 15/15 table, 3/3 page-order;
ASR on the fixture reaches CER 0.0 at RTF 1.23. The vendor wav set is recorded as
MEASURED_NO_TRUTH, not a pass — it ran, and nothing in it matched a declared probe, which is the
honest state rather than a zero scored as 100%.

Evidence: docs/audits/AG11_OCR_ASR_MATRIX_2026-10-07.json (fixture identities hashed, package
versions and served models named, per-cell hits/probes/rates).
Instrument: scripts/audit/ag11_ocr_asr_matrix.py — cells are declared up front, so a cell that
cannot be measured is reported as not_run rather than quietly dropped.
Gate: tests/workflow-assistance/test_ag11_matrix_evidence.py re-derives every rate from
hits/probes, refuses a score on a MEASURED_NO_TRUTH cell, requires RTF to follow its own seconds,
and rejects a summary that disagrees with its cells. 10 tests, 8 negative controls.

The register row keeps its original NOT_RUN sentence and records the closure after it.
Real-user audio and multi-page PDF OCR stay UNKNOWN: no such material exists in the repo.
…ontract breach closed

The head carries the config-compiler conformance fix, the governance vacuity guard and the
authority-index reference gate, so a verdict there covers those records rather than an ancestor.
…s currency

They sat under docs/current and two live documents called them "本轮", so a reader at the
front door was pointed at some past round's record as current procedure; one of those links
had not resolved since the 2026-09-17 convergence. The reference gate now judges four
surfaces and distinguishes a path from a backticked term, and the frozen files keep their
pre-convergence paths on purpose -- ten of them -- because correcting a dated record to
today's names would make it describe a moment that never happened.
… currency

Measured per file: ten pre-convergence paths remain inside the frozen records and are
deliberately neither rewritten nor guarded, which is written down instead of left implicit.
… carrying both

That head is the first one to contain the register-shape gate and the widened reference
checker alongside the records they were measured against.
…s them, and gate it

Containment by a ref is the test, not whether the object parses: three of these hashes
return HTTP 422 from GitHub and exist only on this machine, yet every existence check
called them resolvable. 34 other pins sit on the open PR branch and are legitimate today,
which is why an earlier "not an ancestor of main" count of 40 would have failed a healthy
register.
…uld have used

The earlier 13-tokens/24-rows figure came from judging reachability as "ancestor of
origin/main", which also swept up 34 pins that sit on this open branch and are perfectly
checkable today.
The watcher decided "changed" from row counts, status tallies and aggregates, and an
acquire_lease moves none of them: measured, writing holder/token/checkpoint onto a live
task left the canonical fingerprint byte-identical. The witness is now content over the
mutable table plus a PRAGMA data_version fast path, and acquire/heartbeat/release finally
stamp updated_at, which they silently did not before.
The first version of this fix was caught by its own gate: acquire_lease did not stamp
updated_at either, so the witness stayed put until it did.
…ck guards

A watcher that reads only when another connection commits cannot prove its own readback
works, cannot notice a publish failure to retry, and never enters the read whose blocking
shutdown the guards in test_sidecar_v3_snapshot.py hold. The content was never measured as
a cost worth trading those for; the witness and the updated_at stamps stay, the loop goes
back to reading every tick.
…the retraction

The ledger entry names my own regression: gating the watcher's content read behind
PRAGMA data_version made a watcher that never reads unable to fail, so the shutdown
guard could not see it enter the read, a raising readback could not push it to STALE,
and a failed publish was never retried because the cursor advanced first. ERR-199's
remaining_boundary gets a dated correction rather than a silent rewrite -- the two
halves that still hold (newest_changes as the witness, updated_at stamped by
acquire/heartbeat/release) stay asserted, the fast path does not.

The two republished audits that carry ERR-200's binding are committed; the citation
and tool-inventory regenerations were timestamp-only and are discarded, not smuggled
into a record commit.
…ale from my own pin rewrite

The integration job has been red at d76deef and cdfd27d with BLUEPRINT_COVERAGE_FAIL while the
canonical local gate printed PASS on the same tree: verify_blueprint_coverage.py ran in no local
gate at all. The projection restates register cells verbatim, so my 12 dangling-pin corrections
made it stale, and the only reader that could see that was CI -- a freshness rule a writer cannot
run before pushing is a rule that honest edits will break.

Regenerated with --write (rows=87, anchors_resolve=true) and wired into VERIFY_ORDER as
blueprint-projection. Falsified rather than assumed: the new gate passes on the regenerated
projection, and editing one word of a U19 cell makes it exit 1.
…e caught it on my own run

The governance batch went red at 951dbcc in exactly the way this branch is supposed to fail:
test_quality_gate_runner_is_canonical_and_just_is_optional pins VERIFY_ORDER as a tuple AND retypes a
long prefix of the runner's `verify: Run ...` line, so adding blueprint-projection had to be written
in two places and the second one was missed.

The tuple stays a hand-written pin -- that is the point of it, a gate can only join by someone
accepting the order. The second assertion now derives from VERIFY_ORDER instead of copying it, which
pins more than before (the whole line, not a truncated prefix) and stops disagreeing with itself on
every addition.
…current pages went unjudged

Widened to every tracked markdown under docs/current (38 surfaces, 340 references, broken=0) and fixed
what that found. Sixty references did not resolve: eleven pointed at scripts a reader would run and
watch fail (python scripts/workflow/sync_codex_global_assets.py and friends, dead since the 2026-09
convergence, invisible to the gate because they sit inside fenced blocks), sixteen had lost a root
prefix, thirteen named a client home or an installed app in bare form, twenty-two were a dated page
writing today's tree, and three named things that never existed in any ref.

Client-home references now use the notation this repo already standardises ($HERMES_HOME/, ~/,
<dshInstallRoot>/, <Cognitive-Loop-OS>/) instead of growing a global allowlist, because the string-keyed
list hid a stale README entry once already; DECLARED_NON_PATHS is still exactly one entry and a test
says so. The DSH adapter page gained a dated SUPERSEDED banner: its whole identity is the community
build AGENTS.md retired, and five dead pointers were the symptom, not the disease.

Two rule changes rather than two excuses: a fenced-map entry is judged as a path only in file or
directory shape, which is what stopped origin/main and /interrupt reading as missing files, and the run
refuses to pass below a measured reference floor so a collapsed extractor cannot claim the cleanest
tree yet. Verified the loosening excused nothing: no originally-broken literal survives as a bare
backticked path.

Also: apps/token-monitor/src-tauri/target/ was not ignored by anything (the handoff asserted it was, and
the only /target/ rule in the repo is observer-scoped), so the build output is now ignored and the
sentence says what is true. The live-command audit's own floor was pinned to its debt count, so fixing
one reference made the guard red for doing its job; the floor is now a not-blind floor, with the
measured debt noted beside it.
…out to be false

Two new rows. DOC-REFERENCE-WIDENING-20261008 closes the widened scan with its measured before and
after. REGISTER-CELL-TRUTH-20261008 registers what a full literal-by-literal audit of the register
found and does not quietly rewrite: 21 observer paths missing their apps/observer root, and two claims
that are not a path problem at all -- AG-20 credits a generator named generate_source_registry.py that
appears in no added or deleted path in any ref, and U03-PARITY-20261007 says seven observer JS
consumers are still to be retired while the very commit the same cell cites deleted all seven. Those
two need the owner's intent, not my guess.

The tool-inventory regeneration was timestamp-only and is not smuggled into this commit.
…id now have a local route

Measured: CI invokes 64 python operands across two workflows; 11 of them were reachable from no local
route at all -- eight scripts/ci verifiers, a generator whose committed output CI regenerates, and the
WebView2 E2E. The blueprint projection was this same class two commits ago: red in CI, PASS locally,
and the writer had no way to know because the rule they broke was a rule they could not run.

verify_ci_check_reachability.py decides reachability from repository text (named gate, a test module
that names it, or the module being a discovered test itself) and refuses a run whose declared
exemptions have quietly become reachable -- which is how I found that three of my own draft
exemptions were wrong, and why the list is now two entries with measured reasons rather than five with
guesses. Steps carrying working-directory: are resolved before judging, so apps/observer/scripts/
u19_webview_e2e.py is not mistaken for a root-level path.

tests/ci/test_ci_invoked_verifiers_have_a_local_route.py then runs the eight verifiers here and checks
the committed contracts.ts against a regeneration whose bytes are restored afterwards. All five tests
pass; the reachability run goes green by being satisfied, not by being declared away.
…ed itself before shipping

The REGISTER-CELL-TRUTH row's first draft said "seven JS consumers deleted" and "STALE_PREFIX 21". Per
literal, the measured numbers are six deleted (run_all_tests.js was modified by that commit and is still
tracked) and 22 literals across 11 lines, and AG-20 already carries a dated correction saying its
generator script does not exist -- so the remaining fact there is that no script produces the registry,
which is an owner decision about a producer, not an unmarked lie. The row says what the commands showed
and records that its own opening claim was rewritten after verification.
…state freshness mode is locally reachable

Two reds that were mine, both found by running things rather than by reading my own notes.

The witness rolled up count, fencing token, checkpoint length, holder length and MAX(updated_at), so the
only column a heartbeat necessarily moves -- lease_expires_at -- was not witnessed at all. My renewal
test passed by accident: it relied on updated_at, and on this host _now() returned a single value across
2000 back-to-back calls (~1 ms granularity), so the same statement pair can write byte-identical rows.
Red at 9fb20f2 in the full batch, green standalone. The expiry is now part of tasks_state, and the test
freezes the store clock before the lease is taken so the renewal can change nothing else -- which is how
the blindness shows rather than being asserted away.

CURRENT_STATE_FRESHNESS_FAIL source-digest-mismatch was red in CI at 6e6d4fb and red locally, and the
canonical gate said PASS: the CI step runs generate_current_state.py --check-current, while the local
route for that script is a test module that imports it and never runs that mode. A route naming a file
without its mode is not a route. The mode now runs in the local batch, the projection is regenerated,
and the failure names the moved sources via git history instead of printing a hash pair -- measured: it
pointed straight at config/skill-provenance.yaml and the skill I edited today.
Both are my own defects from this round, recorded with the command that showed them: the witness that
never looked at lease_expires_at behind a test that only passed because the host clock was coarse
(~0.95 ms, two distinct values in 2000 calls), and the freshness mode that CI ran and no local route
executed -- which my own new reachability rule counted as covered, because it judged the file rather
than the flag. ERR-204's phase was first written as CI_VERIFICATION and the ledger verifier refused it;
the enum is LOCAL_VERIFICATION, EVIDENCE_CLOSEOUT, LOCAL_IMPLEMENTATION, AUDIT_ONLY, and the check caught
my invention rather than the record standing.
… so CI's verdict at one SHA was luck

run_forever writes the collector rows inside run_once() and only stamps worker_loop after that tick
returns, which makes a health table holding just the collector a legal intermediate state. The test polled
until the table was non-empty and then asserted worker_loop was in it -- so the outcome depended on where
the poll landed in that window. Measured: the push-triggered work-lab-gate run at 048e616 (tree bd3a678)
was success while the pull-request-triggered run at the identical SHA and tree failed exactly there
("'worker_loop' not found in {'healthy_collector'}"). I had already stamped eight verifiedCommit fields
from the green run; they are withdrawn, because one red run at a head means the head is not verified.

The wait is now on the set the sidecar itself declares as expected, with the deadline raised to 5s and the
failure naming which collectors stayed silent. Falsified rather than assumed: suppressing the worker_loop
write makes the test fail in 5.3s and name the gap; the module is green 5 runs in a row and 41 tests pass
across the four sidecar/worker modules.
…two anchors pointed at the wrong line

A literal-by-literal audit of the register's path claims (686 distinct literals, resolution against
git ls-files only) turned up assertions that are simply false today:

- AG-06o calls references/frontend-baseline-contract.md "confirmed absent from the repository source".
  It is tracked, at packages/client-neutral-core/skills/software-development/windows-development-environment/
  references/, and the very skill whose deployment the row audits cites it at SKILL.md:130-131.
- AG-14 records b10.css at 21,002 B and l10b-shell.css at 2,907 B "verified against main@cd4daa83";
  git cat-file -s at that commit says 20,469 B and 2,832 B, and l10b-shell.css is 22,012 B today.
- U03-PARITY-20261007 asserts the observer JS suites are "退役仍未做" while the commit the same cell
  cites, 891c871, deleted all six of them (run_all_tests.js was modified and survives, and its own
  header comment records the removal). Two anchors in that row are wrong too: required_groups.json has
  no web/scripts/*.js glob anywhere and its :22 group runs node tests/run_all_tests.js, and the
  observer-no-business-write watch list lives at run_quality_gate.py:1381 and protects frontend/src,
  not apps/observer/web/.

The dated corrections keep the original judgement as evidence instead of erasing it. 22 root-prefixed
path literals across 11 rows were rewritten, each replacement re-checked with
git ls-files --error-unmatch (20/20 tracked), and the table shape (170 rows, wrong_width=0) plus pin
reachability (67 pins) still pass. AG-20's claim that a script named generate_source_registry.py
generated a tracked registry stands as an owner ask: no script in the repo produces that file.
…uced them

ERR-205 keeps the flake honest: one trigger green and one red at the same SHA and tree, and the
verifiedCommit stamps I took from the green run are recorded as withdrawn rather than quietly re-taken.
ERR-206 names the four git commands behind the register corrections and states plainly that no machine
guard yet covers the register's factual claims -- applying the docs rule to it yields 61 unresolved out
of 321, and nearly all of those are legitimate client-home, build-output, cross-project or deleted-file
mentions, so gating it needs the per-row exemption design, which is the next open step rather than a
claim of closure.
…p-alive connection, in process

ERR-187 recorded that this test could not be written without a subprocess boot because the handler class
is nested inside the serving function. That was never checked and it is false: create_server is module
level and binds port 0, so the test speaks raw bytes to it in the same process. The boundary stayed open
because the excuse for not writing it was never measured.

One connection, two requests: a POST that must be refused with 405 and must not echo its own body, then a
GET that must begin with a clean status line -- which is the symptom a non-draining refusal actually
produces, invisible to any client that only parses statuses. The control case serves the sloppiness on
purpose: a refusal that never reads its body makes the second response something other than
HTTP/1.1 200 OK, so the assertion above is not green because nothing can fail it. Three consecutive runs
stable; ERR-187's boundary now carries the dated closure instead of the false obstacle.
…is repo's convention and my omission

The repoint tool rewrote 330 lines of the ledger and the diff looked alarming until it was compared
semantically: records head=204 working=204, semantic changes in pre-existing records=0, added=[] and
removed=[]. Every one of those lines was \uXXXX escapes becoming readable CJK. The escape form was mine --
my append scripts used json.dumps' default while the ten ledger writers already in scripts/audit all pass
ensure_ascii=False -- so this commit restores the house convention rather than introducing one, and the
next append must keep it or the same phantom diff returns.
… for everything, now pinned

Measured on this checkout: `git check-ignore -v --no-index services/` exits 0 and names `.gitignore:44`,
which is an empty line, and it does the same for docs/current/ and for a directory that does not exist.
With a file inside the directory the answer is honest -- services/orchestration/sidecar.py exits 1 while
.project-local/runs/x.json exits 0. So "is this path ignored?" is only a question you may ask about a file,
and a claim built on a trailing-slash query is not evidence.

The guard pins the trap, asserts the file-shaped probe still discriminates (including the Cargo target rule
added today and apps/observer/web, which is deleted rather than ignored), and scans tracked call sites for
directory-shaped operands. The first scanner used a regex over a text window starting at the word
check-ignore -- which in every real call site is mid-list, so the quotes paired off by one and it reported
the wrong literals; the planted case now proves the ast-based reader catches build/ and ignores a runtime
concatenation it honestly cannot judge.
…ster guard designed with numbers

Every Actions run at 970fef4 -- both triggers of work-lab-gate plus the production gates -- completed
success, so ten records finally carry a verifiedCommit that means what it says. 970fef4 contains all ten
records, and the ancestry and ledger-present checks are the stamper's own, not my reading.

ERR-207 records the check-ignore probe trap and the off-by-one in the first version of my own guard for
it. Three register rows: the sidecar 405 byte-level proof, the probe-shape rule, and
REGISTER-PATH-GUARD-20261008 carrying the measured 61-unresolved breakdown and the design decision
(declarations live inside the row they excuse, because a code-side map goes stale before the prose does and
hides the exemption from the reader). That row also states a gate hole I found while designing it: the
refs>0 requirement and the reference floor only apply to the default widened scan, so a --index run can
still pass on an extractor that matched nothing -- the register is bootstrap step 6 and belongs in the
navigation set.
…ich is what a broken extractor looks like

The capability floor I added this morning only applied to the default widened scan, so
`--index <one prose page>` printed refs=0 broken=0 and exited 0. A page legitimately making no tree claim
stays fine inside the widened scan, where the other surfaces carry the floor; a whole run that examined
nothing is not a pass. Found while designing the register's own guard, which is exactly the kind of hole
that design work is for.
The ledger requires a non-zero exit_code, and the vacuity I am recording is precisely a run that exited 0
after judging nothing. The record now says which number is which: the defect's code was 0, the 1 on file is
the refusal measured after the fix on the same command. The REGISTER-PATH-GUARD row is closed halfway the
same way -- the --index hole it named is fixed, the in-row declarations, the nine genuine corrections, the
280 floor and pulling the register into the navigation set remain unimplemented and stay written as such.
…he head is green

Three runs exist at that SHA -- work-lab-gate on push and on pull_request, plus the production gates --
and all three concluded success, which is the condition ERR-205 itself set after one trigger lied about
the head last time.

Measured remaining binding debt, stated rather than implied by the stamping: 185 PASS records, 94 with no
fixedCommit and 91 with no verifiedCommit, and the triage reports bindableByBirth=0 for that population.
Those are pre-lifecycle records, so closing them is an owner/evidence decision, not a command I can run.
…g its own dead names in the row

Widening the reference gate to docs/current stopped one short of the surface the authority chain actually
tells a reader to navigate first: the live open-task register sat outside the scanned root with 65
unresolved references out of 330, while the default run exited 0. It joins as EXTRA_SURFACES under the
strict refs>0 rule and its own measured floor (280 against 333 observed).

Exemptions are declared in the row that needs them -- `[no-tree-claim <CODE> ref=<name>]` over nine reason
codes -- rather than in a checker-side map or a sidecar, because a map keyed by row ID goes stale before the
prose does (the register gained ~13 dated rows today) and hides the excuse from the reader; the string-keyed
global allowlist already cost this repo a hidden stale README entry. Every declaration is re-falsified each
run: a declared name that resolves in the tracked tree fails, a name the surface never asked about fails,
an unknown code fails, and DELETED / NEVER_EXISTED are settled against git history. That last check needed
three corrections while being written: --diff-filter=D misses the R100/R081 renames, a row names the tail of
a path so the suffix form must be tried, and a directory is proven by its contents -- `-- '*web/'` finds
nothing while `-- '*web/*'` finds five commits.

Reconciled the 65: 14 literals rewritten to unique tracked successors (each proved with git ls-files before
use, including the two archived-AGENTS.md and the sync_hermes_workflow_assets.py real home), 40 declarations
covering 50 occurrences, plus two shape rules (an elided `…` and any backslash-containing literal are no
longer read as paths, which retires the regex `/\bCPU\b/` false positive). Measured: targets=39 refs=690
broken=0 declared_in_row=50, 28 tests in the ci module pass including four planted-failure cases.
active-authority-index.md's claim that the gate covers "four current surfaces" is corrected in the same
commit, and the design row's own quoted syntax no longer matches the extractor it describes.
…uard that measures it

The row now states what was built rather than what was designed, including two self-inflicted traps found
on the way: the design prose quoted the declaration syntax in its matching form, so the guard flagged its own
documentation until the example moved to placeholder spelling, and my apply script's "[no-tree-claim"
substring check reported that prose as an existing declaration. Reconciled totals: register refs 350,
broken 0, declared in-row 50 across 40 tokens; the whole scan reads targets=39 refs=691 broken=0.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant