From a70b3d8c0b6e04619bf0c0a60a9c83978a2a7584 Mon Sep 17 00:00:00 2001 From: Sollan Systems Date: Sat, 25 Jul 2026 23:06:28 -0400 Subject: [PATCH 1/2] docs(scoreboard): correct the Spec Kit row, refresh stars, note the PRPs rename MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Spec Kit ships a first-party workflow engine our fixture never reached. At the pinned commit 3f7392a, src/specify_cli/workflows/base.py:28 defines a RunStatus enum (created/running/paused/completed/failed/aborted) that engine.py:484 atomically persists to .specify/workflows/runs/RUN_ID/state.json alongside an append-only log.jsonl. A machine reading that file can tell a failed run from a completed one. That falsifies the generalization this board drew in two places: the Spec Kit row's "no machine-readable done flag", and the field-summary's claim that "failed, and here is the machine-readable reason" is never expressible. Both are corrected in place and dated, with the narrower finding stated instead: what state.json gives you is a status a run writes about itself, not evidence a third party can re-check that the completed was earned. The correction is also a live demonstration of the method's limit — a fixture built from documentation cannot see machinery the documentation does not surface. The row now says so and invites the same treatment for any other row. Also refreshes star counts (608,137 combined, was 599,034 on 2026-07-09) and notes that Wirasm/PRPs-agentic-eng now resolves to Wirasm/prp. --- docs/gap-reports/scoreboard.md | 59 ++++++++++++++++++++++++++-------- 1 file changed, 46 insertions(+), 13 deletions(-) diff --git a/docs/gap-reports/scoreboard.md b/docs/gap-reports/scoreboard.md index 1d0463a..5dd2f7b 100644 --- a/docs/gap-reports/scoreboard.md +++ b/docs/gap-reports/scoreboard.md @@ -7,7 +7,9 @@ > pinned to the harness commit named in its row. No version-general claims > about any project are made or implied: every statement is checkable against > one directory and one SHA. -> **Date:** 2026-07-09. +> **Date:** 2026-07-09; corrected 2026-07-25 (Spec Kit workflow-engine fairness +> note, refreshed star counts, PRPs repository rename). Corrections are made in +> place and dated, never silently. > **Reproduce any row:** `python3 -m loop inspect examples/` ## What this measures — and what it does not @@ -73,18 +75,19 @@ answer to the question this scoreboard asks.† | Harness | Stars· | Pinned | Fixture | Score | Verdict | Terminals | |---|---:|---|---|---:|---|---:| | *(calibration)* loop-engineer contract | — | this repo | [`flaky-test-triage`](../../examples/flaky-test-triage/) | **90** | strong | 7/7 | -| Superpowers ([obra/superpowers](https://github.com/obra/superpowers)) | 251k | fixture× | [`superpowers-run`](../../examples/superpowers-run/) | **12** | weak | 0/7 | -| Spec Kit ([github/spec-kit](https://github.com/github/spec-kit)) | 119k | `3f7392a` | [`spec-kit-run`](../../examples/spec-kit-run/) | **12** | weak | 0/7 | +| Superpowers ([obra/superpowers](https://github.com/obra/superpowers)) | 261k | fixture× | [`superpowers-run`](../../examples/superpowers-run/) | **12** | weak | 0/7 | +| Spec Kit ([github/spec-kit](https://github.com/github/spec-kit)) | 124k | `3f7392a` | [`spec-kit-run`](../../examples/spec-kit-run/) | **12** | weak | 0/7 | | CCPM ([automazeio/ccpm](https://github.com/automazeio/ccpm)) | 8.3k | `7d7e462` | [`ccpm-run`](../../examples/ccpm-run/) | **12** | weak | 0/7 | -| BMAD-METHOD ([bmad-code-org/BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD)) | 50k | `49069b8` | [`bmad-run`](../../examples/bmad-run/) | **0** | weak | 0/7 | +| BMAD-METHOD ([bmad-code-org/BMAD-METHOD](https://github.com/bmad-code-org/BMAD-METHOD)) | 51k | `49069b8` | [`bmad-run`](../../examples/bmad-run/) | **0** | weak | 0/7 | | Task Master ([eyaltoledano/claude-task-master](https://github.com/eyaltoledano/claude-task-master)) | 28k | `c0c98d3` | [`task-master-run`](../../examples/task-master-run/) | **0** | weak | 0/7 | -| OpenSpec ([Fission-AI/OpenSpec](https://github.com/Fission-AI/OpenSpec)) | 60k | `93e27a7` | [`openspec-run`](../../examples/openspec-run/) | **0** | weak | 0/7 | -| ruflo ([ruvnet/ruflo](https://github.com/ruvnet/ruflo), né claude-flow) | 64k | `7ef4d4e` | [`ruflo-run`](../../examples/ruflo-run/) | **0** | weak | 0/7 | -| Agent OS ([buildermethods/agent-os](https://github.com/buildermethods/agent-os)) | 5.0k | `cae8e66` | [`agent-os-run`](../../examples/agent-os-run/) | **0** | weak | 0/7 | -| PRPs ([Wirasm/PRPs-agentic-eng](https://github.com/Wirasm/PRPs-agentic-eng)) | 2.2k | `ada2f5b` | [`prp-run`](../../examples/prp-run/) | **0** | weak | 0/7 | +| OpenSpec ([Fission-AI/OpenSpec](https://github.com/Fission-AI/OpenSpec)) | 63k | `93e27a7` | [`openspec-run`](../../examples/openspec-run/) | **0** | weak | 0/7 | +| ruflo ([ruvnet/ruflo](https://github.com/ruvnet/ruflo), né claude-flow) | 66k | `7ef4d4e` | [`ruflo-run`](../../examples/ruflo-run/) | **0** | weak | 0/7 | +| Agent OS ([buildermethods/agent-os](https://github.com/buildermethods/agent-os)) | 5.1k | `cae8e66` | [`agent-os-run`](../../examples/agent-os-run/) | **0** | weak | 0/7 | +| PRPs ([Wirasm/prp](https://github.com/Wirasm/prp), formerly PRPs-agentic-eng) | 2.2k | `ada2f5b` | [`prp-run`](../../examples/prp-run/) | **0** | weak | 0/7 | | *(calibration)* an unstructured agent loop | — | this repo | [`naive-loop`](../../examples/naive-loop/) | **0** | weak | 0/7 | -·Stars as reported by the GitHub API on 2026-07-09, for scale only. +·Stars as reported by the GitHub API on 2026-07-25 (608,137 combined), for +scale only. ×The Superpowers row predates this scoreboard's SHA-pinning: its fixture was built from the layout documented in [`superpowers.md`](superpowers.md) (evaluated 2026-07-08) and carries no pinned commit. @@ -132,8 +135,34 @@ Spec-driven discipline: a numbered `specs/NNN-slug/` folder with a constitution, requirements checklist, and phased tasks. It proves the front of the loop well — verifiable Success Criteria, a Constitution Check, a requirements-quality checklist — the surface the contract layer composes with. -What it structurally cannot prove is how the work ended: no typed terminal, no -held-out gate, no machine-readable done flag. +What the *fixture's* layout cannot prove is how the work ended: no typed +terminal, no held-out gate, no evidence a third party can re-check. + +> **Correction (2026-07-25) — the strongest fairness note on this board.** +> Spec Kit ships a first-party workflow engine that our fixture does not model, +> and it *does* persist a typed run status. At the pinned commit `3f7392a`, +> `src/specify_cli/workflows/base.py:28` defines +> `class RunStatus(str, Enum)` with `created / running / paused / completed / +> failed / aborted`, and `workflows/engine.py` atomically writes it to +> `.specify/workflows/runs//state.json` (line 484) alongside an +> append-only `log.jsonl` (line 560), transitioning to `FAILED` on step failure +> and `PAUSED` on interruption. **A machine reading that file can tell a failed +> run from a completed one.** That is more than eight of the nine layouts here +> offer, and the row's 0/7 should be read with it in view. +> +> The fixture models Spec Kit's documented slash-command flow, which is the +> surface most users drive, so the score stands for what it measures — but the +> generalization this scoreboard drew from it does not. What `state.json` gives +> you is a status a run writes *about itself*. What it does not give you is +> evidence a third party can re-check: nothing in it lets a reader verify the +> `completed` was earned. That narrower gap is the real finding; the broader +> one was wrong. +> +> This is also a live demonstration of the method's own limit. A fixture built +> from documentation cannot see machinery the documentation does not surface. +> If another row has a comparable first-party mechanism we missed, name the file +> and the pinned commit and it gets the same treatment: re-run, corrected in +> place, in public. Fairness notes: `/speckit.implement` runs a real checkbox PASS/FAIL gate and halts on FAIL — but it is a natural-language instruction that writes no file. @@ -272,12 +301,16 @@ success-criteria signal reads. ## What the whole field is missing -Across nine popular harnesses — ~590k stars of them — the same three +Across nine popular harnesses — ~610k stars of them — the same three structural gaps repeat, independent of methodology, maturity, or how much verification machinery each ships: 1. **No typed terminal taxonomy (0/7, every row).** "Done" is always - expressible; *"failed, and here is the machine-readable reason"* never is. + expressible. A machine-readable *failure* almost never is — the one + exception on this board is Spec Kit's workflow engine, which persists a + typed `failed`/`aborted` status (see the correction in its row). Even there + the status is self-written: no row on this board pairs a terminal with + evidence a third party can re-check. 2. **No held-out gate.** Every verification surface that exists is run by the same agent that claims success, and most persist nothing. Nothing prevents the claim from being wrong — and nothing measures how often it is (no row From 08d5df8c540570a3551d3c87f7becf46e2a6d2e5 Mon Sep 17 00:00:00 2001 From: Sollan Systems Date: Sat, 25 Jul 2026 23:06:49 -0400 Subject: [PATCH 2/2] =?UTF-8?q?docs(adr):=20ADR=200002=20=E2=80=94=20attes?= =?UTF-8?q?t=20the=20verdict=20in=20CI,=20never=20sign=20in=20the=20kernel?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Discharges the decision ADR 0001 deferred in its Non-goals ("Those require separate decisions"). ADR 0001 remains Accepted and unamended; its non-equivalence statement is restated here as a standing limit. The seam: the kernel disposes on contents, the CI lane notarizes context, neither is the other. The kernel may hash. It must never sign, never establish authenticity, and never read an environment variable. Decisions: - loop verdict emits a loop-engineer/verdict@1 PREDICATE BODY, not an in-toto Statement. actions/attest builds the Statement from subject + predicate-type + predicate; handing it a full Statement would nest a Statement inside a predicate and write a malformed attestation to a permanent public log. - Subject is the chain head alone. The source commit stays out of the predicate: the kernel reads no environment today, and the Fulcio cert already binds the source repository digest more strongly than the kernel could assert it. - predicateType is urn:loop-engineer:verdict:1 — no vendor host, no org, no repo name, because the URI is immutable in a public log and a rename must not be able to orphan it. - Authenticity (gh attestation verify) and semantic agreement (loop verdict --compare) are separate checks; --compare reports signature_checked: false on every path and no flag may flip it. - Ships as two slices: 4a emission, 4b consumption. 4b changes what the hard gate accepts and earns its own review. - Human review required on loop/**, action.yml, and .github/workflows/**. Rejected: SLSA VSA (its verificationResult is a PASSED/FAILED bit and verifiedLevels is required — it would collapse seven typed terminals and imply a build-level evaluation we did not perform), in-toto SVR (pre-1.0, stringly typed properties), in-toto test-result, in-kernel cryptography, local DSSE/HMAC, attesting on pull requests, and a multi-subject Statement. Standing limits are stated plainly rather than footnoted, including the one the adversarial review corrected: opt-in anchor auto-resolution detects an unattested rewrite for exactly one run. An actor who can land a commit lets CI mint a genuine attestation over the rewritten chain and rolls the anchor forward. Opt-in changes consent, not that property. The Slice 4a plan lands alongside: nine TDD tasks covering the schema, the projection, the verified-evidence digest set, CLI wiring, the boundary made mechanical, the action step, the repo's own attestation job, normative §23, and the governance change. Five open items are flagged for resolution during execution rather than guessed here. Baseline measured for the plan at HEAD 0025acc in a fresh worktree: 1305 passed / 18 skipped with pyyaml+jsonschema+pytest. --- docs/adr/0002-ci-attested-verdict.md | 270 ++++ .../2026-07-25-slice4a-verdict-emission.md | 1232 +++++++++++++++++ 2 files changed, 1502 insertions(+) create mode 100644 docs/adr/0002-ci-attested-verdict.md create mode 100644 docs/superpowers/plans/2026-07-25-slice4a-verdict-emission.md diff --git a/docs/adr/0002-ci-attested-verdict.md b/docs/adr/0002-ci-attested-verdict.md new file mode 100644 index 0000000..d822945 --- /dev/null +++ b/docs/adr/0002-ci-attested-verdict.md @@ -0,0 +1,270 @@ +# ADR 0002: Attest the verdict in CI; never sign in the kernel + +- **Status:** Accepted +- **Date:** 2026-07-25 +- **Decision owners:** Loop Engineer maintainers +- **Relationship to ADR 0001:** Discharges the decision deferred in ADR 0001 + § Non-goals. ADR 0001 remains Accepted and unamended; its non-equivalence + statement is restated below as a standing limit. + +## Context + +Slice 1 gave every run a SHA-256 hash chain with an externally anchorable head, +and said plainly what it does not buy: without an external anchor, a worker that +can rewrite `.loop/` can recompute the chain over the rewritten history, and the +result verifies clean. `loop doctor --expect-chain-head` is the control for +that, and today the expected value is a number a human remembers. Nothing +carries run N's head to run N+1. + +What is missing is not more hashing. It is a record of *head-at-time-T* that is +durable, externally timestamped, and not mintable at will by the process under +judgment. GitHub Actions can produce one: an OIDC token exchanged for a +short-lived Fulcio certificate binds the repository, the commit, the workflow +file and ref, the trigger, and the runner environment, and the resulting bundle +is logged publicly and append-only. + +Taking that capability introduces a real hazard to the architecture. Signing +requires key material, a network call, a trust root, and a vendor identity — +every one of which ADR 0001 keeps out of the portable layer. The decision below +is primarily about where the boundary runs. + +## Decision + +### 1. The kernel projects a verdict; it never signs one + +A new verb, `loop verdict `, emits a `loop-engineer/verdict@1` +document: a projection over `doctor_report()`, the terminal record, and the +chain-bound evidence digests that satisfy the strict verified-evidence bar. It +is a pure function of local state, built with `json` and `hashlib`, reusing +`loop.chain.canonical_json` and `_strict_evidence_failure` rather than +re-deriving either. + +The kernel emits the **predicate body only**. It does not construct the in-toto +Statement. `actions/attest` builds the Statement from `subject-*` + +`predicate-type` + `predicate-path`; handing it a complete Statement would nest +a Statement inside a predicate and write a malformed attestation to a permanent +public log. + +### 2. The signer lane owns the envelope, the subject, and all cryptography + +`action.yml` writes the predicate to `$RUNNER_TEMP`, passes it to +`actions/attest` with `subject-name: loop-chain-head` and +`subject-digest: sha256:`, and declares the resulting attestation URL as +an action output. The calling workflow's job supplies `id-token: write` and +`attestations: write`; a composite action cannot declare its own `permissions:`, +so the step must fail with a legible message when they are absent rather than +surfacing a raw OIDC 403. + +The subject is the chain head alone. The source commit is **not** carried in the +predicate: the kernel reads no environment variable today, and making it read +`GITHUB_SHA` would put a vendor identifier inside the portable layer. The Fulcio +certificate already binds the source repository digest more strongly than the +kernel could assert it. + +### 3. Predicate identity is `urn:loop-engineer:verdict:1` + +Backed by `schemas/verdict.schema.json` with the matching `$id`, following the +convention every existing schema uses. The predicateType is written immutably +into a public log and cannot change without a major version bump, so it names no +vendor host, no organization, and no repository — a rename must not be able to +orphan it. + +### 4. Authenticity and agreement are separate checks, in that order + +`gh attestation verify` establishes *this was signed by that workflow at that +time*. `loop verdict --compare ` establishes *these claims +agree with local state* — attested head equals the local chain head, attested +evidence digests resolve and match, attested terminal state equals the local +terminal record. + +`--compare` reports `signature_checked: false` on every code path, and no flag +may flip it. It accepts a bare `verdict@1` predicate only; unwrapping any +vendor envelope happens outside `loop/`, and a `gh`-shaped envelope is rejected +with a typed error rather than best-effort parsed. + +Disagreement is a typed failure. An absent, unverifiable, or ignored attestation +is never a promotion — consistent with in-toto's monotonicity principle, that +ignoring an attestation must never convert a denial into an approval. + +### 5. The work ships as two slices + +**4a — emission.** Schema, `loop verdict`, the `attest` action input (default +off), a repo CI job restricted to `push` on the default branch, docs, and the +purity tests. Write-only and reversible by ignoring it. + +**4b — consumption.** `--compare`, anchor auto-resolution, and the signer-trust +policy. This changes what the hard gate *accepts*, which is a different risk +class and belongs in its own review. + +Anchor auto-resolution ships opt-in and default off. When enabled, the resolve +step pins `--signer-workflow`, requires a `push` trigger, and an +explicitly-supplied `expect-chain-head` always wins over a resolved one. The +signer-identity and trigger policy is evaluated by a tested kernel-side function +over already-extracted JSON — still no crypto and no network — because a trust +policy written in YAML fails silently and open. + +### 6. Governance changes with it + +Human review is required on `loop/**`, `action.yml`, and `.github/workflows/**`. +Everything else stays autonomous. This is the only control that addresses the +largest hazard in the design, and it is a governance decision, not a code one. + +## Governing rule + +ADR 0001's rule stands: **agents propose; the kernel disposes.** This ADR adds +the seam that makes it survive contact with cryptography: + +**The kernel disposes on contents; the CI lane notarizes context; neither is the +other.** The kernel may hash. It must never sign, and it must never establish +authenticity. + +## Immediate consequences + +1. `loop/verdict.py` imports nothing outside the standard library and `loop.*`. + No module under `loop/` may reference a signing stack, key material, or an + OIDC token; no module under `loop/` reads an environment variable. Both are + asserted mechanically, not intended. +2. The predicate carries digests, enums, and issue **codes** only — never + free-text detail strings, workspace paths, or task titles. Everything in it + is public, append-only, and permanent; adding a field is a one-way door. + An allowlist test enforces the field set. +3. `schemas/verdict.schema.json` is not a contract artifact and does not enter + `_SCHEMA_FILES`; it follows the `evidence.py` precedent of its own schema id + and loader. +4. Section 23 is **appended** to `reference/repo-os-contract.md`. No new file + enters `reference/` — `evals/cases/structural.json` pins that list at eight. +5. Attestation happens on `push` to the default branch only. The PR gate remains + the existing, unattested `loop doctor` hard gate. Attestation is a post-merge + notarization of what landed, not a pre-merge approval. +6. No new runtime dependency. This is the constraint that rules out in-kernel + signature verification, which would require X.509 chain validation, a TUF + trust root, and Rekor inclusion proofs. + +## Standing limits + +ADR 0001's non-equivalence statement stands unchanged: **this does not make +structural evidence equivalent to cryptographic attestation.** The signature +attests context — which repository, which commit, which workflow, which trigger, +which runner, at what time. It never attests correctness. A signed verdict over +a weak gate is a signed weak gate. + +**The worker can edit the verifier.** An agent with ordinary merge rights can +land a change to `loop/**`, `action.yml`, or the gate workflow and then mint a +perfectly genuine attestation for the loosened gate. Decision 6 is the control; +signing is not. `gh attestation verify` says this outright: only the certificate +and the verified timestamps are unforgeable by the workflow that produced the +attestation — the predicate is user-controllable metadata. + +**Auto-resolution detects an unattested rewrite for exactly one run.** An actor +who can land a commit lets CI run once, which mints a genuine attestation over +the rewritten chain, and the next run resolves forward to it. Opt-in changes +operator consent; it does not change that property. The value of decision 5's +resolve step is that it makes the anchoring assumption explicit and revocable, +not that it eliminates self-satisfaction. + +**Rekor makes independent historical audit possible; nothing here performs it.** +That is a property of Sigstore's public log, available with or without this +work. The resolve step reads the most recent matching attestation and looks no +further back. + +**Fabricated history remains indistinguishable.** `compute_event_hash` draws its +preimage entirely from the record's own fields, so a chain fabricated wholesale +at commit-authoring time is byte-valid. The chain proves order and +non-tampering relative to an anchor; it does not prove the events happened when +claimed. + +**Self-hosted runners void the guarantee.** The runner-environment claim is +self-reported. `--deny-self-hosted-runners` is mandatory on the verify side, not +advisory. + +**Fork pull requests cannot be attested.** A fork's `pull_request` run receives +no OIDC token. Same-repo branches — how this repo's own agents push — do receive +one, so "forks cannot sign" is not a mitigation available to us. + +**Most attestations are never verified.** An attestation nothing gates on is +decoration. If 4b is never enabled, this work ships a publication surface, not a +gate, and should be described that way. + +### Limits that become tests + +Following the project's habit of shipping defeat conditions as passing tests: +`loop verdict` refuses, typed, on a workspace with no event store or terminal +record; `--compare` reports `signature_checked: false` across every branch; +head disagreement fails with a typed code and agreement passes as the negative +control; predicate conformance is machine-pinned; the field allowlist holds; +import purity and the no-signing / no-environment assertions hold repo-wide; and +an insider-mint case — two attestations, both from the pinned signer workflow, +the second covering a chain rewritten between runs — is asserted as an accepted, +documented limitation rather than left as prose. + +## Non-goals + +No signature verification, key material, network access, or trust root inside +`loop/`, ever. No local DSSE or HMAC signing — that remains permanently +rejected. No per-event or per-evidence attestations; one verdict per run. No +registry push or storage records. No GitHub Enterprise Server or private +Sigstore instance path. No attestation lifecycle or retention management. No +policy engine deciding who may attest beyond documented `--signer-workflow` +guidance. + +## Rejected alternatives + +**SLSA Verification Summary Attestation.** Its `verificationResult` is a +PASSED/FAILED enum and `verifiedLevels` is required and typed to SLSA results. +Adopting it would collapse seven typed terminal states to a bit and imply a +build-level evaluation we did not perform. The mapping from `verdict@1` to a VSA +or in-toto SVR predicate is total, so a second attestation in a standard shape +can be added later without a `verdict@1` version bump. + +**in-toto SVR.** Purpose-built for policy verdicts, but its payload is a list of +property strings; 64-hex digests would become stringly-typed tokens, which is +the lossy re-encoding this project refuses elsewhere. It is also pre-1.0, and +pinning our hard gate to someone else's pre-1.0 version risk is not a trade we +want. + +**in-toto Test Result.** Models individually-named tests with a single +pass/warn/fail. That is CI-runner output, not a multi-criterion contract verdict. + +**Cryptography in the kernel.** Requires a trust root, X.509 validation, and +network access; violates the dependency constraint and the boundary this ADR +exists to draw. + +**A `github.com` or owned-domain predicateType URI.** The first bakes a vendor +host, organization, and repository name into the portable contract's type +identifier forever, and a rename orphans it permanently. The second creates an +indefinite domain-renewal obligation whose failure mode is a dead type URI in an +append-only log. + +**Attesting on pull requests.** An unreviewed branch would mint a signed verdict +under the repository's own identity before any review — the strongest form of the +confused-deputy problem this ADR is trying not to create. + +**A multi-subject Statement (chain head and commit).** Not expressible through +`actions/attest`, which accepts at most one of `subject-path` / `subject-digest` +/ `subject-checksums`. The commit is cert-bound regardless, and asserting a +weaker second copy invites a reader to trust it. + +## Open verification items + +These must be established before the Slice 4 plan relies on them; none changes +the decisions above. + +1. **Subject digest algorithm.** The chain head is a SHA-256 over a synthesized + event preimage, not over retrievable bytes, so the `sha256` DigestSet key + licenses an inference — fetch, re-hash, compare — that is false here. A + namespaced key (the `gitCommit` / `dirHash` precedent) is correct in in-toto + terms, but `actions/attest`'s `subject-digest` input may accept only + `sha256:`. Resolve by experiment; if the input constrains us, the subject + name and the §23 preimage spec must carry the disambiguation instead. +2. **Attestation retention and index durability.** Rekor log permanence and + GitHub's attestation index are different properties. Auto-resolution queries + the latter. Establish the retention window from GitHub's own documentation + before enabling 4b; do not assume it from Rekor. +3. **`--signer-digest` as a hard pin.** If it invalidates on every legitimate + workflow edit it is unusable in practice. Needs a two-revision experiment. +4. **Fork-PR OIDC unavailability** is corroborated but not smoke-tested. Decision + 5 does not depend on it; any future design that attests on PRs must test it + first. +5. **`create-storage-record` and `push-to-registry` defaults.** Pass both + explicitly as false so the two-permission claim is true by construction rather + than by assumed default. diff --git a/docs/superpowers/plans/2026-07-25-slice4a-verdict-emission.md b/docs/superpowers/plans/2026-07-25-slice4a-verdict-emission.md new file mode 100644 index 0000000..270d334 --- /dev/null +++ b/docs/superpowers/plans/2026-07-25-slice4a-verdict-emission.md @@ -0,0 +1,1232 @@ +# Slice 4a — `verdict@1` Emission Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Give the kernel a `loop verdict` verb that projects local run state into a `loop-engineer/verdict@1` predicate document, and give `action.yml` an opt-in step that hands that predicate to `actions/attest` for keyless signing — so a run's chain head becomes a signed, externally timestamped statement instead of a number an operator remembers. + +**Architecture:** A new pure leaf module `loop/verdict.py` composes `doctor_report()`, the terminal record, and the chain-bound evidence digests that pass the existing strict verified-evidence bar into one canonical JSON document. It emits the **predicate body only** — `actions/attest` constructs the in-toto Statement from `subject-*` + `predicate-type` + `predicate-path`, so the kernel never builds an envelope, never signs, never verifies a signature, and never reads an environment variable. `action.yml` gains an `attest` input, default off. + +**Tech Stack:** Python 3.10+ stdlib only (`json`, `hashlib`, `importlib.metadata`). No new runtime dependencies. Tests via pytest under `uv run`. CI via `actions/attest` in a composite-action step. + +**Baseline (fresh worktree, main @ `0025acc` / v0.11.0):** **1305 passed / 18 skipped** with `--with pyyaml --with jsonschema --with pytest`. All zero-regression claims are against this number. + +> **Baseline measurement trap.** Running the suite via `uv run --project ` installs the wheel, which materializes `loop/_bundle/` and makes `scripts/test_resources.py::test_repo_checkout_resolves_to_repo_dirs` fail *correctly*. Measure from inside the worktree (`bash -c 'cd && uv run --with … pytest -q … scripts'`), never with `--project`. + +**Source of decisions:** `docs/adr/0002-ci-attested-verdict.md`. Where this plan and the ADR disagree, the ADR wins. This plan implements **4a only** — emission. `--compare`, anchor auto-resolution, and signer-trust policy are Slice 4b and are explicitly out of scope. + +## Global Constraints + +- **Zero new runtime dependencies.** `loop/` imports stdlib only; `pyproject.toml` `[project.optional-dependencies]` stays exactly `yaml`, `schemas`, `dev`. +- **The kernel never signs and never establishes authenticity.** No module under `loop/` may reference a signing stack (sigstore, cosign, fulcio, rekor, DSSE), key material, or an OIDC token. +- **The kernel reads no environment variable.** `grep -rn "environ\|getenv" --include=*.py loop/` returns zero matches today and must still return zero after this slice. +- **The kernel emits a predicate, never a Statement.** No `_type`, no `subject`, no `predicateType` key is produced by `loop/`. +- **Predicate content is digests, enums, and issue codes only.** No free-text detail strings, no workspace paths, no task titles. `run_id` is the single operator-controlled string and is allowlisted deliberately (Task 5). +- **`schemas/verdict.schema.json` is not a contract artifact.** It does not enter `SCHEMA_IDS` (`loop/contract.py:34-39`) — that tuple is manifest/state/tasks/terminal only. Follow the `evidence.py` precedent: own schema-id constant, own loader. +- **No new file in `reference/`.** `evals/cases/structural.json` pins `reference_filenames` at 8. Section 23 is **appended** to `reference/repo-os-contract.md`. +- **No version bumps.** Tasks 1–8 land as one feature PR. Version surfaces (`pyproject.toml`, `.claude-plugin/plugin.json`, README, `scripts/test_docs_version.py`, CHANGELOG) move only in a separate release-cut PR. +- **`main` is protected** — PR required, force-push and deletion blocked, no bypass actors. All work lands via PR. +- Repo env quirks: no system pytest. The Bash deny-list blocks `rm`, bare `cd`, `VAR=` prefixes, `timeout`, `printf`, `source` — use `git -C `, absolute paths, and `bash -c '…'`. + +**Canonical test command** (run from the repo root): + +```bash +uv run --with pyyaml --with jsonschema --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts +``` + +--- + +## Existing interfaces this slice consumes + +Read these before starting. Every signature is verbatim from HEAD `0025acc`. + +| Symbol | Location | Signature / behavior | +|---|---|---| +| `doctor_report` | `loop/contract.py:1093` | `(target, *, mode=None, expect_chain_head=None) -> dict`. Returns `{**validate_contract(target, mode=mode), "event_store": …, "issues": [...], "ok": bool}`. Report keys include `ok`, `issues`, `lifecycle`, `schemas_checked`, `validation_mode`. | +| `_strict_evidence_failure` | `loop/contract.py:846` | `(entry, paths, bound) -> str \| None`. **The ONE definition of the verified-evidence bar.** `None` means the entry may back `Succeeded`; a string names which of four sub-checks failed. Reuse it — never restate the bar. | +| `canonical_json` | `loop/chain.py:27` | `(value) -> str`. `json.dumps(sort_keys=True, separators=(",",":"), ensure_ascii=False, allow_nan=False)`. Raises `ChainHashError`. | +| `ChainHashError` | `loop/chain.py:23` | `ValueError` subclass. | +| `EVIDENCE_SCHEMA_ID` | `loop/evidence.py:35` | `"loop-engineer/evidence@1"` — the naming convention to mirror. | +| `verify_bundle_is_green` | `loop/evidence.py:42` | `(bundle) -> bool`. The repo's ONE green-marker rule. | +| `_load_evidence_schema` | `loop/evidence.py:60` | The own-loader pattern to mirror for verdict. | +| `schemas_dir` | `loop/_resources.py:34` | `() -> Path`. Bundle-first, repo-relative fallback. | +| `resolve_loop_paths` | `loop/paths.py` | `(target) -> LoopPaths` with `.workspace`, `.loop_dir`. | +| `_COMMANDS` / `_READ_COMMANDS` / `_USAGE` / `_HELP` | `loop/__main__.py:15,19,21,23` | CLI wiring. `metrics` (line 26 of `_HELP`) is the precedent for a flag-bearing command. | + +--- + +## File structure + +| File | Change | Responsibility | +|---|---|---| +| `schemas/verdict.schema.json` | **create** | The `loop-engineer/verdict@1` predicate shape. Sole identity of the artifact. | +| `loop/verdict.py` | **create** | Pure projection: `build_verdict()`, `VERDICT_SCHEMA_ID`, `PREDICATE_TYPE`, `VerdictError`. stdlib + `loop.*` only. | +| `loop/__main__.py` | modify | Register `verdict` in `_COMMANDS`, `_READ_COMMANDS`, `_USAGE`, `_HELP`; one dispatch block. | +| `action.yml` | modify | New `attest` input (default `false`), `predicate-path` write step, `actions/attest` step, `attestation-url` output. | +| `.github/workflows/attest.yml` | **create** | Push-to-default-branch job that attests a tracked `examples/*` contract. | +| `reference/repo-os-contract.md` | modify | Append §23 (normative `verdict@1` + conformance vectors). | +| `CHANGELOG.md` | modify | Unreleased entry. | +| `scripts/test_verdict.py` | **create** | Projection, refusal, and conformance tests. | +| `scripts/test_verdict_purity.py` | **create** | The boundary made mechanical: no signing, no env, no Statement, field allowlist. | +| `scripts/test_verdict_cli.py` | **create** | CLI wiring, exit codes, stdout shape. | +| `.github/CODEOWNERS` | **create** | Human review on the three gate-defining paths (ADR 0002 decision 6). | + +--- + +## The predicate shape (binding) + +Every task builds toward exactly this document. `null` is legal for `chain.head`, `chain.sequence`, and `tool.version`; every other key is always present. + +```json +{ + "schema": "loop-engineer/verdict@1", + "run_id": "coverage-repair", + "tool": { "name": "loop-engineer", "version": "0.11.0" }, + "doctor": { + "ok": true, + "validation_mode": "jsonschema", + "issue_codes": [], + "schemas_checked": ["loop-engineer/manifest@1", "loop-engineer/state@1"] + }, + "chain": { + "head": "9f2c…64hex", + "sequence": 41, + "unchained_prefix": 0 + }, + "terminal": { + "state": "Succeeded", + "completion_policy": "all_required_verified_evidence", + "false_completion": false + }, + "evidence": [ + { + "digest": "a1b2…64hex", + "code_digest": "c3d4…64hex", + "policy_digest": "e5f6…64hex" + } + ] +} +``` + +**Why each field earns a slot in a permanent public log:** + +- `run_id` — the only operator-controlled string. Deliberately allowlisted so a verdict can be correlated to a run; documented in §23 as public, so run ids must not embed sensitive text. +- `doctor.issue_codes` — **codes only, sorted, de-duplicated.** Never `message`, never `path`. This is the field the allowlist test exists to protect. +- `chain.head` — the subject the signer lane binds. `null` when the store has no chained events. +- `terminal.completion_policy` — distinguishes `all_required` from `all_required_verified_evidence`; without it a reader cannot tell what `Succeeded` meant. +- `evidence[]` — chain-bound entries that passed `_strict_evidence_failure`. Digests only; no URIs, because a URI is a workspace path. + +--- + +## Task 1: The schema and the module skeleton + +**Files:** +- Create: `schemas/verdict.schema.json` +- Create: `loop/verdict.py` +- Test: `scripts/test_verdict.py` + +**Interfaces:** +- Consumes: `loop/_resources.py::schemas_dir()`. +- Produces: `VERDICT_SCHEMA_ID = "loop-engineer/verdict@1"`, `PREDICATE_TYPE = "urn:loop-engineer:verdict:1"`, `class VerdictError(ValueError)`, `_load_verdict_schema() -> dict`. + +- [ ] **Step 1: Write the failing test** + +```python +# scripts/test_verdict.py +import json + + +def test_schema_id_and_predicate_type_are_pinned(): + from loop.verdict import PREDICATE_TYPE, VERDICT_SCHEMA_ID + + assert VERDICT_SCHEMA_ID == "loop-engineer/verdict@1" + assert PREDICATE_TYPE == "urn:loop-engineer:verdict:1" + + +def test_schema_file_declares_the_matching_id(): + from loop._resources import schemas_dir + from loop.verdict import VERDICT_SCHEMA_ID + + schema = json.loads((schemas_dir() / "verdict.schema.json").read_text(encoding="utf-8")) + assert schema["$id"] == VERDICT_SCHEMA_ID + assert schema["$schema"] == "https://json-schema.org/draft/2020-12/schema" + + +def test_verdict_schema_is_not_a_contract_artifact(): + # SCHEMA_IDS is the contract-object tuple (manifest/state/tasks/terminal). + # verdict@1 is a projection, not a contract object, and must stay out of it. + from loop.contract import SCHEMA_IDS + from loop.verdict import VERDICT_SCHEMA_ID + + assert VERDICT_SCHEMA_ID not in SCHEMA_IDS +``` + +- [ ] **Step 2: Run test to verify it fails** + +Run: `uv run --with pyyaml --with jsonschema --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts/test_verdict.py -v` +Expected: FAIL with `ModuleNotFoundError: No module named 'loop.verdict'` + +- [ ] **Step 3: Write the schema** + +```json +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "$id": "loop-engineer/verdict@1", + "title": "Loop Engineer Verdict @1", + "description": "A CI-attestable projection of a loop run's doctor verdict, chain head, terminal record, and verified-evidence digests. Emitted as an in-toto predicate body; the signer constructs the Statement.", + "type": "object", + "additionalProperties": false, + "required": ["schema", "run_id", "tool", "doctor", "chain", "terminal", "evidence"], + "properties": { + "schema": { "const": "loop-engineer/verdict@1" }, + "run_id": { "type": "string", "minLength": 1, "maxLength": 256 }, + "tool": { + "type": "object", + "additionalProperties": false, + "required": ["name", "version"], + "properties": { + "name": { "const": "loop-engineer" }, + "version": { "type": ["string", "null"], "maxLength": 64 } + } + }, + "doctor": { + "type": "object", + "additionalProperties": false, + "required": ["ok", "validation_mode", "issue_codes", "schemas_checked"], + "properties": { + "ok": { "type": "boolean" }, + "validation_mode": { "type": "string", "maxLength": 64 }, + "issue_codes": { + "type": "array", + "items": { "type": "string", "pattern": "^[a-z0-9_]{1,64}$" } + }, + "schemas_checked": { + "type": "array", + "items": { "type": "string", "maxLength": 128 } + } + } + }, + "chain": { + "type": "object", + "additionalProperties": false, + "required": ["head", "sequence", "unchained_prefix"], + "properties": { + "head": { "type": ["string", "null"], "pattern": "^[0-9a-f]{64}$", "maxLength": 64 }, + "sequence": { "type": ["integer", "null"], "minimum": 0 }, + "unchained_prefix": { "type": "integer", "minimum": 0 } + } + }, + "terminal": { + "type": "object", + "additionalProperties": false, + "required": ["state", "completion_policy", "false_completion"], + "properties": { + "state": { + "enum": ["Succeeded", "FailedUnverifiable", "FailedBlocked", "FailedBudget", + "FailedSafety", "FailedSpecGap", "AbortedByHuman"] + }, + "completion_policy": { "type": ["string", "null"], "maxLength": 64 }, + "false_completion": { "type": "boolean" } + } + }, + "evidence": { + "type": "array", + "items": { + "type": "object", + "additionalProperties": false, + "required": ["digest", "code_digest", "policy_digest"], + "properties": { + "digest": { "type": "string", "pattern": "^[0-9a-f]{64}$", "maxLength": 64 }, + "code_digest": { "type": ["string", "null"], "pattern": "^[0-9a-f]{64}$", "maxLength": 64 }, + "policy_digest": { "type": ["string", "null"], "pattern": "^[0-9a-f]{64}$", "maxLength": 64 } + } + } + } + } +} +``` + +> `maxLength` accompanies every `pattern` deliberately. jsonschema's `pattern` is `re.search` semantics, so a bare `^[0-9a-f]{64}$` accepts a trailing newline — the exact hole closed in PR #89. + +- [ ] **Step 4: Write the module skeleton** + +```python +# loop/verdict.py +"""Project a loop run into a loop-engineer/verdict@1 predicate body. + +This module builds a document. It NEVER signs one, never verifies a signature, +never constructs an in-toto Statement, and never reads an environment variable. +The signer lane (action.yml -> actions/attest) owns the envelope, the subject, +and every cryptographic operation. See docs/adr/0002-ci-attested-verdict.md. +""" + +from __future__ import annotations + +import json +from typing import Any + +from ._resources import schemas_dir + +VERDICT_SCHEMA_ID = "loop-engineer/verdict@1" +PREDICATE_TYPE = "urn:loop-engineer:verdict:1" + + +class VerdictError(ValueError): + """A verdict cannot be projected from this workspace.""" + + +def _load_verdict_schema() -> dict[str, Any]: + return json.loads((schemas_dir() / "verdict.schema.json").read_text(encoding="utf-8")) +``` + +- [ ] **Step 5: Run tests to verify they pass** + +Run: `uv run --with pyyaml --with jsonschema --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts/test_verdict.py -v` +Expected: 3 passed + +- [ ] **Step 6: Commit** + +```bash +git add schemas/verdict.schema.json loop/verdict.py scripts/test_verdict.py +git commit -m "feat(verdict): add verdict@1 schema and module skeleton" +``` + +--- + +## Task 2: Project doctor, chain, and terminal + +**Files:** +- Modify: `loop/verdict.py` +- Test: `scripts/test_verdict.py` + +**Interfaces:** +- Consumes: `VerdictError` (Task 1); `loop.contract.doctor_report`; `loop.paths.resolve_loop_paths`. +- Produces: `build_verdict(target, *, mode=None) -> dict` — the full predicate minus `evidence`, which Task 3 fills. Callers rely on the exact key set in "The predicate shape (binding)". + +- [ ] **Step 1: Write the failing test** + +```python +# append to scripts/test_verdict.py +import pytest + + +def _scaffold_terminal(tmp_path, name, state="Succeeded", policy="all_required"): + """A doctor-clean scaffold advanced to a terminal record.""" + from loop.contract import scaffold_contract # existing scaffold entry point + + target = tmp_path / name + scaffold_contract(target) + terminal = { + "schema": "loop-engineer/terminal@1", + "state": state, + "iteration_id": 1, + "false_completion": False, + "completion_policy": {"mode": policy}, + "evidence": [], + } + (target / ".loop" / "terminal_state.json").write_text( + json.dumps(terminal), encoding="utf-8") + return target + + +def test_build_verdict_projects_doctor_terminal_and_chain(tmp_path): + from loop.verdict import build_verdict + + target = _scaffold_terminal(tmp_path, "projected") + verdict = build_verdict(target) + + assert verdict["schema"] == "loop-engineer/verdict@1" + assert set(verdict) == {"schema", "run_id", "tool", "doctor", "chain", "terminal", "evidence"} + assert verdict["terminal"]["state"] == "Succeeded" + assert verdict["terminal"]["completion_policy"] == "all_required" + assert verdict["terminal"]["false_completion"] is False + assert set(verdict["doctor"]) == {"ok", "validation_mode", "issue_codes", "schemas_checked"} + assert set(verdict["chain"]) == {"head", "sequence", "unchained_prefix"} + assert verdict["tool"]["name"] == "loop-engineer" + + +def test_issue_codes_are_sorted_deduplicated_and_carry_no_detail(tmp_path): + from loop.verdict import build_verdict + + target = _scaffold_terminal(tmp_path, "with-issues") + (target / "scripts" / "verify-fast").unlink() + (target / "scripts" / "verify-full").unlink() + + codes = build_verdict(target)["doctor"]["issue_codes"] + assert codes == sorted(set(codes)) + assert all(isinstance(c, str) and " " not in c for c in codes), codes + + +def test_build_verdict_refuses_a_workspace_with_no_terminal_record(tmp_path): + from loop.contract import scaffold_contract + from loop.verdict import VerdictError, build_verdict + + target = tmp_path / "no-terminal" + scaffold_contract(target) + + with pytest.raises(VerdictError, match="no terminal record"): + build_verdict(target) + + +def test_build_verdict_refuses_a_nonexistent_target(tmp_path): + from loop.verdict import VerdictError, build_verdict + + with pytest.raises(VerdictError): + build_verdict(tmp_path / "does-not-exist") +``` + +- [ ] **Step 2: Run tests to verify they fail** + +Run: `uv run --with pyyaml --with jsonschema --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts/test_verdict.py -v` +Expected: FAIL with `ImportError: cannot import name 'build_verdict'` + +> If `scaffold_contract` is not the exported scaffold name at HEAD, run `uv run --with pyyaml python3 -B -c "import loop.contract as c; print([n for n in dir(c) if 'scaffold' in n])"` and use the real one in the helper. Do not invent an API. + +- [ ] **Step 3: Implement the projection** + +```python +# add to loop/verdict.py +from importlib import metadata +from pathlib import Path + +from .contract import doctor_report +from .paths import resolve_loop_paths + + +def _tool_version() -> str | None: + try: + return metadata.version("loop-engineer") + except metadata.PackageNotFoundError: + return None + + +def _terminal_record(paths) -> dict[str, Any]: + path = paths.loop_dir / "terminal_state.json" + if not path.is_file(): + raise VerdictError( + "no terminal record: a verdict projects a finished run " + f"({path.name} is absent)") + try: + data = json.loads(path.read_text(encoding="utf-8")) + except (OSError, json.JSONDecodeError) as exc: + raise VerdictError(f"terminal record is unreadable: {exc}") from exc + if not isinstance(data, dict): + raise VerdictError("terminal record is not an object") + return data + + +def build_verdict(target: str | Path, *, mode: str | None = None) -> dict[str, Any]: + """Project local run state into a verdict@1 predicate body. + + Pure over the workspace: no environment, no network, no signing. + """ + try: + paths = resolve_loop_paths(target) + except (OSError, ValueError) as exc: + raise VerdictError(f"cannot resolve a loop workspace at {target}: {exc}") from exc + + report = doctor_report(paths.workspace, mode=mode) + terminal = _terminal_record(paths) + + store = report.get("event_store") or {} + chain = store.get("chain") or {} + head = chain.get("head") or {} + + policy = terminal.get("completion_policy") + policy_mode = policy.get("mode") if isinstance(policy, dict) else None + + return { + "schema": VERDICT_SCHEMA_ID, + "run_id": str(store.get("run_id") or paths.workspace.name), + "tool": {"name": "loop-engineer", "version": _tool_version()}, + "doctor": { + "ok": bool(report.get("ok")), + "validation_mode": str(report.get("validation_mode") or "unknown"), + "issue_codes": sorted({ + str(i.get("code")) for i in report.get("issues", []) + if isinstance(i, dict) and i.get("code") + }), + "schemas_checked": sorted(str(s) for s in report.get("schemas_checked", [])), + }, + "chain": { + "head": head.get("event_hash"), + "sequence": head.get("sequence"), + "unchained_prefix": int(chain.get("unchained_prefix") or 0), + }, + "terminal": { + "state": terminal.get("state"), + "completion_policy": policy_mode, + "false_completion": bool(terminal.get("false_completion")), + }, + "evidence": [], + } +``` + +- [ ] **Step 4: Run tests to verify they pass** + +Run: `uv run --with pyyaml --with jsonschema --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts/test_verdict.py -v` +Expected: all pass + +- [ ] **Step 5: Confirm the `event_store` key names against reality** + +Run: + +```bash +uv run --with pyyaml --with jsonschema python3 -B -m loop doctor examples/coverage-repair | python3 -c "import json,sys;print(json.dumps(json.load(sys.stdin).get('event_store'), indent=2))" +``` + +If `run_id`, `chain.head.event_hash`, `chain.head.sequence`, or `chain.unchained_prefix` are named differently, fix `build_verdict` to match the real shape and re-run Step 4. Do not adapt the tests to a wrong projection. + +- [ ] **Step 6: Commit** + +```bash +git add loop/verdict.py scripts/test_verdict.py +git commit -m "feat(verdict): project doctor, chain, and terminal into verdict@1" +``` + +--- + +## Task 3: The verified-evidence digest set + +**Files:** +- Modify: `loop/verdict.py` +- Test: `scripts/test_verdict.py` + +**Interfaces:** +- Consumes: `build_verdict` (Task 2); `loop.contract._strict_evidence_failure`. +- Produces: `build_verdict` now returns a populated `evidence` list of `{digest, code_digest, policy_digest}` objects, sorted by `digest`. + +The predicate carries **only** chain-bound evidence that satisfies the strict bar — not everything the terminal record cites. A cited-but-unverified entry in a signed document would be exactly the false completion this kernel exists to refuse. + +- [ ] **Step 1: Write the failing test** + +```python +# append to scripts/test_verdict.py + +def test_evidence_carries_only_entries_that_pass_the_strict_bar(tmp_path, monkeypatch): + """A cited entry that fails the strict bar must not reach the predicate.""" + import loop.verdict as verdict_mod + from loop.verdict import build_verdict + + target = _scaffold_terminal(tmp_path, "mixed-evidence", + policy="all_required_verified_evidence") + terminal_path = target / ".loop" / "terminal_state.json" + terminal = json.loads(terminal_path.read_text(encoding="utf-8")) + terminal["evidence"] = [".loop/evidence/good.json", ".loop/evidence/bad.json"] + terminal_path.write_text(json.dumps(terminal), encoding="utf-8") + + # Patch the bar where verdict.py BINDS it, not where it is defined — + # patching loop.contract is a verified no-op under from-import binding. + def fake_bar(entry, paths, bound): + return None if str(entry).endswith("good.json") else "digest_mismatch" + + monkeypatch.setattr(verdict_mod, "_strict_evidence_failure", fake_bar) + monkeypatch.setattr(verdict_mod, "_evidence_digests", + lambda entry, paths: {"digest": "a" * 64, + "code_digest": None, + "policy_digest": None}) + + evidence = build_verdict(target)["evidence"] + assert len(evidence) == 1, evidence + assert evidence[0]["digest"] == "a" * 64 + + +def test_evidence_entries_carry_digests_only(tmp_path): + """No URIs, no paths — a URI is a workspace path and this document is public.""" + from loop.verdict import build_verdict + + target = _scaffold_terminal(tmp_path, "digest-only") + for entry in build_verdict(target)["evidence"]: + assert set(entry) == {"digest", "code_digest", "policy_digest"} + + +def test_evidence_is_sorted_by_digest(tmp_path, monkeypatch): + """Canonical output: the same run always projects the same bytes.""" + import loop.verdict as verdict_mod + from loop.verdict import build_verdict + + target = _scaffold_terminal(tmp_path, "sorted") + terminal_path = target / ".loop" / "terminal_state.json" + terminal = json.loads(terminal_path.read_text(encoding="utf-8")) + terminal["evidence"] = ["b.json", "a.json"] + terminal_path.write_text(json.dumps(terminal), encoding="utf-8") + + monkeypatch.setattr(verdict_mod, "_strict_evidence_failure", + lambda entry, paths, bound: None) + monkeypatch.setattr( + verdict_mod, "_evidence_digests", + lambda entry, paths: {"digest": ("b" if "b.json" in str(entry) else "a") * 64, + "code_digest": None, "policy_digest": None}) + + digests = [e["digest"] for e in build_verdict(target)["evidence"]] + assert digests == sorted(digests) +``` + +- [ ] **Step 2: Run tests to verify they fail** + +Run: `uv run --with pyyaml --with jsonschema --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts/test_verdict.py -v` +Expected: FAIL — `_evidence_digests` does not exist; `evidence` is empty. + +- [ ] **Step 3: Implement** + +```python +# add to loop/verdict.py imports +from .contract import _strict_evidence_failure, doctor_report + +# add to loop/verdict.py +def _evidence_digests(entry: object, paths) -> dict[str, Any] | None: + """Read the evidence@1 record a terminal entry cites and lift its digests. + + Returns None when the record cannot be read as an object — the strict bar + has already vetted it, so this is a defensive read, not a second gate. + """ + record_path = paths.workspace / str(entry) + try: + record = json.loads(record_path.read_text(encoding="utf-8")) + except (OSError, json.JSONDecodeError, ValueError): + return None + if not isinstance(record, dict): + return None + verified_by = record.get("verified_by") + verified_by = verified_by if isinstance(verified_by, dict) else {} + digest = record.get("digest") + if not isinstance(digest, str): + return None + return { + "digest": digest, + "code_digest": verified_by.get("code_digest"), + "policy_digest": verified_by.get("policy_digest"), + } + + +def _verified_evidence(terminal: dict[str, Any], paths, bound) -> list[dict[str, Any]]: + entries = terminal.get("evidence") + if not isinstance(entries, list): + return [] + collected = [] + for entry in entries: + if _strict_evidence_failure(entry, paths, bound) is not None: + continue + digests = _evidence_digests(entry, paths) + if digests is not None: + collected.append(digests) + return sorted(collected, key=lambda d: d["digest"]) +``` + +Then replace `"evidence": []` in `build_verdict` with: + +```python + "evidence": _verified_evidence(terminal, paths, _bound_evidence(paths)), +``` + +- [ ] **Step 4: Wire the chain-bound map** + +`_strict_evidence_failure`'s third parameter is the chain-bound evidence map. Find how `loop/contract.py` builds it for its own strict check: + +```bash +grep -n "_strict_evidence_failure(" loop/contract.py +``` + +Reuse that exact producer as `_bound_evidence(paths)` — import it if it is a named function, or lift the two or three lines verbatim with a comment naming the source line. **Do not re-derive the binding.** If the producer needs the doctor report, thread it through rather than calling `doctor_report` twice. + +- [ ] **Step 5: Run tests to verify they pass** + +Run: `uv run --with pyyaml --with jsonschema --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts/test_verdict.py -v` +Expected: all pass + +- [ ] **Step 6: Verify against a real contract** + +```bash +uv run --with pyyaml --with jsonschema python3 -B -c " +from loop.verdict import build_verdict +import json; print(json.dumps(build_verdict('examples/flaky-test-triage'), indent=2))" +``` + +Expected: a populated predicate. Confirm `evidence[].digest` values are real 64-hex digests from that contract, not empty. + +- [ ] **Step 7: Commit** + +```bash +git add loop/verdict.py scripts/test_verdict.py +git commit -m "feat(verdict): carry only chain-bound evidence that passes the strict bar" +``` + +--- + +## Task 4: CLI wiring + +**Files:** +- Modify: `loop/__main__.py:15,19,21,23` +- Test: `scripts/test_verdict_cli.py` + +**Interfaces:** +- Consumes: `build_verdict` (Tasks 2–3). +- Produces: `python3 -m loop verdict [--mode basic|strict|release] ` writing canonical JSON to stdout, exit 0 on success and exit 2 with a typed one-line stderr message on `VerdictError`. + +- [ ] **Step 1: Write the failing test** + +```python +# scripts/test_verdict_cli.py +import json +import subprocess +import sys + + +def _run(*args, cwd=None): + return subprocess.run([sys.executable, "-B", "-m", "loop", *args], + capture_output=True, text=True, cwd=cwd) + + +def test_verdict_emits_canonical_json_and_exits_zero(repo_root): + proc = _run("verdict", "examples/flaky-test-triage", cwd=repo_root) + assert proc.returncode == 0, proc.stderr + doc = json.loads(proc.stdout) + assert doc["schema"] == "loop-engineer/verdict@1" + + +def test_verdict_output_is_byte_stable(repo_root): + a = _run("verdict", "examples/flaky-test-triage", cwd=repo_root) + b = _run("verdict", "examples/flaky-test-triage", cwd=repo_root) + assert a.stdout == b.stdout + + +def test_verdict_emits_no_statement_envelope(repo_root): + """The kernel emits a predicate body. actions/attest builds the Statement.""" + doc = json.loads(_run("verdict", "examples/flaky-test-triage", cwd=repo_root).stdout) + for forbidden in ("_type", "subject", "predicateType", "predicate"): + assert forbidden not in doc, f"{forbidden} belongs to the signer lane" + + +def test_verdict_fails_loud_without_a_terminal_record(tmp_path, repo_root): + _run("scaffold", str(tmp_path / "fresh"), cwd=repo_root) + proc = _run("verdict", str(tmp_path / "fresh"), cwd=repo_root) + assert proc.returncode == 2 + assert proc.stdout == "" + assert "Traceback" not in proc.stderr + assert "no terminal record" in proc.stderr + + +def test_verdict_appears_in_usage_and_help(repo_root): + proc = _run("--help", cwd=repo_root) + assert "verdict" in proc.stdout +``` + +Add a `repo_root` fixture to the file if `scripts/conftest.py` does not already provide one: + +```python +import pathlib +import pytest + + +@pytest.fixture +def repo_root(): + return pathlib.Path(__file__).resolve().parent.parent +``` + +- [ ] **Step 2: Run tests to verify they fail** + +Run: `uv run --with pyyaml --with jsonschema --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts/test_verdict_cli.py -v` +Expected: FAIL — `verdict` is not a known command. + +- [ ] **Step 3: Wire the CLI** + +In `loop/__main__.py`, add `"verdict"` to both `_COMMANDS` (line 15) and `_READ_COMMANDS` (line 19), add it to the `_USAGE` alternation (line 21), and add a `_HELP` usage line beside `metrics` plus a command description: + +``` + {_PROG} verdict [--mode basic|strict|release] +``` + +``` + verdict Project the run into a loop-engineer/verdict@1 predicate body for + CI attestation. Emits the predicate ONLY — the signer constructs + the in-toto Statement. Never signs and never verifies a signature. +``` + +Then add the dispatch block, following the `metrics` flag-pre-parsing precedent: + +```python + if command == "verdict": + from .verdict import VerdictError, build_verdict + from .chain import canonical_json + + try: + document = build_verdict(target, mode=mode) + except VerdictError as exc: + print(f"error: {exc}", file=sys.stderr) + return 2 + print(canonical_json(document)) + return 0 +``` + +- [ ] **Step 4: Run tests to verify they pass** + +Run: `uv run --with pyyaml --with jsonschema --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts/test_verdict_cli.py -v` +Expected: all pass + +- [ ] **Step 5: Run the full suite — `_USAGE` and `_HELP` are asserted elsewhere** + +Run: `uv run --with pyyaml --with jsonschema --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts` +Expected: **1305 + new tests passed / 18 skipped**, zero failures. Existing CLI-surface tests that pin the usage string will need their fixtures updated — update the *fixture* to include `verdict`, never weaken the assertion. + +- [ ] **Step 6: Commit** + +```bash +git add loop/__main__.py scripts/test_verdict_cli.py +git commit -m "feat(cli): add the loop verdict verb" +``` + +--- + +## Task 5: The boundary, made mechanical + +**Files:** +- Create: `scripts/test_verdict_purity.py` + +**Interfaces:** +- Consumes: `loop/verdict.py`, the whole `loop/` package. +- Produces: nothing. This task exists so ADR 0002's boundary cannot decay into an intention. + +- [ ] **Step 1: Write the tests** + +```python +# scripts/test_verdict_purity.py +import ast +import json +import pathlib +import subprocess +import sys + +REPO = pathlib.Path(__file__).resolve().parent.parent +LOOP = REPO / "loop" + +_SIGNING_TOKENS = ("sigstore", "cosign", "fulcio", "rekor", "dsse", + "private_key", "PRIVATE KEY", "ACTIONS_ID_TOKEN", + "id_token", "oidc") + + +def test_kernel_never_references_a_signing_stack(): + """ADR 0002: the kernel may hash; it must never sign.""" + offenders = [] + for path in LOOP.rglob("*.py"): + text = path.read_text(encoding="utf-8").lower() + for token in _SIGNING_TOKENS: + if token.lower() in text: + offenders.append(f"{path.relative_to(REPO)}:{token}") + assert offenders == [], offenders + + +def test_kernel_reads_no_environment_variable(): + """Zero matches at HEAD 0025acc; this keeps it that way. A verdict that + read GITHUB_SHA would put a vendor identifier inside the portable layer.""" + offenders = [] + for path in LOOP.rglob("*.py"): + text = path.read_text(encoding="utf-8") + if "os.environ" in text or "getenv" in text: + offenders.append(str(path.relative_to(REPO))) + assert offenders == [], offenders + + +def test_verdict_module_imports_only_stdlib_and_loop(): + tree = ast.parse((LOOP / "verdict.py").read_text(encoding="utf-8")) + third_party = [] + for node in ast.walk(tree): + if isinstance(node, ast.ImportFrom): + if node.level == 0 and node.module and node.module.split(".")[0] not in sys.stdlib_module_names: + third_party.append(node.module) + elif isinstance(node, ast.Import): + for alias in node.names: + if alias.name.split(".")[0] not in sys.stdlib_module_names: + third_party.append(alias.name) + assert third_party == [], third_party + + +def _predicate(): + proc = subprocess.run( + [sys.executable, "-B", "-m", "loop", "verdict", "examples/flaky-test-triage"], + capture_output=True, text=True, cwd=REPO) + assert proc.returncode == 0, proc.stderr + return json.loads(proc.stdout) + + +def test_predicate_field_allowlist_holds(): + """Everything here is public, append-only, and permanent. Adding a field is + a one-way door — this test is the door.""" + doc = _predicate() + assert set(doc) == {"schema", "run_id", "tool", "doctor", "chain", "terminal", "evidence"} + assert set(doc["doctor"]) == {"ok", "validation_mode", "issue_codes", "schemas_checked"} + assert set(doc["chain"]) == {"head", "sequence", "unchained_prefix"} + assert set(doc["terminal"]) == {"state", "completion_policy", "false_completion"} + for entry in doc["evidence"]: + assert set(entry) == {"digest", "code_digest", "policy_digest"} + + +def test_predicate_carries_no_free_text_but_run_id(): + """run_id is the ONE operator-controlled string, allowlisted deliberately. + Any other prose in the document is a leak into a permanent public log.""" + doc = _predicate() + strings = [] + + def walk(node, path): + if isinstance(node, dict): + for k, v in node.items(): + walk(v, f"{path}.{k}") + elif isinstance(node, list): + for i, v in enumerate(node): + walk(v, f"{path}[{i}]") + elif isinstance(node, str): + strings.append((path, node)) + + walk(doc, "") + for path, value in strings: + if path == ".run_id": + continue + # Whitespace is the proxy for prose. Every other string in the document + # is a digest, an enum, a snake_case issue code, or a slash-separated + # schema id — none of which contain a space. + assert " " not in value, f"free text at {path}: {value!r}" + + +def test_predicate_validates_against_its_own_schema(): + jsonschema = __import__("jsonschema") + from loop.verdict import _load_verdict_schema + + jsonschema.validate(_predicate(), _load_verdict_schema()) +``` + +- [ ] **Step 2: Run tests** + +Run: `uv run --with pyyaml --with jsonschema --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts/test_verdict_purity.py -v` + +Expected: all pass. **If `test_kernel_never_references_a_signing_stack` or `test_kernel_reads_no_environment_variable` fails, that is a real finding — fix the code, not the test.** If a token produces a false positive (e.g. the substring `oidc` inside an unrelated identifier), narrow the token, and add a comment naming the false positive so the narrowing is not mistaken for weakening. + +- [ ] **Step 3: Prove the tests have teeth (mutation probe)** + +Temporarily add `import os; os.environ.get("GITHUB_SHA")` to `loop/verdict.py`, re-run, confirm `test_kernel_reads_no_environment_variable` FAILS, then revert. Temporarily add a `"debug"` key to the returned dict, re-run, confirm `test_predicate_field_allowlist_holds` FAILS, then revert. A gate that cannot fail is not a gate. + +- [ ] **Step 4: Commit** + +```bash +git add scripts/test_verdict_purity.py +git commit -m "test(verdict): make the ADR 0002 boundary mechanical" +``` + +--- + +## Task 6: The action step + +**Files:** +- Modify: `action.yml` + +**Interfaces:** +- Consumes: the `loop verdict` CLI (Task 4); the existing `chain-head` step (`action.yml:85-107`), which already extracts `event_store.chain.head.event_hash` into the `chain-head` output. +- Produces: action input `attest` (default `"false"`); outputs `attestation-url`, `attestation-id`. + +- [ ] **Step 1: Add the input** + +Append to `inputs:` in `action.yml`: + +```yaml + attest: + description: >- + Emit a loop-engineer/verdict@1 predicate and attest it keylessly via + GitHub OIDC. Requires the CALLING JOB to declare `id-token: write` and + `attestations: write` — a composite action cannot declare its own + permissions. Default false. The predicate is PUBLIC and permanent for + public repositories. + required: false + default: "false" +``` + +- [ ] **Step 2: Add the outputs** + +Append to `outputs:`: + +```yaml + attestation-url: + description: "URL of the attestation created by this run ('' when attest is false)." + value: ${{ steps.attest.outputs.attestation-url }} + attestation-id: + description: "ID of the attestation created by this run ('' when attest is false)." + value: ${{ steps.attest.outputs.attestation-id }} +``` + +- [ ] **Step 3: Add the predicate + permission-precheck step** + +Insert after the existing `chain head (anchor surface)` step: + +```yaml + - name: verdict predicate + id: verdict + if: ${{ inputs.attest == 'true' }} + shell: bash + env: + LOOP_PATH: "${{ inputs.path }}" + run: | + # Fail with a legible message rather than a raw OIDC 403 from the attest + # step: a composite action cannot declare permissions, so the CALLING + # job must. ACTIONS_ID_TOKEN_REQUEST_URL is present only when the job + # declared id-token: write. + if [ -z "${ACTIONS_ID_TOKEN_REQUEST_URL:-}" ]; then + echo "::error::attest: true requires the calling job to declare" \ + "'permissions: { id-token: write, attestations: write }'." \ + "A composite action cannot declare its own permissions." + exit 1 + fi + loop verdict "$LOOP_PATH" > "${RUNNER_TEMP}/verdict.json" + echo "predicate-path=${RUNNER_TEMP}/verdict.json" >> "$GITHUB_OUTPUT" + + - name: attest verdict + id: attest + if: ${{ inputs.attest == 'true' }} + uses: actions/attest@v4 + with: + subject-name: loop-chain-head + subject-digest: sha256:${{ steps.chain-head.outputs.chain-head }} + predicate-type: urn:loop-engineer:verdict:1 + predicate-path: ${{ steps.verdict.outputs.predicate-path }} + push-to-registry: false +``` + +> **Verify before merging** (ADR 0002 open items 1 and 5): confirm `actions/attest@v4` is the current major, confirm `subject-digest` accepts the `sha256:` prefix form with a separate `subject-name`, and confirm whether `create-storage-record` exists as an input and what it defaults to. If it exists, pass it explicitly as `false`. Run `gh api repos/actions/attest/contents/action.yml --jq '.content' | base64 -d` and read the real input list. Do not ship an input name from memory. + +- [ ] **Step 4: Guard the empty-head case** + +A run with no chained events produces an empty `chain-head`, and `subject-digest: sha256:` is malformed. Change the `attest verdict` step's condition to: + +```yaml + if: ${{ inputs.attest == 'true' && steps.chain-head.outputs.chain-head != '' }} +``` + +and add a step that warns when attestation was requested but skipped: + +```yaml + - name: attest skipped (no chained events) + if: ${{ inputs.attest == 'true' && steps.chain-head.outputs.chain-head == '' }} + shell: bash + run: echo "::warning::attest requested but the store has no chained events; nothing to attest." +``` + +- [ ] **Step 5: Lint the workflow syntax** + +```bash +uv run --with pyyaml python3 -B -c " +import yaml, pathlib +doc = yaml.safe_load(pathlib.Path('action.yml').read_text(encoding='utf-8')) +print('inputs:', sorted(doc['inputs'])) +print('outputs:', sorted(doc['outputs'])) +print('steps:', [s.get('name') for s in doc['runs']['steps']])" +``` + +Expected: `attest` in inputs; `attestation-url`, `attestation-id` in outputs; the new steps in order. + +- [ ] **Step 6: Commit** + +```bash +git add action.yml +git commit -m "feat(action): opt-in keyless attestation of the verdict predicate" +``` + +--- + +## Task 7: The repo's own attestation job + +**Files:** +- Create: `.github/workflows/attest.yml` + +**Interfaces:** +- Consumes: the composite action from Task 6. +- Produces: a real attestation on every push to the default branch, over a **tracked** `examples/*` contract. + +Never point this at the live gitignored `.loop/` — CI runs on a fresh checkout where it does not exist. + +- [ ] **Step 1: Write the workflow** + +```yaml +name: attest + +on: + push: + branches: [main] + +permissions: + contents: read + +jobs: + verdict: + # Push-to-default-branch only, by ADR 0002 decision 5: attesting on a PR + # would mint a signed verdict under the repository's identity before review. + runs-on: ubuntu-latest + permissions: + contents: read + id-token: write + attestations: write + steps: + - uses: actions/checkout@v5 + - id: gate + uses: ./ + with: + path: examples/flaky-test-triage + attest: "true" + - name: record + run: | + echo "chain head: ${{ steps.gate.outputs.chain-head }}" >> "$GITHUB_STEP_SUMMARY" + echo "attestation: ${{ steps.gate.outputs.attestation-url }}" >> "$GITHUB_STEP_SUMMARY" +``` + +- [ ] **Step 2: Confirm the checkout action major matches the repo's other workflows** + +```bash +grep -rn "actions/checkout@" .github/workflows/ +``` + +Use whatever major `ci.yml` uses. A mismatched pin is a dependabot PR waiting to happen. + +- [ ] **Step 3: Commit** + +```bash +git add .github/workflows/attest.yml +git commit -m "ci: attest the verdict for a tracked example contract on push to main" +``` + +- [ ] **Step 4: Note for the PR body** + +This job cannot be validated before merge — `push: branches: [main]` fires only after landing. State that explicitly in the PR description, and check the first post-merge run. If it fails, the fix is a follow-up PR, not a revert: the action input defaults to `false`, so nothing else is affected. + +--- + +## Task 8: Normative documentation + +**Files:** +- Modify: `reference/repo-os-contract.md` (append §23) +- Modify: `CHANGELOG.md` + +**Interfaces:** +- Consumes: everything above. +- Produces: the normative spec third parties implement against. + +- [ ] **Step 1: Append §23** + +Add after §22, matching the numbered-section style of §16 and §17. It must cover: + +1. The predicate shape, field by field, with the exact JSON from "The predicate shape (binding)" above as a machine-pinned conformance vector. +2. `PREDICATE_TYPE = urn:loop-engineer:verdict:1` and why it names no vendor host (ADR 0002 decision 3). +3. **The subject is the chain head**, and — critically — that the chain head is a SHA-256 over a synthesized event preimage (§16), **not** over retrievable bytes. A consumer must never conclude "fetch bytes, re-hash, compare." +4. The separation: the kernel emits the predicate; the signer builds the Statement. +5. That `run_id` is operator-controlled and lands in a permanent public log, so run ids must not embed sensitive text. +6. The honest limits, copied in substance from ADR 0002's Standing Limits — especially that a worker with merge rights can loosen the gate and then mint a genuine attestation for it. + +- [ ] **Step 2: Verify the reference-file count is unchanged** + +```bash +uv run --with pyyaml python3 -B scripts/self_eval.py +``` + +Expected: 13/13 PASS. If `reference_filenames` fails, a new file was created in `reference/` — move the content into `repo-os-contract.md` §23 instead. + +- [ ] **Step 3: Add the CHANGELOG entry** + +Under an `## [Unreleased]` heading, describing: the `loop verdict` verb, the `verdict@1` schema, the opt-in `attest` action input, and — in the project's established voice — what the attestation does **not** buy. Name the one-run detection latency and the worker-can-edit-the-verifier path. Do not use the word "tamper-proof". + +- [ ] **Step 4: Run the full suite** + +Run: `uv run --with pyyaml --with jsonschema --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts` +Expected: **1305 + new tests passed / 18 skipped**, zero failures. + +- [ ] **Step 5: Run the structural-fallback leg** + +Run: `uv run --with pyyaml --with pytest python3 -B -m pytest -q -p no:cacheprovider scripts` +Expected: zero failures. `test_predicate_validates_against_its_own_schema` (Task 5) imports `jsonschema` directly and must be skipped rather than failed in this leg — mark it `@pytest.mark.skipif` on the import, following how existing schema tests handle the fallback mode. + +- [ ] **Step 6: Commit** + +```bash +git add reference/repo-os-contract.md CHANGELOG.md +git commit -m "docs: normative verdict@1 section and honest limits" +``` + +--- + +## Task 9: The governance change (ADR 0002 decision 6) + +**Files:** +- Create: `.github/CODEOWNERS` + +**Interfaces:** +- Consumes: nothing. +- Produces: a review requirement on the three paths that define what the gate checks. + +This is the control for the largest limit in ADR 0002 — an agent with ordinary +merge rights can loosen the gate and then mint a genuine attestation for it. No +code fixes that. This slice is the first PR that touches all three paths, so the +requirement lands with it. + +- [ ] **Step 1: Write CODEOWNERS** + +``` +# Changes to what the gate checks require human review (ADR 0002, decision 6). +# Everything else in this repo stays autonomous. +/loop/ @SollanSystems +/action.yml @SollanSystems +/.github/workflows/ @SollanSystems +/.github/CODEOWNERS @SollanSystems +``` + +- [ ] **Step 2: Verify GitHub parses it** + +After pushing the branch: + +```bash +gh api "repos/SollanSystems/loop-engineer/codeowners/errors" --jq '.errors' +``` + +Expected: `[]`. A malformed CODEOWNERS fails silently — GitHub simply enforces +nothing — so this check is the deliverable, not the file. + +- [ ] **Step 3: Commit** + +```bash +git add .github/CODEOWNERS +git commit -m "chore: require human review on the gate-defining paths" +``` + +- [ ] **Step 4: OPERATOR ACTION — the ruleset half** + +**CODEOWNERS alone enforces nothing.** The `main protection` ruleset currently +requires a PR with **0 approvals**, so a code-owner file without a matching rule +is decoration. The operator must set the ruleset to require review from Code +Owners (Settings → Rules → `main protection` → "Require review from Code +Owners", and raise required approvals to 1). + +Do not attempt this from an agent session — branch-protection edits are operator +territory, and a half-applied change is worse than none. Record in the PR body +whether it has been applied; until it is, ADR 0002's decision 6 is documented but +not in force, and the Standing Limits section is the operative truth. + +--- + +## Definition of done + +- [ ] Full suite green in both dependency legs, against the 1305/18 baseline, measured from **inside** a fresh worktree (not `--project`). +- [ ] `uv run --with pyyaml python3 -B scripts/self_eval.py` → 13/13. +- [ ] `uv run --with pyyaml python3 -B scripts/validate_frontmatter.py` → 9/9. +- [ ] The two mutation probes in Task 5 Step 3 were run and both gates failed as predicted. +- [ ] `grep -rn "environ\|getenv" --include=*.py loop/` → zero matches. +- [ ] `python3 -m loop verdict examples/flaky-test-triage` emits a predicate with no `_type`, `subject`, or `predicateType` key. +- [ ] `actions/attest` input names were read from the live action definition, not from this plan. +- [ ] PR body states that `.github/workflows/attest.yml` is unvalidatable pre-merge and names the first post-merge run as the check. +- [ ] `gh api repos/SollanSystems/loop-engineer/codeowners/errors` → `[]`. +- [ ] PR body records whether the ruleset half of Task 9 has been applied by the operator. +- [ ] No version bump in this PR. + +## Out of scope — Slice 4b + +`loop verdict --compare`, anchor auto-resolution from a prior attestation, and signer-trust policy evaluation. Do not add a `--compare` flag, a `--sign` flag, or any bundle-parsing code in this slice. If a task seems to need one, the task is wrong. + +## Open items to resolve during execution + +Carried from ADR 0002. None blocks starting; each blocks the task that touches it. + +1. **Subject digest algorithm** (Task 6). The chain head is not a hash of retrievable bytes, so the `sha256` DigestSet key licenses a false inference. A namespaced key is correct in in-toto terms but `actions/attest`'s `subject-digest` may accept only `sha256:`. Resolve by experiment. If the input constrains us, §23 must carry the disambiguation instead. +2. **`create-storage-record` / `push-to-registry` defaults** (Task 6). Pass both explicitly so the two-permission claim is true by construction. +3. **`scaffold_contract` name** (Task 2 Step 2). Confirm the real scaffold entry point before writing the test helper. +4. **`event_store` key names** (Task 2 Step 5). Confirm `run_id`, `chain.head.event_hash`, `chain.head.sequence`, `chain.unchained_prefix` against real doctor output. +5. **The chain-bound evidence map producer** (Task 3 Step 4). Reuse `loop/contract.py`'s, never re-derive.