feat(test): gate contract claims against the built binary (#907) - #997
Conversation
A contract can state a load-bearing rule that nothing executes, and the implementation then diverges silently. That is #907: cooklang-compatibility.md says the `{` closing a multiword name "must touch the name" and spells out the exact failure if it does not, while the pinned parser scans to the first `{` on the line. No test, fixture, or gate noticed. test/contract-claims.txt pairs a verbatim quote from a normative contract with a black-box check against the installed binary. The gate fails on three conditions: the quote no longer appears in the contract (prose drift), an `enforced` claim is violated, or a recorded `known-divergence` no longer reproduces — which fails the moment the upstream fix lands and forces the record to be reclassified instead of left stale. Ships six records, including the open Cooklang divergence as a `known-divergence`. `--audit` lists contracts whose rules still have no quoted tether; that list is a worklist, not a defect list, since many contracts carry dedicated test steps of their own. 🤖 Generated with Codebuff Co-Authored-By: Codebuff <noreply@codebuff.com>
The fragment convention keys the filename to the PR number so release assembly sorts correctly; the PR now exists. Generated with Codebuff 🤖 Co-Authored-By: Codebuff <noreply@codebuff.com>
|
SummaryThe run covers normal build and contract-validation flows, version and artifact consistency, expected failure handling, audit behavior, and adversarial registry cases such as missing or duplicate rules, along with a recipe-parsing edge case. Core happy paths pass, but the safeguards around validating the validation rules themselves are not fully reliable. Merge with caution — the PR introduces medium-severity weaknesses in the validation gate that can misclassify unavailable checks and accept duplicate coverage, reducing confidence that CI will detect incomplete or misleading contract coverage. The unrelated recipe-parsing issue is a flag for later rather than a merge driver. Tests run by ItoAdditional Findings DetailsThese findings are unrelated to the current changes but were observed during testing. 🟡 Cooklang names absorb later braces
Evidence PackageTip Reply with @itoqa to send us feedback on this test run. |
| # pinned Oliver parser instead scans to the first `{` on the line, so a braced | ||
| # word later in the sentence is absorbed into the name and the prose between | ||
| # them is dropped from the rendered step. | ||
| contract_check_cooklang_name_termination_adjacency() { |
There was a problem hiding this comment.
Unavailable recipe probe is reported as resolved
What failed: The check could not run the recipe probe, but the gate presented that execution failure as proof that the Cooklang problem was resolved. A failed probe must not be treated as a successful resolution.
Impact · Steps · Stub / mock · Analysis · Why this is likely a bug
- Severity: Medium
- Impact: CI can report that the known Cooklang problem is fixed when the recipe probe did not run. This can mislead maintainers and block or misdirect release checks until the result is reviewed.
- Steps to Reproduce:
- Make the recipe-scale command unavailable and run the contract-claim gate with the Cooklang record still marked as a known divergence.
- Observe that the probe reports it exited with an error, then the gate prints that the divergence no longer reproduces and exits with a failed record.
- Restore the command and run the same check again; it reports the known misparse still reproducing.
- Stub / mock content: The run used a disposable unavailable-oracle setup for the first probe, then restored the command for comparison. No application mocks or route stubs were used.
- Code Analysis: The production repository path is the newly added shell contract gate. In test/contract-claims.checks.sh, contract_check_cooklang_name_termination_adjacency() runs the installed binary at lines 164-172, captures its exit status, and returns nonzero at lines 174-177 when recipe-scale cannot execute. In scripts/test-contract-claims.sh, finish_record() stores that nonzero result as check_rc=1 at lines 167-201, then the known-divergence branch at lines 215-224 treats every nonzero check_rc as the divergence no longer reproducing and emits the resolution instructions. There is no distinction between an observable contract-violating result and an unavailable oracle. The targeted fix is to propagate an execution-error outcome separately and only enter the resolution branch when the probe completed and its output contradicted the known divergence.
- Why this is likely a bug: The recorded unavailable run explicitly says the divergence could not be probed, while the same restored run reports the documented misparse still reproducing. The source confirms that the check returns failure for inability to execute, and the gate then assigns that failure the semantic meaning of a resolved divergence. That is a false regression result in the new test gate, not a setup-only failure: it can make CI claim that an upstream parser defect disappeared when no successful parser observation was made. A small result-state distinction in the check/gate is sufficient; no parser rewrite is needed.
Relevant code
test/contract-claims.checks.sh:164-177
contract_check_cooklang_name_termination_adjacency() {
...
out="$("$BORIS" recipe-scale --input "$corpus/in" --id adjacency --factor 2 --cooklang 2>&1)"
rc=$?
...
if [[ "$rc" -ne 0 ]]; then
cc_detail "recipe-scale exited $rc, so the divergence could not be probed:"
...
return 1
fiscripts/test-contract-claims.sh:203-224
known-divergence)
if [[ "$check_rc" -eq 0 ]]; then
warn "$r_id (known divergence $r_issue — still reproduces)"
...
else
failmsg "$r_id (known-divergence) — the divergence no longer reproduces"
...
failed=$((failed + 1))
fiEvidence Package
Copy prompt for an agent
Ito QA identified the following failure during automated PR testing. Please investigate and propose a fix.
**Medium severity — Unavailable recipe probe is reported as resolved**
**What failed:** The check could not run the recipe probe, but the gate presented that execution failure as proof that the Cooklang problem was resolved. A failed probe must not be treated as a successful resolution.
- **Impact:** CI can report that the known Cooklang problem is fixed when the recipe probe did not run. This can mislead maintainers and block or misdirect release checks until the result is reviewed.
- **Steps to reproduce:**
1. Make the recipe-scale command unavailable and run the contract-claim gate with the Cooklang record still marked as a known divergence.
2. Observe that the probe reports it exited with an error, then the gate prints that the divergence no longer reproduces and exits with a failed record.
3. Restore the command and run the same check again; it reports the known misparse still reproducing.
- **Stub / mock content:** The run used a disposable unavailable-oracle setup for the first probe, then restored the command for comparison. No application mocks or route stubs were used.
- **Code analysis:** The production repository path is the newly added shell contract gate. In test/contract-claims.checks.sh, contract_check_cooklang_name_termination_adjacency() runs the installed binary at lines 164-172, captures its exit status, and returns nonzero at lines 174-177 when recipe-scale cannot execute. In scripts/test-contract-claims.sh, finish_record() stores that nonzero result as check_rc=1 at lines 167-201, then the known-divergence branch at lines 215-224 treats every nonzero check_rc as the divergence no longer reproducing and emits the resolution instructions. There is no distinction between an observable contract-violating result and an unavailable oracle. The targeted fix is to propagate an execution-error outcome separately and only enter the resolution branch when the probe completed and its output contradicted the known divergence.
- **Why this is likely a bug:** The recorded unavailable run explicitly says the divergence could not be probed, while the same restored run reports the documented misparse still reproducing. The source confirms that the check returns failure for inability to execute, and the gate then assigns that failure the semantic meaning of a resolved divergence. That is a false regression result in the new test gate, not a setup-only failure: it can make CI claim that an upstream parser defect disappeared when no successful parser observation was made. A small result-state distinction in the check/gate is sufficient; no parser rewrite is needed.
**Relevant code:**
`test/contract-claims.checks.sh:164-177`
~~~bash
contract_check_cooklang_name_termination_adjacency() {
...
out="$("$BORIS" recipe-scale --input "$corpus/in" --id adjacency --factor 2 --cooklang 2>&1)"
rc=$?
...
if [[ "$rc" -ne 0 ]]; then
cc_detail "recipe-scale exited $rc, so the divergence could not be probed:"
...
return 1
fi
~~~
`scripts/test-contract-claims.sh:203-224`
~~~bash
known-divergence)
if [[ "$check_rc" -eq 0 ]]; then
warn "$r_id (known divergence $r_issue — still reproduces)"
...
else
failmsg "$r_id (known-divergence) — the divergence no longer reproduces"
...
failed=$((failed + 1))
fi
~~~| normalize "$(cat "$1")" | ||
| } | ||
|
|
||
| r_id=""; r_contract=""; r_section=""; r_status=""; r_issue=""; r_quote=""; r_check="" |
There was a problem hiding this comment.
Duplicate registry entries pass the gate
What failed: The gate accepted two identical claim records and counted both of them as coverage.
Impact · Steps · Stub / mock · Analysis · Why this is likely a bug
- Severity: Medium
- Impact: A registry with duplicate claim IDs can pass the gate and count the same claim twice, hiding missing contract coverage from maintainers. This can let an incomplete contract set merge until the duplicate is removed.
- Steps to Reproduce:
- From the repository root, copy the complete frontmatter.parent-legacy-keys-rejected record in test/contract-claims.txt and append the copy as a second record with the same claim ID.
- Run bash scripts/test-contract-claims.sh with the installed Boris binary available.
- Check the output and exit status: the duplicate record runs twice, the summary reports seven total records, and the command exits 0 instead of reporting a duplicate-ID error.
- Stub / mock content: No stubs, mocks, or bypasses were applied for this test in the recorded run.
- Code Analysis: The PR-added runner initializes per-record fields and aggregate counters at scripts/test-contract-claims.sh:91-95, but it has no collection of previously seen claim IDs. In the parser at lines 232-253, every
claim id:line calls finish_record for the preceding record and then assigns the new ID; the same ID is therefore treated as a new independent record. finish_record validates the fields, tethered quote, and executable check, but never compares r_id with earlier IDs. The exact duplicate mutation consequently reaches the normal success path and increments records/enforced again. The smallest fix is to keep a Bash-3.2-compatible list of accepted IDs (or check each new ID against a newline-delimited seen-ID string) and increment failed plus emit a duplicate-ID diagnostic before dispatching the duplicate record; then return a nonzero result for the registry. - Why this is likely a bug: The registry format describes each claim ID as a stable address for one record, and the gate is specifically intended to prevent missing contract coverage from being hidden by bad registry data. The exact duplicate run demonstrates that this is not only a theoretical gap: the command exits successfully and reports the same claim twice. A focused uniqueness check in the newly added parser is sufficient; no broader redesign is needed.
Relevant code
scripts/test-contract-claims.sh:91-95
r_id=""; r_contract=""; r_section=""; r_status=""; r_issue=""; r_quote=""; r_check=""
records=0
enforced=0
divergent=0
failed=0scripts/test-contract-claims.sh:232-253
while IFS= read -r line || [[ -n "$line" ]]; do
case "$line" in
'claim id: '*)
finish_record
r_id="${line#claim id: }"
;;
...
esac
done < "$REGISTRY"
finish_recordEvidence Package
Copy prompt for an agent
Ito QA identified the following failure during automated PR testing. Please investigate and propose a fix.
**Medium severity — Duplicate registry entries pass the gate**
**What failed:** The gate accepted two identical claim records and counted both of them as coverage.
- **Impact:** A registry with duplicate claim IDs can pass the gate and count the same claim twice, hiding missing contract coverage from maintainers. This can let an incomplete contract set merge until the duplicate is removed.
- **Steps to reproduce:**
1. From the repository root, copy the complete frontmatter.parent-legacy-keys-rejected record in test/contract-claims.txt and append the copy as a second record with the same claim ID.
2. Run bash scripts/test-contract-claims.sh with the installed Boris binary available.
3. Check the output and exit status: the duplicate record runs twice, the summary reports seven total records, and the command exits 0 instead of reporting a duplicate-ID error.
- **Stub / mock content:** No stubs, mocks, or bypasses were applied for this test in the recorded run.
- **Code analysis:** The PR-added runner initializes per-record fields and aggregate counters at scripts/test-contract-claims.sh:91-95, but it has no collection of previously seen claim IDs. In the parser at lines 232-253, every `claim id:` line calls finish_record for the preceding record and then assigns the new ID; the same ID is therefore treated as a new independent record. finish_record validates the fields, tethered quote, and executable check, but never compares r_id with earlier IDs. The exact duplicate mutation consequently reaches the normal success path and increments records/enforced again. The smallest fix is to keep a Bash-3.2-compatible list of accepted IDs (or check each new ID against a newline-delimited seen-ID string) and increment failed plus emit a duplicate-ID diagnostic before dispatching the duplicate record; then return a nonzero result for the registry.
- **Why this is likely a bug:** The registry format describes each claim ID as a stable address for one record, and the gate is specifically intended to prevent missing contract coverage from being hidden by bad registry data. The exact duplicate run demonstrates that this is not only a theoretical gap: the command exits successfully and reports the same claim twice. A focused uniqueness check in the newly added parser is sufficient; no broader redesign is needed.
**Relevant code:**
`scripts/test-contract-claims.sh:91-95`
~~~bash
r_id=""; r_contract=""; r_section=""; r_status=""; r_issue=""; r_quote=""; r_check=""
records=0
enforced=0
divergent=0
failed=0
~~~
`scripts/test-contract-claims.sh:232-253`
~~~bash
while IFS= read -r line || [[ -n "$line" ]]; do
case "$line" in
'claim id: '*)
finish_record
r_id="${line#claim id: }"
;;
...
esac
done < "$REGISTRY"
finish_record
~~~…ims-gate revert: remove the contract-claims gate (#997)

Agent Completion Report
Status: complete
Branch and Worktree:
feat/contract-claims-gate./(primary checkout)Commit and PR:
300c1245(head;791e1a0eis the gate,300c1245the fragment rename)main(this PR is feat(test): gate contract claims against the built binary (#907) #997)Linked Issues (auto-close convention):
Refs Cooklang name-termination: contract's
{-adjacency guard is not enforced (@salt into the {bowl}misparses) #907This PR does not close the issue. It records the divergence as a
known-divergenceclaim and makes it fail the moment the upstream fix lands;the defect itself is still open and its fix lives in Oliver.
Changed Files:
test/contract-claims.txt— the claim registry (six records)test/contract-claims.checks.sh— the named black-box checksscripts/test-contract-claims.sh— the gate runnerbuild.zig—test-contract-claimsstep, wired intotesttest/README.md— registry documentation and the three failure conditionsdocs/changelog.d/997-contract-claims-gate.md— changelog fragmentPreserved Unrelated Files:
stash@{0}entrybelonging to
fix/issue-close-declared-intentand thet3code/cli-reference-currentbranch are untouched. Localmainwas notused or moved. The separate rendered-search title fix is on its own branch
and is not included here; this branch is verified green without it.
Implementation Summary:
implementation then diverges silently. The Cooklang adapter contract says
the
{closing a multiword name must touch the name, and spells out theexact failure if it does not — an unrelated braced word later in a sentence
being absorbed into the name and the prose between them deleted. The pinned
parser scans to the first
{on the line and had never enforced it.test/contract-claims.txtpairs a verbatim quote from a normative contractwith a black-box check against the installed binary. Three failure
conditions: the quote no longer appears in the contract (prose drift), an
enforcedclaim is violated, or a recordedknown-divergenceno longerreproduces.
known-divergencepasses while the bug is present, so the upstream fix fails the gate and
forces the record to be reclassified instead of being left stale. Neither
direction can rot quietly.
enforced(legacyparentEntry/parent_entryrejection,
EPARENTMISSINGfor an absent parent,EUSAGEexit 2, pure I/Oexit 3, and the version pin reusing
scripts/test-version-pin.shverbatimvia
check: script:<path>), plus the Cooklang divergence.check that re-derived what the compiler does would pass while the compiler
was wrong. The
parentEntrycheck asserts the absence ofEPARENTMISSING,which is what makes it a test of "no silent alias" rather than of "this
corpus fails".
--auditlists contracts with no quoted tether, framed as a worklist: 52 of57 today. Untethered is not untested —
publication-claims.md,scanner.md,watch-mode.mdand others carry dedicated test steps of their own. Anearlier draft of that output said "no executable claim", which overclaimed,
and was corrected before landing.
Known Gaps:
remaining work, not a completed survey.
rendered-search.mdis deliberately left untethered. Its title rule is notreachable black-box — the compiler CLI exposes no search flag — and tethering
it would have required either a nested build or asserting structure and
calling it behavior. Its own unit test still guards it.
one or two
borisinvocations). Acceptable now; a registry of dozens ofrecords would want batching.
Exact Commands Run:
zig build testzig build test-contract-claimsbash scripts/test-contract-claims.sh --auditzig fmt --check build.ziggit diff --check origin/main...feat/contract-claims-gateExact Gate Results:
zig build test: pass — exit 0 on this branch alone, with the newcontract-claimsstep running inside it andsrc/search_index.zigstill atits original revision
zig build test-contract-claims: pass —5 enforced, 1 known divergence(s), 6 record(s) totalbash scripts/test-contract-claims.sh --audit: pass —52 of 57 normative contracts have no record in this registryexit 2→exit 9in the
EUSAGErow ofdocs/contracts/diagnostics.md) made the gate exit 1,name
cli.usage-error-exit-2, and print the stale quote; the other tworecords in that file stayed green. Restored,
git diffclean.the binary does not produce made the gate exit 1 with
the implementation does not honor this contract claim, naming the record,contract, and check. Restored.
zig fmt --check build.zig: passgit diff --check origin/main...feat/contract-claims-gate: cleanDeterminism Result:
fixed corpora in a per-run scratch tree under
.zig-cache/contract-claims,which is removed on exit. No content enumeration order is relied on; the
audit walk sorts by filename through the shell glob.
Generated Artifacts:
.zig-cache/contract-claims/scratch trees only (gitignored, removed onexit). No
dist/,rag/, orsource-rag/output was committed.Blockers and Next Card:
html-output.md,validation.md,scanner.md— and, separately, land the Cooklang adjacencyguard upstream so the recorded divergence can be flipped to
enforced.