Repository navigation
test: exercise the soak evidence contract at long parameterization - #73
Conversation
Two test-only modules. No production file is touched, and the non-short refusal
in scripts/soak_mock_stack.py is deliberately left exactly as it is.
test_soak_long_profile_contract.py drives full synthetic sample series at the
real 12h and 72h profiles, read from PROFILES rather than retyped, through both
evaluate_resources and the fault-correlation validator. Each rule carries a
negative: an identity change within an epoch, a skipped replacement epoch, and a
recovery over the ceiling must each be rejected. A validator that cannot be made
to fail proves nothing, so the rejections are the point.
test_soak_runner_rehearsal.py pins that the refusal follows the profile's own
identity rather than its lookup key. A compressed rehearsal preserving the 12h
fault structure, substituted into PROFILES under the "short" key, is still
refused with exit 3 and creates no evidence directory. That is defence in depth:
rebinding the mapping is not a way to launder a long-duration run past the gate.
The refusal returns before the platform-specific runner is imported, so this
runs on every platform rather than being skipped off POSIX.
Recorded for a later, separately reviewed change rather than fixed here: the
runner has no seam for a selected profile. soak_mock_stack_runner.py calls
soak.profile("short") directly and writes "profile": "short" into the manifest
unconditionally, so the multi-fault, multi-epoch mechanics the 12h profile
describes have never been executed by anything. Closing that is the activation
work, and it belongs behind a sealed short-profile PASS on a clean commit.
Controls: tests/scripts 126 passed, 94 skipped. ruff check and
ruff format --check clean on both files.
|
@codex review This is a draft and deliberately has no check set yet: under the review-before-CI rule it earns its checks once the review is clean, since a review costs no runner minutes and a full set costs about 28 jobs. Two things I would value your eye on specifically:
|
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c7c67c180c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| identity_changed = copy.deepcopy(samples) | ||
| identity_changed[-1]["roles"]["engine"]["pid"] += 1 | ||
| assert any("identity changed within epoch" in error for error in soak.evaluate_resources(identity_changed, profile)) |
There was a problem hiding this comment.
Make the 72-hour resource thresholds load-bearing
If evaluate_resources stops enforcing the aggregate slope checks or accidentally applies the 12-hour ceilings to the 72-hour profile, this suite remains green because every resource counter in _samples is constant and the only resource-series negative changes a PID. This leaves the stricter 72-hour RSS and descriptor limits unguarded; add mutations with slopes between the 72-hour and 12-hour ceilings so the profile-specific rules demonstrably reject them.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
| skipped_epoch = copy.deepcopy(samples) | ||
| event = profile.events[0] | ||
| for sample in skipped_epoch: | ||
| if sample["elapsed_s"] > event.at_s: | ||
| sample["roles"][event.target]["epoch"] += 1 |
There was a problem hiding this comment.
Add a negative for the complete long fault schedule
If _validate_faults stops enforcing the exact missing/duplicated/reordered schedule, all of these tests still pass: the positive fixture supplies a valid schedule, while this mutation changes only the first replacement epoch and the recovery mutation changes only the first ledger entry. Consequently, the defining six- and ten-fault long schedules are not load-bearing; delete, reorder, or duplicate a later fault and assert the schedule rejection.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
| rehearsal = _compressed_12h_rehearsal() | ||
| monkeypatch.setitem(soak.PROFILES, "short", rehearsal) |
There was a problem hiding this comment.
Isolate profile name in the refusal rehearsal
If main changes to refuse based on duration, event count, warm-up, or another long-profile attribute instead of selected.name, this test still passes because _compressed_12h_rehearsal() differs from the short profile in all of those dimensions. It therefore proves only that the mapping key is not the sole criterion, not that the refusal follows the selected profile's name as claimed; use the short profile with only name replaced for this pin and keep fault-structure compression as a separate assertion.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
| excessive_recovery = copy.deepcopy(faults) | ||
| excessive_recovery[0]["recovery_s"] = soak.RECOVERY_CEILING_S + soak.SAMPLE_INTERVAL_S | ||
| assert any( | ||
| "recovery exceeded ceiling" in error for error in soak._validate_faults(excessive_recovery, profile, samples) |
There was a problem hiding this comment.
Exercise recovery readiness and target-specific health failures
If _validate_faults stops requiring ready, engine bridge-data resumption, or a newer assistant H3 heartbeat, the new long-profile tests remain green because every generated fault supplies the successful values and the sole recovery negative changes only the numeric duration. The existing assert_replacement unit test exercises a different helper, so it cannot protect these validator branches; mutate each target's recovery authority fields here and assert their specific rejection.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
| It also records, as an executable statement rather than prose, what the runner | ||
| still lacks: a reviewed seam that passes the SELECTED profile through to | ||
| ``_PosixSoakRunner``. Today ``soak_mock_stack_runner`` calls | ||
| ``soak.profile("short")`` directly and writes ``"profile": "short"`` into the | ||
| manifest unconditionally, so the multi-fault, multi-epoch mechanics that the 12h |
There was a problem hiding this comment.
Describe the untested runner seam as prose
This module does not make an executable statement about _PosixSoakRunner: no test imports or invokes the runner or checks its hard-coded profile and manifest, and the refusal test deliberately returns before the runner import. A future runner seam could therefore be added or its hard-coding changed without affecting this module, so describing the gap as executable evidence overstates what was collected; either label it as explanatory prose or add a focused runner-level assertion.
AGENTS.md reference: AGENTS.md:L16-L18
Useful? React with 👍 / 👎.
| def _samples(profile: soak.SoakProfile) -> list[dict[str, object]]: | ||
| role_indexes = {role: index for index, role in enumerate(soak.ROLES)} | ||
| result = [] | ||
| for elapsed_s in range(0, int(profile.duration_s) + 1, int(soak.SAMPLE_INTERVAL_S)): |
There was a problem hiding this comment.
Reject truncated long-profile sample series
If validate_sample_series stops using profile.duration_s or regresses to a short-profile duration constant, these tests remain green because _samples always generates through the exact endpoint and none of the negatives removes the tail. The existing short-profile truncation test cannot detect a 72-hour profile being accepted after only 12 hours, so add a truncated 12-hour/72-hour series and assert the profile-duration rejection.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
| "python", | ||
| 100, | ||
| 1, | ||
| 2, | ||
| True, |
There was a problem hiding this comment.
Make later-epoch recovery envelopes fail under mutation
If evaluate_resources checks only the first restart or compares later recoveries against the wrong baseline, this long-profile suite remains green because RSS, thread, and descriptor values stay flat across every generated epoch. The existing resource test has only one synthetic engine transition, so it does not guard the repeated recovery path that is unique to the six- and ten-fault profiles; raise the first counters of a later epoch above each recovery envelope and assert the later-epoch errors.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
The pair is generated from the index, so adding two Python files moves it. I pushed this branch without regenerating, which is what reddened the docs gate and the remaining partition.
Re-derived as a set difference both ways rather than edited to match: master holds 694 paths, this head holds 696, the two that entered are exactly the new test modules, and none left.
…al variable Closes the six review findings, all of the same shape: a production rule that could be deleted with this suite staying green. The 72h resource thresholds are now load-bearing: a series whose growth the 12h ceilings would accept and the 72h ceilings reject is asserted rejected, so the profile-specific limit is what fails it rather than any limit at all. The fault schedule gains negatives for a missing, duplicated and reordered schedule, and for recovery readiness and target-specific health, where every generated case previously succeeded. A truncated sample tail is now rejected, where the fixture previously always ran to the exact endpoint so duration_s was not load-bearing. The refusal pin now isolates its variable. It previously compared a compressed rehearsal that differed from the activated profile in duration, event count and warm-up as well as name, so a main() that refused on duration would have kept it green. It now renames the short profile and changes nothing else, so only the name can account for the refusal. The rehearsal module's docstring no longer claims coverage of the runner seam it does not execute. That claim belongs in the pull request body, not in a docstring, and it is there. Controls: tests/scripts 129 passed, 94 skipped, up from 126. ruff check and ruff format --check clean; both files verified UTF-8, no byte-order mark, no carriage returns.
|
@codex please review at head All six findings are closed. Every one was the same shape — a production rule that could be deleted with this suite staying green — and that framing is what made them findable, so thank you for it. 1. The 72h resource thresholds are now load-bearing. A series whose growth the 12h ceilings would accept and the 72h ceilings reject is asserted rejected, so the profile-specific limit is what fails it rather than any limit at all. Previously every resource counter in the fixture was constant. 2 and 3. The fault schedule and per-target health gain their negatives. Missing, duplicated and reordered schedules, plus recovery readiness and target-specific health, where every generated case previously succeeded. 4. A truncated sample tail is rejected. The fixture always ran to the exact endpoint, so 5. The refusal pin now isolates its variable, and your point here corrected a claim I had made elsewhere. It previously compared a compressed rehearsal differing from the activated profile in duration, event count and warm-up as well as name — so a 6. The docstring no longer claims runner coverage the module does not have. That claim belongs in this pull request body rather than in a docstring, and it is stated below. The finding for the separate activation change, unchanged and unaddressed here on purpose: Also corrected since the last head, and it was my error rather than a finding: the branch went red on the docs gate and the Controls at this head: |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: fe810883bc
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| for sample in samples: | ||
| elapsed_s = float(sample["elapsed_s"]) | ||
| growth = int((elapsed_s - profile.warmup_s) * 0.1 / 3600) if elapsed_s >= profile.warmup_s else 0 | ||
| for role in ("launcher", "engine", "bridge"): | ||
| sample["roles"][role]["descriptors"] += growth |
There was a problem hiding this comment.
Exercise the 72-hour RSS slope ceiling
Fresh evidence at fe810883: the corrective mutation changes only descriptors, while every generated rss_bytes value remains constant at 100. Consequently, replacing the 72-hour RSS ceiling (1 MiB/hour) with the 12-hour ceiling (4 MiB/hour) still leaves this suite green; the pre-existing RSS test uses 5 MiB/hour and therefore cannot distinguish those profiles either. Add an RSS mutation between the two ceilings and assert the 72-hour RSS-specific rejection.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
The previous round made the descriptor slope load-bearing and left its sibling untouched: every generated rss_bytes stayed constant, so replacing the 72-hour RSS ceiling with the 12-hour one left the suite green. The new negative isolates the ceiling rather than merely tripping some limit. Both per-profile limits are read from the module, the test asserts the 72-hour limit really is the tighter of the two, and the injected growth is derived as the midpoint between them rather than hardcoded. The unmutated fixture is then asserted clean, the 72-hour evaluation is asserted to report the aggregate RSS slope error, and the 12-hour evaluation is asserted NOT to. That pairing is what proves the profile-specific ceiling is what fails the series. The general lesson, recorded in the layer: closing a "rule not load-bearing" finding on one rule does not close it on its siblings. The fields the fixture holds constant are where the remaining ones hide. Controls: tests/scripts 130 passed, 94 skipped, up from 129. ruff check and ruff format --check clean.
|
@codex please review at head Your finding is closed, and it was the same defect as the round before it, surviving in a sibling I had not enumerated. The previous round made the descriptor slope load-bearing and left the RSS slope alone: every generated The new negative isolates the ceiling rather than merely tripping some limit:
That last pairing is the point: it proves the profile-specific ceiling is what fails the series, not any ceiling at all. The lesson I have taken from this pair of findings, and written into our layer: closing a "rule not load-bearing" finding on one rule does not close it on its siblings, and the fields the fixture holds constant are where the remaining ones hide. If there is a third such field in this module that I still have not enumerated, that is the finding I most want from you. Controls at this head: |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f966c44b29
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| growth = int((elapsed_s - profile.warmup_s) * 0.1 / 3600) if elapsed_s >= profile.warmup_s else 0 | ||
| for role in ("launcher", "engine", "bridge"): | ||
| sample["roles"][role]["descriptors"] += growth |
There was a problem hiding this comment.
Exercise the per-role descriptor slope rule
Each role grows by only 0.1 descriptors/hour, below the 72-hour per-role ceiling of 0.25, while the three roles aggregate to 0.3 and trip only the aggregate check. Consequently, deleting both per-role descriptor-slope branches in evaluate_resources, or applying the 12-hour ceiling there, leaves every assertion green. Add a mutation whose single-role slope is derived between the two profile limits and assert the role/epoch-specific rejection.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
| errors_72h = soak.evaluate_resources(samples, profile_72h) | ||
| errors_12h = soak.evaluate_resources(samples, profile_12h) | ||
| assert "aggregate RSS slope exceeded profile limit" in errors_72h | ||
| assert "aggregate RSS slope exceeded profile limit" not in errors_12h |
There was a problem hiding this comment.
Make the per-role 72-hour RSS ceiling load-bearing
When the 72-hour per-role RSS check accidentally uses the 12-hour ceiling, this new midpoint series still passes the suite: the aggregate check continues to emit the asserted message, while no assertion requires launcher epoch 0 RSS slope exceeded profile limit; the existing 12-hour per-role test also remains green because its 5 MiB/hour input exceeds that substituted ceiling. Assert the launcher-specific error for 72 hours and its absence for 12 hours so this separate use of the profile limit is protected.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
| "python", | ||
| 100, | ||
| 1, |
There was a problem hiding this comment.
Make the long-profile warm-up boundary load-bearing
Because _samples keeps every resource counter flat throughout startup, all of these tests still pass if evaluate_resources ignores profile.warmup_s or reuses the short profile's 180-second boundary instead of the long profiles' 600-second boundary; the slope mutations begin only at 600 seconds and remain over their ceilings across the remaining hours. That regression would include startup transients that the long-profile contract explicitly excludes and could falsely fail otherwise valid evidence, so add a resource spike confined between the short and long warm-up boundaries and assert that the long evaluation remains clean.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
| "scheduled_s": float(event.at_s), | ||
| "observed_s": float(event.at_s), |
There was a problem hiding this comment.
Make the fault schedule tolerance fail under mutation
Every generated record makes observed_s exactly equal to scheduled_s, and no existing fault test changes that relationship, so removing the SAMPLE_INTERVAL_S tolerance check from _validate_faults or accidentally validating the scheduled value twice leaves the suite green. This would allow evidence from a substantially late fault injection to qualify; mutate a later long-profile record's observed time beyond the tolerance without changing its scheduled time and assert the schedule-tolerance rejection.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
| "pre_pid": 10 + index + before * 100, | ||
| "pre_started_ns": 100 + index + before * 100, | ||
| "recheck_pid": 10 + index + before * 100, | ||
| "recheck_started_ns": 100 + index + before * 100, |
There was a problem hiding this comment.
Exercise the immediate identity recheck
The helper always copies the same identity into the pre-fault and immediate-recheck fields, while the existing wrong_pre negative changes both fields together and therefore exercises only correlation with the preceding sample. If _validate_faults stops requiring the recheck identity to equal the pre-fault identity, every test remains green even though the receipt can no longer prove that the process identity observed before signaling was still current; mutate only a recheck field and assert failed immediate identity recheck.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
| excessive_recovery = copy.deepcopy(faults) | ||
| excessive_recovery[0]["recovery_s"] = soak.RECOVERY_CEILING_S + soak.SAMPLE_INTERVAL_S | ||
| assert any( | ||
| "recovery exceeded ceiling" in error for error in soak._validate_faults(excessive_recovery, profile, samples) |
There was a problem hiding this comment.
Isolate recovery-time correlation from the ceiling
The sole recovery-time negative sets the recorded value above RECOVERY_CEILING_S, which currently produces both the asserted ceiling error and the unasserted sample-history mismatch. Removing the independent check that the receipt's recovery_s equals the elapsed time of the observed replacement therefore leaves this test green, allowing a within-ceiling but fabricated recovery duration to qualify; use a different in-ceiling value from the five-second sample history and assert recovery does not match sample history.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
Landed by the batch lander. The lane's own report and the coordinator's verification are recorded on the pull request.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 43ed1051dd
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if soak.PROFILES["short"].warmup_s <= sample["elapsed_s"] < profile.warmup_s: | ||
| sample["roles"]["launcher"]["rss_bytes"] += 100 * 1024 * 1024 |
There was a problem hiding this comment.
Make the warm-up mutation fail with the short boundary
Fresh evidence at 43ed105: evaluating this fixture with the 72-hour profile copied to use the short profile's 180-second warmup_s still returns []. The added 100 MiB plateau begins at 180 seconds and drops before 600 seconds, so its fitted growth and robust slope are non-positive and the test remains green if evaluate_resources regresses to the short warm-up boundary. Use a rising transient confined to 180–600 seconds that fails when included but is clean under the correct 600-second boundary.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
Landed by the batch lander. The lane's own report and the coordinator's verification are recorded on the pull request.
Landed by the batch lander. The lane's own report and the coordinator's verification are recorded on the pull request.
The test contrasts two warm-up lengths: a ramp that finishes inside the long profile's warm-up is discarded by it, but is visible to a profile with a short warm-up and must trip the 50 MiB per-role growth cap. The ramp was added only INSIDE the window, so resident memory fell back to baseline afterwards. Growth is measured from the first post-warm-up sample to the last, so a bump that comes back down nets to zero and trips nothing. The assertion therefore saw an empty error list and the test failed while the evaluator was behaving correctly. The ramp is now held at its peak after the window, which is what makes the two warm-up lengths disagree and is the property the test set out to pin. tests/scripts: 130 passed, 94 skipped.
|
Codex Review: Didn't find any major issues. You're on a roll. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0f64fedac4
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| short_warmup = replace(profile, warmup_s=soak.PROFILES["short"].warmup_s) | ||
| assert "launcher epoch 0 RSS growth reached 50 MiB" in soak.evaluate_resources(warmup_ramp, short_warmup) | ||
| errors = soak.evaluate_resources(samples, profile) | ||
| assert "aggregate descriptor slope exceeded profile limit" in errors |
There was a problem hiding this comment.
Prove the 12-hour descriptor ceiling accepts the midpoint
If the 12-hour descriptor ceiling is accidentally replaced with the stricter 72-hour value, every test still passes: both midpoint mutations are evaluated only with profile bound to PROFILES["72h"], while the 12-hour positive cases keep descriptors flat. That regression would falsely reject valid 12-hour evidence, so generate corresponding 12-hour aggregate and per-role midpoint series and assert that their descriptor-slope errors are absent.
AGENTS.md reference: AGENTS.md:L370-L379
Useful? React with 👍 / 👎.
| assert soak.evaluate_resources(samples, profile) == [] | ||
| assert soak._validate_faults(_faults(profile), profile, samples) == [] |
There was a problem hiding this comment.
Exercise long profiles through the evidence-sealing entry point
These assertions invoke the two validators directly, so they do not prove that Evidence._build_ledger selects and passes the manifest's long profile on the production sealing path. If that wiring is changed to use short or otherwise drops the selected profile, the existing short sealing tests and every new long-profile test remain green while 12-hour or 72-hour evidence is judged against the wrong thresholds and schedule; add a long-profile sealing-path test rather than validating only these stand-ins.
AGENTS.md reference: AGENTS.md:L380-L386
Useful? React with 👍 / 👎.
Two test-only modules that close the evidence-contract half of the milestone-1 long-soak blocker. No production file is touched, and the non-short refusal in
scripts/soak_mock_stack.pyis left exactly as it is.Opened as a draft per the review-before-CI rule: it earns its check set after a clean review.
Why these, and not an activation change
The refusal says the runner and the evidence contract are validated only for the short profile. Those are two different objects and only one of them is actually short-only.
The contract is already partly exercised at long parameterization —
test_soak_mock_stack.py:197drives a full 51,841-point 72h series through the bounded slope estimator, and:264uses the 12h thresholds. This finishes that half. The runner is the real gap, and it is not addressed here.What each module does
test_soak_long_profile_contract.pydrives full synthetic sample series at the real12hand72hprofiles — read fromPROFILES, never retyped — throughevaluate_resourcesand the fault-correlation validator. Each rule carries a negative: an identity change within an epoch, a skipped replacement epoch, and a recovery over the ceiling must each be rejected. A validator that cannot be made to fail proves nothing, so the rejections are the substance.test_soak_runner_rehearsal.pypins that the refusal follows the profile's own identity, not its lookup key. A compressed rehearsal preserving the 12h fault structure, substituted intoPROFILESunder the"short"key, is still refused with exit 3 and creates no evidence directory. Rebinding the mapping is not a way to launder a long-duration run past the gate.That module began life as a deliberately-failing characterization of the missing seam. I did not land it that way — a knowingly-red node is a broken build, not a characterization — so it was inverted to pin the true statement adjacent to it. It also needs no POSIX gate, because the refusal returns before the platform-specific runner is imported, so it runs everywhere rather than being skipped on Windows.
The finding this leaves for a separate, reviewed change
The runner has no seam for a selected profile.
scripts/soak_mock_stack_runner.py:3387callssoak.profile("short")directly with its own_RunnerFoundationError, and:3419writes"profile": "short"into the manifest unconditionally;soak_mock_stack.py::mainnever passes the selected profile through at all. So the multi-fault, multi-epoch mechanics the 12h profile describes have never been executed by anything —shortruns two faults,12hruns six.One consequence worth stating because it bears on the threat model: deleting the CLI refusal would not produce fake long evidence. It would produce a short run honestly labelled short, because the runner hardcodes the profile one layer below. The system fails safe there.
Activation is a separate change and belongs behind a sealed short-profile PASS on a clean commit. The nightly soak has failed six consecutive nights (2026-08-09 through 08-14) on the missing Qt libraries that #72 fixes, so no sealed short PASS exists yet.
Controls
tests/scripts— 126 passed, 94 skipped.ruff checkandruff format --checkclean on both files.Authored with AI assistance; every measurement above was re-run by the coordinator rather than taken from a worker's report.