Skip to content

test: exercise the soak evidence contract at long parameterization - #73

Merged
test1card merged 14 commits into
masterfrom
test/soak-long-profile-contract
Aug 15, 2026
Merged

test1card merged 14 commits into
masterfrom
test/soak-long-profile-contract

Conversation

@test1card

Copy link
Copy Markdown
Owner

Two test-only modules that close the evidence-contract half of the milestone-1 long-soak blocker. No production file is touched, and the non-short refusal in scripts/soak_mock_stack.py is left exactly as it is.

Opened as a draft per the review-before-CI rule: it earns its check set after a clean review.

Why these, and not an activation change

The refusal says the runner and the evidence contract are validated only for the short profile. Those are two different objects and only one of them is actually short-only.

The contract is already partly exercised at long parameterization — test_soak_mock_stack.py:197 drives a full 51,841-point 72h series through the bounded slope estimator, and :264 uses the 12h thresholds. This finishes that half. The runner is the real gap, and it is not addressed here.

What each module does

test_soak_long_profile_contract.py drives full synthetic sample series at the real 12h and 72h profiles — read from PROFILES, never retyped — through evaluate_resources and the fault-correlation validator. Each rule carries a negative: an identity change within an epoch, a skipped replacement epoch, and a recovery over the ceiling must each be rejected. A validator that cannot be made to fail proves nothing, so the rejections are the substance.

test_soak_runner_rehearsal.py pins that the refusal follows the profile's own identity, not its lookup key. A compressed rehearsal preserving the 12h fault structure, substituted into PROFILES under the "short" key, is still refused with exit 3 and creates no evidence directory. Rebinding the mapping is not a way to launder a long-duration run past the gate.

That module began life as a deliberately-failing characterization of the missing seam. I did not land it that way — a knowingly-red node is a broken build, not a characterization — so it was inverted to pin the true statement adjacent to it. It also needs no POSIX gate, because the refusal returns before the platform-specific runner is imported, so it runs everywhere rather than being skipped on Windows.

The finding this leaves for a separate, reviewed change

The runner has no seam for a selected profile. scripts/soak_mock_stack_runner.py:3387 calls soak.profile("short") directly with its own _RunnerFoundationError, and :3419 writes "profile": "short" into the manifest unconditionally; soak_mock_stack.py::main never passes the selected profile through at all. So the multi-fault, multi-epoch mechanics the 12h profile describes have never been executed by anything — short runs two faults, 12h runs six.

One consequence worth stating because it bears on the threat model: deleting the CLI refusal would not produce fake long evidence. It would produce a short run honestly labelled short, because the runner hardcodes the profile one layer below. The system fails safe there.

Activation is a separate change and belongs behind a sealed short-profile PASS on a clean commit. The nightly soak has failed six consecutive nights (2026-08-09 through 08-14) on the missing Qt libraries that #72 fixes, so no sealed short PASS exists yet.

Controls

  • tests/scripts — 126 passed, 94 skipped.
  • ruff check and ruff format --check clean on both files.
  • Both files verified UTF-8, no byte-order mark, no carriage returns, and parsing.

Authored with AI assistance; every measurement above was re-run by the coordinator rather than taken from a worker's report.

Two test-only modules. No production file is touched, and the non-short refusal
in scripts/soak_mock_stack.py is deliberately left exactly as it is.

test_soak_long_profile_contract.py drives full synthetic sample series at the
real 12h and 72h profiles, read from PROFILES rather than retyped, through both
evaluate_resources and the fault-correlation validator. Each rule carries a
negative: an identity change within an epoch, a skipped replacement epoch, and a
recovery over the ceiling must each be rejected. A validator that cannot be made
to fail proves nothing, so the rejections are the point.

test_soak_runner_rehearsal.py pins that the refusal follows the profile's own
identity rather than its lookup key. A compressed rehearsal preserving the 12h
fault structure, substituted into PROFILES under the "short" key, is still
refused with exit 3 and creates no evidence directory. That is defence in depth:
rebinding the mapping is not a way to launder a long-duration run past the gate.
The refusal returns before the platform-specific runner is imported, so this
runs on every platform rather than being skipped off POSIX.

Recorded for a later, separately reviewed change rather than fixed here: the
runner has no seam for a selected profile. soak_mock_stack_runner.py calls
soak.profile("short") directly and writes "profile": "short" into the manifest
unconditionally, so the multi-fault, multi-epoch mechanics the 12h profile
describes have never been executed by anything. Closing that is the activation
work, and it belongs behind a sealed short-profile PASS on a clean commit.

Controls: tests/scripts 126 passed, 94 skipped. ruff check and
ruff format --check clean on both files.
@test1card

Copy link
Copy Markdown
Owner Author

@codex review

This is a draft and deliberately has no check set yet: under the review-before-CI rule it earns its checks once the review is clean, since a review costs no runner minutes and a full set costs about 28 jobs.

Two things I would value your eye on specifically:

  1. Are the three rejection cases the right ones? Identity-change-within-epoch, skipped replacement epoch, and recovery-over-ceiling. If a rule in evaluate_resources or the fault-correlation validator has no negative here, it is unguarded and I would rather know now.
  2. Is the rehearsal pin honest about what it proves? It asserts the refusal follows selected.name rather than the mapping key. It does not prove endurance and does not claim to. If the docstring overstates that, say so.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: c7c67c180c

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +100 to +102
identity_changed = copy.deepcopy(samples)
identity_changed[-1]["roles"]["engine"]["pid"] += 1
assert any("identity changed within epoch" in error for error in soak.evaluate_resources(identity_changed, profile))

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make the 72-hour resource thresholds load-bearing

If evaluate_resources stops enforcing the aggregate slope checks or accidentally applies the 12-hour ceilings to the 72-hour profile, this suite remains green because every resource counter in _samples is constant and the only resource-series negative changes a PID. This leaves the stricter 72-hour RSS and descriptor limits unguarded; add mutations with slopes between the 72-hour and 12-hour ceilings so the profile-specific rules demonstrably reject them.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Comment on lines +104 to +108
skipped_epoch = copy.deepcopy(samples)
event = profile.events[0]
for sample in skipped_epoch:
if sample["elapsed_s"] > event.at_s:
sample["roles"][event.target]["epoch"] += 1

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Add a negative for the complete long fault schedule

If _validate_faults stops enforcing the exact missing/duplicated/reordered schedule, all of these tests still pass: the positive fixture supplies a valid schedule, while this mutation changes only the first replacement epoch and the recovery mutation changes only the first ledger entry. Consequently, the defining six- and ten-fault long schedules are not load-bearing; delete, reorder, or duplicate a later fault and assert the schedule rejection.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Comment on lines +73 to +74
rehearsal = _compressed_12h_rehearsal()
monkeypatch.setitem(soak.PROFILES, "short", rehearsal)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Isolate profile name in the refusal rehearsal

If main changes to refuse based on duration, event count, warm-up, or another long-profile attribute instead of selected.name, this test still passes because _compressed_12h_rehearsal() differs from the short profile in all of those dimensions. It therefore proves only that the mapping key is not the sole criterion, not that the refusal follows the selected profile's name as claimed; use the short profile with only name replaced for this pin and keep fault-structure compression as a separate assertion.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Comment on lines +114 to +117
excessive_recovery = copy.deepcopy(faults)
excessive_recovery[0]["recovery_s"] = soak.RECOVERY_CEILING_S + soak.SAMPLE_INTERVAL_S
assert any(
"recovery exceeded ceiling" in error for error in soak._validate_faults(excessive_recovery, profile, samples)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Exercise recovery readiness and target-specific health failures

If _validate_faults stops requiring ready, engine bridge-data resumption, or a newer assistant H3 heartbeat, the new long-profile tests remain green because every generated fault supplies the successful values and the sole recovery negative changes only the numeric duration. The existing assert_replacement unit test exercises a different helper, so it cannot protect these validator branches; mutate each target's recovery authority fields here and assert their specific rejection.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Comment on lines +9 to +13
It also records, as an executable statement rather than prose, what the runner
still lacks: a reviewed seam that passes the SELECTED profile through to
``_PosixSoakRunner``. Today ``soak_mock_stack_runner`` calls
``soak.profile("short")`` directly and writes ``"profile": "short"`` into the
manifest unconditionally, so the multi-fault, multi-epoch mechanics that the 12h

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Describe the untested runner seam as prose

This module does not make an executable statement about _PosixSoakRunner: no test imports or invokes the runner or checks its hard-coded profile and manifest, and the refusal test deliberately returns before the runner import. A future runner seam could therefore be added or its hard-coding changed without affecting this module, so describing the gap as executable evidence overstates what was collected; either label it as explanatory prose or add a focused runner-level assertion.

AGENTS.md reference: AGENTS.md:L16-L18

Useful? React with 👍 / 👎.

def _samples(profile: soak.SoakProfile) -> list[dict[str, object]]:
role_indexes = {role: index for index, role in enumerate(soak.ROLES)}
result = []
for elapsed_s in range(0, int(profile.duration_s) + 1, int(soak.SAMPLE_INTERVAL_S)):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reject truncated long-profile sample series

If validate_sample_series stops using profile.duration_s or regresses to a short-profile duration constant, these tests remain green because _samples always generates through the exact endpoint and none of the negatives removes the tail. The existing short-profile truncation test cannot detect a 72-hour profile being accepted after only 12 hours, so add a truncated 12-hour/72-hour series and assert the profile-duration rejection.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Comment on lines +40 to +44
"python",
100,
1,
2,
True,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make later-epoch recovery envelopes fail under mutation

If evaluate_resources checks only the first restart or compares later recoveries against the wrong baseline, this long-profile suite remains green because RSS, thread, and descriptor values stay flat across every generated epoch. The existing resource test has only one synthetic engine transition, so it does not guard the repeated recovery path that is unique to the six- and ten-fault profiles; raise the first counters of a later epoch above each recovery envelope and assert the later-epoch errors.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

The pair is generated from the index, so adding two Python files moves it. I
pushed this branch without regenerating, which is what reddened the docs gate
and the remaining partition.
Re-derived as a set difference both ways rather than edited to match: master
holds 694 paths, this head holds 696, the two that entered are exactly the new
test modules, and none left.
…al variable

Closes the six review findings, all of the same shape: a production rule that
could be deleted with this suite staying green.

The 72h resource thresholds are now load-bearing: a series whose growth the 12h
ceilings would accept and the 72h ceilings reject is asserted rejected, so the
profile-specific limit is what fails it rather than any limit at all.

The fault schedule gains negatives for a missing, duplicated and reordered
schedule, and for recovery readiness and target-specific health, where every
generated case previously succeeded. A truncated sample tail is now rejected,
where the fixture previously always ran to the exact endpoint so duration_s was
not load-bearing.

The refusal pin now isolates its variable. It previously compared a compressed
rehearsal that differed from the activated profile in duration, event count and
warm-up as well as name, so a main() that refused on duration would have kept it
green. It now renames the short profile and changes nothing else, so only the
name can account for the refusal.

The rehearsal module's docstring no longer claims coverage of the runner seam it
does not execute. That claim belongs in the pull request body, not in a
docstring, and it is there.

Controls: tests/scripts 129 passed, 94 skipped, up from 126. ruff check and
ruff format --check clean; both files verified UTF-8, no byte-order mark, no
carriage returns.
@test1card

Copy link
Copy Markdown
Owner Author

@codex please review at head fe810883bca4e45cbb3292aa2160f0db9f608c98.

All six findings are closed. Every one was the same shape — a production rule that could be deleted with this suite staying green — and that framing is what made them findable, so thank you for it.

1. The 72h resource thresholds are now load-bearing. A series whose growth the 12h ceilings would accept and the 72h ceilings reject is asserted rejected, so the profile-specific limit is what fails it rather than any limit at all. Previously every resource counter in the fixture was constant.

2 and 3. The fault schedule and per-target health gain their negatives. Missing, duplicated and reordered schedules, plus recovery readiness and target-specific health, where every generated case previously succeeded.

4. A truncated sample tail is rejected. The fixture always ran to the exact endpoint, so profile.duration_s was not load-bearing.

5. The refusal pin now isolates its variable, and your point here corrected a claim I had made elsewhere. It previously compared a compressed rehearsal differing from the activated profile in duration, event count and warm-up as well as name — so a main that refused on duration would have kept it green, and my summary calling it a pin on selected.name was wrong. It now renames the short profile and changes nothing else, so only the name can account for the refusal.

6. The docstring no longer claims runner coverage the module does not have. That claim belongs in this pull request body rather than in a docstring, and it is stated below.

The finding for the separate activation change, unchanged and unaddressed here on purpose: scripts/soak_mock_stack_runner.py:3387 calls soak.profile("short") directly and :3419 writes "profile": "short" into the manifest unconditionally, while soak_mock_stack.py::main never passes the selected profile through. So the multi-fault, multi-epoch mechanics the 12h profile describes have never been executed — short runs two faults, 12h runs six. Nothing in this pull request executes against _PosixSoakRunner, and it no longer pretends to.

Also corrected since the last head, and it was my error rather than a finding: the branch went red on the docs gate and the remaining partition because I opened it without regenerating the derived pair. The pair is generated from the index, so adding two Python files moves it. The pair is regenerated to a fixed point and committed, and the changed-Python count was re-derived as a set difference both ways — master 694 paths, this head 696, the two that entered are exactly these modules, none left.

Controls at this head: tests/scripts — 129 passed, 94 skipped, up from 126. tests/docs — 67 passed with TRUSTED_BASE_SHA bound as CI binds it. ruff check and ruff format --check clean, and both files verified UTF-8 with no byte-order mark and no carriage returns.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: fe810883bc

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +126 to +130
for sample in samples:
elapsed_s = float(sample["elapsed_s"])
growth = int((elapsed_s - profile.warmup_s) * 0.1 / 3600) if elapsed_s >= profile.warmup_s else 0
for role in ("launcher", "engine", "bridge"):
sample["roles"][role]["descriptors"] += growth

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Exercise the 72-hour RSS slope ceiling

Fresh evidence at fe810883: the corrective mutation changes only descriptors, while every generated rss_bytes value remains constant at 100. Consequently, replacing the 72-hour RSS ceiling (1 MiB/hour) with the 12-hour ceiling (4 MiB/hour) still leaves this suite green; the pre-existing RSS test uses 5 MiB/hour and therefore cannot distinguish those profiles either. Add an RSS mutation between the two ceilings and assert the 72-hour RSS-specific rejection.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

The previous round made the descriptor slope load-bearing and left its sibling
untouched: every generated rss_bytes stayed constant, so replacing the 72-hour
RSS ceiling with the 12-hour one left the suite green.

The new negative isolates the ceiling rather than merely tripping some limit.
Both per-profile limits are read from the module, the test asserts the 72-hour
limit really is the tighter of the two, and the injected growth is derived as
the midpoint between them rather than hardcoded. The unmutated fixture is then
asserted clean, the 72-hour evaluation is asserted to report the aggregate RSS
slope error, and the 12-hour evaluation is asserted NOT to. That pairing is what
proves the profile-specific ceiling is what fails the series.

The general lesson, recorded in the layer: closing a "rule not load-bearing"
finding on one rule does not close it on its siblings. The fields the fixture
holds constant are where the remaining ones hide.

Controls: tests/scripts 130 passed, 94 skipped, up from 129. ruff check and
ruff format --check clean.
@test1card

Copy link
Copy Markdown
Owner Author

@codex please review at head f966c44b29abd6c54a94a5b7091b410ec572abf3.

Your finding is closed, and it was the same defect as the round before it, surviving in a sibling I had not enumerated.

The previous round made the descriptor slope load-bearing and left the RSS slope alone: every generated rss_bytes stayed constant at 100, so replacing the 72-hour ceiling with the 12-hour one still left the suite green. Exactly as you said.

The new negative isolates the ceiling rather than merely tripping some limit:

  • both per-profile RSS limits are read from the module, never retyped;
  • the test first asserts the 72-hour limit really is the tighter of the two, so the premise is checked rather than assumed;
  • the injected growth is derived as the midpoint between the two limits, so it follows the profiles if they change;
  • the unmutated fixture is asserted clean;
  • the 72-hour evaluation is asserted to report aggregate RSS slope exceeded profile limit, and the 12-hour evaluation is asserted not to.

That last pairing is the point: it proves the profile-specific ceiling is what fails the series, not any ceiling at all.

The lesson I have taken from this pair of findings, and written into our layer: closing a "rule not load-bearing" finding on one rule does not close it on its siblings, and the fields the fixture holds constant are where the remaining ones hide. If there is a third such field in this module that I still have not enumerated, that is the finding I most want from you.

Controls at this head: tests/scripts — 130 passed, 94 skipped, up from 129. tests/docs — 67 passed with TRUSTED_BASE_SHA bound as CI binds it. ruff check and ruff format --check clean. The derived pair is regenerated as the last commit.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: f966c44b29

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +128 to +130
growth = int((elapsed_s - profile.warmup_s) * 0.1 / 3600) if elapsed_s >= profile.warmup_s else 0
for role in ("launcher", "engine", "bridge"):
sample["roles"][role]["descriptors"] += growth

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Exercise the per-role descriptor slope rule

Each role grows by only 0.1 descriptors/hour, below the 72-hour per-role ceiling of 0.25, while the three roles aggregate to 0.3 and trip only the aggregate check. Consequently, deleting both per-role descriptor-slope branches in evaluate_resources, or applying the 12-hour ceiling there, leaves every assertion green. Add a mutation whose single-role slope is derived between the two profile limits and assert the role/epoch-specific rejection.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Comment on lines +154 to +157
errors_72h = soak.evaluate_resources(samples, profile_72h)
errors_12h = soak.evaluate_resources(samples, profile_12h)
assert "aggregate RSS slope exceeded profile limit" in errors_72h
assert "aggregate RSS slope exceeded profile limit" not in errors_12h

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make the per-role 72-hour RSS ceiling load-bearing

When the 72-hour per-role RSS check accidentally uses the 12-hour ceiling, this new midpoint series still passes the suite: the aggregate check continues to emit the asserted message, while no assertion requires launcher epoch 0 RSS slope exceeded profile limit; the existing 12-hour per-role test also remains green because its 5 MiB/hour input exceeds that substituted ceiling. Assert the launcher-specific error for 72 hours and its absence for 12 hours so this separate use of the profile limit is protected.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Comment on lines +40 to +42
"python",
100,
1,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make the long-profile warm-up boundary load-bearing

Because _samples keeps every resource counter flat throughout startup, all of these tests still pass if evaluate_resources ignores profile.warmup_s or reuses the short profile's 180-second boundary instead of the long profiles' 600-second boundary; the slope mutations begin only at 600 seconds and remain over their ceilings across the remaining hours. That regression would include startup transients that the long-profile contract explicitly excludes and could falsely fail otherwise valid evidence, so add a resource spike confined between the short and long warm-up boundaries and assert that the long evaluation remains clean.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Comment on lines +65 to +66
"scheduled_s": float(event.at_s),
"observed_s": float(event.at_s),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make the fault schedule tolerance fail under mutation

Every generated record makes observed_s exactly equal to scheduled_s, and no existing fault test changes that relationship, so removing the SAMPLE_INTERVAL_S tolerance check from _validate_faults or accidentally validating the scheduled value twice leaves the suite green. This would allow evidence from a substantially late fault injection to qualify; mutate a later long-profile record's observed time beyond the tolerance without changing its scheduled time and assert the schedule-tolerance rejection.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Comment on lines +67 to +70
"pre_pid": 10 + index + before * 100,
"pre_started_ns": 100 + index + before * 100,
"recheck_pid": 10 + index + before * 100,
"recheck_started_ns": 100 + index + before * 100,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Exercise the immediate identity recheck

The helper always copies the same identity into the pre-fault and immediate-recheck fields, while the existing wrong_pre negative changes both fields together and therefore exercises only correlation with the preceding sample. If _validate_faults stops requiring the recheck identity to equal the pre-fault identity, every test remains green even though the receipt can no longer prove that the process identity observed before signaling was still current; mutate only a recheck field and assert failed immediate identity recheck.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Comment on lines +114 to +117
excessive_recovery = copy.deepcopy(faults)
excessive_recovery[0]["recovery_s"] = soak.RECOVERY_CEILING_S + soak.SAMPLE_INTERVAL_S
assert any(
"recovery exceeded ceiling" in error for error in soak._validate_faults(excessive_recovery, profile, samples)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Isolate recovery-time correlation from the ceiling

The sole recovery-time negative sets the recorded value above RECOVERY_CEILING_S, which currently produces both the asserted ceiling error and the unasserted sample-history mismatch. Removing the independent check that the receipt's recovery_s equals the elapsed time of the observed replacement therefore leaves this test green, allowing a within-ceiling but fabricated recovery duration to qualify; use a different in-ceiling value from the five-second sample history and assert recovery does not match sample history.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Landed by the batch lander. The lane's own report and the coordinator's
verification are recorded on the pull request.
@test1card

Copy link
Copy Markdown
Owner Author

@codex review this pull request at head 43ed105.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 43ed1051dd

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +135 to +136
if soak.PROFILES["short"].warmup_s <= sample["elapsed_s"] < profile.warmup_s:
sample["roles"]["launcher"]["rss_bytes"] += 100 * 1024 * 1024

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make the warm-up mutation fail with the short boundary

Fresh evidence at 43ed105: evaluating this fixture with the 72-hour profile copied to use the short profile's 180-second warmup_s still returns []. The added 100 MiB plateau begins at 180 seconds and drops before 600 seconds, so its fitted growth and robust slope are non-positive and the test remains green if evaluate_resources regresses to the short warm-up boundary. Use a rising transient confined to 180–600 seconds that fails when included but is clean under the correct 600-second boundary.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Landed by the batch lander. The lane's own report and the coordinator's
verification are recorded on the pull request.
Landed by the batch lander. The lane's own report and the coordinator's
verification are recorded on the pull request.
The test contrasts two warm-up lengths: a ramp that finishes inside the long
profile's warm-up is discarded by it, but is visible to a profile with a short
warm-up and must trip the 50 MiB per-role growth cap.

The ramp was added only INSIDE the window, so resident memory fell back to
baseline afterwards. Growth is measured from the first post-warm-up sample to
the last, so a bump that comes back down nets to zero and trips nothing. The
assertion therefore saw an empty error list and the test failed while the
evaluator was behaving correctly.

The ramp is now held at its peak after the window, which is what makes the two
warm-up lengths disagree and is the property the test set out to pin.

tests/scripts: 130 passed, 94 skipped.
@test1card

Copy link
Copy Markdown
Owner Author

@codex review this pull request at head 0f64fed.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. You're on a roll.

Reviewed commit: 0f64fedac4

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@test1card
test1card marked this pull request as ready for review August 15, 2026 21:50
@test1card
test1card merged commit 9433d3d into master Aug 15, 2026
36 of 37 checks passed
@test1card
test1card deleted the test/soak-long-profile-contract branch August 15, 2026 21:52

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0f64fedac4

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

short_warmup = replace(profile, warmup_s=soak.PROFILES["short"].warmup_s)
assert "launcher epoch 0 RSS growth reached 50 MiB" in soak.evaluate_resources(warmup_ramp, short_warmup)
errors = soak.evaluate_resources(samples, profile)
assert "aggregate descriptor slope exceeded profile limit" in errors

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Prove the 12-hour descriptor ceiling accepts the midpoint

If the 12-hour descriptor ceiling is accidentally replaced with the stricter 72-hour value, every test still passes: both midpoint mutations are evaluated only with profile bound to PROFILES["72h"], while the 12-hour positive cases keep descriptors flat. That regression would falsely reject valid 12-hour evidence, so generate corresponding 12-hour aggregate and per-role midpoint series and assert that their descriptor-slope errors are absent.

AGENTS.md reference: AGENTS.md:L370-L379

Useful? React with 👍 / 👎.

Comment on lines +89 to +90
assert soak.evaluate_resources(samples, profile) == []
assert soak._validate_faults(_faults(profile), profile, samples) == []

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Exercise long profiles through the evidence-sealing entry point

These assertions invoke the two validators directly, so they do not prove that Evidence._build_ledger selects and passes the manifest's long profile on the production sealing path. If that wiring is changed to use short or otherwise drops the selected profile, the existing short sealing tests and every new long-profile test remain green while 12-hour or 72-hour evidence is judged against the wrong thresholds and schedule; add a long-profile sealing-path test rather than validating only these stand-ins.

AGENTS.md reference: AGENTS.md:L380-L386

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant