Conversation
The diagnostic paid for itself in one roundAt the pushed head One round ago the refusal said only: Four separate conditions had to hold and the message named none of them. It now That is the third time on this branch that fixing the diagnostic before What it means
From the launcher log the tree contains at least the launcher, the ZMQ bridge Why this matters beyond the soak: a process the launcher cannot classify is Also visible in the same log, and not yet a findingThe launcher receives its shutdown receipt from the engine and the engine What is running nowA worker names the unclassified process — identifier, parent, and command line, The plausible sources, none established: a Qt platform helper on a headless |
Master is RED on the target platform, and this branch is what fixes itMeasured on Ubuntu 22.04.5 LTS, Python 3.14.6, in clean worktrees cut from the
The failure master carriesThe child loads the system
The library the child needs is present in the environment and is not the one it Why this branch fixes itMaster computes the controlled That change was made for a different reason — a stock virtual environment What it means for the queueEvery open pull request inherits this red node, including ones that touch no It is therefore not an acceptance to be written per branch. It is one fix, here. |
Landed by the batch lander. The lane's own report and the coordinator's verification are recorded on the pull request.
|
@codex review Head is |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Landed by the batch lander. The lane's own report and the coordinator's verification are recorded on the pull request.
|
@codex review Head is |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
The soak RUNS. It sampled, it injected a fault, and the fault recovered.Measured at the pushed head
It then stopped at: Where the barriers have gone
The round now running is about the class, not the instanceThe refusal says a child did not recover and does not say which. That is the So this round names this child AND sweeps both soak modules for every refusal Diagnostic only. No bound moves, no condition is relaxed. Recorded, not chasedThe launcher also repeats at shutdown: That happens at |
Landed by the batch lander. The lane's own report and the coordinator's verification are recorded on the pull request.
|
@codex review Head is |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
The soak found a real endurance defect: a faulted engine never comes backMeasured at the pushed head This is the first finding on this branch that is about the product rather than What was measuredThe soak killed the engine on schedule at 185 seconds. It exited. Nothing In the same run, the bridge was faulted and did come back: So replacement exists for one role and either does not exist, or did not fire, What it costsThe objective is a week of continuous thermal-conductivity running without data This is the roadmap's own Also in the log, and possibly the same storyIf a killed engine puts the launcher into a permanent hold by design, then the What a worker is doingEstablishing, with lines, whether an engine replacement path exists at all; It is explicitly told to STOP and report rather than improvise if this turns out The sixty-second ceiling will not be raised to make this pass. |
The engine restart path exists, and a deliberate permanent HOLD blocks itThe worker did not fix this, and was right not to. Its finding, with lines:
Measured, not inferred: And its own conclusion:
What this means for the objective, stated plainlyThe So a sealed short-profile PASS is not reachable until the launcher can restart The bridge, faulted in the same run, came back: The three ways out, and all three are the owner's
Option 3 is the honest description of today. It should be known before anyone What the soak has now proven it can doNine barriers cleared, each named by evidence and each fix verified by reverting The instrument works. What it found is a real endurance limit. |
# Conflicts: # docs/CLAIM_CORRECTIONS.md # docs/architecture-montana-important.svg # docs/current_candidate_metrics.md # governance/agent_preventions_baseline.json # tests/docs/test_docs_freshness.py
|
@codex review Head is |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
@codex review Head under review: Full check set green with zero pending, and |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
Root cause of the short soak's
|
The endurance run on the laboratory machine: one barrier removed, a second one found at masterAll of this is measured on the laboratory WSL image, Ubuntu 22.04.5, in worktrees cut from the native Linux clone. The launcher barrier is behind usThe theme barrier reported here on 2026-08-19 is fixed by #86, and fixing it took two steps rather than one, because the first step moved the failure earlier instead of removing it:
Adding a directory to the isolated configuration broke the fixture seal, which walks an exact topology: a fixed set of files plus one empty directory. That is the seal doing its job, so the seal is what had to learn about the theme directory. The packs are now sealed the way the files are — identity, no links, exact mode 0o600, content hashed, and the directory re-listed after sealing so a file appearing mid-seal is a refusal rather than an unsealed byte. The second barrier is at MASTER, not on the branchRunning the same short soak at So it is not this branch, and it matters more than the barrier above: the storage layer cannot be imported inside the sealed exact-six child on the target machine. The endurance run stops before the program starts. What it is NOT — five hypotheses eliminated by measurement
Import order was also eliminated: Tools, so this is repeatable rather than a story: What is leftThe remaining difference between the passing probes and the failing child is the sealed snapshot itself: the child runs inside the extracted archive, with One more fact worth pinning: an earlier run of this soak at |
This draft is the roadmap item for the laboratory week, and its Ubuntu failure is a collision with #87Measured 2026-08-20, so that whoever picks this up does not re-derive it. Why it matters more than its draft status suggests. The soak's So the week-long guarantee is not blocked by a leak we have not found. It is blocked by having no instrument that can look for one, and this pull request is that instrument. What is already done, and does not need doing again. The judging half is on master and is profile-generic: every slope check reads the profile's own limit behind an The Ubuntu failure, opened rather than read from its nameBoth failing This branch carries its own #87 solves the same problem the opposite way. It resolves the directory, sets The owner's standing rule decides between them, and it is not close:
So this branch should adopt #87's approach in that area rather than the reverse, and the two will conflict there. Worth knowing before the rebase rather than during it. The Windows failure is NOT the same thing, and I am not claiming it is
and the candidate artifact carries no Present state, measured
The CI hang that was failing every Linux job for the whole of 2026-08-19 is fixed and merged as #92, so a rebase onto current master starts from a working gate. That was not true when this branch last ran. |
Follow-up, because the strict check is not wrong everywhere — only where the gate runsI said the strict On the laboratory machine — Ubuntu 22.04.5, the miniforge environment the soak worktrees actually use: Every condition passes there. So the design is not unusable in the laboratory — it is unusable on the hosted runner, which is where a required check runs, and that alone blocks this pull request. That narrows the repair rather than widening it. It is not a case of the check being wrong about what a safe library root is; it is a case of a refusal being the wrong response to a root it does not recognise. The owner's standing rule says exactly that:
which is the shape #87 already has: resolve the directory, set One thing I have not measured and am not guessing at: why the runner's prefix fails the check. The traceback proves it does; which of the five conditions it trips would need the probe run inside a job, and that costs a round. |
Purpose
Make the existing whole-stack soak runner usable for bounded 12-hour, 72-hour, and 168-hour Ubuntu 22.04 laboratory-readiness evidence.
This pull request adds no laboratory completion claim. It prepares the evidence path so a later run can measure continuity, persistence, resource growth, restart behavior, and final verified-OFF state on one frozen commit.
Why this is needed
The repository defined long profiles, but the production runner refused them. The prior path could not produce the evidence needed for a week-long thermal-conductivity campaign.
The correction keeps validation bounded. It verifies the exact runtime library set, avoids loading the full 168-hour sample history into memory, checks native-library identity across restart boundaries, and aligns the laboratory checklist with Ubuntu 22.04 profiles.
Exact candidate evidence
Candidate: 2198fe4
On the exact candidate in CryoDAQ-Lab-Ubuntu-22.04:
The skipped tests remain environment-gated. No physical instrument, dummy-load, 12-hour, 72-hour, or 168-hour result is claimed here.
Merge boundary
Keep this pull request draft until the exact-head Codex verdict is clean and every required hosted check succeeds. A later physical or duration run must use one frozen commit and retain its own immutable artifacts.