Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 5 additions & 2 deletions deploy/agent/codex/Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -17,8 +17,11 @@ RUN apt-get update -qq && apt-get install -y -qq --no-install-recommends \
ca-certificates git openssh-client \
&& rm -rf /var/lib/apt/lists/*

# Pinned to the version the adapter's event parser was measured against (RUN-TOPOLOGY §10).
RUN npm i -g @openai/codex@0.146.0 && npm cache clean --force
# Pinned. The event parser was first measured on 0.146.0 (RUN-TOPOLOGY §10); raised to 0.156.1 on
# 2026-09-23 because the model list is read from this CLI, and 0.146.0 predates the gpt-6 models. Every
# flag the adapter uses was re-checked on 0.156.1; the --json event shape is re-proved by the next paid
# run (UNVERIFIED section B).
RUN npm i -g @openai/codex@0.156.1 && npm cache clean --force

# The mount points exist, owned by the agent user, so a fresh named volume mounted there inherits
# that ownership; without this the volume is root's and the agent cannot write its own workspace.
Expand Down
16 changes: 10 additions & 6 deletions docs/UNVERIFIED.md
Original file line number Diff line number Diff line change
Expand Up @@ -413,13 +413,16 @@ so in the outage this exists for it fails too.
What the operator sees depends on WHEN the outage began, and the two are not the same:

- **After the prompt was stored** — the row is PROMPTED and carries an expiry, so the screen counts down
and then shows a code that has plainly run out.
- **Before it** — the prompt is the only thing that ever writes {@code expires_at}, and the worker does
not check whether that send arrived either. The row stays PENDING with **no countdown at all**: the
screen says "starting the sign-in", indefinitely.
and then shows a code that has plainly run out. Two minutes past that expiry the orchestrator closes the
row itself as expired (`HarnessSignIns.resendUnclaimed`, 2026-09-23).
- **Before it** — the row stays PENDING. Since 2026-09-23 (review of PR #168) the orchestrator re-sends
an unanswered start every 30 seconds, and closes the row as
`sign_in_not_started` if no worker has shown a code within six minutes (a four-minute start window, carried in the start itself, plus two for the code to cross the bus). A worker that cannot
start a unit reports nothing, because another worker may hold it. So the screen no longer says "starting the sign-in" for ever; it ends within
the wait.

Neither resolves itself. Cancelling is the way out, and cancelling publishes before it changes the row,
so it too needs the broker back; until then the screen shows only the cancel's own error.
Both now end on their own, without the broker. What is still lost is the credential itself: a finished
sign-in whose result cannot be delivered is discarded, and the operator signs in again.

**Evidence needed.** None — this is a deliberate trade, not a suspicion. Closing it means a durable
terminal record the screen can read without the broker, which is the same transactional-outbox treatment
Expand Down Expand Up @@ -468,6 +471,7 @@ Each has a runbook mode. None has been run by an operator.
| The whole M1 lifecycle against a real forge | **Mode Q** | Cancel, steer, the watchdog, the push gate and the charge ledger have only ever met a WireMock LLM and a local origin |
| Corporate-only bundle → the failure it produces | Mode R §5 | The documented trap (internal forge works, model API fails) is asserted nowhere; it is the mistake an operator will actually make |
| A private-registry pull | Mode S §4 | Nothing pulls from a private registry in any test. `authFor` and the attachment are unit-tested; the *pull* is not |
| **Codex CLI 0.156.1 in the agent image** (2026-09-23) | none yet | Raised from 0.146.0 so the model list includes the gpt-6 models. Re-checked on 0.156.1: every flag the adapter passes, the API-key login, and the shape of the file it writes (`auth_mode=apikey`). NOT re-checked: the `--json` event stream the usage parser reads, and the device sign-in output. Both need a paid run or a real sign-in, and the first of each proves or breaks them |
| **Runs pinned to the image their model list came from** (M3.5 part M, 2026-09-23) | none yet | Choosing the pin is unit-tested against given daemon answers, and one real-daemon test pins a LOCAL build by its image id. No test pulls a registry image and pins it by its registry digest, and none runs two workers. So "two workers holding different images under one tag run the same one" is argued from the code, not watched. A local-only image is pinned by an id that exists on one daemon only: on a second worker such a run fails to pull — by design, but unobserved |
| **OIDC sessions actually renew instead of re-authenticating** | **Mode J check 11** (2026-09-10) | The bug it fixes needs a real browser, a real Keycloak and **fifteen elapsed minutes**. No suite here has any of the three: there are zero WebSocket client tests, and nothing observes a token reaching its `exp`. `OidcSessionsAreRenewedTest` asserts the four `application.yml` files *say* renewal is on — it cannot assert Quarkus *does* it |

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -26,8 +26,18 @@ public sealed interface HarnessImageCommand {
* @param harness the name the orchestrator dispatches under, echoed back so the answer files itself
* @param image the exact reference the orchestrator will run — the answer describes THIS image, and a
* different tag of the same repository may carry a different CLI and a different list
* @param askedAt when the orchestrator asked. The channels replay from their oldest record, so a
* question can arrive long after it was asked: the worker skips one too old to matter, and the
* answer carries this back so an older answer cannot replace a newer one (review of PR #168).
* Null in a question sent before this existed.
*/
record Describe(String requestId, String harness, String image) implements HarnessImageCommand {
record Describe(String requestId, String harness, String image, java.time.Instant askedAt)
implements HarnessImageCommand {

public Describe(String requestId, String harness, String image) {
this(requestId, harness, image, null);
}

public Describe {
if (requestId == null || requestId.isBlank()) throw new IllegalArgumentException("A request id is required");
if (harness == null || harness.isBlank()) throw new IllegalArgumentException("A harness is required");
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -41,14 +41,57 @@ public sealed interface HarnessSignInCommand {
* @param maxWaitSeconds how long the unit may wait for the operator. The vendor states a 15-minute
* code lifetime; this is the deployment's own ceiling, so a unit cannot outlive the code it is
* waiting on and sit holding a container for ever.
* @param requestedAt when the operator pressed start. The wait is counted from HERE, not from delivery:
* the channel replays from its oldest record and the orchestrator re-sends a start nobody picked
* up, so a start can arrive late or twice. One whose wait has already run out opens no unit, and
* a late one gets only the time left (review of PR #168). Null in a start sent before this existed.
* @param startWithinSeconds how long after the press a unit may still be OPENED, as distinct from how
* long a person may take to approve ({@code maxWaitSeconds}). The orchestrator ends a sign-in that
* shows no code a little after this, so a worker must not open one it could only prompt for once
* that has happened. Carried rather than configured twice, so the two sides cannot drift. Zero in
* a start sent before this existed, which then opens nothing.
*/
record Start(String signInId, String harness, String image, long maxWaitSeconds)
record Start(String signInId, String harness, String image, long maxWaitSeconds, java.time.Instant requestedAt,
long startWithinSeconds)
implements HarnessSignInCommand {

public Start(String signInId, String harness, String image, long maxWaitSeconds) {
this(signInId, harness, image, maxWaitSeconds, null, 0);
}

/** When the code must be on screen by; after this the orchestrator ends the row. Null if unknown. */
public java.time.Instant promptDeadline() {
return requestedAt == null ? null : requestedAt.plusSeconds(startWithinSeconds);
}

/** When the person's time to approve runs out, counted from the press. Null if unknown. */
public java.time.Instant approvalDeadline() {
return requestedAt == null ? null : requestedAt.plusSeconds(maxWaitSeconds);
}

/**
* Whether a unit opened now could still show its code in time: the code must be on screen by the
* end of the start window, and the unit may take {@code toPrompt} to print it.
*/
public boolean mayOpenAt(java.time.Instant now, java.time.Duration toPrompt) {
if (requestedAt == null || startWithinSeconds <= 0) return false;
return !now.plus(toPrompt).isAfter(requestedAt.plusSeconds(startWithinSeconds));
}

/** How long the unit may still wait, counted from the request; the full wait when that is unknown. */
public java.time.Duration remainingWait(java.time.Instant now) {
java.time.Duration full = java.time.Duration.ofSeconds(maxWaitSeconds);
if (requestedAt == null) return full;
java.time.Duration left = java.time.Duration.between(now, requestedAt.plus(full));
return left.isNegative() ? java.time.Duration.ZERO : (left.compareTo(full) > 0 ? full : left);
}

public Start {
if (signInId == null || signInId.isBlank()) throw new IllegalArgumentException("A sign-in id is required");
if (harness == null || harness.isBlank()) throw new IllegalArgumentException("A harness name is required");
if (image == null || image.isBlank()) throw new IllegalArgumentException("An agent image is required");
if (startWithinSeconds < 0 || startWithinSeconds > maxWaitSeconds) throw new IllegalArgumentException(
"A start window of " + startWithinSeconds + " seconds must lie within the wait of " + maxWaitSeconds);
if (maxWaitSeconds <= 0) throw new IllegalArgumentException(
"A sign-in wait of " + maxWaitSeconds + " seconds would end the unit before the operator"
+ " could read the code; the wait is a ceiling, not a switch");
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -68,13 +68,21 @@ enum Status {
* when it could not be reached. A run of this harness uses it rather than the tag,
* so it runs the image these models were read from (review of PR #167). Null in an
* answer sent before pins existed.
* @param askedAt when the question this answers was asked, echoed back. The cache keeps the answer to
* the NEWEST question, so a replayed older answer cannot roll it back. Null in an answer
* sent before this existed.
*/
record Described(String requestId, String harness, String image, Status status, List<Model> models,
String pinnedImage)
String pinnedImage, java.time.Instant askedAt)
implements HarnessImageResult {

public Described(String requestId, String harness, String image, Status status, List<Model> models) {
this(requestId, harness, image, status, models, null);
this(requestId, harness, image, status, models, null, null);
}

public Described(String requestId, String harness, String image, Status status, List<Model> models,
String pinnedImage) {
this(requestId, harness, image, status, models, pinnedImage, null);
}

public Described {
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -105,6 +105,8 @@ record Failed(String signInId, String cause, String detail) implements HarnessSi
public static final String UNIT_FAILED = "sign_in_unit_failed";
/** The CLI signed in, but as an API key rather than a subscription. */
public static final String WRONG_MODE = "sign_in_wrong_mode";
/** No worker picked the sign-in up before its wait ran out, re-sends included. */
public static final String NOT_STARTED = "sign_in_not_started";

public Failed {
if (signInId == null || signInId.isBlank()) throw new IllegalArgumentException("A sign-in id is required");
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
package dev.codespire.contract.command;

import org.junit.jupiter.api.Test;

import java.time.Duration;
import java.time.Instant;

import static org.junit.jupiter.api.Assertions.assertEquals;

/**
* A start's wait is counted from the operator's press, not from delivery, because a start can be
* replayed or re-sent (review of PR #168).
*/
class HarnessSignInStartTest {

private static final Instant PRESSED = Instant.parse("2026-09-23T07:00:00Z");

private static HarnessSignInCommand.Start start(Instant requestedAt) {
return new HarnessSignInCommand.Start("TEST-sign-in", "codex", "TEST-image", 840, requestedAt, 240);
}

/** A unit may be opened only if it can print its code before the start window closes. */
@Test
void aUnitMayBeOpenedOnlyWhileItsCodeCanStillBeShownInTime() {
Duration toPrompt = Duration.ofSeconds(60);
assertEquals(true, start(PRESSED).mayOpenAt(PRESSED.plusSeconds(180), toPrompt));
assertEquals(false, start(PRESSED).mayOpenAt(PRESSED.plusSeconds(181), toPrompt),
"the wait is far from over, but the code would land after the window");
}

@Test
void theDeadlinesAreCountedFromThePress() {
assertEquals(PRESSED.plusSeconds(240), start(PRESSED).promptDeadline());
assertEquals(PRESSED.plusSeconds(840), start(PRESSED).approvalDeadline());
}

@Test
void aStartWithNoWindowOrNoPressOpensNothing() {
assertEquals(false, new HarnessSignInCommand.Start("TEST-sign-in", "codex", "TEST-image", 840, PRESSED, 0)
.mayOpenAt(PRESSED, Duration.ZERO));
assertEquals(false, start(null).mayOpenAt(PRESSED, Duration.ZERO));
}

@Test
void aLateStartGetsOnlyWhatIsLeft() {
assertEquals(Duration.ofSeconds(240), start(PRESSED).remainingWait(PRESSED.plusSeconds(600)));
}

@Test
void aStartWhoseWaitIsGoneGetsNothing() {
assertEquals(Duration.ZERO, start(PRESSED).remainingWait(PRESSED.plusSeconds(900)));
}

/** A clock behind the orchestrator's must not grant more than the full wait. */
@Test
void neverMoreThanTheFullWait() {
assertEquals(Duration.ofSeconds(840), start(PRESSED).remainingWait(PRESSED.minusSeconds(60)));
}

@Test
void aStartFromBeforeTimesExistedGetsTheFullWait() {
assertEquals(Duration.ofSeconds(840), start(null).remainingWait(PRESSED));
}
}
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,19 @@ final class DlqTopics {

private static final Set<String> RUN_RESULT_TYPES = Set.of("RunStarted", "RunFinished", "RunFailed");

/**
* The harness side channels (M3.5 parts F and M). Without these a dead-lettered model-list or
* sign-in record was replayed onto {@code cs.commands}, where nothing reads it (review of PR #168).
*/
static final String HARNESS_IMAGE_COMMANDS = "cs.harness-image-commands";
static final String HARNESS_IMAGE_RESULTS = "cs.harness-image-results";
static final String HARNESS_SIGN_IN_COMMANDS = "cs.harness-sign-in-commands";
static final String HARNESS_SIGN_IN_RESULTS = "cs.harness-sign-in-results";

private static final Set<String> HARNESS_SIGN_IN_COMMAND_TYPES = Set.of("Start", "Cancel");

private static final Set<String> HARNESS_SIGN_IN_RESULT_TYPES = Set.of("Prompted", "Completed", "Failed");

private DlqTopics() {
}

Expand All @@ -74,6 +87,10 @@ static String forType(String type) {
if (RUN_RESULT_TYPES.contains(type)) {
return RUN_RESULTS;
}
if ("Describe".equals(type)) return HARNESS_IMAGE_COMMANDS;
if ("Described".equals(type)) return HARNESS_IMAGE_RESULTS;
if (HARNESS_SIGN_IN_COMMAND_TYPES.contains(type)) return HARNESS_SIGN_IN_COMMANDS;
if (HARNESS_SIGN_IN_RESULT_TYPES.contains(type)) return HARNESS_SIGN_IN_RESULTS;
return COMMANDS;
}
}
Original file line number Diff line number Diff line change
Expand Up @@ -81,13 +81,16 @@ void onStart(@Observes StartupEvent event) {
* because a worker that was down at startup would otherwise leave the cache empty until the next
* orchestrator restart. The answer is cheap: a label read, and a pull only the first time.
*/
@Scheduled(every = "${spire.harness-catalogue-interval:10m}",
// Delayed: the startup ask already covers boot, and a first tick at startup ran before the Kafka
// emitter was connected, logging an injection ERROR on every start (dev stack, 2026-09-23).
@Scheduled(every = "${spire.harness-catalogue-interval:10m}", delayed = "1m",
concurrentExecution = Scheduled.ConcurrentExecution.SKIP)
void refresh() {
for (Map.Entry<String, String> harness : config.agentImage().entrySet()) {
try {
KafkaSends.sendAndAwait(commands, harness.getKey(),
new HarnessImageCommand.Describe(UUID.randomUUID().toString(), harness.getKey(), harness.getValue()),
new HarnessImageCommand.Describe(UUID.randomUUID().toString(), harness.getKey(), harness.getValue(),
Instant.now()),
"describe the image for " + harness.getKey());
} catch (RuntimeException undelivered) {
// The next interval asks again; one lost question is not worth failing startup over.
Expand All @@ -110,17 +113,27 @@ public void record(HarnessImageResult.Described answer) {
return;
}
try (Connection c = dataSource.getConnection(); PreparedStatement ps = c.prepareStatement("""
INSERT INTO harness_catalogue (harness, image, status, models, observed_at, pinned_image)
VALUES (?, ?, ?, ?::jsonb, now(), ?)
INSERT INTO harness_catalogue (harness, image, status, models, observed_at, pinned_image, asked_at)
VALUES (?, ?, ?, ?::jsonb, now(), ?, ?)
ON CONFLICT (harness) DO UPDATE
SET image=excluded.image, status=excluded.status, models=excluded.models, observed_at=now(),
pinned_image=excluded.pinned_image
pinned_image=excluded.pinned_image, asked_at=excluded.asked_at
-- The answer to the NEWEST question wins, not the last one to arrive: the channel
-- replays from its oldest record, and an older answer must not roll the list and its
-- pin back (review of PR #168). An answer with no time predates times, so it can only
-- be a replay and replaces nothing. A row about an image no longer configured is stale
-- whatever its time, so any timed current answer replaces it.
WHERE excluded.asked_at IS NOT NULL
AND (harness_catalogue.image <> excluded.image
OR harness_catalogue.asked_at IS NULL
OR excluded.asked_at >= harness_catalogue.asked_at)
""")) {
ps.setString(1, answer.harness());
ps.setString(2, answer.image());
ps.setString(3, answer.status().name());
ps.setString(4, mapper.writeValueAsString(answer.models()));
ps.setString(5, answer.pinnedImage());
ps.setTimestamp(6, answer.askedAt() == null ? null : java.sql.Timestamp.from(answer.askedAt()));
ps.executeUpdate();
} catch (SQLException | JsonProcessingException failure) {
throw new IllegalStateException("The model catalogue for " + answer.harness() + " could not be stored", failure);
Expand Down
Loading
Loading