Skip to content

M3.5 part M: offer the models a harness runs, with their thinking levels - #167

Merged
artyomsv merged 7 commits into
masterfrom
feat/factory-m35-one-ticket-build
Sep 23, 2026
Merged

artyomsv merged 7 commits into
masterfrom
feat/factory-m35-one-ticket-build

Conversation

@artyomsv

Copy link
Copy Markdown
Owner

Why

The build setup offered codex every model on the price list, Claude and Gemini included, and had no
thinking level at all. Codex itself offers a model and a level together, and each model has its own
levels. The answer must also work on Kubernetes, where nothing can run the image to ask it.

Design: docs/superpowers/specs/2026-09-16-factory-m35-one-ticket-to-a-build-design.md §6A.

What changed

  1. The image says what it runs. deploy/agent/build-codex.sh builds the agent image, reads
    codex debug models, and bakes a trimmed catalogue into the label dev.codespire.agent.models.
    The image contract gains a models clause; spire-agent-image verify reports it.
  2. The run worker reads the label, the orchestrator caches it. New topics
    cs.harness-image-commands / cs.harness-image-results. The run worker reads labels with the same
    authenticated pull it uses for runs. The orchestrator stores the answer in harness_catalogue
    (V79), refreshes on start and every spire.harness-catalogue-interval (10m), and drops an answer
    for an image that is no longer current.
  3. The screen offers that list. The model select shows the harness's models; a model with no
    price is shown but blocked, with the reason. A thinking-level select lists the chosen model's own
    levels. The save refuses a model the harness does not run and a level the model does not offer.
    When the list is not known, models fall back to the price list and no level is offered (§6A.4a).
  4. The level reaches the run. Preparation binding version 3 hashes the level (versions 1 and 2
    keep their exact hashes). ExecuteRun.reasoningEffort carries it; the Codex arm adds
    -c model_reasoning_effort="<level>" only when one was chosen (§6A.4b).

Verification

  • testFast and testServices pass; UI 913 tests across 104 files, tsc and production build pass.
  • Mutation checks: 8 mutants on the level path (one per hop and guard), each killed by its own test.
    The model/level save check: removing it failed 3 tests.

Not in this PR

  • Paying with a Codex subscription (selection, injection, unmetered charging). It waits for a real
    sign-in through the new screen, so the credential shape is measured rather than assumed.

Found by the operator testing part B: the build setup offered every
enabled model in the LLM catalogue, filtered by nothing, so Codex was
offered Claude and Gemini models. The backend agreed and was equally
wrong -- BuildDefaults.save checks that a model is enabled and priced,
and never asks whether the harness can run it. So codex with
claude-opus-5 saves, and dies when the run starts: after an approval,
mid-item, which is the failure this milestone exists to remove.

Measured the same day in the pinned image. The catalogue is not a weaker
version of the harness's list, it is a different list: the reference
image's Codex knows eight models, the catalogue offers six for OpenAI,
and they have TWO in common. Filtering the catalogue by vendor does not
fix that -- it leaves a menu that is still mostly wrong and still missing
everything that works.

`codex debug models` renders the catalogue as JSON, works signed out, and
carries the thinking levels each model allows with its own default. So
the harness answers "which models and which levels", and the LLM
catalogue goes back to what it is good at: what a token costs, which only
an API-key run needs.

The image carries that answer rather than being asked for it. Running the
image to fill in a dropdown would, on Kubernetes, mean scheduling a pod
-- a capability no arm has yet, and one that would then be written twice.
Reading a label is the one thing every runtime must already do, since it
cannot pull otherwise. Trimmed to six fields the value is 864 bytes; the
raw document is 314 KB.

The usual objection to a second copy is drift, and it does not apply: the
value is generated FROM the binary in the image, during the build that
installs it, and sealed into the same artifact.

- A LABEL cannot read a RUN's output, so deploy/agent/build-codex.sh
  builds, asks the binary, then builds again passing the answer as a
  build argument. The second pass is cached but for the label layer.
- Base64, because the value is JSON and has to survive a Dockerfile, a
  shell, `docker inspect` and a Kubernetes manifest without a quote being
  eaten. The verifier decodes it before printing.
- If `codex debug models` ever moves or disappears, the IMAGE BUILD
  fails, naming the script. That is the whole reason it is read during a
  build rather than when somebody opens a page.

`models` joins the declared clauses of the image contract. An image
without it still conforms; the factory then has no model list for it and
says so rather than guessing. The three documents that gave the old
`docker build` line now give the script.
The orchestrator now knows which models each harness can run, and at
which thinking levels, without ever reading an image itself. It has no
container runtime, and on Kubernetes it never will. The run worker must
reach every agent image or no run could start, so it is the one asked:
it reads the image's catalogue label and reports, and the orchestrator
keeps the answer in harness_catalogue (V79). Every screen reads that
table, so no page waits on a runtime, and a deployment whose runtime arm
is unfinished degrades to the last answer rather than to none.

- RunRuntime gains imageLabels(image). On that port and not a new one
  because the registry credential lives there and must reach nothing but
  a pull; the Docker arm reuses the exact authenticated pull a run uses,
  so an image a run could start is an image that can be described. The
  default THROWS rather than answering empty: an empty map is a real
  answer ("declares nothing") and an arm that cannot look is a different
  fact.
- An empty list means four different things -- no catalogue, an
  unreadable one, an unreachable image, or nothing heard yet -- so each
  is named and kept beside the list. A screen that cannot tell them apart
  can only show an empty dropdown.
- One bad entry makes the whole label unreadable rather than yielding the
  entries that parsed: a list with a silent hole looks exactly like a
  complete one.
- An answer describing an image the harness no longer runs is dropped,
  and a stored one is not served once the configuration moves on. A
  different tag may carry a different CLI and a different list.
- The options endpoint returns each harness's offered models beside its
  name, in the vendor's order with its hidden ones left out, so the
  screen makes one call and cannot pair a harness with another's models.

Proved against a real daemon: the reference image's models are read from
metadata alone, and the test asserts no container was started.

Also turns the part C preparation sweep OFF under test, where every
other background loop already was. It had been running every 20 seconds
against whatever items the current test created, writing artifacts and
health rows into the middle of teardowns -- the actual cause of the
foreign-key failures that were being patched one teardown at a time.
The build setup now offers the models the chosen harness's image says it
runs, not every model on the price list. That is the defect the operator
reported: codex was offered Claude and Gemini, and a save accepted one.
Once the list is known, a model the harness cannot run is refused at
save -- here, where it can be fixed -- rather than when the run starts,
after an approval.

The price list keeps its real job. It no longer decides WHICH models are
offered, only whether each can be paid for: a model the harness runs but
nobody has priced is shown, so the operator learns it exists, and
blocked with "no price yet", because an API-key run needs a rate for
every token type the harness reports.

A thinking level is chosen with the model, because the levels belong to
it. Codex itself offers "Model and Effort" together, and the levels
differ per model -- one allows low to ultra, another only low to xhigh --
so the list is rebuilt from the chosen model and checked against that
model's own list at save. No level chosen means the model's own default,
a real choice the vendor publishes, stored as NULL (V80) rather than as a
guessed name. A level chosen for one model is cleared when the model
changes, since the next one may not offer it.

When the harness's list is NOT known the save is not blocked, and that is
a decision: a development stack with no run worker never hears the
answer, and refusing every save would lock the operator out of the build
setup over a reply that has not come. The screen says which of the four
reasons it is and falls back to the price list, as before. A thinking
level is refused in that case, since there is nothing to check it
against. Recorded in the design as 6A.4a.

Mutation-verified: removing the harness check fails three tests.
The build setup saved a thinking level, but nothing carried it to the
run, so every build ran at the model's own default. The level now
travels the whole way:

- WorkPreparation binds it under a new binding version 3, so gates
  opened under versions 1 and 2 keep their exact hashes, and a level
  may not ride on a version that does not hash it.
- ExecuteRun carries it as a nullable reasoningEffort, with an
  atEffort wither; WorkRunAssembly sets it from the approved
  preparation.
- The run worker puts it on HarnessInvocation, and the Codex arm adds
  -c model_reasoning_effort="<level>" only when a level was chosen.

The Codex CLI does not validate the level, so every hop accepts only a
short lower-case word. Eight mutants, one per hop and guard, each fail
exactly the test written for them.
Review of PR #167 found four gaps; this closes three and part of the
fourth.

- A setup was checked against the image only when saved. If the image
  changed afterwards, the task was still prepared and the build still
  dispatched. HarnessCatalogues.refusal now holds the one check, and
  the save, the preparation sweep and the dispatch all ask it. The
  three reasons are named on the item page.
- An image could declare a level such as "x-high": offered, saved,
  then refused at preparation. ThinkingLevel is now the one shape rule,
  used by the declared model, the preparation and the run command, so
  such a label reads as UNREADABLE.
- The build screen kept a level across a harness change, and a saved
  level with no known list had no control on screen yet was sent. The
  level is dropped on a harness change; an unusable one is named, with
  a way back to the model default, and Save waits.
- Image answers were keyed by a random id and handled unordered, so an
  older answer could replace a newer one. They are now keyed by harness
  and handled in order.

Still open: two workers holding different images under one mutable tag.
Two run workers can hold different images under one mutable tag, and
the deployment default is spire-agent-codex:latest. The model list
could be read from one worker's image while the build ran on another.

The run worker now reads the labels and the exact image in one
inspect: the registry digest when the image has one, else the daemon's
image id. The answer carries the pin, the orchestrator stores it with
the list (V81), and every run it sends - item builds, REST dispatches
and /fix runs - uses the pin instead of the tag. An item build takes
its model check and its pin from one read of the cache.

With nothing read yet, runs use the tag as before. A local-only image
pinned by id fails to pull on a worker that does not hold it, which is
the honest outcome. RunRuntime.imageLabels becomes describeImage.
Recorded in UNVERIFIED: no test pulls a registry image or runs two
workers.
Second review of PR #167 found two gaps in the image pin.

- The orchestrator consumed image answers unordered, so an older answer
  queued during an outage could still finish last and overwrite the
  newer list and pin. It now records them one at a time, which keeps
  the per-harness order the worker's key gives.
- Docker reports a Docker Hub image under its short name whatever
  spelling pulled it: docker.io/library/alpine:3.20 reports
  alpine@sha256:... The digest match now compares repositories in
  Docker's canonical form, so such an image is pinned by its portable
  digest instead of a local image id.
@artyomsv
artyomsv merged commit 4c23976 into master Sep 23, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant