M3.5 part M: offer the models a harness runs, with their thinking levels - #167
Merged
Merged
Conversation
Found by the operator testing part B: the build setup offered every enabled model in the LLM catalogue, filtered by nothing, so Codex was offered Claude and Gemini models. The backend agreed and was equally wrong -- BuildDefaults.save checks that a model is enabled and priced, and never asks whether the harness can run it. So codex with claude-opus-5 saves, and dies when the run starts: after an approval, mid-item, which is the failure this milestone exists to remove. Measured the same day in the pinned image. The catalogue is not a weaker version of the harness's list, it is a different list: the reference image's Codex knows eight models, the catalogue offers six for OpenAI, and they have TWO in common. Filtering the catalogue by vendor does not fix that -- it leaves a menu that is still mostly wrong and still missing everything that works. `codex debug models` renders the catalogue as JSON, works signed out, and carries the thinking levels each model allows with its own default. So the harness answers "which models and which levels", and the LLM catalogue goes back to what it is good at: what a token costs, which only an API-key run needs. The image carries that answer rather than being asked for it. Running the image to fill in a dropdown would, on Kubernetes, mean scheduling a pod -- a capability no arm has yet, and one that would then be written twice. Reading a label is the one thing every runtime must already do, since it cannot pull otherwise. Trimmed to six fields the value is 864 bytes; the raw document is 314 KB. The usual objection to a second copy is drift, and it does not apply: the value is generated FROM the binary in the image, during the build that installs it, and sealed into the same artifact. - A LABEL cannot read a RUN's output, so deploy/agent/build-codex.sh builds, asks the binary, then builds again passing the answer as a build argument. The second pass is cached but for the label layer. - Base64, because the value is JSON and has to survive a Dockerfile, a shell, `docker inspect` and a Kubernetes manifest without a quote being eaten. The verifier decodes it before printing. - If `codex debug models` ever moves or disappears, the IMAGE BUILD fails, naming the script. That is the whole reason it is read during a build rather than when somebody opens a page. `models` joins the declared clauses of the image contract. An image without it still conforms; the factory then has no model list for it and says so rather than guessing. The three documents that gave the old `docker build` line now give the script.
The orchestrator now knows which models each harness can run, and at
which thinking levels, without ever reading an image itself. It has no
container runtime, and on Kubernetes it never will. The run worker must
reach every agent image or no run could start, so it is the one asked:
it reads the image's catalogue label and reports, and the orchestrator
keeps the answer in harness_catalogue (V79). Every screen reads that
table, so no page waits on a runtime, and a deployment whose runtime arm
is unfinished degrades to the last answer rather than to none.
- RunRuntime gains imageLabels(image). On that port and not a new one
because the registry credential lives there and must reach nothing but
a pull; the Docker arm reuses the exact authenticated pull a run uses,
so an image a run could start is an image that can be described. The
default THROWS rather than answering empty: an empty map is a real
answer ("declares nothing") and an arm that cannot look is a different
fact.
- An empty list means four different things -- no catalogue, an
unreadable one, an unreachable image, or nothing heard yet -- so each
is named and kept beside the list. A screen that cannot tell them apart
can only show an empty dropdown.
- One bad entry makes the whole label unreadable rather than yielding the
entries that parsed: a list with a silent hole looks exactly like a
complete one.
- An answer describing an image the harness no longer runs is dropped,
and a stored one is not served once the configuration moves on. A
different tag may carry a different CLI and a different list.
- The options endpoint returns each harness's offered models beside its
name, in the vendor's order with its hidden ones left out, so the
screen makes one call and cannot pair a harness with another's models.
Proved against a real daemon: the reference image's models are read from
metadata alone, and the test asserts no container was started.
Also turns the part C preparation sweep OFF under test, where every
other background loop already was. It had been running every 20 seconds
against whatever items the current test created, writing artifacts and
health rows into the middle of teardowns -- the actual cause of the
foreign-key failures that were being patched one teardown at a time.
The build setup now offers the models the chosen harness's image says it runs, not every model on the price list. That is the defect the operator reported: codex was offered Claude and Gemini, and a save accepted one. Once the list is known, a model the harness cannot run is refused at save -- here, where it can be fixed -- rather than when the run starts, after an approval. The price list keeps its real job. It no longer decides WHICH models are offered, only whether each can be paid for: a model the harness runs but nobody has priced is shown, so the operator learns it exists, and blocked with "no price yet", because an API-key run needs a rate for every token type the harness reports. A thinking level is chosen with the model, because the levels belong to it. Codex itself offers "Model and Effort" together, and the levels differ per model -- one allows low to ultra, another only low to xhigh -- so the list is rebuilt from the chosen model and checked against that model's own list at save. No level chosen means the model's own default, a real choice the vendor publishes, stored as NULL (V80) rather than as a guessed name. A level chosen for one model is cleared when the model changes, since the next one may not offer it. When the harness's list is NOT known the save is not blocked, and that is a decision: a development stack with no run worker never hears the answer, and refusing every save would lock the operator out of the build setup over a reply that has not come. The screen says which of the four reasons it is and falls back to the price list, as before. A thinking level is refused in that case, since there is nothing to check it against. Recorded in the design as 6A.4a. Mutation-verified: removing the harness check fails three tests.
The build setup saved a thinking level, but nothing carried it to the run, so every build ran at the model's own default. The level now travels the whole way: - WorkPreparation binds it under a new binding version 3, so gates opened under versions 1 and 2 keep their exact hashes, and a level may not ride on a version that does not hash it. - ExecuteRun carries it as a nullable reasoningEffort, with an atEffort wither; WorkRunAssembly sets it from the approved preparation. - The run worker puts it on HarnessInvocation, and the Codex arm adds -c model_reasoning_effort="<level>" only when a level was chosen. The Codex CLI does not validate the level, so every hop accepts only a short lower-case word. Eight mutants, one per hop and guard, each fail exactly the test written for them.
Review of PR #167 found four gaps; this closes three and part of the fourth. - A setup was checked against the image only when saved. If the image changed afterwards, the task was still prepared and the build still dispatched. HarnessCatalogues.refusal now holds the one check, and the save, the preparation sweep and the dispatch all ask it. The three reasons are named on the item page. - An image could declare a level such as "x-high": offered, saved, then refused at preparation. ThinkingLevel is now the one shape rule, used by the declared model, the preparation and the run command, so such a label reads as UNREADABLE. - The build screen kept a level across a harness change, and a saved level with no known list had no control on screen yet was sent. The level is dropped on a harness change; an unusable one is named, with a way back to the model default, and Save waits. - Image answers were keyed by a random id and handled unordered, so an older answer could replace a newer one. They are now keyed by harness and handled in order. Still open: two workers holding different images under one mutable tag.
Two run workers can hold different images under one mutable tag, and the deployment default is spire-agent-codex:latest. The model list could be read from one worker's image while the build ran on another. The run worker now reads the labels and the exact image in one inspect: the registry digest when the image has one, else the daemon's image id. The answer carries the pin, the orchestrator stores it with the list (V81), and every run it sends - item builds, REST dispatches and /fix runs - uses the pin instead of the tag. An item build takes its model check and its pin from one read of the cache. With nothing read yet, runs use the tag as before. A local-only image pinned by id fails to pull on a worker that does not hold it, which is the honest outcome. RunRuntime.imageLabels becomes describeImage. Recorded in UNVERIFIED: no test pulls a registry image or runs two workers.
Second review of PR #167 found two gaps in the image pin. - The orchestrator consumed image answers unordered, so an older answer queued during an outage could still finish last and overwrite the newer list and pin. It now records them one at a time, which keeps the per-harness order the worker's key gives. - Docker reports a Docker Hub image under its short name whatever spelling pulled it: docker.io/library/alpine:3.20 reports alpine@sha256:... The digest match now compares repositories in Docker's canonical form, so such an image is pinned by its portable digest instead of a local image id.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
The build setup offered codex every model on the price list, Claude and Gemini included, and had no
thinking level at all. Codex itself offers a model and a level together, and each model has its own
levels. The answer must also work on Kubernetes, where nothing can run the image to ask it.
Design:
docs/superpowers/specs/2026-09-16-factory-m35-one-ticket-to-a-build-design.md§6A.What changed
deploy/agent/build-codex.shbuilds the agent image, readscodex debug models, and bakes a trimmed catalogue into the labeldev.codespire.agent.models.The image contract gains a
modelsclause;spire-agent-image verifyreports it.cs.harness-image-commands/cs.harness-image-results. The run worker reads labels with the sameauthenticated pull it uses for runs. The orchestrator stores the answer in
harness_catalogue(V79), refreshes on start and every
spire.harness-catalogue-interval(10m), and drops an answerfor an image that is no longer current.
price is shown but blocked, with the reason. A thinking-level select lists the chosen model's own
levels. The save refuses a model the harness does not run and a level the model does not offer.
When the list is not known, models fall back to the price list and no level is offered (§6A.4a).
keep their exact hashes).
ExecuteRun.reasoningEffortcarries it; the Codex arm adds-c model_reasoning_effort="<level>"only when one was chosen (§6A.4b).Verification
testFastandtestServicespass; UI 913 tests across 104 files,tscand production build pass.The model/level save check: removing it failed 3 tests.
Not in this PR
sign-in through the new screen, so the credential shape is measured rather than assumed.