Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/UNVERIFIED.md
Original file line number Diff line number Diff line change
Expand Up @@ -471,7 +471,7 @@ Each has a runbook mode. None has been run by an operator.
| The whole M1 lifecycle against a real forge | **Mode Q** | Cancel, steer, the watchdog, the push gate and the charge ledger have only ever met a WireMock LLM and a local origin |
| Corporate-only bundle → the failure it produces | Mode R §5 | The documented trap (internal forge works, model API fails) is asserted nowhere; it is the mistake an operator will actually make |
| A private-registry pull | Mode S §4 | Nothing pulls from a private registry in any test. `authFor` and the attachment are unit-tested; the *pull* is not |
| **Codex CLI 0.156.1 in the agent image** (2026-09-23) | none yet | Raised from 0.146.0 so the model list includes the gpt-6 models. Re-checked on 0.156.1: every flag the adapter passes, the API-key login, and the shape of the file it writes (`auth_mode=apikey`). The device sign-in output was then proved by a real sign-in (2026-09-25), and one live `codex exec` on a subscription produced the same `--json` event types and the same five usage buckets. Still NOT re-checked: a multi-turn run with tool calls, which exercises the rest of the parser |
| **Codex CLI 0.156.1 in the agent image** (2026-09-23) | none yet | Raised from 0.146.0 so the model list includes the gpt-6 models. Re-checked on 0.156.1: every flag the adapter passes, the API-key login, and the shape of the file it writes (`auth_mode=apikey`). The device sign-in output was then proved by a real sign-in (2026-09-25), and one live `codex exec` on a subscription produced the same `--json` event types and the same five usage buckets. A live API-key build with 11 tool calls (work item 38, 2026-09-27) was parsed end to end, and its usage was the first with a non-zero `cache_write_input_tokens`. The Codex session log's own `total_tokens` equals input plus output, so the cache write is part of input, not extra to it. The adapter had read it as extra and billed those tokens twice (that run recorded about 43% above its cost); it now subtracts both parts from input. Still NOT re-checked: a run whose parts exceed input on a real vendor, which the adapter degrades to an unreconciled total |
| **Runs pinned to the image their model list came from** (M3.5 part M, 2026-09-23) | none yet | Choosing the pin is unit-tested against given daemon answers, and one real-daemon test pins a LOCAL build by its image id. No test pulls a registry image and pins it by its registry digest, and none runs two workers. So "two workers holding different images under one tag run the same one" is argued from the code, not watched. A local-only image is pinned by an id that exists on one daemon only: on a second worker such a run fails to pull — by design, but unobserved |
| **A build paid by a Codex subscription** (M3.5 part F, 2026-09-25) | none yet | Tests prove the emptied refresh token, one seat per account, seat rotation, the unmetered charge lines and the file the adapter writes. One live `codex exec` accepted a file with its refresh token emptied, three hours after sign-in with the id token expired. No item build has run on a seat yet. NOT observed: two builds using one seat at the same moment (seats are shared by the operator's decision of 2026-09-26; whether the vendor accepts two sessions of one account at once is unknown), a run that outlives the access token (10 days — the seat then needs a new sign-in), whether the vendor's CLI ever tries to refresh mid-run with an empty refresh token, and whether `tokens.account_id` is stable across sign-ins of one account |
| **Run-worker log lines are scrubbed of run secrets** (review of PR #178, 2026-09-25) | none yet | `SecretLogFilter` is unit-tested on a JBoss log record and its console configuration is asserted. Not observed: the filter on the JSON console handler of a running worker. And by design it cannot scrub a unit this process did not start: after a worker restart, the watchdog's log lines about an older unit carry that unit's secrets unscrubbed, as `RunFailures` already states for failure details |
Expand Down
32 changes: 19 additions & 13 deletions docs/factory/RUN-TOPOLOGY.md
Original file line number Diff line number Diff line change
Expand Up @@ -476,20 +476,26 @@ The real event vocabulary, captured from two live runs rather than read from doc
call that used 14064 — a 71% overstatement, worst on the runs that were cheapest.
`TokenUsageMapper.openAi` in spire-llm already subtracts for this reason; `CodexAdapter` follows it.

**Three limits left open, deliberately.** The cumulative-versus-incremental question above is the
first. Second, Codex reports **no total**, so the independent cross-check `TokenUsageMapper`
performs against `totalTokenCount()` is unavailable here — a mis-partition cannot be caught by
arithmetic, only by a contradiction between the vendor's own fields. Third,
`cache_write_input_tokens` is treated as *additional* to input rather than a subset of it, matching
what the name describes and how Anthropic reports the same concept; every run observed so far
reported zero, so measurement has not ruled out the alternative — and the contradiction gate cannot
catch a wrong choice here, because `cacheWrite` is compared against nothing. A run with a non-zero
cache write would settle it.

**What would settle two of the three: one deliberate multi-turn run**, captured with `--json`, its
- **Cache writes are a subset of input too (measured 2026-09-27).** The first run to report a
non-zero `cache_write_input_tokens` (work item 38, an API-key build with 11 tool calls) settled it.
Its Codex session log gives `total_tokens` equal to input plus output, and one turn's
`input_tokens: 14804` is `cached_input_tokens: 14616` + `cache_write_input_tokens: 185` + 3
plain. The adapter had read the cache write as *additional*, as Anthropic reports cache creation;
that billed the run's 14801 cache-write tokens twice and recorded its cost about 43% above what it
was. `CodexAdapter` now subtracts both parts from input, and a turn whose parts add up to more
than input degrades to an unreconciled total.

**Two limits left open, deliberately.** The cumulative-versus-incremental question above is the
first. Second, the `--json` stream carries **no total**, so the independent cross-check
`TokenUsageMapper` performs against `totalTokenCount()` is unavailable here — a mis-partition cannot
be caught by arithmetic, only by a contradiction between the vendor's own fields. The session log
Codex writes in the agent's home does carry `total_tokens`, which is how the cache-write reading
was settled; the adapter does not read that file.

**What would settle the first: one deliberate multi-turn run**, captured with `--json`, its
`turn.completed` lines compared against each other. That is a run to make on purpose, not a thing to
wait for. Until it exists, neither the cumulative reading nor the cache-write reading may be cited
as measured — and this section is the record of which is which.
wait for. Until it exists, the cumulative reading may not be cited as measured — and this section is
the record of which is which.

**Version note.** The plan states its flag set was verified against codex-cli **0.152.0**; the
binary available for this measurement was **0.146.0**. Every flag the adapter uses was re-checked
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -286,13 +286,18 @@ private static boolean failed(JsonNode item) {
* cached would be recorded as 24048 tokens for a call that used 14064, inflated by the cache-hit
* rate and therefore inflated most on the cheapest runs.
*
* <p><b>Two honest limits, recorded rather than papered over.</b> Codex reports no total, so the
* independent cross-check {@code TokenUsageMapper} performs against {@code totalTokenCount()} is
* not available here — a mis-partition cannot be caught by arithmetic. And
* {@code cache_write_input_tokens} is treated as ADDITIONAL to input rather than a subset of it,
* matching what the name describes and how Anthropic reports the same concept; every run
* observed so far reported zero, so the alternative has not been ruled out by measurement. A
* run with a non-zero cache write is what would settle it.
* <p>{@code cache_write_input_tokens} is a subset of input too, like the cached count and unlike
* Anthropic's cache creation. The first run to report one settled it (work item 38, 2026-09-27):
* the Codex session log gave {@code input_tokens=14804, cached_input_tokens=14616,
* cache_write_input_tokens=185, output_tokens=64, total_tokens=14868}, so the total is input plus
* output and the three input parts add up to 14804. Reading the cache write as additional billed
* the run's 14801 cache-write tokens twice, once as input and once as cache write, and recorded
* its cost about 43% above what it was.
*
* <p><b>One honest limit, recorded rather than papered over.</b> The {@code --json} stream this
* parses carries no total, so the independent cross-check {@code TokenUsageMapper} performs
* against {@code totalTokenCount()} is not available here: a mis-partition that keeps every
* part below the headline cannot be caught by arithmetic.
*/
private Optional<RunEvent> usageEvent(JsonNode usage, Instant at) {
if (!usage.isObject()) {
Expand All @@ -309,12 +314,13 @@ private Optional<RunEvent> usageEvent(JsonNode usage, Instant at) {
// once dead-lettered a paid review; here it degrades the turn to UNKNOWN instead.
return Optional.empty();
}
if (cached > input || reasoning > output) {
return Optional.of(new RunEvent.Usage(at, unreconciled(input, cached, output, reasoning)));
long inputParts = cached + cacheWrite;
if (inputParts > input || reasoning > output) {
return Optional.of(new RunEvent.Usage(at, unreconciled(input, inputParts, output, reasoning)));
}

Map<TokenBucket, Long> counts = new EnumMap<>(TokenBucket.class);
put(counts, TokenBucket.INPUT, input - cached);
put(counts, TokenBucket.INPUT, input - inputParts);
put(counts, TokenBucket.CACHED_INPUT, cached);
put(counts, TokenBucket.CACHE_WRITE, cacheWrite);
put(counts, TokenBucket.OUTPUT, output - reasoning);
Expand Down Expand Up @@ -342,14 +348,15 @@ private static void put(Map<TokenBucket, Long> counts, TokenBucket bucket, long
*
* <p>{@code TokenUsageMapper} carries the vendor's OWN total here — the one number it has not
* derived. Codex reports none, so this is derived from the very fields that just failed their
* check, and the direction of the guess matters. Under {@code cached > input} the natural
* reading is that input EXCLUDES cached, making the true total {@code input + cached + output};
* {@code input + output} would then be short by the cached amount. For waste detection,
* understating is the harmful direction, so each side takes the larger of its two figures.
* check, and the direction of the guess matters. Under {@code cached + cacheWrite > input} the
* natural reading is that input EXCLUDES those parts, making the true total
* {@code input + parts + output}; {@code input + output} would then be short by the parts. For
* waste detection, understating is the harmful direction, so each side takes the larger of its
* two figures.
*/
private static UsageReport unreconciled(long input, long cached, long output, long reasoning) {
private static UsageReport unreconciled(long input, long inputParts, long output, long reasoning) {
return UsageReport.of(Map.of(TokenBucket.TOTAL,
Math.max(input, cached) + Math.max(output, reasoning)));
Math.max(input, inputParts) + Math.max(output, reasoning)));
}

/** A field that is absent reads as zero; one that is present but not a number reads as -1. */
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -388,25 +388,42 @@ void anUnknownEnvelopeTypeIsClippedLikeEveryOtherModelControlledField() {
}

@Test
void aNonZeroCacheWriteIsTreatedAsAdditionalToInput() {
// Pins an UNVERIFIED reading on purpose. cache_write is treated as ADDITIONAL to input,
// matching its name and how Anthropic reports the same concept — but every run observed so
// far reported zero, so the subset reading is not ruled out. The contradiction gate cannot
// catch a wrong choice here: cacheWrite is compared against nothing, so if it really is a
// subset the adapter overstates by exactly that amount and degrades to nothing.
//
// Asserting the reading explicitly means changing it is a deliberate act with a failing
// test, rather than a silent drift in an arithmetic nobody re-reads.
void aCacheWriteIsPartOfInputNotAdditionalToIt() {
// Measured, not assumed: the first run to report a cache write (work item 38, 2026-09-27).
// Its Codex session log gave these counts with total_tokens=14868, which is input plus output,
// so cached (14616) + cache write (185) + plain input (3) make up the 14804 input tokens. The
// earlier reading, cache write ADDITIONAL to input, billed the cache-write tokens twice and
// recorded that run about 43% above its cost.
RunEvent event = adapter.parse("""
{"type":"turn.completed","usage":{"input_tokens":100,"cached_input_tokens":10,\
"cache_write_input_tokens":7,"output_tokens":5,"reasoning_output_tokens":0}}""")
{"type":"turn.completed","usage":{"input_tokens":14804,"cached_input_tokens":14616,\
"cache_write_input_tokens":185,"output_tokens":64,"reasoning_output_tokens":0}}""")
.orElseThrow();

UsageReport report = adapter.usage(RunEventSummary.of(List.of(event)));

assertEquals(90L, report.tokens(TokenBucket.INPUT),
"input minus cached only — cache_write is NOT subtracted under the additional reading");
assertEquals(7L, report.tokens(TokenBucket.CACHE_WRITE));
assertEquals(3L, report.tokens(TokenBucket.INPUT), "input minus cached minus cache write");
assertEquals(14_616L, report.tokens(TokenBucket.CACHED_INPUT));
assertEquals(185L, report.tokens(TokenBucket.CACHE_WRITE));
assertEquals(64L, report.tokens(TokenBucket.OUTPUT));
assertEquals(14_868L, report.asMap().orElseThrow().values().stream().mapToLong(Long::longValue).sum(),
"the parts add up to Codex's own total, counting nothing twice");
}

@Test
void cachedPlusCacheWriteAboveInputDegradesToAnUnreconciledTotal() {
// Each part fits under input on its own, but together they are more than input: the counts
// contradict the subset reading. No split can be trusted, so the run is unreconciled rather
// than recorded with an input floored at zero.
RunEvent event = adapter.parse("""
{"type":"turn.completed","usage":{"input_tokens":100,"cached_input_tokens":80,\
"cache_write_input_tokens":30,"output_tokens":10,"reasoning_output_tokens":0}}""")
.orElseThrow();

UsageReport report = adapter.usage(RunEventSummary.of(List.of(event)));

assertEquals(Set.of(TokenBucket.TOTAL), report.asMap().orElseThrow().keySet(),
"no split is offered alongside a TOTAL");
assertEquals(120L, report.tokens(TokenBucket.TOTAL), "max(100, 80 + 30) + max(10, 0)");
}

@Test
Expand Down
Loading