Count Codex cache writes as part of input, not extra to it - #182
Merged
Merged
Conversation
The Codex adapter read cache_write_input_tokens as additional to input, as Anthropic reports cache creation. The first run to report a non-zero cache write (work item 38) disproved that: the Codex session log's own total_tokens is input plus output, and cached + cache write + plain input add up to input_tokens. So the cache-write tokens were billed twice, once as input and once as cache write, and that run's cost was recorded about 43% above what it was. - INPUT is now input minus cached minus cache write. - When cached plus cache write exceed input, the counts contradict the subset reading, and the turn degrades to an unreconciled total. - The test that pinned the old reading is replaced by one built on the real counts from that run, plus one for the contradiction. - docs/UNVERIFIED.md: the Codex row records the multi-turn run and what it settled. The cost already recorded for work item 38 is not rewritten.
RUN-TOPOLOGY still listed cache writes as additional to input, the reading work item 38 disproved and this branch fixes. It now records the measurement, and leaves two limits open instead of three.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Build 2 of the pay-with comparison (work item 38, API key,
gpt-6-sol) was the first run to report a non-zerocache_write_input_tokens. Its recorded cost was $0.197; its real cost is about $0.138.The adapter treated the cache write as additional to input (as Anthropic reports cache creation). The code marked this as unverified. The Codex session log from that run settles it:
total_tokens= input + output (123,068 + 1,383 = 124,451), and per turninput_tokens14,804 = cached 14,616 + cache write 185 + plain 3. So the cache write is a subset of input, and 14,801 tokens were billed twice (as INPUT at the input rate and as CACHE_WRITE).What changed
CodexAdapter.usageEvent:INPUT = input − cached − cacheWrite. Ifcached + cacheWrite > input, the counts contradict the subset reading, and the turn degrades to an unreconciledTOTAL(the larger of input and its parts, plus the larger of output and reasoning).CodexAdapterTest: the test that pinned the old reading is replaced by one on the real counts from that run (parts sum to Codex's total), plus one for the contradiction.docs/UNVERIFIED.md: the Codex CLI row records the multi-turn tool-call run and what it settled.The Anthropic mapping in
TokenUsageMapperis unchanged; there, cache creation is additional to input.Not done: the cost already recorded for work item 38 is left as it is.
Verification
:spire-harness-codex:test55 tests pass;testFastpasses.inputPartswithout the cache write, the guard on cached alone, the unreconciled total from cached alone.