Conversation
…age id Claude Code writes one transcript row per content block of an assistant response (thinking, text, each tool_use). The rows share `message.id`, carry identical prompt-side usage and a non-decreasing `output_tokens`, and all of them already hold the finalized `stop_reason`. The stop_reason gate only catches streaming partials (stop_reason: null), so every one of those rows was summed: tokens-input, tokens-output, tokens-cached and tokens-total over-reported by roughly 2x, and the speed widgets counted one request per row. Hold the latest row of the current message id while scanning and count it when the next message starts or the scan ends. Rows without an id keep the old per-row behaviour, the in-progress streaming row is still counted once, and a compaction boundary that lands while a row is pending keeps it out of the post-compaction context the way it did before. Measured on 23 local transcripts: cache read 3.15G -> 1.84G, output 12.1M -> 5.58M, input 55.1k -> 27.8k. Closes sirmalloc#549 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01StkifHpSLWH4J5TvuP56ur
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #549.
Problem
Claude Code writes one transcript row per content block of an assistant response (
thinking,text, eachtool_use). Those rows sharemessage.id(andrequestId), carry identicalinput_tokens/cache_read_input_tokens/cache_creation_input_tokens, a non-decreasingoutput_tokens, and all of them already hold the finalizedstop_reason. The current gate incollectTokenMetricRecordonly drops streaming partials (stop_reason: null), so every content-block row is summed. As #549 measured,tokens-input,tokens-output,tokens-cachedandtokens-totalover-report by ~2x;context-lengthis unaffected because it reads a single row.collectSpeedMetricRecordhas the same shape: it pushed oneSpeedRequestper row, so the speed widgets counted one request (and one set of output tokens) per content block.Reproduced here on 23 local transcripts, 7,600 API calls, ccstatusline 2.2.30:
60% of messages in those transcripts are multi-row with every row finalized; the remaining 40% are single-row. Within a message id, prompt-side usage was identical and
output_tokensnon-decreasing in every case, as the report states.Fix
Keep the scan single-pass: hold the latest row of the current
message.idand count it when the next message starts or the file ends (flushPendingTokenMetricEntry/flushPendingSpeedRequest). The last row of a message carries its complete usage, so this is exact rather than a heuristic.Preserved behaviour:
message.idare still counted per row;stop_reasongate still applies, and the last in-progress row (stop_reason: null) is still counted once at the end of the scan for live updates;boundaryAfterLastUsagedid before (that flag is folded into the pending entry);stop_reasonfirst appears also drops the pending row, matching the old accumulator reset.TranscriptLine.messagegains an optionalidso the field is typed rather than read throughunknown.Tests
New cases in
jsonl-metrics.test.ts(the two token/speed ones fail onmain):bun test(2354 pass) andbun run lintgreen.🤖 Generated with Claude Code