Skip to content

fix(metrics): count each API response once when its rows share a message id - #595

Open
z1pp090 wants to merge 1 commit into
sirmalloc:mainfrom
z1pp090:fix/dedupe-usage-rows-by-message-id
Open

z1pp090 wants to merge 1 commit into
sirmalloc:mainfrom
z1pp090:fix/dedupe-usage-rows-by-message-id

Conversation

@z1pp090

@z1pp090 z1pp090 commented Sep 20, 2026 •

Copy link
Copy Markdown

Fixes #549.

Problem

Claude Code writes one transcript row per content block of an assistant response (thinking, text, each tool_use). Those rows share message.id (and requestId), carry identical input_tokens / cache_read_input_tokens / cache_creation_input_tokens, a non-decreasing output_tokens, and all of them already hold the finalized stop_reason. The current gate in collectTokenMetricRecord only drops streaming partials (stop_reason: null), so every content-block row is summed. As #549 measured, tokens-input, tokens-output, tokens-cached and tokens-total over-report by ~2x; context-length is unaffected because it reads a single row.

collectSpeedMetricRecord has the same shape: it pushed one SpeedRequest per row, so the speed widgets counted one request (and one set of output tokens) per content block.

Reproduced here on 23 local transcripts, 7,600 API calls, ccstatusline 2.2.30:

metric before after ratio
cache read 3,151,311,722 1,841,844,917 1.71x
cache creation 60,820,618 30,586,715 1.99x
input 55,138 27,766 1.99x
output 12,147,333 5,576,123 2.18x

60% of messages in those transcripts are multi-row with every row finalized; the remaining 40% are single-row. Within a message id, prompt-side usage was identical and output_tokens non-decreasing in every case, as the report states.

Fix

Keep the scan single-pass: hold the latest row of the current message.id and count it when the next message starts or the file ends (flushPendingTokenMetricEntry / flushPendingSpeedRequest). The last row of a message carries its complete usage, so this is exact rather than a heuristic.

Preserved behaviour:

  • rows without a message.id are still counted per row;
  • the stop_reason gate still applies, and the last in-progress row (stop_reason: null) is still counted once at the end of the scan for live updates;
  • a compaction boundary that arrives while a row is pending keeps that row out of the post-compaction context, as boundaryAfterLastUsage did before (that flag is folded into the pending entry);
  • the reset when stop_reason first appears also drops the pending row, matching the old accumulator reset.

TranscriptLine.message gains an optional id so the field is typed rather than read through unknown.

Tests

New cases in jsonl-metrics.test.ts (the two token/speed ones fail on main):

  • one response written as three finalized rows with the same id is counted once, including the context-length pick;
  • rows with no id keep the per-row behaviour;
  • a still-streaming response whose rows share an id is counted once;
  • speed metrics count one request per message id, not per row.

bun test (2354 pass) and bun run lint green.

🤖 Generated with Claude Code

…age id

Claude Code writes one transcript row per content block of an assistant
response (thinking, text, each tool_use). The rows share `message.id`, carry
identical prompt-side usage and a non-decreasing `output_tokens`, and all of
them already hold the finalized `stop_reason`. The stop_reason gate only
catches streaming partials (stop_reason: null), so every one of those rows
was summed: tokens-input, tokens-output, tokens-cached and tokens-total
over-reported by roughly 2x, and the speed widgets counted one request per
row.

Hold the latest row of the current message id while scanning and count it
when the next message starts or the scan ends. Rows without an id keep the
old per-row behaviour, the in-progress streaming row is still counted once,
and a compaction boundary that lands while a row is pending keeps it out of
the post-compaction context the way it did before.

Measured on 23 local transcripts: cache read 3.15G -> 1.84G, output 12.1M ->
5.58M, input 55.1k -> 27.8k.

Closes sirmalloc#549

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01StkifHpSLWH4J5TvuP56ur
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Token widgets over-report ~1.84x: one JSONL entry per content block is counted per API call

1 participant