Skip to content

[high] fix(metrics): count one API response once when the transcript logs it per content block - #591

Open
elhoim wants to merge 1 commit into
sirmalloc:mainfrom
elhoim:fix/token-metrics-block-dedupe
Open

elhoim wants to merge 1 commit into
sirmalloc:mainfrom
elhoim:fix/token-metrics-block-dedupe

Conversation

@elhoim

@elhoim elhoim commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

BLUF

  • Priority: high. Every transcript-derived token number is currently inflated — roughly 2x on today's Claude Code transcripts.
  • Cause: one API response is logged once per content block (thinking, tool_use, …). Each of those lines repeats the same usage, the same message.id, and an increasing apiBlockIndex — and the fallback counted every line.
  • Affects: Tokens Input / Output / Cached / Total, and the speed widgets' requestCount / token totals.
  • Fix: skip continuation blocks — apiBlockIndex > 0 when present, a repeated message.id otherwise — so one response is counted once.
  • Safe: the live context_window path and contextLength are untouched; streamed responses (stop_reason: null then the real count) are still counted exactly once.
  • Verified: 5 new tests, bun run lint clean, bun run build clean, and the patched build's numbers now match an independent per-message tally of real transcripts exactly.

The evidence

Two consecutive lines from a live transcript, one API response:

field line 1 line 2
apiBlockIndex 0 1
message.id msg_011CfAdDY4… msg_011CfAdDY4… (same)
message.stop_reason tool_use tool_use
usage.input_tokens 2 2
usage.output_tokens 288 288
usage.cache_read_input_tokens 23991 23991
content block thinking tool_use

Both lines carry stop_reason, so the existing hasStopReasonField filter keeps both and the response is counted twice.

Rendering the same session's transcript through both builds, side by side:

before   In: 390 | Cached: 30.1M | Out: 108.8k
after    In: 212 | Cached: 16.7M | Out:  59.7k

An independent per-message tally of that transcript (a separate script, deduping on message.id) gives 212 / 16.7M / 59.7k, matching the patched build exactly.

The fix

isUsageContinuationBlock() in src/utils/jsonl-metrics.ts, used by both collectors:

  • Token metrics — a continuation block is not accumulated. lastCountedMessageId is only updated when an entry is actually counted, so the in-flight format (several stop_reason: null lines, then the final line with the real usage under the same id) still counts that final line once.
  • Speed metrics — a continuation block does not push a second request; it stretches the pending request's interval to when the later block landed and adopts its usage (identical for a block repeat, complete for a streamed response). requestCount becomes responses, not lines.

Untouched on purpose:

  • The status-JSON context_window path — it never had this problem.
  • contextLength / most-recent-usage — repeats carry identical usage, so dropping them cannot move those figures.
  • src/utils/jsonl-blocks.ts — it only reads timestamps of usage-bearing lines for block-window activity, so the Block Timer widget is unaffected.

Tests

New cases in src/utils/__tests__/jsonl-metrics.test.ts:

  1. a per-block logged response is counted once (tokens + cache + contextLength);
  2. the message.id fallback, for transcripts without apiBlockIndex;
  3. distinct responses with no id at all still both count (no behavior change for older formats);
  4. a streamed response repeating its id still counts its final line;
  5. speed metrics report one request spanning the prompt to the last block.

bun test: 2360 pass / 3 fail. The failures are pre-existing flaky TUI/subprocess cases in this environment (fetchUsageData 5s subprocess timeout, TerminalWidthMenu, UsageTimezoneEditor); the failing set varies run to run - an unmodified main here fails 7, this branch 3 - and none of them is in metrics or token widgets. bun run lint and bun run build are clean.

Relationship to other PRs

🤖 Generated with Claude Code

Claude Code writes an assistant response to the transcript once per content
block - a `thinking` line, a `tool_use` line, and so on - and every one of
those lines repeats the same `usage` object, the same `message.id` and an
increasing `apiBlockIndex`. The transcript fallback counted each line, so a
two-block response was billed twice: Tokens Input / Output / Cached / Total and
the speed widgets' request count all read roughly 2x on current transcripts.

Continuation blocks are now skipped - `apiBlockIndex` when the transcript
carries it, a repeated `message.id` otherwise - and in the speed path the later
block only stretches the request's interval and takes its final usage, so a
streamed response (`stop_reason: null` lines followed by the real count) is
still counted exactly once.

The live `context_window` path in the status JSON is untouched, and so is
`contextLength`: repeats carry identical usage, so dropping them cannot move
the most-recent-usage figures.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant