Skip to content

[medium] perf(jsonl): skip JSON.parse for transcript lines no collector reads - #640

Open
elhoim wants to merge 1 commit into
sirmalloc:mainfrom
elhoim:perf/transcript-prefilter
Open

elhoim wants to merge 1 commit into
sirmalloc:mainfrom
elhoim:perf/transcript-prefilter

Conversation

@elhoim

@elhoim elhoim commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

BLUF

  • Priority: medium.
  • When no widget needs session duration or speed metrics, the transcript scan only reads a few record kinds, yet it still ran JSON.parse on every line (tool results and attachments are most of the bytes).
  • Lines that contain none of the fixed markers those collectors need, and no escape that could spell one, are now skipped before JSON.parse. The recursive agentId walk (speed widgets only) gets the same guard.
  • A differential test proves the skip is lossless: old and new analysis are deep-equal on generated corpora across every option set, and on 160 runs over real transcripts.
  • Measured on a real 60 MB transcript: -19% render CPU with token widgets, -10% with all widgets. The synthetic corpus shows -5%.

Details

With neither session duration nor speed metrics requested, each collector reads only these records:

Collector Records it reads Raw-text marker
token metrics message.usage, compact boundary "usage", "compact_boundary"
compaction stats subtype: "compact_boundary" "compact_boundary"
thinking effort string content with the effort/model stdout prefix <local-command-stdout>Set (shared start of both prefixes, now exported from jsonl-metadata.ts)
session name type: "custom-title" "custom-title"
agent-id walk (speed only) any agentId string "agentId"

Soundness:

  • Every marker is printable ASCII. JSON has no short escape for those characters, so the only other way to spell one is a –~ escape. Any line matching /\\u00[2-7]/ is always parsed.
  • Other escapes can never form a marker. That includes \u001b from ANSI-colored tool output, which accounts for about 30% of the bytes in the real transcript measured here.
  • On a skipped line, every enabled collector is provably a no-op.
  • Session duration and speed metrics read the timestamp of every record, so the filter turns off when either is requested. Only the agentId walk stays guarded in that case.

Tests (jsonl-metrics-prefilter.test.ts):

  • The test generates seeded corpora with:
    • usage with/without stop_reason, sidechain, and API-error records
    • compact boundaries with every metadata variant
    • effort/model stdout with ANSI escapes
    • custom titles
    • agentId nested in arrays, empty, and spelled agentId
    • usage, compact_boundary, custom-title
    • odd whitespace, malformed lines, arrays and scalars, CRLF, BOM, and subagent files
  • Each corpus is compared with a copy whose records carry an extra unread field containing every marker and a A. That copy forces the unfiltered path on every line. The comparison covers all 16 option sets: speed, compaction, effort, name.
  • Mutation check: removing the escape guard, or any single marker, makes the test fail.

Overlap: #591 (ours) and #622 change token counting in src/utils/jsonl-metrics.ts. #634 (ours) adds includeTokenMetrics there. This PR only touches the scan loop and adds the marker helpers, so conflicts would be mechanical. The two perf PRs are independent: #634 skips the scan when nothing reads it, and this PR makes the scans that remain cheaper.

Measurements

node dist/ccstatusline.js. Base and patched arms were interleaved round-robin in one run with 20 passes plus a node -e 0 control. CPU is user+sys including children, in ms. Settings use gitCacheTtlSeconds: 0. "Token widgets" means the default widgets plus tokens-total, tokens-input, compaction-counter, and session-name. "All widgets" is every widget, including speed and session clock.

Arm CPU median CPU p90 Wall median
control node -e 0 90 133 125
base, token widgets, real 60 MB transcript 3389 4181 3617
patched, token widgets, real 60 MB transcript 2758 3298 2928
base, all widgets, real 60 MB transcript 4996 6040 5384
patched, all widgets, real 60 MB transcript 4513 5329 4833
base, token widgets, synthetic 50 MB 2278 2744 2188
patched, token widgets, synthetic 50 MB 2171 2730 2310

Load1 during the run was min 10.7, median 17.1, max 20.2 on 6 cores. The host was loaded, so compare ratios. The saving depends on how many bytes sit in marker-free lines. The synthetic corpus has short tool results and no escapes, so it gains less than a real session.

Equivalence:

  • getTranscriptAnalysis on main and on this branch matched on 160 runs: 10 real transcripts, some with subagents, times 16 option sets.
  • stdout of base and patched dist is byte-identical on 99 config × payload pairs.

Checks

  • bun run lint: clean.
  • bun run build: OK.
  • bun test on the jsonl, metrics, compaction, speed, and prefilter suites: 74 pass. The full suite has the same host-load flakes as unmodified main on this machine: fetchUsageData error handling, custom command capture, and TUI menu timeouts. None of them touch this code.

🤖 Generated with Claude Code

When neither session duration nor speed metrics are requested, the
transcript scan only reads usage records, compact boundaries, effort
stdout, and custom-title records. Each of these carries a fixed
printable-ASCII marker ("usage", "compact_boundary",
"<local-command-stdout>Set ", "custom-title"). JSON can only spell those
characters differently with a  -~ escape. A line with neither
a marker nor such an escape therefore cannot change any enabled
collector, and it is skipped before JSON.parse. The recursive agentId
walk (speed widgets only) gets the same guard with the "agentId" marker.
ANSI \u001b escapes in tool output do not trigger the guard.

A differential test builds seeded corpora with escaped keys and values,
odd whitespace, malformed lines, CRLF, BOM, nested and escaped agentIds,
and subagent files. It compares each against a copy that forces the full
parse on every line, over every option set. Dropping any marker or the
escape guard makes it fail.

Measured (node dist, 20 interleaved passes, load1 ~17 on 6 cores), CPU
median per render on a real 60 MB Claude Code transcript: token widgets
3389 -> 2758 ms (-19%), all widgets 4996 -> 4513 ms (-10%). On the
synthetic 50 MB corpus: -5%. Analysis output is identical on 160 real
transcript x option runs, and stdout is byte-identical on 99
config/payload pairs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant