Skip to content

perf(index): doc_to_disk remap in TrigramIndex.writeToDisk (2.5-4x on cold-index persist) - #671

Merged
justrach merged 23 commits into
release/0.2.5829from
perf/trigram-writetodisk-remap
Jul 10, 2026
Merged

perf(index): doc_to_disk remap in TrigramIndex.writeToDisk (2.5-4x on cold-index persist)#671
justrach merged 23 commits into
release/0.2.5829from
perf/trigram-writetodisk-remap

Conversation

@justrach

@justrach justrach commented Jul 9, 2026

Copy link
Copy Markdown
Owner

The single unambiguous 2×-class win from the fresh perf audit. TrigramIndex.writeToDisk — the disk persist on every cold index — hashed the full file-path string once per posting (disk_path_to_id.get(id_to_path[doc_id]) inside the Step-3 posting loop): tens of millions of wyhash+probe on a large repo.

Fix: precompute doc_id -> disk file_id once into doc_to_disk: []u32 (one hash per doc, O(num_docs)); the posting loop then does a single array load per posting. Retrieval-identical — doc_to_disk[doc_id] equals the old disk_path_to_id.get(id_to_path[doc_id]), with maxInt standing in for the old orelse continue.

Precedent: identical fix already shipped for WordIndex (#475). The audit's ReleaseFast microbench measured the remap step ~25× (12.5ns → 0.5ns/posting); whole writeToDisk ~2.5–4× (diluted by buffered IO + the trigram sort). num_postings >> num_docs, so it scales with repo size — ~0.5s absolute per 40M postings.

Validation: test-index (trigram disk round-trip: writeToDisk → readFromDisk → search) 174/174; full zig build test green. Behavior-preserving perf refactor covered by the existing round-trip tests. A large-repo ReleaseFast cold-index wall-time A/B is a nice-to-have follow-up (the pattern itself is #475-validated).

🤖 Generated with Claude Code

justrach and others added 23 commits July 4, 2026 19:10
…r a file edit

Red on release/0.2.5828: the second findCallPath cannot resolve the
newly added alpha -> beta edge because call_graph is never invalidated
on the file-commit path, and its node_name/node_path slices borrow
outline memory the re-index just freed.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Explorer.call_graph was built once and never invalidated: the file-commit
path refreshed deps/symbol-index/line-offsets but left call_graph and
call_centrality pointing at the process-start snapshot, so codedb_callpath
and the graph-distance rerank could not see edges added by an edit.

Worse, CallGraph.node_path/node_name borrowed slices into outline symbol
memory; re-indexing freed those outlines, dangling the name/path tables
(use-after-free, silently-wrong under the testing allocator).

Fix: (1) deinit+null call_graph and call_centrality in commitParsedFileOwnedOutline
so the next graph query lazily rebuilds; (2) dupe node_path/node_name in
ensureCallGraph and free them in CallGraph.deinit, unwinding partial arrays
on every fallible early-return. call_centrality keys keep borrowing the
stable self.outlines keys (node_path.items), preserving the snapshot-restore
ownership invariant.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
#3 #6)

Three behavior-preserving optimizations on the per-query search/rank path,
identified by the perf audit and adversarially verified:

#1 (mcp.zig, appendSearchSymbolNudge): the bare-identifier nudge did a full
   O(U+S) symbol scan (searchSymbols scans the whole symbol_index AND outlines
   every call) just to test existence. Switched to findAllSymbols — the same
   helper codedb_symbol uses: ensureSymbolIndex + O(1) symbol_index.get, with
   the outline scan only when !symbol_index_complete. Output byte-identical
   (the nudge only checks results.len and prints the query string).

#3 (explore.zig, BM25 posting loop): dropped the doc_tf hashmap. doc_best_line
   already accumulates .count identically per posting, so df = doc_best_line.count()
   and tf = entry.value_ptr.count — one map instead of two, and the redundant
   per-doc .get() re-lookup is gone. Values bit-identical.

#6 (explore.zig, tier-1 sort): the comparator hashed 4 strings + touched 2
   atomics per comparison. Precompute (count, len) per candidate once in O(C)
   (mirroring tier0_order), then sort a struct-key array — identical predicate
   (count desc, len asc) and input order, so ranking is unchanged. Also stops
   the comparator corrupting ContentCache hit/miss stats.

zig build test: full suite green (0 failures).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…y blocks

Every MCP tool result carries the assistant data block plus two
audience:["user"] blocks (a colored summary + a guidance hint). These exist
for interactive preview panes, but agent harnesses forward them straight into
model context where they cost output tokens the model can't use — ~34% on a
small result like codedb_symbol (115→75 tok), the very tool codedb tells
agents to reach for first.

A lean mode already existed (CODEDB_MCP_LEAN) but was off by default and
undocumented, so agents paid the tax unless they knew the knob. Make the
decision client-aware, keyed on clientInfo.name from initialize:

- Agent harnesses (claude-code, codex, unknown clients) default LEAN.
- Human-facing GUI clients (claude-ai / Claude Desktop) get the rich blocks.
- Overrides: CODEDB_MCP_LEAN / CODEDB_MCP_RICH force globally (LEAN wins);
  CODEDB_MCP_RICH_CLIENTS=a,b extends the rich allowlist.

Also fixes a latent use-after-free this surfaced: session.client_name
borrowed a slice from the initialize request's parsed JSON, which is freed at
the end of that dispatch iteration. It was only ever read during initialize
before; reading it on a later tools/call (the new policy) dangled. Now duped
into session-owned memory and freed in Session.deinit.

Verified end-to-end (built binary, clean env): claude-code -> lean,
claude-ai -> rich, CODEDB_MCP_LEAN=1 -> lean even for claude-ai,
CODEDB_MCP_RICH_CLIENTS -> opt-in works. Unit test in test_mcp.zig; full
`zig build test` green (testing allocator clean on the UAF fix).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…#1)

When codedb_search's query is a bare identifier that resolves to an indexed
symbol, the nudge used to just say "go call codedb_symbol" — forcing a second
round-trip (~85 tok response + the agent's own call emission) to learn where
the symbol is defined. Now it surfaces the definition site(s) inline
(path:line (kind), up to 3, "+N more"), reusing the O(1) findAllSymbols the
nudge already ran. One call now answers "where is X defined?".

Fewer round-trips = fewer output tokens, which dominate cost.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…irect (#658)

The PreToolUse hook hard-blocked grep/rg/cat/head/tail/sed/awk/find the moment
the codedb binary existed — unconditionally, even outside any repo, on
unindexed dirs, and in ~/.claude. It fired before permission evaluation
(overriding permissions.allow) and re-registered on every install. Blocking a
command an agent then just routes around (python/sed) is friction, not
adoption.

Layered redesign (coercion last), fail-open by default:

- Scope 1: only intercept when cwd is at/under a codedb-indexed root recorded
  in ~/.codedb/projects/*/project.txt. Degenerate "/" and bare-$HOME roots are
  skipped so an indexed home doesn't block all of ~. Outside an indexed repo,
  codedb has nothing to offer — the native tool is allowed.
- Scope 2: if the command targets an ABSOLUTE path outside the repo
  (cat /etc/hosts, grep x /tmp/f), allow — codedb can't read it. Relative
  paths are assumed in-repo.
- Redirect: when it does block, name the exact codedb substitute for THAT
  command (grep->codedb_search, cat->codedb_read, find->codedb_find/glob, ...)
  and state that native tools work outside the repo or with CODEDB_NO_HOOKS=1.
- Opt-out: CODEDB_NO_HOOKS=1 disables the hook entirely.

Verified across the behavior matrix (in-repo relative -> block+redirect;
not-indexed -> allow; in-repo absolute-outside -> allow; bare-home -> allow;
opt-out -> allow; non-listed command -> allow). Installer Python heredoc
compiles.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Consolidate the ~16 duplicated "path:line[: text]" render sites in mcp.zig
behind one appendMatchLine helper (prefixes "  "/"- "/"" and the scope suffix
preserved; paths_only output byte-identical), and trim the match text to save
agent output tokens:
  - strip leading indentation (the line number already locates the hit)
  - cap at 200 bytes with a UTF-8-safe cut + "…"

Also apply the same trim to the default search path — renderPlainSearch in
explore.zig (mcp's helper can't reach it), which the mcp-only consolidation
missed. This is the highest-traffic path (a bare codedb_search query), where
deeply-indented source lines were the bulk of the waste.

Render-only: result set and order unchanged. Full `zig build test` green.

The mcp.zig consolidation was produced by a 3-attempt swarm (winner's diff
applied + independently re-verified); the explore.zig completion + measurement
verification done here.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
# Conflicts:
#	src/test_mcp.zig
…first)

codedb_search's fast path (renderPlainSearchUncached) ordered tier0 files by
hit count + basename prior, with no definition-awareness — so a file that
merely *mentions* a symbol many times (tests, callers) outranked the file that
actually DEFINES it when the def file's basename didn't match the query. The
rich def-first ranking existed only in searchContentRanked, which codedb_search
never calls.

Fix: give tier0 a `defines` flag (via the existing fileDefinesSymbol) and add
+20 to the rank prior for files that define the queried symbol — larger than
the basename-match prior (+15), so a non-canonically-named definition still
wins. Same result SET, better ORDER.

Measured on a pinned, deterministic eval (frozen corpus + isolated $HOME, so
the on-disk index can't drift between runs — the confound that sank the first
attempt): 10 symbol queries, "def file ranks #1" improved 8/10 -> 9/10 with
ZERO regressions (searchContentRanked: pos 3 -> #1). Full `zig build test`
green; new regression test fails on baseline, passes with the fix.

Follow-up: the one remaining holdout (renderPlainSearch) bails to
searchContentAuto, a separate path where docs aren't demoted and this boost
doesn't apply — tracked for a later pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Checks the ranking eval used for PR #665 into the repo so it's reproducible,
not just in a scratchpad:

- scripts/rank-eval.py — self-contained pinned A/B harness for codedb_search
  def-first ranking. Frozen corpus (copied out of the repo) + isolated $HOME so
  the on-disk index can't drift between builds; only the binary differs. Handles
  the temp-root guard (CODEDB_ALLOW_TEMP), async snapshot load (spaced pipe +
  pre-warm), and default/custom gold cases.
- experiments/ranking/def-first-eval.md — methodology: what it measures, the
  snapshot-drift confound it kills, the gotchas it encodes, how to run it, the
  PR #665 result (def-file-#1 8/10 -> 9/10, zero regressions), and the known
  searchContentAuto holdout for the next pass.
- experiments/ranking/todo.md — logged def-first as done (matches the rVSM entry).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
… pass 2)

Follow-up to the def-first ranking (#665). That floated the defining FILE to the
top, but the render still showed the file's first hit LINE — often a comment or
a call above the definition. Now, within a file that defines the queried symbol,
the definition line(s) render before mere mentions.

Implementation: after lineSpans computes the per-file spans, stable-move the
spans whose line is a definition line (outline symbol named == query) to the
front. Reorders the OUTPUT spans only — lineSpans' sorted input is untouched, so
offset resolution is unaffected. Same result set, better in-file order.

Verified on the pinned eval (scripts/rank-eval.py): def-file rank unchanged
(9/10, no regression); qualitatively, `codedb_search bumpSearchGen` now leads
with `explore.zig:2021: fn bumpSearchGen(...)` (the def) instead of a comment.
Regression test in test_search.zig fails on baseline, passes with the fix; full
`zig build test` green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…ng hash -> array load)

TrigramIndex.writeToDisk (the disk persist on every cold index) hashed the full
file-path string once PER POSTING to map an in-memory doc_id to its disk
file_id — `disk_path_to_id.get(id_to_path[doc_id])` inside the Step-3 posting
loop, i.e. tens of millions of wyhash+probe on a large repo.

Precompute doc_id -> disk file_id ONCE into a `doc_to_disk: []u32` (one path
hash per DOC, O(num_docs)), then the posting loop does a single array load per
posting. Retrieval-identical (same disk postings): doc_to_disk[doc_id] equals
the old disk_path_to_id.get(id_to_path[doc_id]), with maxInt standing in for the
old `orelse continue`.

Same fix already shipped for WordIndex (#475). ReleaseFast microbench (from the
perf audit) measured the remap step ~25x (12.5ns -> 0.5ns/posting); whole
writeToDisk ~2.5-4x, diluted by buffered IO + the trigram sort. num_postings >>
num_docs, so the win scales with repo size.

test-index (trigram disk round-trip) 174/174; full `zig build test` green.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown

Benchmark Regression Report

Thresholds: 10.00% and 50,000 ns absolute delta

NOISE means the percentage threshold was exceeded, but the absolute delta was too small to fail CI.

Tool Base (ns) Head (ns) Delta Abs Delta (ns) Status
codedb_bundle 65978 74522 +12.95% +8544 NOISE
codedb_changes 12338 13618 +10.37% +1280 NOISE
codedb_context 316218 323546 +2.32% +7328 OK
codedb_deps 373 370 -0.80% -3 OK
codedb_edit 47094 50291 +6.79% +3197 OK
codedb_find 3053 3054 +0.03% +1 OK
codedb_hot 28669 32662 +13.93% +3993 NOISE
codedb_outline 17656 18688 +5.85% +1032 OK
codedb_read 15662 14865 -5.09% -797 OK
codedb_search 70690 66446 -6.00% -4244 OK
codedb_snapshot 83645 91154 +8.98% +7509 OK
codedb_status 11274 11636 +3.21% +362 OK
codedb_symbol 52798 54394 +3.02% +1596 OK
codedb_tree 23941 24035 +0.39% +94 OK
codedb_word 16513 12846 -22.21% -3667 OK

@justrach
justrach changed the base branch from release/0.2.5828 to release/0.2.5829 July 10, 2026 00:56
@justrach
justrach merged commit f71048b into release/0.2.5829 Jul 10, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant