kv-unified: per-message tail cache units, recipe hysteresis certificate, cached compile metadata - #119
Conversation
The tail was one opaque ('tail','tail') layout unit. Every append shifted
its position (messages leaving the tail became new raw units in front of
it), so the end-of-tail cache marker fell outside the identical prefix and
an unchanged layout was priced as the whole tail recomputed on every turn
— ~100k tokens on Sill — although the wire bytes were identical. The false
churn was constant across candidates and cancelled in the floor
normalization, so selection was unaffected, but churn/cacheFloor were wrong
and the hysteresis certificate's zero-churn precondition never held in
steady state (3 of 50 solves certified in a 40-turn trial).
Tail chunks now render as raw units keyed by chunk id — the identity they
keep after sliding into the middle — in renderLayout, both solver storages
and the terminal evaluator; tail-message cache markers map to that unit.
Tail tokens no chunk accounts for (synthetic inputs) keep the opaque block,
so token totals are unchanged.
Trial on a copy of Sill's store, 30 simulated turns, certificate on:
22 of 23 unchanged turns certified, solve 0.3 s (was 4-6 s). Full suite
824/824; oracle-equivalence tests unchanged.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…1 extensions Adds `hysteresisCertificate` to KvUnifiedConfig (the adapter already spreads the config into solver options). The certificate no longer declines when an appended leaf has a non-raw option: hysteresis keeps the accepted layout under its best extension, so every cut that fixes each accepted leaf at its accepted level is enumerated (cap 256 → decline), scored exactly with the zero-floor witnesses, and the best is certified against the tangent bound. Tests: the PR's "ambiguous extension" test now asserts agreement with the exact oracle; new tests cover fresh-L1 extensions across 24 varied runs. 30 simulated turns on a copy of Sill's store with the flag set in the recipe: 24/25 unchanged turns certified (0.4 s solve), 24 L1 mints. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ross turns Profile of certified kv-unified turns put ~1.5 s per compile outside the solver. mergeAdjacentBodyGroupRaw called store.get() (chronicle fetch + blob resolve) once per raw entry to read bodyGroupId/shardIndex: ~0.7 s on a 75k-message store. It now indexes the cached getAll() listing (sharded messages only). rankLatentDemand rebuilt a forest and ran two what-if solves per candidate on every compile; its ranking depends only on the summary-root runs and the budget, so it is cached on the strategy and reused until a mint or merge changes the roots (trailing raw roots — the messages appended each turn — are excluded from the key). 30 simulated turns on a copy of Sill's store, certificate on: 23/24 unchanged turns certified, compile median 1.16 s, 73% of turns under 1.5 s (the rest are real transitions at 4-6 s). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Review of this PR (GPT-6 Sol, high, own clone at The rebase is clean: 1. Medium: the latent-demand cache reuses stale rankings ( The cache key is the summary roots + It's the right idea that trailing raw roots don't change the candidates, but they do change the scores. Reproduced with a probe on the new test's 8-chunk fixture:
Options: cache only the candidate list (which really does depend only on the roots) and re-score each turn, or add every scoring input to the key. Either way, a test that compares a cached run with a fresh run on the same roots but different inputs would guard it. 2. Low: the certificate doc is out of date ( It still says new leaves must be raw-only (foldable ones decline) and that successful results have one candidate. Not checked by the reviewer: the replay and timing numbers in the PR body, which come from an external corpus. |
|
Follow-up to the review above: a benchmark of the latent-demand ranking cache ( Setup. Our 60-turn replay of the synthetic corpus built from the Sill store, with a fake summarizer so the strategy's own L1/merge pipeline runs. Recipe:
A shadow mode also ran a fresh ranking on every cache hit and compared the
Why.
Suggestion: drop the ranking cache and keep the shard-metadata index from the same commit. Ranking is free when there's nothing to rank, and it costs time only when summaries pile up, which is when a stale answer does harm. Checked, not a problem: what-ifs keep the certificate on. On certified turns every merge scores 0. With the what-if certificate forced off, the full solves give the same 0 on the same turns (the carried layout wins by hysteresis either way) and the same requests, while taking ~3.5 s more per certified turn. So the what-if certificate is correct and worth keeping. Not covered: the real Sill/KR stores, and a live summarizer with real latency. |
|
Post-merge second pass (Claude + GPT-6 Sol at xhigh) on Checked and fine
1. Latent-demand ranking cache: widening the key won't fix it, so drop it. This adds to finding 1 of the earlier review.
2. The adapter-level marker test only exercises the fallback. 3. Nit: Not re-checked: the corpus replay numbers in the PR body. |
This is @antra-tess's
fix/kv-unified-tail-unitsbranch, opened as a PR from a fork so it can be reviewed. The 3 commits are hers (author kept), put on top of currentmain.What it does
4de34bd→7797b51): the raw tail is rendered as one cache unit per message, so appending a message no longer rewrites the whole tail in the provider cache.hysteresisCertificate(d3e4f65→b4bc9ae): the certificate becomes a recipe option, and fresh-L1 extensions can be certified.bad4b4e→db8fb06): shard metadata is cached and the latent-demand ranking is reused while the summary roots are unchanged.Rebase onto main
One conflict, in
test/adaptive/kv-unified-policy.test.ts: main (#110, #118) and this branch both added tests at the same spot. Kept both.git range-diffshows no code change in the three commits, only line offsets.npm test: 851/851.Bench
Measured before #118 merged, on this branch merged with #110's head (
ad9bae6), replaying the synthetic corpus built from the real Sill store (corpus-20260921: 75,717 messages, 3,406 summaries), 20 turns, one message appended per turn.rebuildChunks0.24 s,getTree0.11 s. Those are the next things to trim.🤖 Generated with Claude Code
https://claude.ai/code/session_01UN6d9TeXZDWNMxQ3dT5uR8