Skip to content

kv-unified: per-message tail cache units, recipe hysteresis certificate, cached compile metadata - #119

Merged
antra-tess merged 3 commits into
anima-research:mainfrom
theaspirational:feat/kv-unified-tail-units
Sep 25, 2026
Merged

antra-tess merged 3 commits into
anima-research:mainfrom
theaspirational:feat/kv-unified-tail-units

Conversation

@theaspirational

Copy link
Copy Markdown
Contributor

This is @antra-tess's fix/kv-unified-tail-units branch, opened as a PR from a fork so it can be reviewed. The 3 commits are hers (author kept), put on top of current main.

What it does

  • Per-message tail cache units (4de34bd → 7797b51): the raw tail is rendered as one cache unit per message, so appending a message no longer rewrites the whole tail in the provider cache.
  • Recipe-level hysteresisCertificate (d3e4f65 → b4bc9ae): the certificate becomes a recipe option, and fresh-L1 extensions can be certified.
  • Cached compile metadata (bad4b4e → db8fb06): shard metadata is cached and the latent-demand ranking is reused while the summary roots are unchanged.

Rebase onto main

One conflict, in test/adaptive/kv-unified-policy.test.ts: main (#110, #118) and this branch both added tests at the same spot. Kept both. git range-diff shows no code change in the three commits, only line offsets.

npm test: 851/851.

Bench

Measured before #118 merged, on this branch merged with #110's head (ad9bae6), replaying the synthetic corpus built from the real Sill store (corpus-20260921: 75,717 messages, 3,406 summaries), 20 turns, one message appended per turn.

layout and score vs #110 quiet-turn cache rewrite time
certificate off same on 20/20 turns 0 (was ~120k tokens) mean 6.8 → 6.1 s
certificate on same on 20/20 turns 0 quiet-turn median 5.5 → 1.9 s
  • With the certificate on, 14 of 15 quiet turns were certified. The one it declined (turn 5) fell back to the full solve and still picked the same layout.
  • Memory: heap after GC is about 360 MB with the certificate on or off.
  • A certified turn still takes about 1.7 s: certificate check 0.57 s, forest rebuild 0.30 s, rebuildChunks 0.24 s, getTree 0.11 s. Those are the next things to trim.

🤖 Generated with Claude Code

https://claude.ai/code/session_01UN6d9TeXZDWNMxQ3dT5uR8

antra-tess and others added 3 commits September 25, 2026 20:53
The tail was one opaque ('tail','tail') layout unit. Every append shifted
its position (messages leaving the tail became new raw units in front of
it), so the end-of-tail cache marker fell outside the identical prefix and
an unchanged layout was priced as the whole tail recomputed on every turn
— ~100k tokens on Sill — although the wire bytes were identical. The false
churn was constant across candidates and cancelled in the floor
normalization, so selection was unaffected, but churn/cacheFloor were wrong
and the hysteresis certificate's zero-churn precondition never held in
steady state (3 of 50 solves certified in a 40-turn trial).

Tail chunks now render as raw units keyed by chunk id — the identity they
keep after sliding into the middle — in renderLayout, both solver storages
and the terminal evaluator; tail-message cache markers map to that unit.
Tail tokens no chunk accounts for (synthetic inputs) keep the opaque block,
so token totals are unchanged.

Trial on a copy of Sill's store, 30 simulated turns, certificate on:
22 of 23 unchanged turns certified, solve 0.3 s (was 4-6 s). Full suite
824/824; oracle-equivalence tests unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…1 extensions

Adds `hysteresisCertificate` to KvUnifiedConfig (the adapter already spreads
the config into solver options). The certificate no longer declines when an
appended leaf has a non-raw option: hysteresis keeps the accepted layout
under its best extension, so every cut that fixes each accepted leaf at its
accepted level is enumerated (cap 256 → decline), scored exactly with the
zero-floor witnesses, and the best is certified against the tangent bound.
Tests: the PR's "ambiguous extension" test now asserts agreement with the
exact oracle; new tests cover fresh-L1 extensions across 24 varied runs.

30 simulated turns on a copy of Sill's store with the flag set in the
recipe: 24/25 unchanged turns certified (0.4 s solve), 24 L1 mints.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ross turns

Profile of certified kv-unified turns put ~1.5 s per compile outside the
solver. mergeAdjacentBodyGroupRaw called store.get() (chronicle fetch +
blob resolve) once per raw entry to read bodyGroupId/shardIndex: ~0.7 s on
a 75k-message store. It now indexes the cached getAll() listing (sharded
messages only). rankLatentDemand rebuilt a forest and ran two what-if
solves per candidate on every compile; its ranking depends only on the
summary-root runs and the budget, so it is cached on the strategy and
reused until a mint or merge changes the roots (trailing raw roots — the
messages appended each turn — are excluded from the key).

30 simulated turns on a copy of Sill's store, certificate on: 23/24
unchanged turns certified, compile median 1.16 s, 73% of turns under 1.5 s
(the rest are real transitions at 4-6 s).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@theaspirational

Copy link
Copy Markdown
Contributor Author

Review of this PR (GPT-6 Sol, high, own clone at db8fb06): changes needed, 2 findings, both in the original commits, not the rebase.

The rebase is clean: git range-diff shows the three changes and both test additions intact, and npm test passes 851/851.

1. Medium: the latent-demand cache reuses stale rankings (db8fb06, src/adaptive/strategies/kv-unified.ts:126-130)

The cache key is the summary roots + totalBudget + mergeThreshold + maxCandidates. On a hit it returns the old evaluations and produced. But the scores in them also depend on the policy, the chunks, the presentation, the recall costs and the solver options. None of those are in the key. The live adapter keeps this cache across compiles (autobiographical.ts), so this affects real turns.

It's the right idea that trailing raw roots don't change the candidates, but they do change the scores. Reproduced with a probe on the new test's 8-chunk fixture:

  • Change budgetOverLambda from 1,000,000 to 0, same roots and budget. The cached result still requests the L2, with improvement 22,156. A fresh solve requests nothing (improvement 0).
  • Append one 60-token raw chunk. The cached improvement stays 22,156; fresh gives 75,496.
  • Switch engine to leaf. Cached approximate is true; fresh is false. So the fix(kv-unified): exact token key in the leaf engine (#109); tests from the #110 review #118 flag can go stale too.

Options: cache only the candidate list (which really does depend only on the roots) and re-score each turn, or add every scoring input to the key. Either way, a test that compares a cached run with a fresh run on the same roots but different inputs would guard it.

2. Low: the certificate doc is out of date (docs/kv-unified-hysteresis-certificate.md:44-46, 56-57)

It still says new leaves must be raw-only (foldable ones decline) and that successful results have one candidate. b4bc9ae now enumerates foldable extensions (kv-unified-certificate.ts:137-190) and can return several.

Not checked by the reviewer: the replay and timing numbers in the PR body, which come from an external corpus.

@theaspirational

Copy link
Copy Markdown
Contributor Author

Follow-up to the review above: a benchmark of the latent-demand ranking cache (db8fb06), to help pick a fix.

Setup. Our 60-turn replay of the synthetic corpus built from the Sill store, with a fake summarizer so the strategy's own L1/merge pipeline runs. Recipe: tokenBucketSize 50k, budgetHighRatio 0.85, certificate on, adoptEpsilon 2,000. Options, switched by an env var in a local build only:

  • stale: the PR as is
  • off: no ranking cache (reference)
  • keyplus: key also covers policy, engine, buckets and root recall costs
  • ttl5: keyplus plus a full re-rank every 5 turns

A shadow mode also ran a fresh ranking on every cache hit and compared the produced lists.

summarizer candidates stale keyplus / ttl5 off
keeps up (4 calls/turn, 700 words) 0 on all 60 turns ranking ~0 ms, same layouts same same
lags (1 call/turn, 2,000 words) 1–3 on 23 turns no faster: ~99 s ranking total, same layouts same same
stops after turn 15 (turns 16–21) 2 per turn 2.7 s/turn, 4 of 6 hits wrong same as stale 9.4 s/turn

Why.

  • Candidates need 6+ same-level roots side by side. The pipeline merges at 6, so while it keeps up there is nothing to rank.
  • While it lags, each new L1 changes the key, so the cache misses on exactly the expensive turns.
  • It only hits when the summarizer stalls. There it replayed the turn-15 answer ("no merges") while a fresh ranking asked for 1–2 L2 merges worth 293–1,107. The run hit OverBudgetError at turn 22 (the fake summarizer was stopped, so nothing could act on the requests).
  • Widening the key doesn't help: the drift comes from appended messages, not settings.

Suggestion: drop the ranking cache and keep the shard-metadata index from the same commit. Ranking is free when there's nothing to rank, and it costs time only when summaries pile up, which is when a stale answer does harm.

Checked, not a problem: what-ifs keep the certificate on. On certified turns every merge scores 0. With the what-if certificate forced off, the full solves give the same 0 on the same turns (the carried layout wins by hysteresis either way) and the same requests, while taking ~3.5 s more per certified turn. So the what-if certificate is correct and worth keeping.

Not covered: the real Sill/KR stores, and a live summarizer with real latency.

@antra-tess
antra-tess merged commit 0975648 into anima-research:main Sep 25, 2026
5 checks passed
@Anarchid

Copy link
Copy Markdown
Contributor

Post-merge second pass (Claude + GPT-6 Sol at xhigh) on db8fb06, whose tree is identical to 0975648 on main. No new correctness defects in the per-message tail or the certificate extension. There are three small follow-ups on top of the earlier review; none is urgent.

Checked and fine

  • Tail units on the production path: selectAdaptive makes one picker chunk per tail message (id = message id, sequence = message index), and tailTokens is the same per-message sum (autobiographical.ts:7593-7599, 7731-7744). So in production the opaque tail residual is always empty. The receipt records every chunk, tail included (7815-7835), so only messages appended since the last accepted presentation count as extension. Merged shard entries map their marker to the last shard's unit.
  • The certificate picks the same carried cut as the full solve (the best-scoring cut that matches the presentation, with new leaves free). Requiring zero churn on every enumerated cut keeps its floors equal to the global ones. A randomized in-memory sweep (400 cases on a 7-chunk forest) certified 127, with 0 selection or score differences from the exact solver.
  • npm test passes 851/851 on the merged tree.

1. Latent-demand ranking cache: widening the key won't fix it, so drop it. This adds to finding 1 of the earlier review.

  • Not branch-safe. kvUnifiedLatentDemandCache lives on the strategy instance. A branch switch re-initializes that instance in place (clearBranchMirrors() at autobiographical.ts:1693, then loadPersistedState()), and neither clears the cache. The key is built from summary ids (kv-unified.ts:127), and those come from a per-branch counter (autobiographical.ts:2006). So sibling branches forked at the same counter value can mint the same ids over different chunks. A wider key (the "keyplus" or "ttl5" options in the benchmark above) would still hit across branches unless it includes the branch identity or the cache is cleared on a switch. A hit returns produced ranges computed on the other branch. If their endpoint ids don't resolve here, enqueueMergeForRange treats every source as in range (autobiographical.ts:4041) and merges the first N unparented summaries: the Demand-path merges (enqueueMergeForRange, produce level ≥ 2) still bypass the strict-adjacency grammar — can mint a non-contiguous L2 that kv-unified then rejects forever #95 path. This is from reading the code; I did not run it end to end.
  • Never hits when productionBudgetTokens is set. The shadow pick (autobiographical.ts:7917-7920) builds its KvUnifiedStrategy with the same single-slot cache but a different totalBudget. Live and shadow compiles keep overwriting each other's entry.
  • A test asserts the bug. test/adaptive/kv-unified-policy.test.ts:1360-1364 asserts that the cached evaluations are reused after appending one 60-token raw chunk. That is exactly the case the earlier probe showed to be wrong (22,156 cached vs 75,496 fresh). The fix will need to invert this test.

2. The adapter-level marker test only exercises the fallback. test/kv-unified-strategy-integration.test.ts:182-194 still passes the pre-PR layout with a single ('tail','tail') unit. So the tail marker resolves through the new 'tail' fallback (autobiographical.ts:8701), a branch production no longer takes. The new tail-unit and certificate tests use synthetic PickerInputs. One test through selectAdaptive → reconcileKvUnifiedMarkerIndices with per-message tail units, including a sharded message in the tail, would cover the path this PR actually changed.

3. Nit: mergeAdjacentBodyGroupRaw calls store.getAll() a second time (autobiographical.ts:8743). On a plain MessageStore this is cached and free. Behind a viewFilter or auxiliary views, each call re-filters or re-sorts the whole history (message-view.ts:31, 74-77). selectAdaptive already holds that listing (7566), so passing it in avoids the extra pass. It's still a net win over the old per-entry get().

Not re-checked: the corpus replay numbers in the PR body.

antra-tess added a commit that referenced this pull request Sep 26, 2026
#121)

fix(kv-unified): drop the latent-demand ranking cache; #119 follow-ups
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants