feat: transient compression holds for provisional messages (for AF #159) - #120
antra-tess wants to merge 2 commits into
Conversation
Adds holdCompression/releaseCompression/getCompressionHolds and an addMessage `holdCompression` option. Autobiographical/Knowledge stop the compressible region before the earliest held message (keeping its tool_use paired) and skip already-closed chunks that reach the hold boundary, so a placeholder later replaced via editMessage is never summarized and the final content compresses after release. In-memory only; unused = unchanged. Needed by anima-research/agent-framework#159 (tool-result guard, review finding #3). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Reviewed head
Validation: the six compression-hold tests pass; Integration note: AF #159 does not yet acquire or release these holds. Running its current — Review by Codex. |
Review (Codex) on #120: - After resetHeadWindow the head window can sit past the hold boundary and compressChunkHierarchical/executeMerge/generateTransitionSummary emitted it as prompt context, leaking the held placeholder. Head context now ends at the hold boundary (stepped back onto the paired tool_use). - rebuildChunks rescanned the timeline per uncompressed chunk. The boundary is computed once per pass (holdBlockedIds), and the view's hasCompressionHolds() fast path skips scanning entirely when nothing is held. Live checks remain for async drains (tick). Regressions: reviewer's moved-head repro; scan counter (0 with no holds, bounded per compile with holds). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Unblocks anima-research/agent-framework#159 (tool-result guard), review finding #3: the guard stages a tool_result as a ~20-token placeholder via
addMessageand later swaps in the real output viaeditMessage. Edits never reachderivedentries or the strategy, and Autobiographical dedupes by message id, so if the placeholder is summarized while pending, the accepted output never reaches compressed memory. The baseline test in this PR reproduces that.Design
The smallest correct fix is a transient compression hold:
addMessage(..., { holdCompression: true })places the hold beforeonNewMessagefires, so the ingress rebuild can't chunk the message.holdCompression(ids),releaseCompression(ids)(idempotent; releasing a sharded add releases every shard), andgetCompressionHolds().isCompressionHeld(id)predicate. Autobiographical'sgetRecentWindowStartis clamped to the earliest held message, stepping back one so the tool_use stays paired with its tool_result. The held message, everything after it, and its tool_use therefore stay in the raw protected tail and never enter a chunk. Knowledge inherits this throughgetCompressibleMessages.holdCompressioncall, and that reaches the hold boundary, is not queued or compressed until release. This also covers chunks after the hold whose lead-in would carry the placeholder. It also covers the kv-stable/kv-unified demand path (enqueueL1ForRange→tick).Edits: I did not add an
onMessageEditedhook. The hold ensures nothing derived exists for a held message, so an edit made while held, followed by release, is correct. Invalidating summaries after an edit would mean re-minting and re-merging, which carries much more risk.editMessageis now documented as "edit while held, then release".Strategies: Autobiographical and Knowledge honor the hold. Passthrough and windowed-passthrough don't compress, so they're unaffected. The folding solvers (kv-stable, kv-unified, flat-profile, oldest-first) only fold existing summaries, and none can cover a held message. Merges are unaffected for the same reason.
Invariants: With no holds the predicate returns false everywhere, so chunking, requests, cache markers and compiled context are unchanged (the existing suite passes untouched). The tool_use/tool_result pairing is kept by the clamp. While a hold is active, the raw tail simply starts earlier; after release it goes back to the token boundary.
Tests
test/compression-hold.test.tsruns a realAutobiographicalStrategyagainst a recording membrane:holdCompression: the same outcome after chunks had already closed.These were red before the fix (type errors first, then 3/6 behavioral failures with a stub API). Result:
npm test868 pass / 0 fail / 0 skipped.Open questions
holdCompression()after the fact can't un-chunk a persisted chunk record. It only defers that chunk's compression, which is enough for correctness. Callers should prefer the addMessage option.🤖 Generated with Claude Code