Skip to content

fix(compress): refuse compressions that do not shrink the context - #643

Open
Mukller wants to merge 1 commit into
Tarquinen:devfrom
Mukller:fix/compression-size-guard
Open

Mukller wants to merge 1 commit into
Tarquinen:devfrom
Mukller:fix/compression-size-guard

Conversation

@Mukller

@Mukller Mukller commented Oct 1, 2026

Copy link
Copy Markdown

fix(compress): refuse compressions that do not shrink the context

Closes #573

Problem

Across one 454-message session, repeated range compressions entered a
self-reinforcing loop: each new block re-absorbed the previous block's (bN)
placeholder plus a thin new tail of raw messages, and produced a summary larger
than the content it replaced
. The on-disk summary grew monotonically to ~68k
tokens, 71 blocks were created, and 738,738 prune tokens were burned.

Two things allowed it:

  1. No size guard at all. A compression that destroys the original content and
    replaces it with something bigger was accepted and persisted. The model was
    never told the compression was pointless, so it retried, and because each
    attempt wrapped the previous block again, the block grew every round.

  2. The cost of a range was under-counted. selectionTokenCost did not exist;
    only the raw messages were considered. A b10..m0190 range looks almost free
    when what it actually absorbs is block 10's whole summary, which is precisely
    the range shape the runaway growth used.

Change

  • selectionTokenCost(selection, messagesState) — raw message tokens in the range
    plus summaryTokens of every compressed block the range absorbs.
  • assertCompressionShrinksContext(summaryTokens, inputTokens, boundaryLabel, idFormat) —
    throws when the summary is not smaller than the content it replaces, naming the
    boundary and telling the caller to widen the range or leave the content alone.
    Applied in both compress/range.ts and compress/message.ts, before
    applyCompressionState, so a refusal leaves no partial state.

The refusal is what breaks the loop: the model gets a real error instead of a
successful-but-useless compression, and the original content is still there.

The floor, and why it exists

The guard only engages at or above 2,000 input tokens. Below that it stays out
of the way entirely. Compressing three short messages into a slightly longer
summary is a normal, deliberate choice, and tokenizer noise on very small texts
makes the comparison unreliable — an early version of this guard with no floor
broke 17 existing tests that compress tiny fixtures.

The runaway growth being guarded against is orders of magnitude above that: the
smallest pathological block in the report was already ~1.6k tokens on the way to
~68k. tests/compression-size-guard.test.ts pins both sides of the floor.

Verification

tests/compression-size-guard.test.ts, seven cases. Five fail on the current
implementation.

  • a range whose summary is bigger than its content is refused, and no block is
    created
    (asserted on state.prune.messages)
  • a range with a genuine gain is still applied
  • message mode applies the same guard and names the message id
  • selectionTokenCost counts absorbed block summaries (1_000 + 40_000)
  • the guard fires at equal and at worse-than-original, and not below the floor
  • the boundary is rendered in the session's id format (@4..6@ for compact)
  • a small range is left alone

Full suite 126 passing, tsc --noEmit clean, prettier --check clean.

Closes Tarquinen#573

## Problem

Across one 454-message session, repeated range compressions entered a
self-reinforcing loop: each new block re-absorbed the previous block's `(bN)`
placeholder plus a thin new tail of raw messages, and produced a summary **larger
than the content it replaced**. The on-disk summary grew monotonically to ~68k
tokens, 71 blocks were created, and 738,738 prune tokens were burned.

Two things allowed it:

1. **No size guard at all.** A compression that destroys the original content and
   replaces it with something bigger was accepted and persisted. The model was
   never told the compression was pointless, so it retried, and because each
   attempt wrapped the previous block again, the block grew every round.

2. **The cost of a range was under-counted.** Only the raw messages were
   considered. A `b10..m0190` range looks almost free when what it actually
   absorbs is block 10's whole summary, which is precisely the range shape the
   runaway growth used.

## Change

- `selectionTokenCost(selection, messagesState)` — raw message tokens in the range
  **plus** `summaryTokens` of every compressed block the range absorbs.
- `assertCompressionShrinksContext(summaryTokens, inputTokens, boundaryLabel, idFormat)` —
  throws when the summary is not smaller than the content it replaces, naming the
  boundary and telling the caller to widen the range or leave the content alone.
  Applied in both `compress/range.ts` and `compress/message.ts`, before
  `applyCompressionState`, so a refusal leaves no partial state.

The refusal is what breaks the loop: the model gets a real error instead of a
successful-but-useless compression, and the original content is still there.

## The floor, and why it exists

The guard only engages at or above **2,000 input tokens**. Below that it stays out
of the way entirely. Compressing three short messages into a slightly longer
summary is a normal, deliberate choice, and tokenizer noise on very small texts
makes the comparison unreliable — an early version of this guard with no floor
broke 17 existing tests that compress tiny fixtures.

The runaway growth being guarded against is orders of magnitude above that: the
smallest pathological block in the report was already ~1.6k tokens on the way to
~68k. `tests/compression-size-guard.test.ts` pins both sides of the floor.

## Verification

`tests/compression-size-guard.test.ts`, seven cases. Five fail on the current
implementation.

- a range whose summary is bigger than its content is refused, and **no block is
  created** (asserted on `state.prune.messages`)
- a range with a genuine gain is still applied
- message mode applies the same guard and names the message id
- `selectionTokenCost` counts absorbed block summaries (`1_000 + 40_000`)
- the guard fires at equal and at worse-than-original, and not below the floor
- the boundary is rendered in the session's id format (`@4..6@` for compact)
- a small range is left alone

Full suite 126 passing, `tsc --noEmit` clean, `prettier --check` clean.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant