Skip to content

Compaction that always makes progress, and runs that end on their own terms - #2

Merged
ascorblack merged 10 commits into
mainfrom
compaction-contract
Sep 26, 2026
Merged

ascorblack merged 10 commits into
mainfrom
compaction-contract

Conversation

@ascorblack

Copy link
Copy Markdown
Contributor

What changed

Two groups of fixes, ten commits, each self-contained: every commit passes the
full suite, ruff, mypy --strict and the leak scan on its own, so the branch
can be read — and bisected — one commit at a time.

How a run ends (four commits)

  • The wind-down notice ("tools are withdrawn, write your answer") is removed
    from history once the drive reaches a terminal state, and written back only
    for a resumed run that is still wound down.
  • A run stopped by its wall clock is told about the time limit
    (soft_stop_notice_text_deadline), and a run whose model keeps returning
    reasoning with no answer and no tool call enters a new model_no_progress
    cause (soft_stop_notice_text_model_no_progress) instead of the
    provider-error one.
  • The retry decision honours the retryable flag of the classification an
    adapter attaches (ClassifiedLike carries it); a permanent
    LLMProviderError no longer enters the wind-down. The fallback chain is
    still tried first.
  • A failed round completes "on its preserved answer" only when that answer was
    written in the round that failed, not anywhere in untagged history.

Compaction (six commits)

  • Adjacent short units are grouped and summarised as one
    (compaction_summary_group_max_tokens, 6,000; 0 disables).
  • The gate sizes the whole prompt: it adds what the last request carried
    besides the history (system messages and tool definitions, calibrated). The
    figure rides the snapshot.
  • A compaction the pre_compact hook refuses returns the run to RUNNING
    (compaction_refused_by_hook) instead of leaving it in COMPACTING.
  • The compaction contract: a pass ends at or below
    compaction_target_ratio (0.6) of the trigger, or at a floor, whatever the
    summariser does. Written up in docs/compaction.md, mirrored in
    docs/ru/compaction.md.

Why

How a run ends

  • The notice outlived its run. When the wind-down's own final turn failed
    (the outage that caused the wind-down also killed the answer), the notice
    was the last thing in the session. The next run obeyed it: it reported what
    it had not finished and called nothing, although every tool was back.
  • The notice named the wrong cause. A deadline read as "reached its
    budget", and a model stuck in reasoning read as "the model endpoint failed".
    The model repeats the notice to the user, so the user was told that an
    endpoint that had answered every request had failed.
  • A final refusal was retried and then wound down. A reply the adapter
    classified as final under a coarse reason such as server_error went
    through the backoff. The wind-down then sent one more request, with the same
    history, to the endpoint that had just refused it, and was refused the same
    way. The run failed a turn later, and its cause read as a wind-down.
  • A failed run completed on the previous run's reply. The seed tag is set
    by the executor, not by every host. A host that hands the engine a session's
    earlier turns untagged therefore had a run whose every request was refused
    complete "successfully", with nothing written and no error shown.

Compaction

The failure that led here was a long-running autonomous loop. A pass failed
the run on the retry budget every time the loop woke. Three things combined:

  1. The gate compared the history with a whole-prompt trigger. The trigger
    and the emergency cliff are whole-prompt sizes. With a large tool surface
    and system prompt (153 tools, 56k of a 256k window), the history reached
    the trigger only once the request was past the provider's ceiling. The
    request fit then clipped the output cap turn after turn, down to a few
    hundred tokens in a replay, and compaction first ran on a refusal.
  2. Short rounds could not be compacted by any tier. Tier 2 skipped every
    unit below compaction_summary_min_unit_tokens, and the fold takes only
    summaries and operator turns. The loop's history was 160 units of 300 to
    1,400 tokens under a 1,500-token floor, and not one unit was eligible. Each
    pass freed zero to a hundred tokens.
  3. The summary carrier lost data and failed often. It was a single JSON
    string capped at 1,024 characters. Verbatim identifiers did not survive it,
    and a reply the model left unterminated, or wrote as prose, was discarded
    as "not JSON".

The compaction contract

docs/compaction.md states twelve clauses, each checked by a test in
tests/unit/runtime/test_compaction_contract.py. In short:

  • Progress. An opened pass ends with the whole prompt at or below its
    trigger, or at the floor. A pass that could change nothing is not opened.
    The retry budget is spent only when the history is already at the floor
    and a tier raised.
  • Tier order. Each tier runs only while the prompt is above its target:
    1. mask old tool outputs by size and by age, with a pointer to the stored
      original and the lines the output said only once;
    2. summarise spans of adjacent units;
    3. fold old summaries;
    4. the floor: remove the oldest spans with no model call, raw before
      summarised, leaving one digest written by code.
  • A failed summary is a normal outcome. Transport errors, timeouts, empty
    replies, replies too large and replies not smaller than their input all
    count as normal. The summariser's input is bounded and each call has a
    deadline.
  • Facts come from state, not from the model. A ledger message is rebuilt
    by code on every pass and is never shown to the summariser. It holds the
    operator's words verbatim, files by tool role, values of recognisable
    shape with their line, failed calls and the latest plan.
  • The carrier is plain text under five fixed headings, requested with
    complete_text. The reader accepts JSON, unterminated JSON, a reply with
    no headings and a truncated reply. Sections are clamped at line
    boundaries.
  • Instructions are isolated. The instruction goes in the system role and
    the material inside <transcript> tags. Each carrier opens by saying the
    runtime wrote it.
  • What is never touched: the task's first turn, the ledger, reference
    blocks, the keep tail and the in-flight batch. Tool pairs stay whole.
  • Carriers carry no timestamps, so recorded requests still replay.

Evidence (aggregate)

Replay of the failing history. A real QueryEngine ran on the recorded
shape: 357 messages, the real system prompt, 153 tool definitions, a 256k
window. The agent was a double that makes one small tool call per turn, for
150 turns. The provider was a double that counts the prompt the way the live
provider did and refuses any request that does not fit. Every summariser call
went to a real model endpoint, four endpoints in all.

core outcome provider refusals passes max prompt summariser calls identifiers kept
before the fixes failed on call 76 2 75 190k 0 —
2.0.0a21 + gate + grouping 150 turns 0 2 155.8k 28–29 237–257 / 354
this branch 150 turns 0 2 155.8k 21–22 278–318 / 354

Summariser faults injected. Two unterminated JSON replies, two empty
replies, a 300-second hang and a 502. Without the contract, the pass waited
out the hang (303 s). With it, the call deadline cut the hang, both
unterminated replies were recovered, and the first pass ended below the
trigger without reaching the floor.

Planted facts. A synthetic agent session of about 150k tokens in six
parts had 14 scored values planted across it: ports, a path, a UUID, UTC
times, a URL, a key id and a non-Latin name. A forced pass ran after each
part, with 40k and 24k windows, on four endpoints, three runs per cell. The
same endpoint then answered from the compacted context.

recall context left (40k window) context left (24k window)
without the contract 0.76–0.93, down to 0.57 in a single run 16.9k–20.6k 9.2k–13.3k
this branch 0.929 in every one of the 24 runs 11.9k–12.9k 5.8k–6.0k

The value missed everywhere, in both arms, is a count with a unit ("19
shards"). The model answers "19": the carrier holds the value, and the loss
happens in the answer.

What changes for a consumer

This is a pre-release, and these changes are intentional. They do not follow
the "new behaviour defaults off" convention, because the old defaults are
what failed.

  • Masking by age starts at once. Tool outputs of at least
    compaction_mask_min_tokens, older than the most recent
    compaction_mask_keep_recent_results (8), become placeholders. The
    placeholder points at the stored original.
  • The ledger takes up to compaction_ledger_max_tokens of the window once
    it fills.
  • Summaries are longer. The output budget is
    compaction_summary_ratio of a span instead of 1,024 characters.
  • Removed constants. Eight constants that sized a JSON string in words
    are removed; the changelog lists them. Setting one now fails validation,
    as for any unknown field.
  • Retry-budget semantics. A failing summariser no longer exhausts the
    retry budget. Exhaustion now means the floor was unavailable.
  • Retry decisions. An adapter that sets retryable=False on a
    classification now stops the retry, whatever the reason says.
  • The summariser is called through complete_text, which is already on
    ILLMProvider, instead of complete_structured. An adapter whose
    complete_text is a stub must implement it.

To get closer to the previous behaviour without code changes, raise
compaction_mask_keep_recent_results, move compaction_target_ratio towards
1, lower compaction_ledger_max_tokens, or set
compaction_summary_group_max_tokens=0.

Checklist

  • uv run pytest . passes: 3657 passed, 139 skipped, coverage 92.88%.
    Every commit also passes on its own.
  • uv run ruff check . is clean
  • uv run mypy --strict is clean (341 source files, no path argument)
  • uv run bandit -r protocore -q -c pyproject.toml is clean
  • Behaviour changes have a test that fails without the change:
    test_soft_stop.py, test_provider_refusals.py,
    test_history_run_boundary.py, test_compaction_small_units.py,
    test_compaction_gate_overhead.py, test_compaction_retry_budget.py and
    test_compaction_contract.py (24 contract tests, with a planted-facts
    fixture and a shape-only long-loop fixture)
  • New tunables are RuntimeConstants fields with a description. The
    compaction defaults change on purpose; see above.
  • Docs updated: docs/compaction.md and its Russian mirror,
    docs/architecture.md in both languages, and the changelog

ascorblack added 10 commits September 26, 2026 15:41
The notice telling the model its tools are withdrawn stayed in history after
the run ended. When the wind-down's final turn itself failed, the next turn
read it as its instruction and gave up without calling a tool. It is removed
when a drive reaches a terminal state and put back only for a resumed run
that is still wound down.
…ey are

A run stopped by its wall clock was told it had reached its budget, and a run
whose model kept returning reasoning with neither an answer nor a tool call
was wound down under the provider-error cause, whose notice says the model
endpoint failed. The model reads the notice literally and passed that on to
the operator as the reason the run stopped, although the endpoint had
answered every request.

The deadline now has its own text naming the time limit, and the empty-round
policy enters a new model_no_progress cause with a text that says the
model's output failed and the requests did consume tokens and time. Each
text is a constant, and a blank one falls back to the general notice.
…d down

Whether to try the same endpoint again read the reason of the classification
an adapter attached to the error and ignored its retryable flag, although the
classification contract carries both and the adapter set the flag as its
answer to exactly that question. A reply the adapter read as final under a
coarse reason such as server_error was retried through the backoff. The flag
now decides when the classification sets one; without it the reason decides
as before.

A permanent provider error also no longer enters the wind-down. Its one turn
is a request to the endpoint that has just refused the run, with the same
history and one more message, and it is refused the same way: the run failed
a turn later, and its cause read as a wind-down rather than as the
provider's refusal. The fallback chain is still tried first, since another
model may serve the request.
Completing a failed run on its "preserved answer" required substantive prose
after the run's latest work, read from every message not tagged as seeded.
The seed tag is set by the executor and not by every host, so on a host that
hands the engine a session's earlier turns untagged, the previous run's reply
counted: a run whose every request was refused completed with nothing
written, and the operator was shown no reply and no error.

The answer must now have been written in the round that failed, the span
after the last message a caller put in. The helper that finds that span is
shared with the check that decides whether a run produced anything a
wind-down could report on.
Tier 2 skipped every unit below compaction_summary_min_unit_tokens, and the
fold takes only summaries and operator turns. A run that works in many
rounds of one short tool call and a short result each therefore built a
history in which no tier had anything to do, however large it grew. Over a
live history of that shape, 160 units of 300 to 1,400 tokens under a
1,500-token floor, not one unit was eligible: every proactive pass freed
nothing and the third one failed the run on the retry budget.

Adjacent small units are now joined and summarised as one unit. Units join
only when nothing lies between them, when each is contiguous in itself and
when they share a seed provenance, so tool pairing stays whole and the
replacement carries exactly one provenance. A group stops before it would
pass compaction_summary_group_max_tokens (6,000 by default; 0 turns grouping
off) and is sent only when it clears the floor a single unit must clear. It
is keyed by its first anchor, so the failure census and the dedup set treat
it as that unit.
The compaction trigger is a whole-prompt size: the lower of the configured
ratio of the window and the largest prompt the provider accepts, less a
turn's headroom. The emergency cliff is one too. The gate held the history's
estimate alone against both. On a deployment whose system prompt and tool
definitions take a large share of the window — 153 tools and a long system
prompt, 56k of a 256k window — the history reached the trigger only once
the request as a whole was well past the provider's ceiling. Proactive
compaction then never ran first: the request fit clipped the output cap
turn after turn, down to a few hundred tokens, and compaction ran only after
a refusal.

The engine now records, at each dispatch, the heuristic size of what the
request carried besides the history (its system messages and tool
definitions), and the gate adds it at the current calibration. Before a
run's first request the figure is zero and the gate reads the history
alone, as it did. It rides the snapshot, so a resumed run's first gate is
sized the same way.
The refusal returned from inside COMPACTING, so the run stayed in that state
for the rest of the turn. It now returns to RUNNING and says why.
…ser does

A long-running loop gated on its whole prompt, could shrink only the history,
found no unit worth a summariser call and failed after 75 passes that freed a
few tokens each. The summariser itself was a single JSON string under a
1,024-character ceiling: it lost verbatim identifiers, and a reply the model
did not close, or wrote as prose, was thrown away as "not JSON".

A pass is now a cascade, each tier running only while the whole prompt is
above its target (compaction_target_ratio of the trigger):

- masking: old tool outputs become a placeholder that names the tool, keeps
  the lines the output said only once and points at the stored original;
- summarising: adjacent units are joined into spans and each span is written
  up as plain text under five fixed headings, requested with complete_text
  under a system-role instruction, read tolerantly and clamped section by
  section at line boundaries; runs of old summaries are then folded;
- the floor: when those stop short of the trigger, the oldest spans are
  removed without a model, raw before summarised, leaving one digest.

Before any tier removes something it records operator words, files touched,
exact values of recognisable shape, failed calls and the latest plan in a
ledger message that code rebuilds on every pass and the summariser never
sees. The summariser's input is bounded, each call has a deadline, and every
failure kind is a normal outcome. compaction_completed reports each tier and
the pass's outcome.
What a pass guarantees, in clauses a test can check: progress to the trigger
or the floor, the tier order, a failed summary as a normal outcome, facts from
state, the carrier, instruction isolation, re-injection, the gate,
reversibility, what is never removed, whole tool pairs and replayable
requests; with the budget arithmetic, the failure states and what a completed
pass reports.
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
🔒 Security Review ✅ Completed 2026-09-26T11:02:14.510952Z 0d6eafd PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@ascorblack
ascorblack merged commit 5517492 into main Sep 26, 2026
7 checks passed
@ascorblack
ascorblack deleted the compaction-contract branch September 26, 2026 11:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant