Repository navigation
Compaction that always makes progress, and runs that end on their own terms - #2
Merged
Merged
Conversation
added 10 commits
September 26, 2026 15:41
The notice telling the model its tools are withdrawn stayed in history after the run ended. When the wind-down's final turn itself failed, the next turn read it as its instruction and gave up without calling a tool. It is removed when a drive reaches a terminal state and put back only for a resumed run that is still wound down.
…ey are A run stopped by its wall clock was told it had reached its budget, and a run whose model kept returning reasoning with neither an answer nor a tool call was wound down under the provider-error cause, whose notice says the model endpoint failed. The model reads the notice literally and passed that on to the operator as the reason the run stopped, although the endpoint had answered every request. The deadline now has its own text naming the time limit, and the empty-round policy enters a new model_no_progress cause with a text that says the model's output failed and the requests did consume tokens and time. Each text is a constant, and a blank one falls back to the general notice.
…d down Whether to try the same endpoint again read the reason of the classification an adapter attached to the error and ignored its retryable flag, although the classification contract carries both and the adapter set the flag as its answer to exactly that question. A reply the adapter read as final under a coarse reason such as server_error was retried through the backoff. The flag now decides when the classification sets one; without it the reason decides as before. A permanent provider error also no longer enters the wind-down. Its one turn is a request to the endpoint that has just refused the run, with the same history and one more message, and it is refused the same way: the run failed a turn later, and its cause read as a wind-down rather than as the provider's refusal. The fallback chain is still tried first, since another model may serve the request.
Completing a failed run on its "preserved answer" required substantive prose after the run's latest work, read from every message not tagged as seeded. The seed tag is set by the executor and not by every host, so on a host that hands the engine a session's earlier turns untagged, the previous run's reply counted: a run whose every request was refused completed with nothing written, and the operator was shown no reply and no error. The answer must now have been written in the round that failed, the span after the last message a caller put in. The helper that finds that span is shared with the check that decides whether a run produced anything a wind-down could report on.
Tier 2 skipped every unit below compaction_summary_min_unit_tokens, and the fold takes only summaries and operator turns. A run that works in many rounds of one short tool call and a short result each therefore built a history in which no tier had anything to do, however large it grew. Over a live history of that shape, 160 units of 300 to 1,400 tokens under a 1,500-token floor, not one unit was eligible: every proactive pass freed nothing and the third one failed the run on the retry budget. Adjacent small units are now joined and summarised as one unit. Units join only when nothing lies between them, when each is contiguous in itself and when they share a seed provenance, so tool pairing stays whole and the replacement carries exactly one provenance. A group stops before it would pass compaction_summary_group_max_tokens (6,000 by default; 0 turns grouping off) and is sent only when it clears the floor a single unit must clear. It is keyed by its first anchor, so the failure census and the dedup set treat it as that unit.
The compaction trigger is a whole-prompt size: the lower of the configured ratio of the window and the largest prompt the provider accepts, less a turn's headroom. The emergency cliff is one too. The gate held the history's estimate alone against both. On a deployment whose system prompt and tool definitions take a large share of the window — 153 tools and a long system prompt, 56k of a 256k window — the history reached the trigger only once the request as a whole was well past the provider's ceiling. Proactive compaction then never ran first: the request fit clipped the output cap turn after turn, down to a few hundred tokens, and compaction ran only after a refusal. The engine now records, at each dispatch, the heuristic size of what the request carried besides the history (its system messages and tool definitions), and the gate adds it at the current calibration. Before a run's first request the figure is zero and the gate reads the history alone, as it did. It rides the snapshot, so a resumed run's first gate is sized the same way.
The refusal returned from inside COMPACTING, so the run stayed in that state for the rest of the turn. It now returns to RUNNING and says why.
…ser does A long-running loop gated on its whole prompt, could shrink only the history, found no unit worth a summariser call and failed after 75 passes that freed a few tokens each. The summariser itself was a single JSON string under a 1,024-character ceiling: it lost verbatim identifiers, and a reply the model did not close, or wrote as prose, was thrown away as "not JSON". A pass is now a cascade, each tier running only while the whole prompt is above its target (compaction_target_ratio of the trigger): - masking: old tool outputs become a placeholder that names the tool, keeps the lines the output said only once and points at the stored original; - summarising: adjacent units are joined into spans and each span is written up as plain text under five fixed headings, requested with complete_text under a system-role instruction, read tolerantly and clamped section by section at line boundaries; runs of old summaries are then folded; - the floor: when those stop short of the trigger, the oldest spans are removed without a model, raw before summarised, leaving one digest. Before any tier removes something it records operator words, files touched, exact values of recognisable shape, failed calls and the latest plan in a ledger message that code rebuilds on every pass and the summariser never sees. The summariser's input is bounded, each call has a deadline, and every failure kind is a normal outcome. compaction_completed reports each tier and the pass's outcome.
What a pass guarantees, in clauses a test can check: progress to the trigger or the floor, the tier order, a failed summary as a normal outcome, facts from state, the carrier, instruction isolation, re-injection, the gate, reversibility, what is never removed, whole tool pairs and replayable requests; with the budget arithmetic, the failure states and what a completed pass reports.
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
Two groups of fixes, ten commits, each self-contained: every commit passes the
full suite,
ruff,mypy --strictand the leak scan on its own, so the branchcan be read — and bisected — one commit at a time.
How a run ends (four commits)
from history once the drive reaches a terminal state, and written back only
for a resumed run that is still wound down.
(
soft_stop_notice_text_deadline), and a run whose model keeps returningreasoning with no answer and no tool call enters a new
model_no_progresscause (
soft_stop_notice_text_model_no_progress) instead of theprovider-error one.
retryableflag of the classification anadapter attaches (
ClassifiedLikecarries it); a permanentLLMProviderErrorno longer enters the wind-down. The fallback chain isstill tried first.
written in the round that failed, not anywhere in untagged history.
Compaction (six commits)
(
compaction_summary_group_max_tokens, 6,000; 0 disables).besides the history (system messages and tool definitions, calibrated). The
figure rides the snapshot.
pre_compacthook refuses returns the run toRUNNING(
compaction_refused_by_hook) instead of leaving it inCOMPACTING.compaction_target_ratio(0.6) of the trigger, or at a floor, whatever thesummariser does. Written up in
docs/compaction.md, mirrored indocs/ru/compaction.md.Why
How a run ends
(the outage that caused the wind-down also killed the answer), the notice
was the last thing in the session. The next run obeyed it: it reported what
it had not finished and called nothing, although every tool was back.
budget", and a model stuck in reasoning read as "the model endpoint failed".
The model repeats the notice to the user, so the user was told that an
endpoint that had answered every request had failed.
classified as final under a coarse reason such as
server_errorwentthrough the backoff. The wind-down then sent one more request, with the same
history, to the endpoint that had just refused it, and was refused the same
way. The run failed a turn later, and its cause read as a wind-down.
by the executor, not by every host. A host that hands the engine a session's
earlier turns untagged therefore had a run whose every request was refused
complete "successfully", with nothing written and no error shown.
Compaction
The failure that led here was a long-running autonomous loop. A pass failed
the run on the retry budget every time the loop woke. Three things combined:
and the emergency cliff are whole-prompt sizes. With a large tool surface
and system prompt (153 tools, 56k of a 256k window), the history reached
the trigger only once the request was past the provider's ceiling. The
request fit then clipped the output cap turn after turn, down to a few
hundred tokens in a replay, and compaction first ran on a refusal.
unit below
compaction_summary_min_unit_tokens, and the fold takes onlysummaries and operator turns. The loop's history was 160 units of 300 to
1,400 tokens under a 1,500-token floor, and not one unit was eligible. Each
pass freed zero to a hundred tokens.
string capped at 1,024 characters. Verbatim identifiers did not survive it,
and a reply the model left unterminated, or wrote as prose, was discarded
as "not JSON".
The compaction contract
docs/compaction.mdstates twelve clauses, each checked by a test intests/unit/runtime/test_compaction_contract.py. In short:trigger, or at the floor. A pass that could change nothing is not opened.
The retry budget is spent only when the history is already at the floor
and a tier raised.
original and the lines the output said only once;
summarised, leaving one digest written by code.
replies, replies too large and replies not smaller than their input all
count as normal. The summariser's input is bounded and each call has a
deadline.
by code on every pass and is never shown to the summariser. It holds the
operator's words verbatim, files by tool role, values of recognisable
shape with their line, failed calls and the latest plan.
complete_text. The reader accepts JSON, unterminated JSON, a reply withno headings and a truncated reply. Sections are clamped at line
boundaries.
the material inside
<transcript>tags. Each carrier opens by saying theruntime wrote it.
blocks, the keep tail and the in-flight batch. Tool pairs stay whole.
Evidence (aggregate)
Replay of the failing history. A real
QueryEngineran on the recordedshape: 357 messages, the real system prompt, 153 tool definitions, a 256k
window. The agent was a double that makes one small tool call per turn, for
150 turns. The provider was a double that counts the prompt the way the live
provider did and refuses any request that does not fit. Every summariser call
went to a real model endpoint, four endpoints in all.
Summariser faults injected. Two unterminated JSON replies, two empty
replies, a 300-second hang and a 502. Without the contract, the pass waited
out the hang (303 s). With it, the call deadline cut the hang, both
unterminated replies were recovered, and the first pass ended below the
trigger without reaching the floor.
Planted facts. A synthetic agent session of about 150k tokens in six
parts had 14 scored values planted across it: ports, a path, a UUID, UTC
times, a URL, a key id and a non-Latin name. A forced pass ran after each
part, with 40k and 24k windows, on four endpoints, three runs per cell. The
same endpoint then answered from the compacted context.
The value missed everywhere, in both arms, is a count with a unit ("19
shards"). The model answers "19": the carrier holds the value, and the loss
happens in the answer.
What changes for a consumer
This is a pre-release, and these changes are intentional. They do not follow
the "new behaviour defaults off" convention, because the old defaults are
what failed.
compaction_mask_min_tokens, older than the most recentcompaction_mask_keep_recent_results(8), become placeholders. Theplaceholder points at the stored original.
compaction_ledger_max_tokensof the window onceit fills.
compaction_summary_ratioof a span instead of 1,024 characters.are removed; the changelog lists them. Setting one now fails validation,
as for any unknown field.
retry budget. Exhaustion now means the floor was unavailable.
retryable=Falseon aclassification now stops the retry, whatever the reason says.
complete_text, which is already onILLMProvider, instead ofcomplete_structured. An adapter whosecomplete_textis a stub must implement it.To get closer to the previous behaviour without code changes, raise
compaction_mask_keep_recent_results, movecompaction_target_ratiotowards1, lower
compaction_ledger_max_tokens, or setcompaction_summary_group_max_tokens=0.Checklist
uv run pytest .passes: 3657 passed, 139 skipped, coverage 92.88%.Every commit also passes on its own.
uv run ruff check .is cleanuv run mypy --strictis clean (341 source files, no path argument)uv run bandit -r protocore -q -c pyproject.tomlis cleantest_soft_stop.py,test_provider_refusals.py,test_history_run_boundary.py,test_compaction_small_units.py,test_compaction_gate_overhead.py,test_compaction_retry_budget.pyandtest_compaction_contract.py(24 contract tests, with a planted-factsfixture and a shape-only long-loop fixture)
RuntimeConstantsfields with a description. Thecompaction defaults change on purpose; see above.
docs/compaction.mdand its Russian mirror,docs/architecture.mdin both languages, and the changelog