Skip to content

fix(compaction): write the handoff note in our own words and keep it on overflow - #6622

Open
Hmbown wants to merge 2 commits into
mainfrom
feat/original-compaction-handoff
Open

Hmbown wants to merge 2 commits into
mainfrom
feat/original-compaction-handoff

Conversation

@Hmbown

@Hmbown Hmbown commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

When Codewhale makes room, it leaves a note for the next turn. That note's prompt and header were taken from Codex's templates, and the header told the model that "another language model" had done the earlier work. This PR replaces both with original text, gives the note a fixed layout, and fixes two ways the note or the user's own words could be lost.

What changed

The handoff text (crates/tui/src/compaction.rs)

  • The prompt asks for fixed headings in this order: Objective, User direction, Permissions and limits, Done, Changed files, Still running, Verification left, Open questions, Next action, Reference. An empty heading says None.
  • It asks the model to quote user corrections exactly, record only permissions the user actually stated, keep verified work apart from assumed work, and quote the last failure verbatim.
  • A fold-in rule says what to do with an earlier note: carry forward what is still true, update what changed, and drop what is finished.
  • HANDOFF_SECTIONS is one list shared by the first request and the quality retry.
  • "Still running" covers only shell commands, servers and ports. Agents are left out, because the agent topology message remains the only source of agent state.
  • The header starts with "Codewhale handoff note." It says what is actually kept above the note: the recent user messages and the last steps of the current round. Long tool output there is shortened and marked, the oldest kept message may be shortened, and earlier steps of the round exist only in the note, so the model is told to rerun a command or reread a file when the full output matters. A test (summary_header_matches_what_replacement_history_keeps) ties this wording to what replacement_messages and bound_last_round keep. The closing says the note grants no new permissions.
  • The "ported from Codex" comments are removed. Comments about matching Codex behaviour stay, because they describe how the code works rather than copied text.

Detection

  • A new checkpoint is recognised only by its structure: the header prefix followed by the provenance block.
  • Substring matching stays for the two legacy markers only.
  • A user message that quotes "Codewhale handoff note" is now kept, stays where it was on restore, and keeps its place on the wire. Before, any message containing the marker sentence was dropped from the replacement history.
  • Checkpoints saved with the old header still pass the wire check, keep their position on restore, and are replaced, not stacked, by the next pass.

Overflow retry

  • drop_oldest_history_messages never drops the newest checkpoint or the trailing instruction.
  • When nothing else can be dropped, the pass returns the provider's context-window error. It does not write a new note without the previous one.

Where this differs from the approved design

  1. Legacy wire check uses the marker, not the full old header. The design kept LEGACY_V2_SUMMARY_HEADER, the full copied paragraph, for the wire check. This PR checks starts_with(LEGACY_V2_COMPACTION_SUMMARY_MARKER) plus the provenance block. The provenance block is what makes the check safe, so the copied paragraph no longer needs to be in the source. Only the marker sentence stays, as a detection constant next to LEGACY_COMPACTION_SUMMARY_MARKER.
  2. The carrier fallback is legacy-only. The design said substring matching was safe in strip_summary_text and extract_compaction_summary, because they read only engine-authored carriers. They do not: a runtime-thread carrier is the host's base system prompt with the summary section merged in. A project instruction that mentions "Codewhale handoff note" would have cut the system prompt short at that point. New carriers are always wrapped in the begin/end delimiters, so the bare-marker fallback now matches only the legacy markers. A test covers host text that quotes the new phrase.

Verification

  • node crates/tui/src/compaction/validate_survival_contract.mjs: 12 fixtures passed on origin/main before the change, 14 after. The two new fixtures are a legacy-v2 checkpoint replaced by the new note, and a user quote of the new marker that must be kept. This check is structural only, not a quality eval.

  • CARGO_BUILD_JOBS=3 scripts/dev-cargo.sh test -p codewhale-tui --lib -- <filter>:

    Filter Passed
    compaction 159
    chat:: 147
    checkpoint 64
    restore 161
    last_round 18
    runtime_handoff 19
    survival_contract 5
    worker_compacts_past_its_context_window 1
  • New tests:

    • a user quote of the new marker is not a checkpoint (retain, restore, wire);
    • a legacy-v2 checkpoint with provenance is recognised and replaced;
    • legacy single-block markers still count as checkpoints;
    • on overflow retry, the previous note and the instruction are always kept, results orphaned next to the kept note are removed, and the retry stops cleanly when nothing else can go;
    • both prompts carry every heading in order, the fold-in rule and the focus line;
    • the note opens with the marker and ends with the closing;
    • the header describes what the replacement history keeps: a long round loses its earlier steps and long tool output is shortened with a marker.
  • cargo fmt --all -- --check: ok.

  • clippy on codewhale-tui with the CI flags (--all-targets --all-features --locked -D warnings): clean.

  • scripts/check-lexicon.py: 363 findings before and 363 after, none in compaction.rs. The two stale allowlist entries are removed.

  • check-blocking-calls-budget.py, check-dead-code-budget.py, check-command-crate-boundaries.py and split/module_graph.py --check all exit 0.

  • CHANGELOG gates:

    • sync-changelog.sh and the web derive scripts ran;
    • check-versions.sh --range-audit-advisory reports OK;
    • check-contributor-credit.py v0.10.0 reports every contributor credited;
    • vitest lib/public-copy.test.ts: 6 passed.
  • No paid provider calls were made.

Not done / risks

  • No offline quality eval exists, so this PR makes no claim that summaries are better. The headings are requested but not enforced: the validator still only rejects corrupt output.
  • Background shells are still described by the model in the note. A typed shell-job message is a separate follow-up.
  • A very large session whose summary request can fit nothing but the previous note and the instruction now fails that pass. Before, it summarized without the note. This change is intended.

No-Issue: original compaction handoff text

🤖 Generated with Claude Code

@Hmbown Hmbown added this to the v0.10.1 milestone Sep 26, 2026
Copilot AI lite review requested due to automatic review settings September 26, 2026 09:12

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

CodeWhale Bot and others added 2 commits September 26, 2026 19:44
…on overflow

The note Codewhale leaves when it makes room used a prompt and header
taken from Codex's templates, and the header said "another language
model" had done the earlier work. The note had no fixed layout and no
rule to carry an earlier note forward.

- New prompt with fixed headings (Objective, User direction, Permissions
  and limits, Done, Changed files, Still running, Verification left, Open
  questions, Next action, Reference), a fold-in rule for an earlier note,
  and one HANDOFF_SECTIONS list shared with the quality retry. "Still
  running" leaves agents out; the agent topology message stays the only
  source of agent state.
- New header ("Codewhale handoff note. ...") that says accurately what is
  kept above (recent user messages and the latest round, the oldest
  possibly shortened), and a closing that grants nothing new.
- Detection: a new checkpoint is recognised only structurally (header
  prefix plus the provenance block). Substring matching is kept for the
  two legacy markers only, so a user message quoting "Codewhale handoff
  note" is no longer dropped. Old checkpoints with provenance still pass
  the wire check and are replaced, not stacked. The bare-marker fallback
  for system-prompt carriers is legacy-only too, so host text quoting the
  new phrase is never truncated.
- Overflow retry: drop_oldest_history_messages never drops the newest
  checkpoint or the instruction; when nothing else can go, the pass fails
  with the provider's context error instead of summarizing without the
  previous note.
- Removed the "ported from Codex" comments and the two stale lexicon
  allowlist entries.

Verification:
- node crates/tui/src/compaction/validate_survival_contract.mjs:
  ok 12 before, ok 14 after (structural fixtures, not a quality eval)
- scripts/dev-cargo.sh test -p codewhale-tui --lib -- <filter>:
  compaction 158 passed; chat:: 147 passed; checkpoint 64 passed;
  restore 161 passed; last_round 17 passed; runtime_handoff 19 passed;
  survival_contract 5 passed; worker_compacts_past_its_context_window 1
- cargo fmt --all -- --check: ok
- clippy -p codewhale-tui --all-targets --all-features with CI flags: clean
- check-lexicon.py: 363 findings before and after, none in compaction.rs
- blocking-calls, dead-code, command-crate-boundaries, module_graph: 0
- changelog gates: sync, web derive, check-versions, contributor credit,
  vitest public-copy 6 passed

No-Issue: original compaction handoff text

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The header claimed the latest round was kept above the note "as they
were". It is not: replacement_messages keeps only the last two tool
exchanges of a long round (plus the user's text), and bound_last_round
shortens last-round tool results over 8 KB with a truncation marker.
A model trusting the old header would assume it still had the whole
round word for word.

The header now says the last steps of the current round are kept, that
long tool output is shortened and marked, that earlier steps of the
round exist only in the note, and to rerun a command or reread a file
when the full output matters.

A new test, summary_header_matches_what_replacement_history_keeps,
builds a three-step round with an oversized last result, checks the
first step is dropped and the last result carries the marker, and checks
the header describes both, so the two cannot drift apart again.

Tests (CARGO_BUILD_JOBS=3 scripts/dev-cargo.sh test -p codewhale-tui --lib):
- compaction: 159 passed, 0 failed
- last_round checkpoint: 82 passed, 0 failed
node validate_survival_contract.mjs: 14 fixtures ok
cargo fmt --check ok; clippy -p codewhale-tui (CI flags) clean;
blocking-calls, dead-code, command-crate-boundaries, module_graph: exit 0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants