Summary
Ephemeral context injections (beforeUser / afterUser, from MCPL beforeInference hooks and module gatherContext) are re-anchored to the latest user-participant message on every compile and never stored. On a provider whose prompt cache only hits at boundaries a previous request wrote, and which writes only at the request tail, this makes every new activation miss everything except the static head (tools + instructions).
Measured on a production agent (gpt-6-astra via the Codex subscription lane, ~170k context, 30-minute heartbeats): inside a tool loop 99.7–99.9 % of input is cached; the first call of every activation gets exactly 40,448 cached tokens (tools + instructions to a 128-token block), whatever the gap since the previous call (observed at 40 s and 26 s too, so it is not cache retention). Each idle heartbeat costs ~130k uncached input for 50–200 output tokens; 42 such calls in 20 h accounted for 5.44M of 5.92M uncached tokens.
The injection in that case was the zulip MCPL server's "Recent messages from Zulip #…" recap (~2.7k tokens, beforeUser), but the mechanism is generic. The subagents module's [Fleet Status] injection (afterUser) will do the same thing the moment a subagent spans an activation boundary.
Mechanism
request N (last call of activation N)
… │ fco "Turn ended" │ RECAP_N │ [trigger N] │ fc │ fco │ … │ fc │ fco ◄ N's cache writes all lie here
request N+1 (first call of activation N+1)
… │ fco "Turn ended" │ [trigger N] │ fc │ fco │ … │ fco │ fc "Turn ended" │ fco │ RECAP_N+1 │ [trigger N+1]
▲ first difference: RECAP_N is gone, so nothing N wrote is reachable
ContextManager.compile() (context-manager src/context-manager.ts) splices beforeUser injections before the last message whose participant is user, afterUser after it, and persists nothing ("inserted before the last user message").
- AF collects the injections at stream start and in
previewActivation (src/framework.ts), authorised per position by src/mcpl/hook-orchestrator.ts.
- At the next activation the injection vanishes from its old slot and reappears at the new tail, so the shared prefix with any earlier request ends where the previous injection used to be. Every cache boundary earlier requests wrote lies after that point.
Anchor bug (independent of caching): normalized tool results travel in user-participant messages, so a mid-activation recompile (stream retry after a socket drop) anchors the injection between a function_call and its function_call_output. Observed once in production; the backend accepted it, but it is a cache break and odd context for the model.
Provider scope
| Lane |
Effect |
Why |
OpenAI Responses, Codex subscription and token-rate API (openai-responses / openai-codex, same membrane adapter) |
Head-only on every activation |
Implicit mode writes only at the request tail; no prompt_cache_key / breakpoints are sent. Same wire shape on both; the API bills it in dollars instead of quota |
| Anthropic, autobiographical strategy |
Bounded loss (≈ two activations of tail) while the previous activation's tool loop is under ~9 rounds; falls to head-only beyond that |
CM's measured-stable-prefix marker sits on the previous trigger message, which precedes the injection; Anthropic's ~20-block backward lookup has to reach it. The floating tool-loop marker (membrane #50) does not help across activations, all its writes are after the injection |
| Anthropic, Passthrough |
Unaffected, already head-only |
No markers |
| Plain block-prefix caches (vLLM/SGLang-style, Gemini implicit, DeepSeek) |
Bounded: only tokens after the old injection slot |
Hit anywhere along the shared prefix, no prior write needed |
Mitigations available today (no code)
- Recipe:
disabledFeatureSets: ["zulip.context"] on the zulip server entry, or MCPL_CONTEXT_HISTORY_SIZE=0.
- Both stop the loss on the OpenAI lanes, at the price of the agent having no channel recap when woken by a mention.
Fix options (for discussion)
- C. Freeze the anchor per activation, and fix the anchor rule. Anchor on the message that triggered the activation, not "last
user-participant message"; reuse the position for retries. Small, fixes the retry split outright. Necessary, not sufficient.
- A. Sticky injections. Once rendered, record the injection at its position (message store or a CM ledger) and re-render it there on later compiles; append a new one only when the namespace's content hash changes. Provider-agnostic, makes history append-only and reproducible; costs growth, needs a compression policy and a story for grant revocation (MCPL §10.8).
- B. Delta injections. Server injects only what is new since a host-supplied cursor. Natural for chat recaps; needs an MCPL spec addition or per-agent server state.
- E. Map CM
cacheBreakpoint to prompt_cache_breakpoint in the OpenAI Responses adapter, so each request also writes at the last stable entry before the injection. Certain on the token-rate API; unknown whether the Codex backend accepts it.
- F. Policy. Skip injections on timer-triggered activations (heartbeats never need the recap), and/or a conhost recipe-validation warning when an
openai-* provider is combined with a server declaring beforeUser/afterUser injection.
Tentative lean: C regardless; A as the general mechanism with B for recaps; E as a provider-level extra where accepted; F.timer as a cheap first step.
Verification
- Per activation: first call's cached count ≈ previous call's input minus its output and a ≤128-token tail.
- Aggregate: head-only calls drop to ~0; the heartbeat agent's daily uncached input falls from ~6–7M to well under 1M.
Related
Summary
Ephemeral context injections (
beforeUser/afterUser, from MCPLbeforeInferencehooks and modulegatherContext) are re-anchored to the latest user-participant message on every compile and never stored. On a provider whose prompt cache only hits at boundaries a previous request wrote, and which writes only at the request tail, this makes every new activation miss everything except the static head (tools + instructions).Measured on a production agent (
gpt-6-astravia the Codex subscription lane, ~170k context, 30-minute heartbeats): inside a tool loop 99.7–99.9 % of input is cached; the first call of every activation gets exactly 40,448 cached tokens (tools + instructions to a 128-token block), whatever the gap since the previous call (observed at 40 s and 26 s too, so it is not cache retention). Each idle heartbeat costs ~130k uncached input for 50–200 output tokens; 42 such calls in 20 h accounted for 5.44M of 5.92M uncached tokens.The injection in that case was the zulip MCPL server's "Recent messages from Zulip #…" recap (~2.7k tokens,
beforeUser), but the mechanism is generic. The subagents module's[Fleet Status]injection (afterUser) will do the same thing the moment a subagent spans an activation boundary.Mechanism
ContextManager.compile()(context-managersrc/context-manager.ts) splicesbeforeUserinjections before the last message whose participant isuser,afterUserafter it, and persists nothing ("inserted before the last user message").previewActivation(src/framework.ts), authorised per position bysrc/mcpl/hook-orchestrator.ts.Anchor bug (independent of caching): normalized tool results travel in
user-participant messages, so a mid-activation recompile (stream retry after a socket drop) anchors the injection between afunction_calland itsfunction_call_output. Observed once in production; the backend accepted it, but it is a cache break and odd context for the model.Provider scope
openai-responses/openai-codex, same membrane adapter)prompt_cache_key/ breakpoints are sent. Same wire shape on both; the API bills it in dollars instead of quotaMitigations available today (no code)
disabledFeatureSets: ["zulip.context"]on the zulip server entry, orMCPL_CONTEXT_HISTORY_SIZE=0.Fix options (for discussion)
user-participant message"; reuse the position for retries. Small, fixes the retry split outright. Necessary, not sufficient.cacheBreakpointtoprompt_cache_breakpointin the OpenAI Responses adapter, so each request also writes at the last stable entry before the injection. Certain on the token-rate API; unknown whether the Codex backend accepts it.openai-*provider is combined with a server declaringbeforeUser/afterUserinjection.Tentative lean: C regardless; A as the general mechanism with B for recaps; E as a provider-level extra where accepted; F.timer as a cheap first step.
Verification
Related
session_idheader; fixed in-turn caching, which exposed this as the remaining loss)