Skip to content

Ephemeral context injections re-anchor on every activation and defeat prompt caching (head-only hits on OpenAI lanes) #171

Description

@Anarchid

Summary

Ephemeral context injections (beforeUser / afterUser, from MCPL beforeInference hooks and module gatherContext) are re-anchored to the latest user-participant message on every compile and never stored. On a provider whose prompt cache only hits at boundaries a previous request wrote, and which writes only at the request tail, this makes every new activation miss everything except the static head (tools + instructions).

Measured on a production agent (gpt-6-astra via the Codex subscription lane, ~170k context, 30-minute heartbeats): inside a tool loop 99.7–99.9 % of input is cached; the first call of every activation gets exactly 40,448 cached tokens (tools + instructions to a 128-token block), whatever the gap since the previous call (observed at 40 s and 26 s too, so it is not cache retention). Each idle heartbeat costs ~130k uncached input for 50–200 output tokens; 42 such calls in 20 h accounted for 5.44M of 5.92M uncached tokens.

The injection in that case was the zulip MCPL server's "Recent messages from Zulip #…" recap (~2.7k tokens, beforeUser), but the mechanism is generic. The subagents module's [Fleet Status] injection (afterUser) will do the same thing the moment a subagent spans an activation boundary.

Mechanism

request N   (last call of activation N)
  … │ fco "Turn ended" │ RECAP_N │ [trigger N] │ fc │ fco │ … │ fc │ fco      ◄ N's cache writes all lie here
request N+1 (first call of activation N+1)
  … │ fco "Turn ended" │ [trigger N] │ fc │ fco │ … │ fco │ fc "Turn ended" │ fco │ RECAP_N+1 │ [trigger N+1]
                       ▲ first difference: RECAP_N is gone, so nothing N wrote is reachable
  1. ContextManager.compile() (context-manager src/context-manager.ts) splices beforeUser injections before the last message whose participant is user, afterUser after it, and persists nothing ("inserted before the last user message").
  2. AF collects the injections at stream start and in previewActivation (src/framework.ts), authorised per position by src/mcpl/hook-orchestrator.ts.
  3. At the next activation the injection vanishes from its old slot and reappears at the new tail, so the shared prefix with any earlier request ends where the previous injection used to be. Every cache boundary earlier requests wrote lies after that point.

Anchor bug (independent of caching): normalized tool results travel in user-participant messages, so a mid-activation recompile (stream retry after a socket drop) anchors the injection between a function_call and its function_call_output. Observed once in production; the backend accepted it, but it is a cache break and odd context for the model.

Provider scope

Lane Effect Why
OpenAI Responses, Codex subscription and token-rate API (openai-responses / openai-codex, same membrane adapter) Head-only on every activation Implicit mode writes only at the request tail; no prompt_cache_key / breakpoints are sent. Same wire shape on both; the API bills it in dollars instead of quota
Anthropic, autobiographical strategy Bounded loss (≈ two activations of tail) while the previous activation's tool loop is under ~9 rounds; falls to head-only beyond that CM's measured-stable-prefix marker sits on the previous trigger message, which precedes the injection; Anthropic's ~20-block backward lookup has to reach it. The floating tool-loop marker (membrane #50) does not help across activations, all its writes are after the injection
Anthropic, Passthrough Unaffected, already head-only No markers
Plain block-prefix caches (vLLM/SGLang-style, Gemini implicit, DeepSeek) Bounded: only tokens after the old injection slot Hit anywhere along the shared prefix, no prior write needed

Mitigations available today (no code)

  • Recipe: disabledFeatureSets: ["zulip.context"] on the zulip server entry, or MCPL_CONTEXT_HISTORY_SIZE=0.
  • Both stop the loss on the OpenAI lanes, at the price of the agent having no channel recap when woken by a mention.

Fix options (for discussion)

  • C. Freeze the anchor per activation, and fix the anchor rule. Anchor on the message that triggered the activation, not "last user-participant message"; reuse the position for retries. Small, fixes the retry split outright. Necessary, not sufficient.
  • A. Sticky injections. Once rendered, record the injection at its position (message store or a CM ledger) and re-render it there on later compiles; append a new one only when the namespace's content hash changes. Provider-agnostic, makes history append-only and reproducible; costs growth, needs a compression policy and a story for grant revocation (MCPL §10.8).
  • B. Delta injections. Server injects only what is new since a host-supplied cursor. Natural for chat recaps; needs an MCPL spec addition or per-agent server state.
  • E. Map CM cacheBreakpoint to prompt_cache_breakpoint in the OpenAI Responses adapter, so each request also writes at the last stable entry before the injection. Certain on the token-rate API; unknown whether the Codex backend accepts it.
  • F. Policy. Skip injections on timer-triggered activations (heartbeats never need the recap), and/or a conhost recipe-validation warning when an openai-* provider is combined with a server declaring beforeUser/afterUser injection.

Tentative lean: C regardless; A as the general mechanism with B for recaps; E as a provider-level extra where accepted; F.timer as a cheap first step.

Verification

  • Per activation: first call's cached count ≈ previous call's input minus its output and a ≤128-token tail.
  • Aggregate: head-only calls drop to ~0; the heartbeat agent's daily uncached input falls from ~6–7M to well under 1M.

Related

Activity

  1. nissa-seru commented on Oct 2, 2026

    @nissa-seru
    Contributor

    Confirmed that the placement defect remains in published context-manager 0.11.0 and upstream CM main 7959602c; AF main 03c31d9 forwards injections without an activation-anchor parameter.

    A local probe using a real ContextManager and temporary Chronicle store produces:

    First compile:
      injection(recap-1), user(trigger-1)
    
    Append assistant(tool_use) and user(tool_result), then recompile:
      user(trigger-1), assistant(tool_use), injection(recap-1), user(tool_result)
    
    Append the next activation's user message, then compile:
      user(trigger-1), assistant(tool_use), user(tool_result), assistant(turn-ended),
      injection(recap-2), user(trigger-2)
    

    The stored messages contain neither recap. This confirms both the historical-prefix change and the insertion between a tool call and its result. It is a compile-shape reproduction, not a live-provider cache measurement.

    The component boundary matters: AF owns activation/retry identity and hook permission lifetime; CM owns placement against rendered source entries. Current ContextInjection has no explicit anchor, and compilation discards source-message IDs when constructing normalized messages. An explicit anchor needs a defined outcome when that source is absent or folded, rather than silently moving to a later user-participant message.

    Nearby CM PR #134 repairs tool pairing inside strategy selection. Its author confirmed that injection placement is outside that change; injections are inserted afterward.

    An anchor repair would address retry/tool-pair placement but would not, by itself, preserve injected prefixes across activations. Sticky historical injections still need the separate retention, compression, and revocation decisions described in this issue. This is an assessment, not an implementation claim.

    — Jasper (OpenAI GPT-6 Astra)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions