Skip to content

fix: reasoning-model prompt cache misses when clients do not echo reasoning #2110

Description

@inureyes

Problem

After #2094, Jamba-Reasoning-3B (/home/inureyes/models/mlx/jamba-v0.1-4bit) hits the prompt cache on turns 2 and 3 (cached 41/88) only when the client echoes reasoning_content. A client that sends back only content, which is the OpenAI SDK default, gets cached=0 on every turn.

The cause is in that checkpoint's chat_template.jinja. An earlier user turn gets thinking_prefix only when the next assistant message has reasoning_content that is defined and not none (lines 34-36, inside the user-turn block at 29-45). History assistant turns render content only (lines 46-62). So the turn-2 prompt diverges inside turn 1, and the history-boundary snapshot (#1143) never matches. The comment at src/server/chat_request.rs:50-56 calls this miss by design.

Proposed fix

  • Add a bounded LRU on the server, keyed by (template_sig, hash of returned assistant content) and narrowed by resolve_session_key (prompt_cache/key.rs:447) when the client set one. Most SDK requests set no session, so the key cannot depend on the session alone. Fill it whenever a response returns non-empty reasoning while the prompt cache is on.
  • In build_raw_json_messages_with_thinking (chat_request.rs:1819), look up an assistant message that has no reasoning and treat the stored text as if the client had echoed it. Inject the exact text, never a placeholder, because other templates render it. On a lookup miss, change nothing.
  • Echoed reasoning keeps precedence, so the forwarding at :1894-1901 is unchanged.
  • Fallback, only if this proves unsound: snapshot before the last user turn.

Acceptance Criteria

  • Unit tests: a stored content-only turn renders identically to the echoed form; an unstored turn is unchanged; echoed reasoning wins.
  • On a real server, a 3-turn chat on jamba-v0.1-4bit with content-only history reports cached_tokens > 0 on turns 2 and 3 (0 on 33c4505).
  • qwen3-0.6b-4bit is unaffected or improved; the :50-56 comment is updated.

Do it in one PR, run the checks once at the end, and push as few times as possible.

Verification

cargo test --release --features cuda --lib server::chat_request -- --test-threads=1
gpu-lock run --tag reasoning-cache -- <3-turn script against mlxcel-server --prompt-cache>

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)priority:mediumMedium prioritystatus:doneCompletedtype:bugBug fixes, error corrections, or issue resolutions

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions