Problem
After #2094, Jamba-Reasoning-3B (/home/inureyes/models/mlx/jamba-v0.1-4bit) hits the prompt cache on turns 2 and 3 (cached 41/88) only when the client echoes reasoning_content. A client that sends back only content, which is the OpenAI SDK default, gets cached=0 on every turn.
The cause is in that checkpoint's chat_template.jinja. An earlier user turn gets thinking_prefix only when the next assistant message has reasoning_content that is defined and not none (lines 34-36, inside the user-turn block at 29-45). History assistant turns render content only (lines 46-62). So the turn-2 prompt diverges inside turn 1, and the history-boundary snapshot (#1143) never matches. The comment at src/server/chat_request.rs:50-56 calls this miss by design.
Proposed fix
- Add a bounded LRU on the server, keyed by
(template_sig, hash of returned assistant content) and narrowed by resolve_session_key (prompt_cache/key.rs:447) when the client set one. Most SDK requests set no session, so the key cannot depend on the session alone. Fill it whenever a response returns non-empty reasoning while the prompt cache is on.
- In
build_raw_json_messages_with_thinking (chat_request.rs:1819), look up an assistant message that has no reasoning and treat the stored text as if the client had echoed it. Inject the exact text, never a placeholder, because other templates render it. On a lookup miss, change nothing.
- Echoed reasoning keeps precedence, so the forwarding at :1894-1901 is unchanged.
- Fallback, only if this proves unsound: snapshot before the last user turn.
Acceptance Criteria
Do it in one PR, run the checks once at the end, and push as few times as possible.
Verification
cargo test --release --features cuda --lib server::chat_request -- --test-threads=1
gpu-lock run --tag reasoning-cache -- <3-turn script against mlxcel-server --prompt-cache>
Problem
After #2094, Jamba-Reasoning-3B (
/home/inureyes/models/mlx/jamba-v0.1-4bit) hits the prompt cache on turns 2 and 3 (cached 41/88) only when the client echoesreasoning_content. A client that sends back onlycontent, which is the OpenAI SDK default, gets cached=0 on every turn.The cause is in that checkpoint's
chat_template.jinja. An earlier user turn getsthinking_prefixonly when the next assistant message hasreasoning_contentthat is defined and not none (lines 34-36, inside the user-turn block at 29-45). History assistant turns render content only (lines 46-62). So the turn-2 prompt diverges inside turn 1, and the history-boundary snapshot (#1143) never matches. The comment atsrc/server/chat_request.rs:50-56calls this miss by design.Proposed fix
(template_sig, hash of returned assistant content)and narrowed byresolve_session_key(prompt_cache/key.rs:447) when the client set one. Most SDK requests set no session, so the key cannot depend on the session alone. Fill it whenever a response returns non-empty reasoning while the prompt cache is on.build_raw_json_messages_with_thinking(chat_request.rs:1819), look up an assistant message that has noreasoningand treat the stored text as if the client had echoed it. Inject the exact text, never a placeholder, because other templates render it. On a lookup miss, change nothing.Acceptance Criteria
cached_tokens > 0on turns 2 and 3 (0 on 33c4505).qwen3-0.6b-4bitis unaffected or improved; the :50-56 comment is updated.Do it in one PR, run the checks once at the end, and push as few times as possible.
Verification