Skip to content

fix: Jamba-Reasoning-3B chat output stays in reasoning_content #2089

Description

@inureyes

Problem

On GB10, the checkpoint at /home/inureyes/models/mlx/jamba-v0.1-4bit served through /v1/chat/completions logs WARN mlxcel::thinking: generation produced tokens but content is empty and all output routed to reasoning_content (src/server/routes/chat.rs:2333-2368), returns empty content, and reports cached=0 every turn (history-boundary snapshot stored at 3314 tokens, next turn diverges at 3277; unchanged after #2082).

That directory name is misleading. Its README says it is mlx-community/AI21-Jamba-Reasoning-3B-4bit (hidden 2560, 28 layers). The tokenizer declares <think>/</think> (ids 541/542), and chat_template.jinja:98-100 always appends <|im_start|>assistant\n<think>\n (no enable_thinking gate). Routing the output to the reasoning channel is therefore expected. The bug is that </think> is never seen or never reached.

Investigation (do first)

Read finish_reason and completion_tokens from the warn line on a default-launch server.

  • finish_reason=length: the model never closed </think> within max_tokens. Fix the budget handling so the client gets usable content, matching how other primed-thinking families behave.
  • finish_reason=stop: </think> was emitted and not recognized. Check how LlamaTokenizerFast decodes id 542, infer_thinking_markers (src/tokenizer/mod.rs:495), is_prompt_primed_open_thinking (chat.rs:2220), prompt_primed_open_thinking (src/reasoning_stream.rs:406), and extract_reasoning_content (chat.rs:2476).

Cache miss

Before blaming rendering, rule out an entry larger than the whole prompt-cache store. Then check rendering: chat_template.jinja:46-62 drops reasoning from earlier assistant turns and keeps <think> only when </think> is in content or reasoning_content is sent, so an empty re-sent turn renders shorter than what was generated. In scope only if it persists once content is non-empty. Otherwise file it separately as a general reasoning-model issue.

Acceptance Criteria

  • A unit test for the branch the investigation selects, failing on current code: length covers the budget behavior; stop covers close-marker recognition with this tokenizer's markers.
  • A real-server 3-turn chat on this checkpoint returns non-empty content on every turn, and turns 2 and 3 report cached_tokens > 0.

Implement in one PR with one verification pass. Run the server under gpu-lock run, building without the lock.

  • Investigation result (PR fix(server): forward echoed reasoning under reasoning_content too #2094): empty content was finish_reason=length only (pinned by a test, documented); the cache miss with non-empty content came from echoed reasoning being forwarded to templates only as reasoning, not reasoning_content. Fixed; real-server 3 turns echoing reasoning_content return non-empty content with cached_tokens 0/41/88. Content-only echo still misses by the template's own rule (documented).

Activity

  1. added
    type:bugBug fixes, error corrections, or issue resolutions
    and removed on Oct 1, 2026
  2. added a commit that references this issue on Oct 2, 2026
    390ad01
  3. added a commit that references this issue on Oct 5, 2026
    9a0be04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority:lowLow prioritystatus:doneCompletedtype:bugBug fixes, error corrections, or issue resolutions

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions