Skip to content

api: add Qwen3.6 tool calling support - #1723

Closed
GenericRikka wants to merge 1 commit into
JustVugg:devfrom
GenericRikka:rikka
Closed

GenericRikka wants to merge 1 commit into
JustVugg:devfrom
GenericRikka:rikka

Conversation

@GenericRikka

Copy link
Copy Markdown
Contributor

Summary

Add native tool-calling support for Qwen3.6 through the OpenAI-compatible API.

The implementation follows Qwen3.6's native chat template for:

  • tool declarations in the system turn
  • assistant <tool_call> blocks
  • consecutive <tool_response> blocks
  • preservation of reasoning for the current query
  • removal of stale reasoning from earlier turns
  • conversion of native tool calls to OpenAI tool_calls

The Qwen tool-call renderer/parser helpers are shared with Qwen3.8 where
the wire format is identical.

Qwen3.6 tokenizer fix

End-to-end testing exposed a separate tokenizer issue affecting Qwen3.6
tool calling.

Qwen3.6's tokenizer.json stores protocol tokens such as:

  • <tool_call> / </tool_call>
  • <tool_response> / </tool_response>
  • <think> / </think>

in added_tokens, outside model.vocab.

The Qwen3.6 decoder previously built its ID-to-piece table only from
model.vocab. As a result, generated added-token IDs were accepted by
the tokenizer but produced no bytes during detokenization. In practice,
this caused native protocol delimiters such as </think> and
<tool_call> to disappear from generated output.

The second commit includes added_tokens in the decoder table and adds
a real-adapter regression test covering the six Qwen3.6 protocol tokens.

Testing

  • Python test suite: 1070 tests passed, 103 skipped
  • git diff --check: clean
  • Qwen3.6 real-adapter tokenizer regression passes against the full
    Qwen3.6-35B-A3B tokenizer
  • End-to-end OpenAI API test produces a native Qwen3.6 tool call which
    is returned as tool_calls with finish_reason: "tool_calls"

The existing Qwen3.6 real-adapter oracle test reports:

qwen36: first generated token differs from the independent oracle

The same failure reproduces unchanged on the upstream base (9d5d05d)
with the same model/reference fixture, so it is unrelated to this
change.

@JustVugg

Copy link
Copy Markdown
Owner

Thank you, and welcome. Two things happened here, so the red checks make sense.

The red checks are not your code. The PR was opened against main; everything here targets dev, so I retargeted it. That cancelled the running workflow (the jobs show as failed with no log), and on dev the branch now conflicts, so GitHub cannot start a new run until it is rebased.

The decoder fix is right and is a real bug on dev: </think> decoded to nothing, so with thinking on the whole answer came back as reasoning_content. I split your commit out as #1724 with your authorship, plus one change: only the non-special added tokens are decoded. Decoding all of them also turns <|im_start|> and <|endoftext|> into text (I saw them in the output of the real 35B), which dev never emitted. That goes into 1.12.1.

The tool-calling half needs a rebase that is not mechanical: dev's render_chat_qwen36 gained two things after the base you started from, the continuation of a trailing assistant turn (#1402, add_generation_prompt=False) and the opt-in prompt-injected tool fallback (#1497, COLI_TOOL_FALLBACK=1). Your renderer replaces the function and would drop both. Please rebase on dev, keep the continuation branch, and make native tool calling the path when tools is present (the fallback can then be removed for qwen36, with its test updated, since native support supersedes it). tests/test_qwen36_chat_template.py against the checkpoint's chat_template.jinja is the check I will look at, with a tools case and a tool-response case. Once that is green it goes in the next release.

@GenericRikka

Copy link
Copy Markdown
Contributor Author

Thank you for looking into this so quickly, and also for already splitting out and fixing the decoder issue in #1724.

And sorry about originally basing the PR on main. That was mostly reflex on my side since a lot of repositories I contribute to use main/master as the development branch, but I should have checked CONTRIBUTING.md before opening the PR.

I have now rebased the tool-calling commit onto dev and resolved the Qwen3.6 renderer changes.

tests/test_qwen36_chat_template.py was also updated to cover the requested tool-call and tool-response cases against Qwen3.6's checkpoint chat_template.jinja, for both thinking modes. The continuation check is still covered as well.

The reference-template test is now fully green:

ok   un turno utente [thinking=True]
ok   sistema piu' utente [thinking=True]
ok   assistant in cronologia [thinking=True]
ok   dichiarazione e chiamata tool [thinking=True]
ok   risposta tool [thinking=True]
ok   un turno utente [thinking=False]
ok   sistema piu' utente [thinking=False]
ok   assistant in cronologia [thinking=False]
ok   dichiarazione e chiamata tool [thinking=False]
ok   risposta tool [thinking=False]
ok   prosecuzione: turno aperto = template(add_generation_prompt=False) senza il <|im_end|> finale

template Qwen3.6: il gateway e' identico al riferimento

Thanks again for the prompt and detailed review.

@JustVugg

Copy link
Copy Markdown
Owner

Thanks for the quick rebase, and the template test with the tool cases in both thinking modes is exactly the check I asked for. CI is red for a real reason though, in the Python suite:

  • tests/test_openai_server.py no longer imports: cannot import name 'parse_qwen38_tool_calls' from 'openai_server'. The rename to parse_qwen_tool_calls needs either the callers in the tests updated or the old name kept as an alias for Qwen3.8.
  • tests/test_openai_tools_fallback_e2e.py fails: Qwen36ToolFallbackE2E.test_tool_call_parsed_back ('stop' != 'tool_calls', then KeyError: 'tool_calls') and Qwen36ToolRefusedByDefault.test_tool_role_refused. Those pin feat(api): opt-in prompt-injected tool translation for OLMoE and Qwen3.6 #1497's behaviour on qwen36. With native tool calling the refusal-by-default test is obsolete, so update it to the new contract, and make the fallback test either run on the native path or be removed for qwen36 with a sentence in the PR saying why.

1.12.1 is being tagged now, so this lands in the next release; once the suite is green I merge it.

Implements the checkpoint's native tool declaration/call/response
protocol and translates it to OpenAI tool_calls.
@JustVugg

Copy link
Copy Markdown
Owner

Thanks for this. Your branch started before #1767 and replaced the whole of render_chat_qwen, so resolving the conflict your way would have dropped #1767's preserve_thinking / last_query logic, and every qwen36 chat request would then fail with a TypeError. We carried your commit onto the current dev in #1794, keeping you as the author. It keeps #1767 and CHAT_FLAVOR, matches the official Qwen3.6 template byte for byte in 40 cases, and passes an end-to-end tool loop. The details are in #1794. Once that is merged this one will be closed as superseded. Please have a look if you have time.

JustVugg added a commit that referenced this pull request Sep 28, 2026
api: Qwen3.6 tool calling (#1723 on current dev, keeps #1767)
@JustVugg

Copy link
Copy Markdown
Owner

Merged as #1794, with your commit and authorship kept, carried onto the current dev together with #1767 and CHAT_FLAVOR. Thanks for the tool calling work; Qwen3.6 now has native tool calls on dev.

@JustVugg JustVugg closed this Sep 28, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants