[Bugfix][Qwen3.5] Keep the MTP drafter stage-local under pipeline parallelism - #636
Merged
Merged
Conversation
…allelism Qwen3.5-family targets with --speculative-config method=mtp and --pipeline-parallel-size greater than one do not boot. Three stacked causes, each hiding the next: 1. Qwen3_5MTP does not implement SupportsPP, so the draft model config check refuses to start the engine. 2. The predictor's forward branches on get_pp_group().is_first_rank, the TARGET model's pipeline position. The drafter only runs on the last rank, took the "receive from the previous stage" path and asserted on intermediate tensors that nobody sends (same defect as 1CatAI#573). 3. With VLLM_QWEN35_MTP_SHARE_IO_WEIGHTS=1 the drafter's embedding is a PPMissingLayer awaiting the target's. Both eagle loaders skip embedding sharing under PP, so the placeholder stayed and the first compile failed on the hidden/embedding shape assertion. Implement SupportsPP, drop both pipeline branches in the predictor, and let the drafter own its embedding when the pipeline has more than one rank. Single-rank behaviour is unchanged. Measured on 2x Quadro RTX 8000, TP1 x PP2, Qwen3.8-27B-NVFP4, MTP k=3: main refuses to start; with this change the engine boots and serves at 61.05 tok/s, acceptance length 2.963, and the 400-token greedy text is byte-identical to the TP2 reference. Co-authored-by: Claude Signed-off-by: Peuqui <peuqui@github.com>
Peuqui
pushed a commit
to Peuqui/1Cat-vLLM
that referenced
this pull request
Sep 14, 2026
- multiproc_executor: PP batch-queue cap removed; 1Cat's own PP spec transport no longer deadlocks DeepSeek-V4 PP5 without it (4 requests, 22-26 tok/s with VLLM_SM70_ASYNC_SCHEDULING_QUEUE_DEPTH=0) - qwen_gdn_linear_attn: dead UpstreamCall adapter (sm75 GDN module gone) - qwen3_5_mtp: replaced by the 1CatAI#636 version (stage-local drafter, own embedding under PP); PP2 MTP text sha 0106659946c064b1 Co-authored-by: Claude Signed-off-by: Peuqui <peuqui@github.com>
This was referenced Sep 19, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Qwen3.5-family targets (Qwen3.5, Qwen3.6, Qwen3.8 dense and MoE) with
--speculative-config method=mtpand--pipeline-parallel-sizegreater than onedo not boot. There are three stacked causes, each hiding the next:
Qwen3_5MTPdoes not implementSupportsPP. The draft model config isverified against the parallel config, so the engine refuses to start:
Qwen3_5MultiTokenPredictor.forwardbranches onget_pp_group().is_first_rank,which is the TARGET model's pipeline position. The drafter only runs on the
last rank (
execute_modelreturns the IntermediateTensors on every earlier rankbefore speculation), so it took the "receive from the previous stage" path and
asserted on intermediate tensors that nobody sends. This is the same defect
[Bugfix][Qwen4Exp] Keep the MTP drafter stage-local under pipeline parallelism #573 fixed for Qwen4Exp.
With the default
VLLM_QWEN35_MTP_SHARE_IO_WEIGHTS=1the drafter builds itsembedding as a
PPMissingLayerand expects the loader to share the target's.Both
load_eagle_model(V2) and_maybe_share_embeddings(V1) deliberatelyskip embedding sharing under PP, because the target's embedding only exists on
the first rank. The placeholder was therefore never replaced, and the first
compile on the last rank failed with
Changes:
Qwen3_5MTPimplementsSupportsPPand forwardsmake_empty_intermediate_tensorsfrom the predictor, asQwen4ExpMTPdoes.and never returns IntermediateTensors for a next stage it does not have.
embed_tokens) instead of waiting for a share that the loaders skip. With asingle pipeline rank nothing changes: sharing stays as before. The LM head is
still shared, since the target's head lives on the drafter's rank.
Single-rank behaviour is unchanged: there
is_first_rankandis_last_rankareboth True, the removed branches were dead, and
world_size == 1keeps the sharedembedding.
Not a duplicate: no open PR here or in vllm-project/vllm touches Qwen3_5MTP under pipeline parallelism; #573 fixed only the Qwen4Exp drafter.
Test Plan
End to end on 2x Quadro RTX 8000 (TP1 x PP2), RadixArk/Qwen3.8-27B-NVFP4,
--speculative-config '{"method":"mtp","num_speculative_tokens":3,"draft_sample_method":"greedy"}',fp16, V2 model runner, 400 tokens at temperature 0, five repetitions, plus the same
prompt on TP2 with a DFlash2 draft as a text reference.
Test Result
(AssertionError on the intermediate tensors, missing SupportsPP).
rank with the dynamo assertion quoted above.
61.10), acceptance length 2.963. The 400-token greedy text is byte-identical to
the TP2 reference (sha256 prefix
0106659946c064b1), so the drafter changesacceptance, never the output.
AI assistance (Claude) was used to write this change and this description. I
reviewed every changed line and ran the tests and measurements above.
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.