Skip to content

[Bugfix][Qwen3.5] Keep the MTP drafter stage-local under pipeline parallelism - #636

Merged
yangzhuxinyzx merged 1 commit into
1CatAI:mainfrom
Peuqui:qwen3-5-mtp-pp
Sep 26, 2026
Merged

yangzhuxinyzx merged 1 commit into
1CatAI:mainfrom
Peuqui:qwen3-5-mtp-pp

Conversation

@Peuqui

@Peuqui Peuqui commented Sep 14, 2026

Copy link
Copy Markdown
Contributor

Purpose

Qwen3.5-family targets (Qwen3.5, Qwen3.6, Qwen3.8 dense and MoE) with
--speculative-config method=mtp and --pipeline-parallel-size greater than one
do not boot. There are three stacked causes, each hiding the next:

  1. Qwen3_5MTP does not implement SupportsPP. The draft model config is
    verified against the parallel config, so the engine refuses to start:

    NotImplementedError: Pipeline parallelism is not supported for this model.
    Supported models implement the `SupportsPP` interface.
    
  2. Qwen3_5MultiTokenPredictor.forward branches on get_pp_group().is_first_rank,
    which is the TARGET model's pipeline position. The drafter only runs on the
    last rank (execute_model returns the IntermediateTensors on every earlier rank
    before speculation), so it took the "receive from the previous stage" path and
    asserted on intermediate tensors that nobody sends. This is the same defect
    [Bugfix][Qwen4Exp] Keep the MTP drafter stage-local under pipeline parallelism #573 fixed for Qwen4Exp.

  3. With the default VLLM_QWEN35_MTP_SHARE_IO_WEIGHTS=1 the drafter builds its
    embedding as a PPMissingLayer and expects the loader to share the target's.
    Both load_eagle_model (V2) and _maybe_share_embeddings (V1) deliberately
    skip embedding sharing under PP, because the target's embedding only exists on
    the first rank. The placeholder was therefore never replaced, and the first
    compile on the last rank failed with

    torch._dynamo.exc.Unsupported: Assertion failed on symbolic shapes
      assert hidden_states.shape[-1] == inputs_embeds.shape[-1]
    

Changes:

  • Qwen3_5MTP implements SupportsPP and forwards
    make_empty_intermediate_tensors from the predictor, as Qwen4ExpMTP does.
  • The predictor's forward drops both pipeline branches: it always embeds locally
    and never returns IntermediateTensors for a next stage it does not have.
  • Under PP the drafter owns its embedding (loaded from the checkpoint's
    embed_tokens) instead of waiting for a share that the loaders skip. With a
    single pipeline rank nothing changes: sharing stays as before. The LM head is
    still shared, since the target's head lives on the drafter's rank.

Single-rank behaviour is unchanged: there is_first_rank and is_last_rank are
both True, the removed branches were dead, and world_size == 1 keeps the shared
embedding.

Not a duplicate: no open PR here or in vllm-project/vllm touches Qwen3_5MTP under pipeline parallelism; #573 fixed only the Qwen4Exp drafter.

Test Plan

.venv/bin/python -m pytest tests/models/qwen3_5/test_mtp_stage_local.py -v
pre-commit run --files <changed files>
pre-commit run mypy-3.10 --hook-stage manual --files <changed files>

End to end on 2x Quadro RTX 8000 (TP1 x PP2), RadixArk/Qwen3.8-27B-NVFP4,
--speculative-config '{"method":"mtp","num_speculative_tokens":3,"draft_sample_method":"greedy"}',
fp16, V2 model runner, 400 tokens at temperature 0, five repetitions, plus the same
prompt on TP2 with a DFlash2 draft as a text reference.

Test Result

  • New tests: 3 passed. With the model file reverted to main: 3 failed
    (AssertionError on the intermediate tensors, missing SupportsPP).
  • main, TP1 x PP2: the engine refuses to start with the NotImplementedError above.
  • With causes 1 and 2 fixed but the embedding still shared: boot fails on the last
    rank with the dynamo assertion quoted above.
  • This PR, TP1 x PP2: boots in 105 s and serves. 61.05 tok/s median (61.00 to
    61.10), acceptance length 2.963. The 400-token greedy text is byte-identical to
    the TP2 reference (sha256 prefix 0106659946c064b1), so the drafter changes
    acceptance, never the output.

AI assistance (Claude) was used to write this change and this description. I
reviewed every changed line and ran the tests and measurements above.


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

…allelism

Qwen3.5-family targets with --speculative-config method=mtp and
--pipeline-parallel-size greater than one do not boot. Three stacked
causes, each hiding the next:

1. Qwen3_5MTP does not implement SupportsPP, so the draft model config
   check refuses to start the engine.
2. The predictor's forward branches on get_pp_group().is_first_rank, the
   TARGET model's pipeline position. The drafter only runs on the last
   rank, took the "receive from the previous stage" path and asserted on
   intermediate tensors that nobody sends (same defect as 1CatAI#573).
3. With VLLM_QWEN35_MTP_SHARE_IO_WEIGHTS=1 the drafter's embedding is a
   PPMissingLayer awaiting the target's. Both eagle loaders skip
   embedding sharing under PP, so the placeholder stayed and the first
   compile failed on the hidden/embedding shape assertion.

Implement SupportsPP, drop both pipeline branches in the predictor, and
let the drafter own its embedding when the pipeline has more than one
rank. Single-rank behaviour is unchanged.

Measured on 2x Quadro RTX 8000, TP1 x PP2, Qwen3.8-27B-NVFP4, MTP k=3:
main refuses to start; with this change the engine boots and serves at
61.05 tok/s, acceptance length 2.963, and the 400-token greedy text is
byte-identical to the TP2 reference.

Co-authored-by: Claude
Signed-off-by: Peuqui <peuqui@github.com>
Peuqui pushed a commit to Peuqui/1Cat-vLLM that referenced this pull request Sep 14, 2026
- multiproc_executor: PP batch-queue cap removed; 1Cat's own PP spec
  transport no longer deadlocks DeepSeek-V4 PP5 without it (4 requests,
  22-26 tok/s with VLLM_SM70_ASYNC_SCHEDULING_QUEUE_DEPTH=0)
- qwen_gdn_linear_attn: dead UpstreamCall adapter (sm75 GDN module gone)
- qwen3_5_mtp: replaced by the 1CatAI#636 version (stage-local drafter, own
  embedding under PP); PP2 MTP text sha 0106659946c064b1

Co-authored-by: Claude
Signed-off-by: Peuqui <peuqui@github.com>
@yangzhuxinyzx
yangzhuxinyzx merged commit 4b8855c into 1CatAI:main Sep 26, 2026
5 of 6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants