Skip to content

[Bugfix][Spec Decode] Ship the speculative round state to non-last PP ranks - #662

Merged
yangzhuxinyzx merged 2 commits into
1CatAI:mainfrom
Peuqui:pp-spec-state-transport-pr
Sep 26, 2026
Merged

yangzhuxinyzx merged 2 commits into
1CatAI:mainfrom
Peuqui:pp-spec-state-transport-pr

Conversation

@Peuqui

@Peuqui Peuqui commented Sep 19, 2026

Copy link
Copy Markdown
Contributor

Purpose

Follow-up to #511 (context: #439). With pipeline parallelism, a drafter and
async scheduling, the non-last ranks never learn what the last rank sampled.
_pp_broadcast_prev_sampled_token_ids asserts a [num_reqs, 1] tensor, which
the speculative sampler does not produce, and the draft token ids that the next
step scatters into input_ids exist on the last rank only. On main with #574
and #636 applied (Qwen3.8 MTP needs both to get this far under PP) the first
speculative step ends on rank 0 with

RuntimeError: Speculative decode scheduled draft input slots, but the worker
has no draft token tensor to scatter.

and the request hangs.

This change ships the round state:

  • The last rank broadcasts the sampled matrix, padded with -1 to the static
    shape [num_reqs, num_spec_tokens + 1] (the sampler emits fewer columns in
    rounds with fewer or no scheduled drafts), and this step's draft token ids
    [num_reqs, num_spec_tokens]. List-form drafts (ngram) travel as zeros; the
    scheduler schedules no GPU-resident spec slots from those, so they are never
    read.
  • A non-last rank derives the next token ids and the accepted counts from the
    matrix with _count_contiguous_spec_tokens, hands them to
    _copy_valid_sampled_token_count, and runs
    _update_states_after_model_execute on the scheduler_output it stashed
    when it returned its intermediate tensors. A missing stash raises instead of
    silently skipping the hybrid-state update.
  • Both payloads go over the gloo cpu_group. An NCCL broadcast on the
    device_group shares the communicator with the pipeline's send/recv. On a
    five-stage pipeline the two interleaved and the first request hung, ranks
    0-2 in the broadcast and ranks 3-4 in irecv. A CPU rendezvous has no
    stream ordering to violate, and the payloads are a few dozen int32. The
    non-speculative [num_reqs, 1] path is unchanged and stays on the
    device_group.
  • The sender now asserts that its row count equals input_batch.num_reqs, the
    number the receiver sizes its buffer from, so a divergence raises instead of
    hanging every rank in an unmatched collective.

The second commit lets the DSpark drafter load its own embedding table from
embed.weight. With PP=1 the proposer replaces it by the shared target
embedding afterwards. Under pipeline parallelism the target embedding lives on
the first stage and the drafter on the last, so the drafter ran on an
uninitialized table; a checkpoint without embed.weight now fails loudly under
PP.

Not covered: #539 stops earlier, in custom_all_reduce.cuh during the memory
profile run, which this does not touch. Qwen3.5-family MTP under PP also needs
#636, and #574 trims the optimistic tokens on every rank; both are independent
of this change and merge cleanly with it.

Test Plan

pytest tests/v1/worker/test_gpu_model_runner_pp_spec.py \
       tests/v1/worker/test_gpu_model_runner.py \
       tests/v1/spec_decode/test_dspark.py -q
pre-commit run --files <4 files>
pre-commit run mypy-3.10 --hook-stage manual --files <3 files>

The new test runs two CPU processes over gloo as the last and a non-last rank
of a PP=2 deployment: a speculative round with a narrower sampler output, a
round without a stashed scheduler_output (must raise after both broadcasts
were consumed, otherwise the last rank would hang on the next collective),
list-form drafts, and the plain [num_reqs, 1] path.

Server run: Qwen3.8-27B-NVFP4 with MTP (num_speculative_tokens=3), pipeline
parallel over two Tesla V100-PCIE-32GB, --enforce-eager, async scheduling on
(the default), this branch merged with #574 and #636.

Test Result

Tesla V100-PCIE-32GB, CUDA_DEVICE_ORDER=PCI_BUS_ID: 53 passed, 2 skipped.
The new test against the unmodified runner of main (b711d53): 1 failed.
pre-commit: all hooks passed; mypy-3.10 manual stage passed.

Server run with this change: four greedy fact prompts correct, a 320-token
greedy generation at 39.9 tok/s, every greedy prompt repeated token for token,
mean acceptance length 2.9 to 3.2 of 4, and 90 short requests, three in
flight, greedy and temperature 1.0 mixed, without an error. The same stack
with this change reverted: the first request hangs with the RuntimeError
quoted above on Worker_PP0.

A fork of this repository has carried the same transport since early
September on a five-stage pipeline (two RTX 8000, three V100) with
DeepSeek-V4-Flash and DSpark, num_speculative_tokens=5; that is where the
NCCL hang was seen and the gloo path came from.

Not a duplicate

gh pr list -R 1CatAI/1Cat-vLLM --state open --search "pipeline parallel speculative"
gh pr list -R 1CatAI/1Cat-vLLM --state open --search "PP spec decode"
gh pr list -R 1CatAI/1Cat-vLLM --state open --search "draft token"
gh pr list -R 1CatAI/1Cat-vLLM --state open --search "async scheduling pipeline"
gh issue list -R 1CatAI/1Cat-vLLM --state open --search "pipeline parallel"
gh issue list -R 1CatAI/1Cat-vLLM --state open --search "draft token tensor"

The hits are my own #574, #636 and #639 (different defects, see above) and
#625, which rejects out-of-range sampled ids in the async output path and does
not move data between ranks.

AI assistance

AI assistance (Claude) was used to port the change from the fork, to write the
test and to draft this text. I have read every changed line and ran the tests
and the server runs above on my own hardware.


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

Peuqui and others added 2 commits September 19, 2026 21:20
DSparkDeepseekV4ForCausalLM has no embedding of its own in the
checkpoint (has_own_embed_tokens = False): with PP=1 the proposer
shares the target's embed_tokens. Under pipeline parallelism it does
not (_maybe_share_embeddings: "will be loaded separately"), because the
target embedding lives on the first stage and the drafter on the last;
but _remap_dspark_name drops every key outside mtp.*, embed.weight
included, so the drafter's VocabParallelEmbedding kept its random
initialisation. The boot succeeds and acceptance collapses to a few
percent.

Map embed.weight onto the drafter's table (with PP=1 the shared target
embedding replaces it afterwards, as before) and fail loudly under PP
when the checkpoint did not supply it.

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Peuqui <peuqui@github.com>
… ranks

With async scheduling, pipeline parallelism and a drafter, the non-last
ranks never learn what the last rank sampled: the [num_reqs, 1] broadcast
asserts on the wider speculative matrix, and the draft token ids the next
step scatters into input_ids exist on the last rank only.

The last rank now broadcasts the sampled matrix, padded with -1 to the
static shape [num_reqs, num_spec_tokens + 1], and this step's draft token
ids. A non-last rank derives the next token ids and the accepted counts
with _count_contiguous_spec_tokens and runs the hybrid-state update on the
scheduler output it stashed for this step.

Both payloads go over the gloo cpu_group. An NCCL broadcast on the
device_group shares the communicator with the pipeline's send/recv; with
five stages the two interleave and the first request hangs.

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Peuqui <peuqui@github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants