Skip to content

[Core][SM70] Repair Qwen4Exp PLE integration from #338 - #374

Merged
yangzhuxinyzx merged 9 commits into
mainfrom
agent/v100-repair-pr338-ple-amd-latest-main-20260828-0327
Aug 28, 2026
Merged

yangzhuxinyzx merged 9 commits into
mainfrom
agent/v100-repair-pr338-ple-amd-latest-main-20260828-0327

Conversation

@yangzhuxinyzx

@yangzhuxinyzx yangzhuxinyzx commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Reconstruct the still-useful parts of #338 on the current public main without restoring its stale NVIDIA implementation:

  • add the ROCm Qwen4Exp model/QSA/PLE/MTP route;
  • add node-local CPU PLE offload for MRV1 and MRV2;
  • retain the latest SM70 NVIDIA QSA, pinned-host PLE, GDN, MTP, and launch policies already on main;
  • preserve the signed [Model][SM70] Support Qwen3.8 Flash Next with PLE offload #338 ancestry with a no-source-change merge.

Audit repairs

  • Wire the current NVIDIA N-gram embedding through PleOffloadLayer while preserving the existing dynamic-shape and SM70 UVA paths.
  • Keep CUDA IPC imports feature-local so disabled PLE offload does not add a CUDA dependency to ROCm/general model loading.
  • Remove the model-name whitelist; admission is based on PLE capability plus supported topology.
  • Repair AMD weight mapping against the current Qwen3.5 API, attach current fused-MoE expert mappings, and fail closed on incomplete MTP expert loads.
  • Keep the newer mainline SM70 QSA launch policy and update the imported regression expectations instead of restoring the older [Model][SM70] Support Qwen3.8 Flash Next with PLE offload #338 policy.
  • Use a native PyTorch Gemma RMSNorm reference on V100 rather than FlashInfer's SM75+ JIT path.
  • Use ForkingPickler symmetrically for tensor IPC and satisfy the forbidden-import gate.
  • Preserve the callable hf_overrides quantization fix while removing a cosmetic global PLE import that would couple ROCm/CPU loading to CUDA bindings.

Test plan and result

  • Repository-wide pre-commit run --all-files: passed (Ruff, format, mypy, lockfiles, SPDX, forbidden imports, configuration gates).
  • PLE worker and NVIDIA PLE focused tests: 34 passed; the pre-existing UVA custom-op test requires a compiled vllm._C and was deselected in this source checkout.
  • MRV2 cache layout, current SM70 QSA launch policy, fused pre-indexer, and HC tests: 20 passed on V100.
  • Existing Qwen4 config, weight loading, model-forward, QSA-cache, and PLE regression set: 52 passed.
  • Post-main-sync PLE plus compressed-KV zeroer regression: 30 passed.
  • AMD model and MTP modules import and compile against current main; 10 ROCm kernel tests are correctly skipped on the V100 host.
  • DCO: all reconstructed and ancestry commits are signed.

No broad end-to-end run was used; validation is source/static plus the smallest focused V100 tests that exercise the changed contracts.

Leonccaa and others added 9 commits August 26, 2026 11:14
Integrate the Qwen3.8 Flash Next model and PLE CPU offload work from vLLM PRs #53896 and #53899, then carry the validated 1Cat SM70 compatibility layer for FP16 activations, hybrid caches, NVFP4 emulation, and top-token sampling.

The included TP4 smoke covers checkpoint loading, PLE offload, prompt prefill, and one autoregressive decode step on four V100 32GB GPUs.

Co-authored-by: huanghaoyan.hhy <huanghaoyan.hhy@alibaba-inc.com>

Co-authored-by: zjy0516 <riverclouds.zhu@qq.com>

Assisted-by: OpenAI Codex
Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
Match the 1Cat mRoPE, FP8 skip, short-conv metadata, and host-to-device helper signatures. Narrow optional FP8 test scales so the upstream pre-commit mypy matrix can validate the imported code.

Assisted-by: OpenAI Codex

Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
Widen the 1Cat hybrid state-shape protocol to the current vLLM shape union and document intentional Qwen4Exp overrides for the older 1Cat Qwen3Next and pipeline-parallel interfaces.

Assisted-by: OpenAI Codex

Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
Apply the pinned Ruff 0.14.0 formatter to the two PR-specific files that differ from the current 1Cat main formatting baseline.

Assisted-by: OpenAI Codex

Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
Use four warps for the single-split sparse QSA prefill profile on exact SM70 devices while preserving the existing launch choices everywhere else. Add a pure policy test for the original profile boundaries and the V100-specific branch.

The corresponding four-V100 TP4 gate improved 1K and 4K prefill by 5.47x and 6.16x with matching deterministic first tokens.

Assisted-by: OpenAI Codex

Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
@yangzhuxinyzx
yangzhuxinyzx marked this pull request as ready for review August 28, 2026 04:00
@yangzhuxinyzx

Copy link
Copy Markdown
Contributor Author

审计与修复完成:以最新 main 重构 #338 的 AMD/PLE 有效部分,保留当前 NVIDIA SM70 快路,修复了 offload 接口、ROCm 导入隔离、AMD 权重/MoE/MTP 兼容、QSA 测试基线和 IPC 安全门禁。V100 定向回归、现有 Qwen4 回归、本地与 GitHub 全量 pre-commit 均通过;锁定 SHA 11abba6 合并。

@yangzhuxinyzx
yangzhuxinyzx merged commit 0ca7111 into main Aug 28, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants