Skip to content

[Model][SM70] Enable Qwen4Exp multimodal inference - #458

Merged
yangzhuxinyzx merged 1 commit into
1CatAI:mainfrom
Leonccaa:fix/qwen4exp-multimodal
Sep 3, 2026
Merged

yangzhuxinyzx merged 1 commit into
1CatAI:mainfrom
Leonccaa:fix/qwen4exp-multimodal

Conversation

@Leonccaa

@Leonccaa Leonccaa commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Remove the provisional guard that rejected full Qwen4Exp image and video inference on the SM70 route.
  • Preserve MRoPE for full multimodal requests while continuing to strip it in language-model-only mode.
  • Replace the removed video-pruning helper in the NVIDIA implementation with the current Qwen3-VL initialization contract.
  • Add a config regression test for multimodal admission and MRoPE preservation.

Scope

This is a Qwen4Exp model-integration fix for the 1Cat V100/SM70 route. It does not change QSA kernels, FP8 KV conversion, or attention dispatch. The same initialization issue affected FP16 KV and calibrated E4M3 QSA KV runs.

Duplicate check

I searched open and closed 1Cat PRs and open vLLM PRs for Qwen4Exp multimodal, MRoPE, and video-pruning fixes. No equivalent PR is open. 1Cat PRs #359 and #374 repaired adjacent Qwen4Exp integration paths but did not enable the multimodal tower or replace this removed helper.

Validation

  • Focused config regressions: 2 passed.
  • Python compile check passed for the changed config, NVIDIA model, and test files.
  • All changed-file pre-commit hooks passed, including Ruff, format, mypy, SPDX, forbidden-import, accelerator API, and configuration checks.
  • End-to-end validation used four Tesla V100 PCIe 32GB GPUs with TP4, MTP0, FP16 activations, eager execution, and prefix caching disabled.
  • Five image cases and four video cases were each repeated twice under both FP16 KV and calibrated E4M3 main QSA KV: 18/18 successful requests and 18/18 semantically correct responses per arm, with 9/9 exact repeat stability and no runtime or request errors.

A full local test-file run was not used as acceptance evidence because this CPU environment lacks torchvision and its global teardown tries to access an unavailable accelerator. The two changed config contracts were rerun directly with the repository cleanup boundary disabled and passed.

AI assistance

OpenAI Codex assisted with implementation, test execution, duplicate checks, and PR preparation. The human submitter reviewed the three-file diff and the reported V100 runtime evidence.

Remove the provisional full-multimodal guard while preserving MRoPE for image and video requests. Align NVIDIA video-pruning state initialization with the current Qwen3-VL contract after the old helper was removed.

Validated on four V100 GPUs with five image and four video cases, each repeated twice, using both FP16 KV and calibrated E4M3 QSA KV.

Assisted-by: OpenAI Codex
Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
@Leonccaa
Leonccaa force-pushed the fix/qwen4exp-multimodal branch from d59cef7 to dd68b9e Compare September 2, 2026 20:39
@Leonccaa Leonccaa changed the title [Model] Enable Qwen4Exp multimodal inference [Model][SM70] Enable Qwen4Exp multimodal inference Sep 2, 2026
@yangzhuxinyzx
yangzhuxinyzx merged commit 18982e3 into 1CatAI:main Sep 3, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants