Skip to content

[Research] Preserve Qwen3.8 Flash Next full-compile checkpoint - #5

Draft
Leonccaa wants to merge 1 commit into
feat/qwen38-flash-next-sm70-corefrom
exp/qwen38-flash-next-full-compile
Draft

Leonccaa wants to merge 1 commit into
feat/qwen38-flash-next-sm70-corefrom
exp/qwen38-flash-next-full-compile

Conversation

@Leonccaa

Copy link
Copy Markdown
Owner

Purpose

Preserve the validated Qwen3.8 Flash Next full-compile path as a fork-internal
research checkpoint. This PR is intentionally based on
feat/qwen38-flash-next-sm70-core, so its diff contains only the compile-graph
work and does not expand the upstream model-support PR.

This is a correctness/fallback checkpoint, not a throughput claim and not an
upstream merge candidate as a single patch.

Validated behavior

  • Four Tesla V100 32 GiB GPUs, TP4, FP16 activations, W4A16 NVFP4/Marlin, PLE
    CPU offload.
  • One dynamic compile range (1, 256) compiled successfully.
  • Mixed prefill/decode captured PIECEWISE and batch-one decode captured FULL.
  • CUDA Graph capture added about 0.05 GiB per GPU.
  • The 8-token correctness gate matched the current eager oracle exactly.
  • The 64-interval benchmark matched the prior no-compile output hash in all
    three repeats:
    261ac6d4ac7500646010f2ae56882b3d0d020ce71f5279de2db1747c54f7cee3.

Performance result

Path Mean decode Median decode
no-compile FULL_DECODE_ONLY 26.433 tok/s 26.427 tok/s
full compile + FULL decode graph 25.803 tok/s 25.735 tok/s

The full-compile route was 2.38% slower by mean throughput. The whole-decoder
opaque boundary restores correctness but prevents Inductor from optimizing
inside the layer, so this patch should remain a reference/fallback while the
opaque scope is reduced incrementally.

Patch structure

  • Dynamic-range warmup no longer treats a size-one capture as a safe warmup
    for a multi-size range.
  • PLE CPU-offload wait, runtime slice, and optional FP8 dequantization are
    kept behind one compile-safe custom op.
  • Qwen4Exp decoder, input embedding, and final HyperConnection mixer use
    caller-owned-output custom-op boundaries for exact compiled execution.
  • Functionalization restores the original output/cache tensors for those
    caller-owned operations.

Checks

  • Compile-range warmup tests: 5 passed.
  • New PLE compile/offload tests: 3 passed.
  • Targeted pre-commit, including Ruff and mypy: passed.

Upstream boundary

Potential upstream changes should be split and justified independently:

  1. The generic dynamic-range size-one warmup fix.
  2. The Qwen4Exp PLE runtime-shape/offload synchronization fix.

The whole-layer opaque wrappers and their functionalization rules stay in this
fork-only research PR until they are narrowed and demonstrate a performance or
maintainability benefit.

Keep Qwen4Exp decoder layers, input embedding, final HyperConnection mixer, and PLE offload synchronization behind stable custom-op boundaries so the dynamic compile graph preserves the eager CUDA custom-op execution order. Warm a non-specializing size for dynamic ranges first observed at one token.

Validated on TP4 V100 with PLE CPU offload and W4A16 NVFP4/Marlin: eager-exact token IDs, dynamic compile range 1-256, and full decode CUDA Graph capture.

Assisted-by: OpenAI Codex

Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant