Skip to content

fix(qsa): preserve fast_topk candidates across buffer overflow - #6

Merged
mochgolf merged 3 commits into
mainfrom
fix/qsa-topk-candidate-overflow
Oct 4, 2026
Merged

mochgolf merged 3 commits into
mainfrom
fix/qsa-topk-candidate-overflow

Conversation

@mochgolf

@mochgolf mochgolf commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

Motivation

When more than 4096 different scores occupy the coarse threshold bucket, JIT fast_topk drops uncached candidates and their histogram contributions. QSA prefill and the independent decode call use this JIT path for block top-k 512, so selection can silently return incorrect scores.

Modifications

  • Retain the cached refinement path when the coarse bucket fits. On overflow, rebuild its full histogram and rescan the complete valid row on each refinement pass, matching the original coarse bucket and every previously selected full-key radix byte.
  • Preserve relative indices, arbitrary valid tie selection, short-row -1 padding, PDL synchronization and the 32KB dynamic candidate allocation. Add a small shared radix-prefix array.
  • Adapt the algorithm from upstream #38144, head e9ba4f06996deb7d7d6fdedd340940a99b258e8c, and correct the misleading capacity comment.
  • Extend the existing registered test with distinct same-bin scores at 4095/4096/4097/16384 entries, positive and negative scores, repeated radix overflow, both supported k values, nonzero starts, padded row strides, empty/short/exact-k rows, and graph replay after in-place score and length changes. Check the exact selected-value multiset, uniqueness and bounds.

Validation

  • git diff --check and Python syntax compilation passed.
  • Fixture geometry and oracle-helper checks passed for k=512/2048 and both signs on a local RTX 4070 Ti SUPER (SM89), Torch 2.8.0+cu128. These checks exercised only Torch fixture/oracle operations, not the repaired kernel.
  • Actual kernel regression execution stopped at import with ModuleNotFoundError: sglang. SGLang, Triton and FlashInfer are absent. No CUDA kernel compilation, pre-fix failure / post-fix pass comparison, CUDA graph kernel execution, target dual-4090 run, sanitizer, end-to-end inference or performance benchmark was completed. This remains a draft pending those checks.
  • Latest main was a3154e0; existing fork PRs Maintenance/prefix cache service #1–Add opt-in NUMA memory interleave with GPU-local CPU affinity #5 contained no duplicate top-k repair. No archived files were changed; no service was restarted, deployed or merged.

Scope and remaining risks

This fixes the Python-shipped JIT kernel only. Deterministic QSA already uses full stable sorting. QSA qsa_fast_topk routes non-512 supported CUDA k to installed sgl_kernel.fast_topk_v2: the AOT source still has the same risk and this PR does not update or rebuild its installed wheel. JIT dsa/kpool_topk_transform.cuh is separately reachable through fast_kpool_topk_transform_fused and also remains unchanged. Upstream #37893, head f36a7588361ec92351e40e3fb59d56f560ad792b, addresses those two paths with a different cached-tail design; porting it needs separate build and GPU qualification.

Run in a compatible source environment:

python test/registered/kernels/ops/elementwise/test_fast_topk.py

CI States

Latest PR Test (Base): ❌ Run #37185123993
Latest PR Test (Extra): ❌ Run #37185123889
Latest PR Test (AMD ROCm 10): ❌ Run #37185124005

Adapt exact full-row radix refinement from sgl-project/sglang#38144
(Leslie360/sglang e9ba4f06996deb7d7d6fdedd340940a99b258e8c).
Add focused capacity-boundary, repeated-overflow and graph replay coverage.
@mochgolf
mochgolf marked this pull request as ready for review October 4, 2026 12:55
@mochgolf
mochgolf merged commit f7ae0f1 into main Oct 4, 2026
94 of 110 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant