Skip to content

QSA: integrate SGLang 35f3c96 and fix FP8/cache/parallel compatibility - #7

Merged
mochgolf merged 574 commits into
mainfrom
maintenance/dsh-upstream-qsa-20261004
Oct 4, 2026
Merged

mochgolf merged 574 commits into
mainfrom
maintenance/dsh-upstream-qsa-20261004

Conversation

@mochgolf

@mochgolf mochgolf commented Oct 4, 2026 •

Copy link
Copy Markdown
Owner

Motivation

Integrate SGLang main at 35f3c96ff4794a4de15daf12caad371084a037ee while keeping the QSA raw-KV offload and private host-prefix cache working with the new cache lifecycle and parallel context. Non-unit FP8 KV scales previously produced incorrect cached-prefix/prefill and reference reads; the removed runner TP fields also prevented model startup after the upstream merge.

Modifications

  • Preserve the two-parent upstream merge 80dc48ddfc over the verified common ancestor 32290dda2c (561 upstream commits).
  • Pass and apply K/V descales on the affected FP8 reads; retain already-restored BF16 compact-gather behavior to avoid double scaling.
  • Migrate checkpoint/claim/free/unpin/release and live TP rank/CPU-group APIs, preserving host-prefix snapshots, DMA leases, Mamba/PLE state, NUMA placement and W4A16 Marlin routing.
  • Guard cache selection and fused KV paths so generic sharing or fusion cannot bypass the private offload contract. Retain PR fix(qsa): preserve fast_topk candidates across buffer overflow #6 fast_topk overflow handling and complete the fused-fallback test fixture.
  • Record the frozen CPU/GPU/API checks and the later local production promotion as separate source-bound validation windows.

Accuracy Tests

The tested runtime is 2fe0731e03; production snapshot 897286b12a and this PR's final commits differ from it only in validation/provenance documentation.

  • CPU runner in two environments: 221 passed, 7 skipped, 24 subtests; JUnit 252 cases, zero failures/errors. Independent live TP rank/group regressions: 8 passed.
  • RTX 4090 SM89 kernel tests: 240 passed, 1 skipped (SM121-only), zero failures. Non-unit FP8 and gather/overflow/graph regressions retain their original tolerances.
  • TP2 service: 7/7 fixed cases, including an exact 262016-token input; 36 B1..8 graph requests, prefix reuse, explicit abort/recovery and flush/reseed passed.
  • Local production 8081 promoted and healthy using the validated isolated dependency environment; rollback profile retained.
  • Three real MCP-driven DSH sessions passed at session effort max, actual wire xhigh, with real reasoning output: pelican SVG independently rendered/reviewed; image 5/5 with byte-equal image payload; IMO 1988 Problem 6 proof independently reviewed. 13 primary model requests, all HTTP 200, no fallback or relay errors. The pelican refined its output within the same first turn; the math problem is known rather than claimed unseen.

Full validation and preserved failure history: VALIDATION.md and PROVENANCE.json. Raw local artifacts, DSH configuration, logs and credentials are not included in this fork.

Speed Tests and Profiling

Functional CUDA graph replay was observed for every B1..8 batch in the fixed service window. No throughput, latency improvement or performance certification is claimed.

Scope and limitations

  • Serving remains OBSERVE=light; full-model strict bitwise/determinism was not validated. DSH max is model reasoning effort.
  • No-host-prefix mode has CPU registry coverage only; AMD/other GPUs, SM121, packaging and this revision's sanitizers were not exercised. The four later upstream commits at the recorded affa261e3d snapshot are not included.
  • Earlier startup/harness failures are retained. The later DSH window also retains one pre-model credential failure and three 64-token auxiliary title requests (two length limits, one client cancellation); these are excluded from the primary-task pass count.
  • Repository-wide lint has inherited formatting debt; a full-tree lint pass is not claimed. Current GitHub checks are recorded below before merge.

Checklist

  • Add and run relevant correctness regressions.
  • Update source-bound validation and provenance.
  • Independently review and exercise the real local model/service.
  • Repository-wide pre-commit/lint clean (inherited debt).
  • Performance benchmark/certification (not part of this maintenance scope).

CI States

Latest PR Test (Base): ❌ Run #37217134865
Latest PR Test (Extra): ❌ Run #37217134654
Latest PR Test (AMD ROCm 10): ❌ Run #37217134818

GitHub checks and merge basis

  • Lint run failed its shebang/executable, isort, ruff-format and clang-format hooks. Python AST/Ruff diagnostics, YAML/TOML, codespell, test registries, Rust formatting and Clippy passed. Of the 39 formatting-modified paths, 38 already appeared in PR fix(qsa): preserve fast_topk candidates across buffer overflow #6's lint record; the remaining path is python/sglang/kernels/ops/gemm/hc_mix.py. This PR does not clean up the repository-wide formatting debt.
  • SMG benchmark compilation check passed.
  • The inherited hardware workflows fail their trigger-label gates, before hardware tests run. The base gate log explicitly reports Missing required label 'run-ci'; other gates likewise require their trigger labels. These results are not counted as executed hardware tests or passing CI. NPU/nightly and image-build jobs that remain pending add no validation coverage.
  • Merge is based on the independent, source-bound CPU/GPU/TP2 service and real DSH qualification described above. No claim of all-green GitHub CI is made. The local 8081 service keeps its tested frozen snapshot and rollback profile.

EdwardXuy and others added 30 commits September 30, 2026 17:17
… (#41856)

Co-authored-by: EdwardXuy <EdwardXuy@users.noreply.github.com>
…ccounting (#40227)

Co-authored-by: Ke Bao <ispobaoke@gmail.com>
…e growth (#41572)

Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
…445)

Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
…#40695)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
… (#40696)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…(11/13) (#40697)

Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nd methodology (#41689)

Co-authored-by: Mick Qian <mickqian@users.noreply.github.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…traction (#41380)

Signed-off-by: Sohom Chakraborty <sohomchakraborty.iitkgp@gmail.com>
Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
…efore preallocation (#41451)

Co-authored-by: Shangming Cai <csmthu@gmail.com>
… match (#41758)

Co-authored-by: merrymercy <merrymercy@users.noreply.github.com>
Co-authored-by: xiezhq-hermann <xiezhq-hermann@users.noreply.github.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
…(#37984)

Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com>
Co-authored-by: Po-Han Huang <pohanh@nvidia.com>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
ruff-format/isort on the two merge-touched files that the project-version
static checks still flagged; no behavior change.
gpu_service_validation is now explicitly marked as the historical f8a6962
result only, with its numbers unchanged and no claim over the Phase-1 commit or
this merge. The Phase-1 service evidence moves to its own block carrying
source_commit bdb935d and its evidence path, marked as externally reported.

README/UPSTREAM/PROVENANCE no longer describe the 172b1b4 record as a
rebase-era baseline. They state the verifiable facts: the fork's actual common
ancestor is 32290dd, 172b1b4 is 52 upstream commits behind it, and those
52 commits are already integrated through 2f06478, so counts taken from the
old record overstate the missing work.
…undle

Production startup on the merged tree failed in QSAHiSparseRuntime.__init__:
ModelRunner no longer carries tp_rank after the parallel-context migration, so
`runner.tp_rank` raised AttributeError on both TP ranks. Read the rank through
the same published context upstream uses (get_parallel().tp_rank) in the runtime
and in QSAHiSparseSingleRequest; no getattr default and no compatibility field
on ModelRunner.

The local audit of QSA construction sites kept runner.token_to_kv_pool,
token_to_kv_pool_allocator, req_to_token_pool, server_args, model_config and
dtype (all still present) and server_args.tp_size/pp_size/max_running_requests/
cuda_graph_backend_decode (still present); tp_rank was the only removed field.

Regression: test/qsa_hisparse/test_parallel_rank_migration.py builds real
runtime/single-request objects with a runner that either lacks tp_rank (the
production condition) or still carries a stale 0, while the published bundle
reports rank 1; it asserts rank 1 and the rank-1 event path. The pre-existing
single-request constructor test now publishes the rank it previously passed
through the removed runner field.
Second production startup failure on the merged tree, both TP ranks:
ModelRunner.init_attention_backends built QSAHiSparseCoordinator with
`self.tp_group.cpu_group`, but ModelRunner no longer carries tp_group after the
parallel-context migration. Take it from the published bundle
(get_parallel().tp_group.cpu_group), the same source upstream uses; the
coordinator keeps the TP CPU-group contract (all_reduce/all_gather over TP).

Bounded audit of our downstream code for parallel fields removed upstream
(added lines only, 46 changed python files): the other references are
get_parallel().attn_tp_size/attn_dcp_size in the backend and the kv-cache
configurator, Scheduler.tp_cpu_group (assigned from get_parallel().tp_group),
runner.server_args.tp_size (still present on ServerArgs), and the runtime and
single-request rank helpers fixed earlier. model_runner.py:1049 was the only
remaining removed field.

Regression test/qsa_hisparse/test_attention_backend_group_migration.py executes
the real init_attention_backends body with the backend/coordinator setup mocked:
a runner without tp_group (production condition) and a runner with a stale group
distinct from the live context; only the live group may reach the coordinator,
and the coordinator is skipped for non-lease pools.
test_block_indices_expand_on_attention_fallback built the backend with
__new__, so the decode path's `self.req_to_token_pool.req_to_token` (and the
metadata row-request lookup) raised AttributeError once the upstream fused-KV
test file entered the GPU suite. Construct the backend fully and provide the
request table plus row_req_pool_indices the path reads.

Assertions and production behavior are unchanged; the oracle still compares the
expanded indices against expand_qsa_block_indices.
…vice review

Docs-only. VALIDATION.md gains the 2026-10-04 section: merge commit
80dc48d over the real ancestor 32290dd, frozen tested source 2fe0731,
CPU 221/7/24 in both environments, the independent 4090 SM89 GPU suite
(240 passed, 1 SM121 skip, 0 failures), the service window (7/7 frozen cases,
B1..8 36 requests with 65536 cached tokens, abort/recovery/reseed/health,
graph evidence) and the confirmed restore (run-summary.json:
candidate_passed=true, production_restored=true). PROVENANCE replaces the
CPU-only pending status with a scoped phase2_independent_review block plus the
latest-main recheck (affa261e, 4 commits, no QSA runtime or dependency overlap,
not merged), keeps the historical f8a6962 and Phase-1 bdb935d blocks separate,
and records the retained startup/harness failures and the explicit not-covered
list. UPSTREAM.md notes why the four later main commits are not chased.
Docs-only follow-up to 7a4a0ea:
- UPSTREAM.md no longer says the GPU/service review is pending; it points at the
  2026-10-04 frozen-source section of VALIDATION.md and keeps the source
  distinction (validated 2fe0731, later doc-only commits).
- VALIDATION/PROVENANCE now record both earlier GPU windows precisely: window 1
  (results/dsh-maintenance-20261004/phase2-final-gpu-tests.txt) 68 failed, 172
  passed, 1 skipped from a PATH without ninja, insufficient free GPU memory
  while the production service ran, and the uninitialized fixture; window 2
  (lab/results/.../phase2-group/gpu-preflight.stdout/.xml) 34 failed, 206
  passed, 1 skipped, of which 33 came from the harness inheriting the production
  SGLANG_QSA_HISPARSE_V3=p2-offload mode plus the same single fixture. The
  fixture is a test-fixture migration gap, not a product runtime API bug, and
  window 2 is a corrected environment run of the same window rather than a
  second environment.
- PROVENANCE.validation.test_environment now carries an explicit
  legacy/historical scope so the old Torch 2.13 / xgrammar 0.2.1 values are not
  read as the current GPU environment; the current environments live in
  phase2_cpu and phase2_independent_review.
- Both documents state that evidence paths are relative to the maintenance
  workspace root /home/zyk/projects/interests/ai-video/qwen, not to this
  checkout, and that the raw artifacts are not committed.
The two earlier GPU runs are distinct execution windows; the point that needed
stating is that the single fixture gap was one root cause, not a separate round.
Remove the sentence claiming window 2 was a re-run of the same window from
VALIDATION.md and from the PROVENANCE history_retained entry; both now end at
the 68/34 counts, their root causes and the final 240-pass result.
…rt window

The frozen-source window (2fe0731) ended with production restored and no
packaging. Add the later window as its own record instead of merging the two:
897286b, which differs from the tested runtime only in PROVENANCE.json,
UPSTREAM.md and VALIDATION.md, was promoted to the local 8081 service, matches
its launch snapshot, reports health 200 and keeps its previous profile for
rollback.

Three standard-local DSH 0.1.7-rc.2 sessions then ran with session effort max
dispatched as wire reasoning_effort=xhigh: 13 primary requests, all HTTP 200
with real reasoning, no fallback route and no relay error. Pelican passed with
10 requests and two self-directed edits inside the same first turn, the image
case passed 5/5 with 660 image tokens and a byte-equal decoded payload, and the
IMO 1988 Problem 6 proof passed an independently predeclared, line-by-line
review.

The retained non-pass records stay visible: the pre-model MISSING_CREDENTIAL
route failure, and the three 64-token title requests (two length-limited, one
client-cancelled BrokenPipe) that are not counted as passes. Limits keep the
window honest: no statistical or performance certification, a known competition
problem rather than an unseen one, local promotion only, still OBSERVE=light, and
identity by Git commits and byte comparison rather than file hashes. Evidence
paths stay relative to the maintenance workspace and raw artifacts, session logs
and credentials are not committed.
…al wording

Five corrections inside the later-window record, no code, test, environment or
service change:

- remove a stray backtick in the 13-primary-request sentence;
- state the environment transition accurately: the promotion did not upgrade the
  previous production environment in place, it switched the service to the
  already-validated upstream-runtime-env runtime path while weights and serving
  configuration stayed as validated;
- relate the two windows by their tested Git revisions 2fe0731 and its
  documentation-only successor 897286b, separated in time, instead of
  claiming they share no snapshot;
- narrow the credential claim: the isolated test instance neither copied nor
  modified the global DSH credential store, and MCP control authentication was
  loaded by the standard wrapper/environment mechanism without being disclosed;
- say that the lab evidence is not included in or committed to this fork and is
  kept under the separate lab's own local evidence Git commit aa33f74, whose raw
  artifacts are not published.

PROVENANCE.json carries the same boundaries in
validation.local_production_promotion and its shared raw-artifacts note.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.