QSA: integrate SGLang 35f3c96 and fix FP8/cache/parallel compatibility - #7
Merged
Merged
Conversation
… (#41856) Co-authored-by: EdwardXuy <EdwardXuy@users.noreply.github.com>
…ccounting (#40227) Co-authored-by: Ke Bao <ispobaoke@gmail.com>
…e growth (#41572) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
…445) Co-authored-by: luoyuan.luo <luoyuan.luo@antgroup.com>
…#40695) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
… (#40696) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…(11/13) (#40697) Co-authored-by: Kangyan Zhou <kangyan.zhou@radixark.ai> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nd methodology (#41689) Co-authored-by: Mick Qian <mickqian@users.noreply.github.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…traction (#41380) Signed-off-by: Sohom Chakraborty <sohomchakraborty.iitkgp@gmail.com> Co-authored-by: Shuwen Wang <47200617+alphabetc1@users.noreply.github.com>
Signed-off-by: jain-ria <riajain@NVIDIA.com>
…efore preallocation (#41451) Co-authored-by: Shangming Cai <csmthu@gmail.com>
… match (#41758) Co-authored-by: merrymercy <merrymercy@users.noreply.github.com> Co-authored-by: xiezhq-hermann <xiezhq-hermann@users.noreply.github.com>
Co-authored-by: Kevin Mi <kevin.mi@radixark.ai>
…(#37984) Co-authored-by: Po-Han Huang (NVIDIA) <53919306+nvpohanh@users.noreply.github.com> Co-authored-by: Po-Han Huang <pohanh@nvidia.com>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai>
Co-authored-by: Mohammad Angkad <mohammad.angkad@radixark.ai> Co-authored-by: Xinyuan Tong <xinyuantong.cs@gmail.com>
…under PP x speculative decoding (#40001)
ruff-format/isort on the two merge-touched files that the project-version static checks still flagged; no behavior change.
gpu_service_validation is now explicitly marked as the historical f8a6962 result only, with its numbers unchanged and no claim over the Phase-1 commit or this merge. The Phase-1 service evidence moves to its own block carrying source_commit bdb935d and its evidence path, marked as externally reported. README/UPSTREAM/PROVENANCE no longer describe the 172b1b4 record as a rebase-era baseline. They state the verifiable facts: the fork's actual common ancestor is 32290dd, 172b1b4 is 52 upstream commits behind it, and those 52 commits are already integrated through 2f06478, so counts taken from the old record overstate the missing work.
…undle Production startup on the merged tree failed in QSAHiSparseRuntime.__init__: ModelRunner no longer carries tp_rank after the parallel-context migration, so `runner.tp_rank` raised AttributeError on both TP ranks. Read the rank through the same published context upstream uses (get_parallel().tp_rank) in the runtime and in QSAHiSparseSingleRequest; no getattr default and no compatibility field on ModelRunner. The local audit of QSA construction sites kept runner.token_to_kv_pool, token_to_kv_pool_allocator, req_to_token_pool, server_args, model_config and dtype (all still present) and server_args.tp_size/pp_size/max_running_requests/ cuda_graph_backend_decode (still present); tp_rank was the only removed field. Regression: test/qsa_hisparse/test_parallel_rank_migration.py builds real runtime/single-request objects with a runner that either lacks tp_rank (the production condition) or still carries a stale 0, while the published bundle reports rank 1; it asserts rank 1 and the rank-1 event path. The pre-existing single-request constructor test now publishes the rank it previously passed through the removed runner field.
Second production startup failure on the merged tree, both TP ranks: ModelRunner.init_attention_backends built QSAHiSparseCoordinator with `self.tp_group.cpu_group`, but ModelRunner no longer carries tp_group after the parallel-context migration. Take it from the published bundle (get_parallel().tp_group.cpu_group), the same source upstream uses; the coordinator keeps the TP CPU-group contract (all_reduce/all_gather over TP). Bounded audit of our downstream code for parallel fields removed upstream (added lines only, 46 changed python files): the other references are get_parallel().attn_tp_size/attn_dcp_size in the backend and the kv-cache configurator, Scheduler.tp_cpu_group (assigned from get_parallel().tp_group), runner.server_args.tp_size (still present on ServerArgs), and the runtime and single-request rank helpers fixed earlier. model_runner.py:1049 was the only remaining removed field. Regression test/qsa_hisparse/test_attention_backend_group_migration.py executes the real init_attention_backends body with the backend/coordinator setup mocked: a runner without tp_group (production condition) and a runner with a stale group distinct from the live context; only the live group may reach the coordinator, and the coordinator is skipped for non-lease pools.
test_block_indices_expand_on_attention_fallback built the backend with __new__, so the decode path's `self.req_to_token_pool.req_to_token` (and the metadata row-request lookup) raised AttributeError once the upstream fused-KV test file entered the GPU suite. Construct the backend fully and provide the request table plus row_req_pool_indices the path reads. Assertions and production behavior are unchanged; the oracle still compares the expanded indices against expand_qsa_block_indices.
…vice review Docs-only. VALIDATION.md gains the 2026-10-04 section: merge commit 80dc48d over the real ancestor 32290dd, frozen tested source 2fe0731, CPU 221/7/24 in both environments, the independent 4090 SM89 GPU suite (240 passed, 1 SM121 skip, 0 failures), the service window (7/7 frozen cases, B1..8 36 requests with 65536 cached tokens, abort/recovery/reseed/health, graph evidence) and the confirmed restore (run-summary.json: candidate_passed=true, production_restored=true). PROVENANCE replaces the CPU-only pending status with a scoped phase2_independent_review block plus the latest-main recheck (affa261e, 4 commits, no QSA runtime or dependency overlap, not merged), keeps the historical f8a6962 and Phase-1 bdb935d blocks separate, and records the retained startup/harness failures and the explicit not-covered list. UPSTREAM.md notes why the four later main commits are not chased.
Docs-only follow-up to 7a4a0ea: - UPSTREAM.md no longer says the GPU/service review is pending; it points at the 2026-10-04 frozen-source section of VALIDATION.md and keeps the source distinction (validated 2fe0731, later doc-only commits). - VALIDATION/PROVENANCE now record both earlier GPU windows precisely: window 1 (results/dsh-maintenance-20261004/phase2-final-gpu-tests.txt) 68 failed, 172 passed, 1 skipped from a PATH without ninja, insufficient free GPU memory while the production service ran, and the uninitialized fixture; window 2 (lab/results/.../phase2-group/gpu-preflight.stdout/.xml) 34 failed, 206 passed, 1 skipped, of which 33 came from the harness inheriting the production SGLANG_QSA_HISPARSE_V3=p2-offload mode plus the same single fixture. The fixture is a test-fixture migration gap, not a product runtime API bug, and window 2 is a corrected environment run of the same window rather than a second environment. - PROVENANCE.validation.test_environment now carries an explicit legacy/historical scope so the old Torch 2.13 / xgrammar 0.2.1 values are not read as the current GPU environment; the current environments live in phase2_cpu and phase2_independent_review. - Both documents state that evidence paths are relative to the maintenance workspace root /home/zyk/projects/interests/ai-video/qwen, not to this checkout, and that the raw artifacts are not committed.
The two earlier GPU runs are distinct execution windows; the point that needed stating is that the single fixture gap was one root cause, not a separate round. Remove the sentence claiming window 2 was a re-run of the same window from VALIDATION.md and from the PROVENANCE history_retained entry; both now end at the 68/34 counts, their root causes and the final 240-pass result.
…rt window The frozen-source window (2fe0731) ended with production restored and no packaging. Add the later window as its own record instead of merging the two: 897286b, which differs from the tested runtime only in PROVENANCE.json, UPSTREAM.md and VALIDATION.md, was promoted to the local 8081 service, matches its launch snapshot, reports health 200 and keeps its previous profile for rollback. Three standard-local DSH 0.1.7-rc.2 sessions then ran with session effort max dispatched as wire reasoning_effort=xhigh: 13 primary requests, all HTTP 200 with real reasoning, no fallback route and no relay error. Pelican passed with 10 requests and two self-directed edits inside the same first turn, the image case passed 5/5 with 660 image tokens and a byte-equal decoded payload, and the IMO 1988 Problem 6 proof passed an independently predeclared, line-by-line review. The retained non-pass records stay visible: the pre-model MISSING_CREDENTIAL route failure, and the three 64-token title requests (two length-limited, one client-cancelled BrokenPipe) that are not counted as passes. Limits keep the window honest: no statistical or performance certification, a known competition problem rather than an unseen one, local promotion only, still OBSERVE=light, and identity by Git commits and byte comparison rather than file hashes. Evidence paths stay relative to the maintenance workspace and raw artifacts, session logs and credentials are not committed.
…al wording Five corrections inside the later-window record, no code, test, environment or service change: - remove a stray backtick in the 13-primary-request sentence; - state the environment transition accurately: the promotion did not upgrade the previous production environment in place, it switched the service to the already-validated upstream-runtime-env runtime path while weights and serving configuration stayed as validated; - relate the two windows by their tested Git revisions 2fe0731 and its documentation-only successor 897286b, separated in time, instead of claiming they share no snapshot; - narrow the credential claim: the isolated test instance neither copied nor modified the global DSH credential store, and MCP control authentication was loaded by the standard wrapper/environment mechanism without being disclosed; - say that the lab evidence is not included in or committed to this fork and is kept under the separate lab's own local evidence Git commit aa33f74, whose raw artifacts are not published. PROVENANCE.json carries the same boundaries in validation.local_production_promotion and its shared raw-artifacts note.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Integrate SGLang main at
35f3c96ff4794a4de15daf12caad371084a037eewhile keeping the QSA raw-KV offload and private host-prefix cache working with the new cache lifecycle and parallel context. Non-unit FP8 KV scales previously produced incorrect cached-prefix/prefill and reference reads; the removed runner TP fields also prevented model startup after the upstream merge.Modifications
80dc48ddfcover the verified common ancestor32290dda2c(561 upstream commits).Accuracy Tests
The tested runtime is
2fe0731e03; production snapshot897286b12aand this PR's final commits differ from it only in validation/provenance documentation.max, actual wirexhigh, with real reasoning output: pelican SVG independently rendered/reviewed; image 5/5 with byte-equal image payload; IMO 1988 Problem 6 proof independently reviewed. 13 primary model requests, all HTTP 200, no fallback or relay errors. The pelican refined its output within the same first turn; the math problem is known rather than claimed unseen.Full validation and preserved failure history: VALIDATION.md and PROVENANCE.json. Raw local artifacts, DSH configuration, logs and credentials are not included in this fork.
Speed Tests and Profiling
Functional CUDA graph replay was observed for every B1..8 batch in the fixed service window. No throughput, latency improvement or performance certification is claimed.
Scope and limitations
OBSERVE=light; full-model strict bitwise/determinism was not validated. DSHmaxis model reasoning effort.affa261e3dsnapshot are not included.Checklist
CI States
Latest PR Test (Base): ❌ Run #37217134865
Latest PR Test (Extra): ❌ Run #37217134654
Latest PR Test (AMD ROCm 10): ❌ Run #37217134818
GitHub checks and merge basis
python/sglang/kernels/ops/gemm/hc_mix.py. This PR does not clean up the repository-wide formatting debt.Missing required label 'run-ci'; other gates likewise require their trigger labels. These results are not counted as executed hardware tests or passing CI. NPU/nightly and image-build jobs that remain pending add no validation coverage.