Repository navigation
wip/tensor-split-fixes: page faults and graph churn under -sm tensor + MTP (issue #105) - #106
briansp2020 wants to merge 1 commit into
Conversation
|
A heads-up before you spend time on this: I found a crash caused by 0002 (the data-pointer graph key), in plain (Claude-assisted, as before.) |
|
Thanks for bearing with me on this one, and apologies: my earlier comment was wrong. The crash I attributed to 0002 turned out to come from our own environment, not from either patch. We run the server with
One more correction: the build I called "stock r12 + 0002" was actually r12 + 0001 + 0002 plus two inactive env switches of ours. So 0002 can come back off the withdrawn list if it's still useful to you; the numbers in the PR description stand ( (Claude-assisted, as before.) |
9d50c52 to
00446c5
Compare
|
Thanks for r15 — the block-06 Meta host-buft fallback is what made I've updated this PR for r15 (rebased onto
With all five, On TODO #41, flagging in case it's useful: these runs keep all experts in VRAM, so the staging ring never activates and we don't hit it. If a second box would help, I'm happy to run your repro here (on our GSQ IQ3_XXS model) with No rush on any of this. (Claude-assisted, as before.) |
…sm tensor + MTP (issue stew675#105) Nine format-patches on top of v16-a55e952b8-r17 (tree 04764deb) and a README with cause, fix, switches and validation: Q8_1 arena retention, data-pointer graph key, recapture after captured memory is freed, split-state cache versions, LLAMA_KV_N_PAD_MIN, unpinned block_out in prefill, compact split-state cache entries, direct conv-state tail copy, planar HC_MIX output. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
00446c5 to
e7de090
Compare
|
Thanks again for r16, and for the TODO #43 write-up - the per-device guard was a neat find. One more round on this PR, flagging it in case it's useful: rebased onto r17 (no conflicts; tested on r16), plus four more small patches (0006-0009), all on the same
Separately, a configuration note rather than a patch: under With all nine patches and p-min 0, No rush on any of it - happy to split these out, rework them, or re-run anything here. (Claude-assisted, as before.) |
This packages the
-sm tensorfixes from #105 as awip/item in the same form as #104, so you can review, re-validate and fold them in however suits you. Updated for r17: rebased ontoc862072(no conflicts), now nine patches; tested on r16.Contents (
wip/tensor-split-fixes/, nothing else touched): nine format-patches on top ofv16-a55e952b8-r17(applied tree04764deb) that apply withgit amafterpatches/00*.patch, plus a README with cause, fix and validation for each.quantize_q8_1page fault)GGML_CUDA_Q8_1_ARENA_FREE_OLD=1GGML_CUDA_GRAPH_KEY_NO_DATA=1GGML_CUDA_GRAPH_MEM_GEN=0GGML_META_SS_VERIFY=1checks hitsLLAMA_KV_N_PAD_MINraises the n_kv padding floor (opt-in, default unchanged)block_outas a prefill graph output (~1.9 GiB of the compute buffer at-ub 2048); the fused combine+norm reads a pool copy if neededLLAMA_HC_PIN_BLOCK_OUT=1GGML_META_SS_VERIFY=1checks hitscont)LLAMA_CONV_TAIL_CONT=1ggml_hc_mix_set_planar(): planar HC_MIX output, noggml_contof the mixed head in the verify band (BF16 CUDA only)LLAMA_HC_MIX_PLANAR=0Results on r16 (2 × R9700, qwen4exp GSQ-RCO IQ3_XXS + MTP n-max 3 p-min 0, all experts in VRAM, 256K,
-ub 2048, stock ROCm 10.0, one session):-sm tensorwith the nine patches vs our production-sm layer: greedy 113.9-114.5 vs 94.6-96.6 t/s, chat 98.7 vs 90.1, 2 concurrent 142-147 vs 122-124, prefill 37k / 155k 2428-2447 / 2022 vs 2032-2054 / 1709, 259.6k prompt prefill / decode 1754 / 47.2 vs 1473 / 30.6. Same as r15 within noise; not re-run on r17.-ub.--spec-draft-p-min 0(configuration, no patch) is part of the tensor result: a constant verify width lets the graph be reused; details in the README.GGML_HIP_GRAPH_FORCE_UPDATE=1or the TheRock nightly that avoids the ROCm 10.0hipGraphExecUpdateleak (hipGraph: kernarg staging never reclaims slots within an exec's lifetime — request slot reuse on exec update ROCm/rocm-systems#10713); with them it is ~6-7 % faster on brand-new prompts, details in the README.The README lists what I didn't measure: 1-GPU and non-RDNA4 layouts, more than 2 GPUs, and host-resident experts with these patches.
Happy to re-run anything on this box or rework the patches if you'd rather they land differently. No rush on our side.
(Claude-assisted, as before.)