Repository navigation
Eval bug: probabilistic CUDA error 'invalid argument' at kernel launch on 7x Volta (sm_70) layer-split, triggered by large prefill + partial prefix reuse #29255
Description
Activity
I'm so sorry to intrude here @yudagua but have you tried the patch I showed? I don't know if it would benefit you in any way.
My apologies if I've over stepped, and thank you to the llama.cpp team and community for making this project possible!
New data points from the same 7x Volta box (runtime 2.47.0 / b11189, driver 582.78, 2026-09-29), same long prompt each time:
-
DeepSeek-V4-Flash-Vision (MXFP4) crashes at ctx=65536 (previously only 261888/262144 tested). The same ~19.2K-token prefill dies at the identical point:
progress = 1.00-> first decode step,invalid argument@ ggml_cuda_kernel_launch (common.cuh), WER ucrtbase.dll 0xc0000409. This rules out a ctx-dependent gridDim.y overflow as the trigger for this signature (at ctx 65536 the block count is ~2x under the 65535 limit). -
MiMo-V2.6-Flash-RL (MXFP4) crashes mid-prefill at ~8.3K tokens (~40% progress). No sparse indexer (hybrid SWA/GA MoE). Same WER fingerprint (ucrtbase.dll 0xc0000409, identical offset), and notably no CUDA error line in the log before death -- fail-fast before any launch is reported.
Affected architectures are now three distinct MoE designs (qwen4exp, dsv4, mimo_v2); the only common factors left are large MoE + 7-way layer split + sm_70. A dense 27B model on the same box runs the same prompt repeatedly with zero crashes.
-
@EnlistedGhost no need to apologize -- I appreciate the follow-up. I did look into it: your patch (PR #28640) targets the FA path for non-standard head sizes (#28661) -- it falls back to cuBLAS when the Q head dim is not 64/128/256. My crashes here are a different signature:
invalid argumentat the genericggml_cuda_kernel_launch, reproduced with flash attention both on and off (see my #29562, deterministic at a fixed prefill token position across 4 runtimes). The models involved all use standard head sizes (128/192/256), so the fallback condition in your patch would never trigger. Keeping it here for anyone else on SM70 hitting the FA/non-standard-head variant -- it may well be the right fix for #28661.Correction to my earlier comment (09-29): the "CPU offload immunity" observation is retracted. New data from the same box, runtime 2.47.0 / b11189, driver 582.78, 2026-09-30:
I restricted LM Studio to 4x PG503 (128 GB) so that strictVramCap forces part of the model onto CPU, then re-ran the same long prompt from the LM chat window for each model. Both crash at the exact same token position as their all-GPU runs:
Model Config GPU layers Crash point (4-GPU + CPU overflow) Crash point (6-GPU all-GPU, earlier data) Qwen3.8-Flash-Next Q8_0 (qwen4exp) ctx 261888, parallel 1, kv unified 27/48 (log: "adjusted from 'max' to '27'") n_tokens = 7,633, progress = 0.39 identical (see #29562) DeepSeek-V4-Flash-Vision MXFP4 ctx 261888, parallel 1, kv unified 29 (log: "adjusted from 'max' to '29'") n_tokens = 19,179, progress = 1.00 (prefill->decode boundary) identical (my 09-29 comment) LM Studio log evidence for the offload state (main.log, at load time):
[2026-09-30 10:09:38] GPU offload layers for model 'mradermacher_qwen3.8-flash-next-uncensored' was adjusted from 'max' to '27' to respect the strict GPU VRAM cap. Resolved GPU config options: Num Offload Layers: 27 / Disabled GPUs: [6,5,4] (4x PG503 active) [2026-09-30 12:07:29] GPU offload layers for model 'deepseek-v4-flash-vision-exp-uncensored' was adjusted from 'max' to '29' ... Resolved GPU config options: Num Offload Layers: 29 / Disabled GPUs: [6,4,5] (4x PG503 active)Both crashes carry the same signature as before:
invalid argumentatggml_cuda_kernel_launch(common.cuh), WER ucrtbase.dll 0xc0000409.Revised conclusion: the crash is a deterministic function of (prompt, model). GPU count, layer split ratio, and CPU offload are all ruled out as variables - with 21 layers on CPU the qwen4exp run still dies at token 7633, and DSV4-Vision still dies at 19179. My earlier observation that CPU participation prevents the crash was not reproduced in these runs and is retracted.
MiMo V2.6 (hybrid SWA/GA) is expected to behave identically; not re-tested separately since the two data points above already close the loop.
Name and Version
Running llama.cpp through LM Studio runtime packs, b10760 through b11026 (runtime 2.31.2 - 2.41.0). All of them crash for me. I have Windows Error Reporting records pinning each crash to the exact engine binary via PE timestamp, so this is not a stale-process artifact.
Operating systems
Windows
GGML backends
CUDA
Hardware
7x Volta sm_70 (1x Quadro GV100 + 6x Tesla PG503-216, 32 GB HBM2 each),
--split-mode layer --tensor-split 0 --main-gpu 0, CUDA graphs disabled viaGGML_CUDA_DISABLE_GRAPHS=1.Models
Reproduces across architectures (Qwen3.8-Flash-Next GGUF and DeepSeek-V4-Flash variants, various quantizations), so I don't think it's model-specific. UBatches 256 and 384 both crash at the same rate on the Qwen model.
Problem description & steps to reproduce
During normal long agent sessions (OpenAI-compatible server, one client, very long contexts with heavy prefix reuse) llama-server dies roughly daily — recently ~6 times a day — with a fail-fast (
0xc0000409in ucrtbase.dll). The last engine line before death is always:The trigger pattern is consistent across three weeks of logs (~100 crashes): a large prefill (tens of thousands of tokens, up to ~112K context) with high partial prefix-cache reuse (
f_keep0.98-0.99), dying either right at the prefill→decode boundary (progress = 1.00) or mid-prefill around a batch boundary. It is probabilistic: replaying the same session mostly works, and a synthetic append-only stress run (growing prefix, ~3K new tokens per round) passed 40/40 rounds once and crashed in round 2 another time. Note it does NOT strictly require >100K context — I also have samples at 56K, 18K, and even a fresh slot crashing on its first 2-message prefill; deep context just raises the probability.What I've ruled out by testing, not guessing:
q8_0K/V): survived 9.5h / ~200K new prefill tokens once, then reproduced the exact same signature. Reduces frequency at best.GGML_CUDA_DISABLE_GRAPHS=1: verified present in the process PEB environment. Helps on ≤2.33 builds; newer builds have launch paths that bypass it and crash anyway.--no-kv-unified, split strategy (spread vs priority-order fill): no effect.Since both
cudaLaunchKernelExand the classic launch path reject the same configuration, I suspect a grid/block dimension goes out of range for some batch shape on sm_70 — similar in spirit to #27901 (gridDim.y overflow), but here it seems to need multi-GPU layer split plus partial slot reuse to show up. Volta is presumably not covered by any CI runner, which might be why this survives.I could not reduce this to a clean
llama-completionrepro yet because the failure needs the prefix-cache reuse pattern of a long agent session; I'm happy to run bisects or test patches against my workload — I have crash logs with exact timestamps, token counts and f_keep values for every occurrence. Related but distinct signatures: #28661 (fattn-mma-f16.cuh:1945, non-standard head dim), #27871/#27901 (GB10 gridDim.y).First bad commit: unknown — every runtime from b10760 to b11026 crashes; the frequency jump around Sep 3/4 suggests #28177 / #28198 amplified an older bug rather than introducing it.
First Bad Commit
No response
Relevant log output
Logs