Skip to content

Eval bug: probabilistic CUDA error 'invalid argument' at kernel launch on 7x Volta (sm_70) layer-split, triggered by large prefill + partial prefix reuse #29255

Description

@yudagua

Name and Version

Running llama.cpp through LM Studio runtime packs, b10760 through b11026 (runtime 2.31.2 - 2.41.0). All of them crash for me. I have Windows Error Reporting records pinning each crash to the exact engine binary via PE timestamp, so this is not a stale-process artifact.

Operating systems

Windows

GGML backends

CUDA

Hardware

7x Volta sm_70 (1x Quadro GV100 + 6x Tesla PG503-216, 32 GB HBM2 each), --split-mode layer --tensor-split 0 --main-gpu 0, CUDA graphs disabled via GGML_CUDA_DISABLE_GRAPHS=1.

Models

Reproduces across architectures (Qwen3.8-Flash-Next GGUF and DeepSeek-V4-Flash variants, various quantizations), so I don't think it's model-specific. UBatches 256 and 384 both crash at the same rate on the Qwen model.

Problem description & steps to reproduce

During normal long agent sessions (OpenAI-compatible server, one client, very long contexts with heavy prefix reuse) llama-server dies roughly daily — recently ~6 times a day — with a fail-fast (0xc0000409 in ucrtbase.dll). The last engine line before death is always:

CUDA error: invalid argument
current device: 0, in function ggml_cuda_kernel_launch at ggml/src/ggml-cuda/common.cuh:<line varies by build: 1676/1710/1716>

The trigger pattern is consistent across three weeks of logs (~100 crashes): a large prefill (tens of thousands of tokens, up to ~112K context) with high partial prefix-cache reuse (f_keep 0.98-0.99), dying either right at the prefill→decode boundary (progress = 1.00) or mid-prefill around a batch boundary. It is probabilistic: replaying the same session mostly works, and a synthetic append-only stress run (growing prefix, ~3K new tokens per round) passed 40/40 rounds once and crashed in round 2 another time. Note it does NOT strictly require >100K context — I also have samples at 56K, 18K, and even a fresh slot crashing on its first 2-message prefill; deep context just raises the probability.

What I've ruled out by testing, not guessing:

  • Engine version. Clean reinstall of runtime 2.33 still crashes (~1/day); 2.34 through 2.41 crash more often (6-10/day). ggml : remove GGML_CUDA_PEER_MAX_BATCH_SIZE #28177 and CUDA: Allow concurrent streams per split for multi-GPU #28198 (merged Sep 3/4) look like amplifiers on the multi-GPU launch path, not the root cause.
  • KV quantization (q8_0 K/V): survived 9.5h / ~200K new prefill tokens once, then reproduced the exact same signature. Reduces frequency at best.
  • GGML_CUDA_DISABLE_GRAPHS=1: verified present in the process PEB environment. Helps on ≤2.33 builds; newer builds have launch paths that bypass it and crash anyway.
  • --no-kv-unified, split strategy (spread vs priority-order fill): no effect.

Since both cudaLaunchKernelEx and the classic launch path reject the same configuration, I suspect a grid/block dimension goes out of range for some batch shape on sm_70 — similar in spirit to #27901 (gridDim.y overflow), but here it seems to need multi-GPU layer split plus partial slot reuse to show up. Volta is presumably not covered by any CI runner, which might be why this survives.

I could not reduce this to a clean llama-completion repro yet because the failure needs the prefix-cache reuse pattern of a long agent session; I'm happy to run bisects or test patches against my workload — I have crash logs with exact timestamps, token counts and f_keep values for every occurrence. Related but distinct signatures: #28661 (fattn-mma-f16.cuh:1945, non-standard head dim), #27871/#27901 (GB10 gridDim.y).

First bad commit: unknown — every runtime from b10760 to b11026 crashes; the frequency jump around Sep 3/4 suggests #28177 / #28198 amplified an older bug rather than introducing it.

First Bad Commit

No response

Relevant log output

Logs

Activity

  1. EnlistedGhost commented on Sep 25, 2026

    @EnlistedGhost

    I'm so sorry to intrude here @yudagua but have you tried the patch I showed? I don't know if it would benefit you in any way.

    My apologies if I've over stepped, and thank you to the llama.cpp team and community for making this project possible!

  2. yudagua commented on Sep 29, 2026

    @yudagua
    Author

    New data points from the same 7x Volta box (runtime 2.47.0 / b11189, driver 582.78, 2026-09-29), same long prompt each time:

    1. DeepSeek-V4-Flash-Vision (MXFP4) crashes at ctx=65536 (previously only 261888/262144 tested). The same ~19.2K-token prefill dies at the identical point: progress = 1.00 -> first decode step, invalid argument @ ggml_cuda_kernel_launch (common.cuh), WER ucrtbase.dll 0xc0000409. This rules out a ctx-dependent gridDim.y overflow as the trigger for this signature (at ctx 65536 the block count is ~2x under the 65535 limit).

    2. MiMo-V2.6-Flash-RL (MXFP4) crashes mid-prefill at ~8.3K tokens (~40% progress). No sparse indexer (hybrid SWA/GA MoE). Same WER fingerprint (ucrtbase.dll 0xc0000409, identical offset), and notably no CUDA error line in the log before death -- fail-fast before any launch is reported.

    Affected architectures are now three distinct MoE designs (qwen4exp, dsv4, mimo_v2); the only common factors left are large MoE + 7-way layer split + sm_70. A dense 27B model on the same box runs the same prompt repeatedly with zero crashes.

  3. yudagua commented on Sep 29, 2026

    @yudagua
    Author

    @EnlistedGhost no need to apologize -- I appreciate the follow-up. I did look into it: your patch (PR #28640) targets the FA path for non-standard head sizes (#28661) -- it falls back to cuBLAS when the Q head dim is not 64/128/256. My crashes here are a different signature: invalid argument at the generic ggml_cuda_kernel_launch, reproduced with flash attention both on and off (see my #29562, deterministic at a fixed prefill token position across 4 runtimes). The models involved all use standard head sizes (128/192/256), so the fallback condition in your patch would never trigger. Keeping it here for anyone else on SM70 hitting the FA/non-standard-head variant -- it may well be the right fix for #28661.

  4. yudagua commented on Sep 30, 2026

    @yudagua
    Author

    Correction to my earlier comment (09-29): the "CPU offload immunity" observation is retracted. New data from the same box, runtime 2.47.0 / b11189, driver 582.78, 2026-09-30:

    I restricted LM Studio to 4x PG503 (128 GB) so that strictVramCap forces part of the model onto CPU, then re-ran the same long prompt from the LM chat window for each model. Both crash at the exact same token position as their all-GPU runs:

    Model Config GPU layers Crash point (4-GPU + CPU overflow) Crash point (6-GPU all-GPU, earlier data)
    Qwen3.8-Flash-Next Q8_0 (qwen4exp) ctx 261888, parallel 1, kv unified 27/48 (log: "adjusted from 'max' to '27'") n_tokens = 7,633, progress = 0.39 identical (see #29562)
    DeepSeek-V4-Flash-Vision MXFP4 ctx 261888, parallel 1, kv unified 29 (log: "adjusted from 'max' to '29'") n_tokens = 19,179, progress = 1.00 (prefill->decode boundary) identical (my 09-29 comment)

    LM Studio log evidence for the offload state (main.log, at load time):

    [2026-09-30 10:09:38] GPU offload layers for model 'mradermacher_qwen3.8-flash-next-uncensored' was adjusted from 'max' to '27' to respect the strict GPU VRAM cap.
    Resolved GPU config options:
      Num Offload Layers: 27 / Disabled GPUs: [6,5,4]   (4x PG503 active)
    
    [2026-09-30 12:07:29] GPU offload layers for model 'deepseek-v4-flash-vision-exp-uncensored' was adjusted from 'max' to '29' ...
    Resolved GPU config options:
      Num Offload Layers: 29 / Disabled GPUs: [6,4,5]   (4x PG503 active)
    

    Both crashes carry the same signature as before: invalid argument at ggml_cuda_kernel_launch (common.cuh), WER ucrtbase.dll 0xc0000409.

    Revised conclusion: the crash is a deterministic function of (prompt, model). GPU count, layer split ratio, and CPU offload are all ruled out as variables - with 21 layers on CPU the qwen4exp run still dies at token 7633, and DSV4-Vision still dies at 19179. My earlier observation that CPU participation prevents the crash was not reproduced in these runs and is retracted.

    MiMo V2.6 (hybrid SWA/GA) is expected to behave identically; not re-tested separately since the two data points above already close the loop.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions