Skip to content

fix(deepseek-v4): the MoE step says why it failed, instead of "failed in MoE" - #1468

Merged
JustVugg merged 1 commit into
devfrom
fix/v4-moe-says-why
Sep 13, 2026
Merged

JustVugg merged 1 commit into
devfrom
fix/v4-moe-says-why

Conversation

@JustVugg

Copy link
Copy Markdown
Owner

For #1464.

hybrid batched block failed in MoE is the generic tail of the layer block. moe_token_pipeline and v4_moe_batch_union can fail in some thirty places (an expert that would not read, an upload that would not land, the GPU expert group refusing, a routing table missing, a plain malloc) and every one returned a bare -1, so the block could not tell an out-of-VRAM card from a bad read over /mnt/c, and neither could the reporter after a day of flags.

Every failure site now records a reason (layer, expert where there is one, the step) through moe_fail(), thread-local and cleared on entry, and both block tails append it:

hybrid batched block failed in MoE: layer 12: the GPU expert group of 6 experts refused (a CUDA allocation or launch failed; check nvidia-smi for free VRAM)
block computation failed in MoE: layer 3: reading expert 137 from the expert store failed

No behaviour change on any success path: only the text in the error buffer differs. All of it lives in the COLI_V4_UNIT_BLOCK_HYBRID unit, which needed its own <stdarg.h>.

Gate: tests/test_v4_moe_reason_source.py pins the shape (no silent -1 left in the two functions except the job markers, reason cleared on entry, both tails read it); it fails on the tree before this commit. The tiny V4 oracle stays token-exact: target long, CLI, serve, prefix reuse.

Not part of 1.11.0 (#1467 is already open); this ships in the next release.

… in MoE"

#1464: an RTX 3080 under WSL2 failed every prompt with

    request failed: hybrid batched block failed in MoE

and a day of flag-flipping could not narrow it, because that message is the
generic tail of the layer block. moe_token_pipeline and v4_moe_batch_union
can fail in some thirty places -- an expert that would not read, an upload
that would not land, the GPU expert group refusing, a routing table missing,
a plain malloc -- and every one of them returned a bare -1. The block could
not tell an out-of-VRAM card from a bad read over /mnt/c, so neither could
the reporter.

Every failure site now records a reason (layer, expert where there is one,
and the step) through moe_fail(), thread-local and cleared on entry so a
stale one cannot outlive the call that set it, and both block tails append
it:

    hybrid batched block failed in MoE: layer 12: the GPU expert group of 6
    experts refused (a CUDA allocation or launch failed; check nvidia-smi
    for free VRAM)

    block computation failed in MoE: layer 3: reading expert 137 from the
    expert store failed

No behaviour changes on any success path; the only difference is the text
in the error buffer. Everything lives in the COLI_V4_UNIT_BLOCK_HYBRID unit,
which needed <stdarg.h> of its own since set_error's unit is a different
object.

tests/test_v4_moe_reason_source.py pins the shape: the only `result = -1`
left in those two functions are the "not finished yet" job markers, both
functions clear the reason on entry, and both block tails read it. It fails
on the tree before this commit. The tiny V4 oracle stays token-exact
(target long, CLI, serve, prefix reuse).
@JustVugg
JustVugg merged commit 7d3e30c into dev Sep 13, 2026
28 checks passed

@skhari1989-cmyk skhari1989-cmyk left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for the support and the resolution.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants