Skip to content

feat: fmt=8 (fp8-e4m3) decode on the kv_b absorb path, CPU and CUDA - #1102

Merged
JustVugg merged 8 commits into
JustVugg:devfrom
monotophic:f8/absorb-fmt8
Sep 23, 2026
Merged

JustVugg merged 8 commits into
JustVugg:devfrom
monotophic:f8/absorb-fmt8

Conversation

@monotophic

@monotophic monotophic commented Aug 19, 2026 •

Copy link
Copy Markdown
Contributor

Authored by Fable 5.1 in Claude Code, analysis in partnership with @monotophic

Context. This PR is one of four independent contributions derived from a single locally-verified working tree (fp8 container support, scoring-evidence tooling, server API hardening, and batched group scoring). This one carries the fp8 container line: the fmt=8 absorb decode path, the mint tool's tests, and their end-to-end test. It stands alone — nothing else in the set needs to land for this to be complete, and merging or declining the others does not affect it. The others will be proposed separately, each with its own evidence.

Checkpoint-faithful FP8 containers stamp kv_b_proj as fmt=8 (raw e4m3 bytes + one f32 scale per 128×128 block). The absorb decode path had no fmt=8 branch: CPU qt_addrow/qt_matvec_rows refused loudly, and the CUDA absorb gate refused before any kernel ran — an FP8 container could load but not decode attention. This adds the fmt=8 branch to both backends, mirroring matmul_fp8's block-scale indexing exactly, plus an explicit named skip (with the absorb path called out) where kv_b GPU sharding legitimately cannot serve fmt=8. The block geometry (FP8_BLOCK, fp8_nblk) moves to a shared header, fp8_format.h, so the CPU branch, the CUDA kernels, and the tests read one definition site.

The dense fmt=8 matmul path already exists on dev; this PR's own new content is the kv_b/absorb decode arms (CPU and CUDA), the shared block-geometry header, and their tests. Platform coverage at this head: CPU is native for both dense and absorb fmt=8. CUDA is native for both and is device-proven end-to-end on GB10: the two-slot batched-serve CUDA-absorb arm passed under CUDA_DENSE=1 + COLI_CUDA_ATTN=1, with the engine's own boot line [CUDA] mode: routed experts + resident dense tensors captured as a positive witness (kv_b CUDA-eligible, absorb path selected); the same test fails, on the same hardware, when that boot line reads resident dense on CPU. Metal has a real, unit-tested fmt=8 decode kernel that no live dispatch path calls, so every fmt=8 matmul on a Metal build continues to fall back to CPU by existing dispatch design — unchanged by this PR; Vulkan has no fmt=8 arm and continues to refuse it at upload, also unchanged. Two other engines in this tree (glm53.c, kimi_k3.c) share the identical kv_b_proj-absorb attention shape and are attachable to the same feature but are not wired to it — a separate follow-on, not a gap here.

Behavioral contract

  • fmt=8 kv_b decodes correctly against the quant.h reference on both backends: on CPU, qt_addrow bit-exact and qt_matvec_rows within float-accumulation-order tolerance; on CUDA, within 1e-3 relative on the GB10 oracle, with the absorb kernels' per-element scaling documented in backend_cuda.cu as the accepted RDNA2 (gfx1030) field report: AMD backend works; greedy decode not token-stable across mixed CPU/GPU expert tiers #510 divergence class. Every other fmt's behavior is unchanged.
  • Unsupported fmts still refuse loudly on both backends (admission widened by exactly {8}).
  • layer_cuda_shard_kvb refuses un-shardable kv_b formats BY NAME (notice + skip; the absorb path serves them) instead of proceeding on a NULL pointer by accident — and the same allowlist closes two pre-existing silent misreads (fmt=5 and fmt=6 computed a wrong row-byte stride with a non-NULL q4).
  • The CUDA fmt=8 LUT-ready flag is cleared on shutdown, and a re-init that names a different device set is refused before any context is rebuilt (dev's own f0f8dfc4), so the device set cannot widen past the last publish without passing through shutdown; the process-wide-flag-vs-per-device-table hole is closed by that pair. The two decisions are pure predicates in backend_cuda.h, called at the real sites, and a host-side test pins the lifecycle with no GPU: shutdown clears, a same-set re-init keeps, a different-set re-init is refused.
  • Everything that worked before still works: the full gate set is green at this head, and the branch passes make check standing alone.

Scope note @JustVugg — observability deliberately left out. During verification we used a temporary log line to prove the absorb branch executes on the serve path, then removed it to keep this change minimal; the end-to-end test now captures the engine's own boot-mode line instead. A permanent, env-gated absorb-path debug line is a reasonable follow-on if you'd find it useful — happy to add it to this PR or a follow-up at your request.

Capstone matrix (evidence as originally re-run on 2026-09-03; the branch was since re-derived onto dev 9e152d49 as ea8a7e4d, and the evidence re-run at that head — per-commit gates, the HIP lane, the CPU decode-identity check on two containers, and the CUDA arms on a GB10 — is in the comment of 22 September)

claim decisive evidence
CPU decode matches the reference validator-written, tree-independent OCP-E4M3FN decoder, 6 disjoint shapes at block boundaries and tails: qt_addrow bit-exact; qt_matvec_rows ≈6e-8 relative (float accumulation order)
CUDA decode matches the reference within the documented tolerance. make cuda-test on GB10: the fmt=8 absorb-kernel oracle, the fmt trap test, and the LUT-lifecycle pin all pass (oa-lane/results/spark1/cells/cuda_test/)
CPU/CUDA indexing agree static host-side sizing probe, 324 shapes × 7 invariants + a 4,097-ordinal row_bytes sweep, with a negative control (wrong ng formula fails as expected)
old head could not serve fmt=8 old binary vs new test on both a CPU host and a CUDA host: BITE CONFIRMED: qt_addrow: unsupported fmt=8 … (oa-lane/results/{strix1,spark1}/cells/cell_old/)
the tests would catch index corruption old test binary vs new colibri.c: 2 FAILED at exactly the fmt=8 refusal pins, fmt=6 pins unchanged; the CUDA sizing probe's negative control fails on a wrong block-count formula
the shard-path hole is closed CUDA=1 shard-refuse guard run on GB10: layer_cuda_shard_kvb refusal probe: ok (oa-lane/results/spark1/cells/shard_kvb_refuse/)
fmt=8 absorb serves end-to-end on CUDA two-slot batched serve on GB10, both slots 200, engine survives, boot witness routed experts + resident dense tensors captured; the negative twin (CPU-dense boot line) fails the same test on the same box (oa-lane/results/spark1/cells/cell_new_cuda_absorb/, …/negative_polarity_ce501dd/). The witness proves CUDA eligibility and path selection; kernel execution follows from the source-verified dispatch chain (no silent-fallback branch between selection and launch) rather than a separate artifact
fmt=8 LUT cannot go stale across re-init test_cuda_lut_gate (host-side, no CUDA toolchain) pins all three edges — shutdown clears, same-set re-init keeps, different-set re-init refused — and the same edges ran on a GB10 in test_fp8_cuda's lifecycle phase
HIP lane sees no new warning HIP syntax gate: warning identity 1090 = 1090 vs the dev baseline (run at the 2026-09-03 head; the delta to that push was Python tests and docs only); re-run at ea8a7e4d, where the lane now also builds and runs the shard-refuse test against the CUDA backend — see the comment of 22 September
Metal non-regression make METAL=1: zero warnings at every commit in the series; metal-test ok under both COLI_METAL_RESSET states
Fuller matrix, review record, and origin accounting

Review record: four bisectable commits (shared header → CPU engine → CUDA engine → tests), each built from make clean and passing the full suite standing alone; each commit reviewed by an independent blind validator and a deep auditor, with fix rounds closed at primary source; the assembled program then passed a program-head deep audit and a final pre-push verification. One review finding is worth naming: an earlier version of the CUDA-absorb end-to-end test went green without ever reaching the CUDA absorb path (the eligibility knob was stripped by the test's own env hygiene). The test now injects the knob, asserts the engine's boot-mode line, and echoes what it observed on pass and fail, so that class of vacuous green cannot recur.

Origin accounting for a re-derive onto current dev: the decode branches and CUDA arms are byte-identical or offset-only carries of the previously reviewed content; two pre-existing defects found during re-derivation are fixed here and disclosed above (the fmt=5/6 shard-allowlist misreads, and a missing fp8_format.h dependency in Makefile.deepseek-v4 that left five DeepSeek-V4 objects stale against a block-size edit); the LUT re-init hole an earlier revision closed on the init side is now closed by dev's own init refusal together with the shutdown-side reset this PR keeps, pinned by the host-side lifecycle test and by the GB10 one; a few stale line-number anchors in docs/FORMATS.md and the repack tool's docstring were corrected while those files were open. The end-to-end test and the refusal canary diverge from the earlier hardened lineage by exactly the reviewed hardening described above.

The mint→load regression test is armed only where torch/numpy/safetensors are installed (it names its skip otherwise); upstream CI does not install torch, so it SKIPs there and runs on our fleet. Compiled-binary digests for the CUDA build varied run-to-run on identical sources (nvcc non-determinism; the gcc build was byte-stable) — disclosed as a reproducibility note, not a behavior.

The fmt=8/fmt=6 scale-byte VRAM accounting interaction disclosed in the original submission was fixed separately in #1100 (merged 2026-08-19), which is in this PR's base; the diagnostic counters are correct here without further change.

Style: changed lines were held to the file's measured local idiom via a diff-scoped consistency check; no surrounding code was reformatted.

Durable vs current-state: decode branches, refusals, and the LUT gate are durable; GB10/sm_121 timings, skip counts, and warning counts are current-state (2026-09-03 at base 387653f; re-derived 2026-09-22 onto base 9e152d49 as ea8a7e4d, evidence in the comment of 22 September).

@monotophic

Copy link
Copy Markdown
Contributor Author

Rebased on 0282193.

More coming to build out support for the FP8 checkpoint-faithful container and full logprobs instrumentation support when using the OpenAI API.

@monotophic

Copy link
Copy Markdown
Contributor Author

Rebased onto dev @ 12b0fa1 (30 commits past the previous base 387653f); head 77d8939 → 0282193, force-with-lease.

What changed, by origin

  • Dev-motion-caused: c/Makefile only — upstream's header-prerequisite work (build: list the headers each engine includes as Makefile prerequisites #1284) and later rules expanded the
    prerequisite lists of colibri$(EXE), cuda-dll, hip-dll and backend_cuda.o; this branch's own addition
    (fp8_format.h in each of those four lists) is re-applied on top of the expanded lists. c/backend_metal.mm merged
    automatically against fix(metal): bit-exact fp8-e4m3 decode via ieee754 exponent bitcast #1346 (bit-exact fp8-e4m3 decode); this branch's one-line comment fix there is unchanged.
  • Pre-existing defects found by re-review: none.
  • Adjacent improvements: none.
  • Contribution content: byte-identical to 77d8939 on every file except the Makefile lines above (diff-of-diffs 0 lines
    over the branch's other 19 files).

Re-verified on 0282193

check result
make check (serial) exit 0 — 743 Python tests OK (43 skipped), all C test binaries OK
c/tests/test_makefile_deps.py (new upstream gate) 2 passed, 7 subtests passed
lint on changed lines, per commit (4 commits) no findings
preflight (preflight_push.sh) against a fresh fetch GO (base unmoved)
CI on the pushed head all passed (operator report)

Preflight's advisory about POSIX-only test constructs refers to c/tests/test_qt_addrow.c (fork/pipe/waitpid at
lines 359–370); that block sits inside #ifndef _WIN32 … #else … #endif (lines 358–390), so the Windows CI lane skips
it by construction — consistent with the green Windows run.

Authored by Claude Fable 5.1 in Claude Code, analysis in partnership with @monotophic

@monotophic

monotophic commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor Author

Authored by Claude Fable 5.1 in Claude Code, analysis in partnership with @monotophic

Rebased onto current dev (87c54a93). No functional change on this branch's side — the 4-commit chain is preserved unsquashed with the same messages.

Files needing conflict resolution, all caused by dev motion since the last rebase:

  • c/quant.h: dev added an AVX2 e4m3/bf16 decode block right where this PR removes FP8_BLOCK/fp8_nblk. Resolved by keeping the whole AVX2 block and moving only the two definition lines out to fp8_format.h, as before. matmul_fp8 and the AVX2 helpers are untouched; FP8_BLOCK/fp8_nblk are each defined exactly once.
  • c/Makefile: dev expanded the qwen38, glm53 and cuda-test prerequisite lists, added four rules that name quant.h and did not exist when this PR was written (deepseek_v41, tests/test_e4m3_vector, tests/test_qwen38_tier_engine, tests/bench_router_select), and (941d5fe) added omp_tune.h to the deepseek_v41, qwen38, and glm53 rules. Resolved by keeping dev's full lists and pairing fp8_format.h right after quant.h wherever a rule names quant.h — since quant.h now includes fp8_format.h, a rule without the pairing is exactly the stale-binary case this PR's Makefile change exists to prevent. tests/test_makefile_deps.py passes on the result.
    Everything else auto-merged with this PR's hunks verbatim (diff against the new base identical to the diff against the old base on the other 19 files).

Re-verified on 619d51fa: make check exit 0 (983 tests, no failures), every one of the 4 commits builds, METAL=1 build with zero warnings, metal-test under both COLI_METAL_RESSET states, the HIP syntax lane replica (rocm 6.2 container), and the branch's fp8 test modules. Lint on changed lines: no findings.

@JustVugg: Checking in on the viability of this PR. Is there interest in merging the fmt=8 absorb decode path? The format itself is already in via the mint tooling; this PR closes the one path that still refuses it. I have a personal interest in having the option to preserve the published checkpoint quality for the resident spine for GLM 5.2/5.3 and also a full checkpoint quality container for a quant degradation benchmark. The quant performance/cost experiments are also the driver for my associated PRs that build out proper instrumentation for the OpenAI API endpoint.

I've got everything I need integrated and working at my end so my container experiments are not blocked by this stuff sitting unmerged, but I would prefer not to be investing in an isolated fork and rebasing regularly to avoid drifting too far from what you are doing here. Steering input is welcome, my hope is that these contributions would be generally useful for anyone wanting to use Colibri for serious work, including model tuning/development and research.

@monotophic

Copy link
Copy Markdown
Contributor Author

Authored by Claude Fable 5.1 in Claude Code, analysis in partnership with @monotophic

Field evidence for the CPU fmt=8 absorb path in this PR, measured on real fmt=8 containers.

Identity gate (2026-09-07, this PR's build 0282193e; spark GB10 class, CPU only, temperature 0, MTP off): direct decode with the absorb path on (ABSORB=1) versus off (ABSORB=0), same binary, same five raw prompts, 64 new tokens each.

container format identical extracted bytes per prompt (on / off)
we-v1 fp8 e4m3 experts, int3 g64 5/5 325/325 · 377/377 · 345/345 · 312/312 · 297/297
wq-v1 fp8 e4m3 experts, int4 g64 5/5 370/370 · 379/379 · 353/353 · 301/301 · 314/314

Control: the dev binary of the day (e1efc687) under the identical ABSORB=1 command on we-v1 exits 1 with qt_addrow: unsupported fmt=8 for the per-row-scale absorb path — so the ABSORB=1 arm reaches the absorb surface, and this PR is what makes it pass. Two-slot serving through openai_server.py on both containers completed 5/5 requests with zero refusals.

Repeated 2026-09-15 with this PR integrated alongside #1353 and #1355–#1357 on dev edd4f48d: we-v1 on aarch64 (spark GB10) 5/5 identical with the same byte counts as above; wq-v1 on x86_64 (Strix Halo, AVX-512 absorb specialisation) 5/5 identical; dev control refused on both. A serve battery on the same integrated build (two KV slots, so the absorb path is forced) returned text identical to this PR's standalone build on the default completion, chat and Anthropic-messages requests.

Not measured: CUDA runtime behaviour of the fmt=8 absorb kernels (no CUDA host in these runs; the kernels are verified by compilation and the HIP syntax lane only). Serving rate on the CPU absorb path was ~0.31 tok/s on a GB10 for we-v1, so these are identity measurements, not throughput ones.

Records: F-5 gate colibri_lab/dispatch/2026-09-07-f5-gate/results/ and integration cells colibri_lab/dispatch/2026-09-15-integ-cells/results/ (available on request).

@monotophic

Copy link
Copy Markdown
Contributor Author

Authored by Claude Fable 5.1 in Claude Code, analysis in partnership with @monotophic

Rebased onto current dev (b5968eb6). No functional change on this branch's side — the 4-commit chain is preserved unsquashed with the same messages.

One file needed conflict resolution, caused by dev motion since the last rebase: c/Makefile, where #1558 added kv_prefix.h to the deepseek_v41 prerequisite list on the same line this PR adds fp8_format.h. Resolved by keeping both. Every other file merged with this PR's hunks verbatim — the diff against the new base is identical to the diff against the old base except for that one line.

Re-verified on f25ac455: make check exit 0 (1002 tests, no failures), every one of the 4 commits builds, METAL=1 build with zero warnings, metal-test under both COLI_METAL_RESSET states, the HIP syntax lane replica (rocm 6.2 container), the branch's fp8 test modules, and tests/test_makefile_deps.py. Lint on changed lines: no findings.

@JustVugg

Copy link
Copy Markdown
Owner

Reviewed, and this is close. The CPU arm mirrors matmul_fp8 exactly, same block index, same scale stride, same partial-tail handling, and test_qt_addrow is the standout of the set: an independent per-element reference decoder rather than the block-batched shape, a partial-column tail with a stated negative control, direct parity against the reference, a NaN pin, and byte-identity re-checks for the other six formats so the guard's placement is pinned too. That is real coverage.

One change before I merge, and it is wording rather than code. The behavioural-contract bullet says fmt=8 kv_b decodes bit-consistently with the reference on both backends. The CUDA side does not, and your own comment in backend_cuda.cu says so plainly: the fp8 accumulation convention is not carried into the absorb kernels and the divergence is an accepted documented class. The capstone matrix is honest about this; the one-line summary is not. Please make the summary match the comment.

Worth knowing for sequencing rather than for you to fix: the roughly seventy lines of Makefile dependency churn will conflict with most other open PRs touching that file, so this wants to land either before them or well after.

@monotophic

Copy link
Copy Markdown
Contributor Author

Authored by Claude Fable 5.1 in Claude Code, analysis in partnership with @monotophic

Agreed, and fixed in the PR body — wording only, no push.

  • The behavioral-contract bullet now states the two backends separately: CPU qt_addrow bit-exact and qt_matvec_rows within float-accumulation-order tolerance; CUDA within 1e-3 relative on the GB10 oracle, with the absorb kernels' per-element scaling named as the accepted RDNA2 (gfx1030) field report: AMD backend works; greedy decode not token-stable across mixed CPU/GPU expert tiers #510 divergence class, matching the backend_cuda.cu comment.
  • While there, the two capstone labels that said "exact" ("CPU decode is exact", "CUDA decode is exact") were the same summary-vs-evidence gap one tier down, so they now read "matches the reference" / "matches the reference within the documented tolerance". The evidence cells are unchanged.

Head is still f25ac455; nothing in the branch changed.

@JustVugg: Understood on the Makefile sequencing. Knowing that this is getting review is sufficient to keep me focused on making sure this is ready to merge when it fits your broader priorities. Thanks!

@monotophic

Copy link
Copy Markdown
Contributor Author

Heads-up on an interaction with #1356, which is in review now: #1356 and this PR both add entries to c/Makefile and will conflict there — only there. We merged them in both orders in a scratch worktree to check: the one conflicted path is c/Makefile in each direction, and c/colibri.c merges cleanly in both orders even though both PRs modify it. So it is the ordinary consequence of two PRs each adding build targets, not an engine-level collision.

Whichever lands second, we re-merge dev and push the resolution within a day. Nothing is needed from you here.

One related note, if this PR lands first, quant.h gains fp8_format.h, and #1356's two new Makefile rules should list it too, exactly as their sibling rules do — a one-line follow-up in whichever order the two land.

Authored by Opus 5 in Claude Code, analysis in partnership with @Monotophic

The CUDA backend needs the same 128-column block edge and block count the
CPU decoder uses. Sharing them through quant.h is not possible: that header
pulls in the whole CPU kernel set, which backend_cuda.cu must not compile.
Move the two definitions into a minimal header both sides include, so there
is one definition of the fp8 block geometry rather than a constant repeated
per backend.

Every prerequisite list whose translation unit reaches the new header gains
it beside quant.h, in both build files. One rule, tests/test_qwen36_dnproj_batch,
names quant.h but is left alone: qwen36.c does not include quant.h, so that
translation unit never reaches fp8_format.h and listing it would declare a
dependency that does not exist.

That claim is now enforced rather than asserted. The existing Makefile
prerequisite test walks only an engine's own #include lines, so a header
reached THROUGH another one is invisible to it -- which is exactly
fp8_format.h's shape, since quant.h includes it and no engine .c does. A new
case closes the include graph for this header: drop fp8_format.h from a rule
whose engine reaches it and the test names that rule. It is scoped to this
header deliberately; the same closure over every header reports seven rules
that predate this change, which the case documents as a separate fix rather
than failing on arrival for reasons this change did not cause.

Pointers to the old location (the fmt=8 entry in docs/FORMATS.md, comments in
colibri.c and backend_metal.mm) now name fp8_format.h. FORMATS.md's source
references drop their line numbers and cite symbols instead: those anchors
went stale twice, and a stale line number reads as a verification that was
not performed.
qt_addrow and qt_matvec_rows had no fmt=8 case, so an fp8-e4m3 kv_b_proj fell
through to the per-row-scale tail, read t->s[row] past the end of a
block-scale array and dereferenced a NULL q4. Both now decode fp8 directly,
taking the block scale from [ceil(O/128), ceil(I/128)] exactly as matmul_fp8
does.

qt_matvec_rows' arm mirrors that kernel statement for statement, including
the f32-partial-per-block, double-across-blocks accumulation. qt_addrow is an
axpy, so it mirrors the block-scale indexing only and folds the coefficient
into the scale, the convention its fmt=4 arm has always used.

A NaN block scale is refused by name, with the block that carried it: it
poisons the whole block and every accumulator downstream of it, and these
functions cannot repair it. A ZERO block scale is not refused -- it is valid
data. Block scales are amax/448, so an all-zero block legitimately carries
one, and decoding it as zeros is the correct answer; unlike on the GPU there
is no table here that could be unwritten, because the CPU decoder reads e4m3
from a compile-time constant. The check costs one branch per block, not per
element, and both properties are pinned: a NaN scale must refuse through
either function, a zero scale must decode to zeros through both.

The format guard below the new arms keeps its condition; only its message
changes, to say that fmt=8 now returns above it.

layer_cuda_shard_kvb gains a format allowlist. fmt 5/6/7 previously computed
an int4 row stride against a group-scaled layout, took a non-NULL q4 and
uploaded garbage silently; they are now refused by name before any pointer or
stride is used, with a notice bounded to one line per format per process.
weight_at and absorb_scale gain fmt=8 branches so the attention absorb path
decodes fp8 on the GPU, with the block-scale geometry of the dense fmt=8
matmul and of matmul_fp8. The surviving 128 literals are pinned to the shared
constant by a file-scope assertion.

The e4m3 table is per-device while the flag that gates fmt=8 uploads is
process-wide, so the two must not drift. coli_cuda_shutdown clears the flag;
coli_cuda_init never writes it. That is sufficient, and this states the
mechanism rather than asserting the conclusion: init will not rebuild contexts
underneath a live device set. A re-init naming the same set returns success
and leaves the contexts -- and the table published to them -- untouched; one
naming a different set is refused before anything is rebuilt. So the set
cannot widen past what the last publish covered without passing through
shutdown.

The two decisions that gate rests on are factored into pure predicates in
backend_cuda.h, which the backend calls at the real sites rather than keeping
a second copy. tests/test_cuda_lut_gate.c pins them, and the resulting state
machine, on a plain CPU build with no GPU and no CUDA toolchain: previously
nothing about this gate ran anywhere except on a CUDA host, which is how
three sites in the tree came to assert init semantics that no longer held.
The lifecycle test in tests/test_fp8_cuda.cu now asserts all three edges,
including that a same-set re-init keeps the flag -- requiring a republish
there would force one that buys nothing.

absorb_scale's own contract comment named fmt=4 as the only format without a
per-row scale; it now names fmt=8's per-block layout as the other exception,
since this change adds that branch inside the same function.
…ession

test_fp8_serve_batch_e2e.py pins the fmt=8 kv_b batched-serve decode end to
end against a real container, and is a named SKIP when none is present. Its
child environment is allowlist-scrubbed: ambient engine knobs are stripped,
read-only model-location settings pass, documented legacy aliases are stripped
alongside their primaries, and the ratified non-absorb arm passes
value-restricted. The CUDA absorb path is opt-in: the test INJECTS the exact
engine bundle that path needs rather than letting it arrive ambiently, then
witnesses the engine's boot report and fails outright on the
resident-dense-on-CPU line, so a misconfigured lane cannot bank a vacuous
green. That witness caught a first attempt which value-allowed the bundle
through the scrub instead of injecting it.

test_fp8_refusal_canary.py runs in CI without a container: it pins that each
absorb function keeps at least one refusal matching the end-to-end matcher --
existential on purpose, since an additional differently-worded guard is new
coverage rather than drift -- and that the matcher stays selective against
both a synthetic near-miss and the real in-tree sibling.

test_e8x4g64_mint_load.py and test_e8x4g64_loader.c pin the container class
through the real conversion CLI and the real C loader at toy scale, plus the
duplicate-tensor-name refusal on a tool-produced container. The C harness
takes a container directory on argv, so it is excluded from the gate list and
driven by the Python test; it gains a Makefile rule anyway, so that it
compiles under the suite's own flags instead of only the driver's hand-copied
list, which is what its no-warnings assertion had been vouching for. The
driver no longer swallows a failed OpenMP probe on macOS: silently compiling a
different binary from the one the Makefile builds is what made that assertion
misleading.

The external-files table in the repack tool now cites symbols instead of line
numbers. Those anchors had rotted twice while the table claimed to be verified
against the current tree, and a stale line number reads as a verification that
was not performed.
@monotophic

Copy link
Copy Markdown
Contributor Author

Rebased onto dev 9e152d49. Two files needed resolution: c/Makefile, in its prerequisite lines, and c/backend_cuda.cu, in two places — one of which changed a decision in this PR.

c/Makefile: the conflicting regions are all the same shape, a prerequisite list both sides edited. Resolving them is mechanical; checking them is the part worth stating. Strip every fp8_format.h back out of the resolved file and what remains differs from dev's Makefile only by what this PR adds on purpose: three new test rules, and one entry in the exclude list for the harness among them that takes an argument. Nothing dev put there was dropped — oracle.h, which #1605 added to 43 of these prerequisite lines while this was in review, is on all 43 at the head.

c/backend_cuda.cu, first place: f0f8dfc4 removed the g_nctx = 0; line that this PR's init-side g_fp8_lut_ready = 0; was anchored to, and closed the hazard that reset existed for — init({0}) → set_lut → init({0,1}) now returns early with "device list change requires shutdown first" instead of rebuilding contexts.

I dropped the init-side reset. Keeping it would have left a comment that is false on the merged tree.

The shutdown-side reset stays, and its comment now names f0f8dfc4's refusal as the mechanism. What it rests on: g_nctx is written in only two places, the one that clears it is immediately followed by the LUT reset with no early return between them, and a device absent from the context list is refused before the upload guard is reached — so the device set can only widen by passing through shutdown.

Until now that argument lived only in comments, and when f0f8dfc4 changed what re-initialization does, two of them went on describing the old behaviour — the note in backend_cuda.h and the CUDA lifecycle test. Both are corrected here. The two decisions behind the gate are now pure predicates that backend_cuda.cu calls at the real sites, and a new test pins them and the resulting state machine on a plain CPU build, with no GPU and no CUDA toolchain. The lifecycle test asserts all three edges, including that a same-set re-init keeps the flag — requiring a republish there would force one that buys nothing.

Second place: a65ea155's if (fmt == 7) block sits directly above this PR's fmt == 8 rewrite. Both kept, unchanged. #1676's SiTU work on the same file merged without conflict.

On the matmul_fp8 comparison, worth restating since #1313 split that kernel in two: the scale row (row/128), the stride (ceil(I/128)) and the handling of a partial block along I are the same in qt_addrow and in both forms. The four-row form also has a tail along O, for when fewer than four rows are left; qt_addrow decodes one row at a time, so it has no equivalent and needs none. #1313 changed row blocking and accumulator interleaving, not scale geometry — so the comparison holds, it only wants naming the tail it means. No code change there.

Two things did change in the CPU arms on re-reading them: a NaN block scale is now refused by name, with the block it came from, instead of propagating through the accumulator; and a zero block scale is explicitly kept valid and decoded as zeros, because block scales are amax/448 and an all-zero block genuinely has one of zero.

Gates: make check 1368 tests, OK (skipped=119), and the same on each of the four commits separately rather than only on the last one. test_qt_addrow, test_cuda_fmt_guard and the new LUT-gate test pass, as do the three Python tests. Diff-scoped lint is clean and merge-tree against 9e152d49 is clean. The HIP syntax lane passes, and it now also builds test_shard_kvb_refuse with CUDA=1 and runs it: that test needs a GPU-capable toolchain but no GPU, so the whole file was a SKIP everywhere it was run before and now executes the guard it was written for.

CPU decode-identity check, re-run on the rebased binary (direct ABSORB=1 against ABSORB=0, same binary, temperature 0, with the dev binary's refusal as the control): 5/5 identical on both containers, and the dev binary refused both.

CUDA, on a GB10 with CUDA 13: the cuda-test suite builds and every test in it passes except test_alloc_footprint_cuda, which fails on this device when built from dev's own tree too — the same three live-measurement checks, with different numbers each run, and one further check that our build happens to pass — so it is this device's allocator against that test's model, not this change. The same five prompts decoded through the CUDA absorb path give bytes identical to the CPU absorb path and to the CPU non-absorb path, 5/5, with the GPU busy only on the CUDA arm; the dev binary under the same CUDA environment exits with the format refusal by name and no CUDA error. The PR's own end-to-end serve test ran against the fp8 container on that box and passed.

One thing is yours to call. c/Makefile moved seven times in the hours this took. #1605 put oracle.h on 43 prerequisite lines, 38 of them lines this PR also edits, and #1676 added to the cuda-test list this PR edits too, so the diff against dev goes stale about as fast as it is made. Do you want it now, or parked until the other Makefile-touching PRs are through?

Authored by Claude Opus 5 in Claude Code, analysis in partnership with @monotophic

…g test comment

Five accuracy items from the r7 audit (AUDIT_1102_r7_ea8a7e4d.md, findings
F1-F5), none touching the decode path.

docs/FORMATS.md's Metal fused-gate paragraph (F1) still named a colibri.c
line number for each of five symbols and cited a non-upstream branch
(kvb/fmt-gate-notice-r4) as its source of truth, even though the "Sources
for all rows" table just below it had already been de-anchored. All five are
now symbol references in the same style as that table, and the branch name
is gone.

repack_fp8_passthrough.py's EXTERNAL-files table (F2) de-anchored its line
numbers everywhere the preamble above it promises, except one: the
Engine.__init__ reference still carried a line number (previously wrong per
r6 F4, and edited to a different wrong number in r7). It is now a bare
symbol reference, matching the rest of the block.

tests/test_e8x4g64_loader$(EXE) (F3) had a rule and a comment claiming it
compiles under the suite's own CFLAGS so its warnings are caught, but the
rule was unreachable from every target -- test-c, test, check, and all --
since TEST_EXCLUDE drops it from TEST_BINS and nothing else names it as a
prerequisite. Added e8x4g64-loader-check, a phony build-only target shaped
like qwen38-tier-engine-check, so `make e8x4g64-loader-check` actually
builds it under those flags on demand. check itself is untouched: the new
target is not a prerequisite of check/test/test-c/all, so this adds no time
to `make check`.

test_qt_addrow.c's block-scale comment (F4) said a ZERO fmt=8 scale is
"refused by name" alongside NaN. It is not: colibri.c is explicit that a
zero scale is valid and must decode to zeros, and
test_fmt8_zero_scale_decodes_to_zeros twenty lines below already asserts
exactly that. The comment now describes only the NaN guard the two calls
below it exercise, and points at that assertion for the zero case.

fp8_format.h (F5) is added beside quant.h/backend_cuda.cu on two of the
three Makefile rules the audit named: tests/test_kimi_cuda_expert$(EXE)
(kimi_k3.c includes quant.h directly, which pulls in fp8_format.h) and
tests/bench_cuda_resident_batch$(EXE) (backend_cuda.cu -- a separate
translation unit built into the same binary -- includes fp8_format.h
directly, a miss no quant.h-based framing could see). The third,
tests/test_qwen36_dnproj_batch$(EXE), is deliberately left alone: full
preprocessing under both its build variants (plain and the
-DCOLI_DNPROJ_REAL_CUDA alternate) confirms qwen36.c never reaches quant.h
or fp8_format.h at all, exactly as 6826c00's own commit message already
documented when it introduced fp8_format.h. The audit's framing of this
rule as a regression does not hold up under that check; the pre-existing
quant.h entry on that rule is itself already a no-op and is not this
change's to fix.

Not in scope: the fmt=8 NaN exit(1)-inside-omp policy, disclosed in the r7
thread note and left for a later round.
@monotophic

Copy link
Copy Markdown
Contributor Author

Pushed two small commits on top of the previous head. A review pass I ran after the last push found five accuracy problems in the documentation and build rules, none in code paths: docs/FORMATS.md still carried line-number anchors, a branch name from my fork, and preambles promising line numbers it no longer gives, the repack tool's docstring cited a stale line number, the test_e8x4g64_loader rule was reachable from no target while its comment said its warnings were caught (it now builds via make e8x4g64-loader-check, still outside check), a comment in test_qt_addrow.c said the opposite of what its assertions check, and two rules that reach fp8_format.h did not list it. No test logic or product code changed; the test-file edit is comment-only.

One thing worth naming rather than changing: the NaN block-scale refusal exits the process from inside the OpenMP region, where matmul_fp8 would let the NaN propagate. That is deliberate, since a poisoned accumulator is a wrong answer that looks like a right one, but it is a policy choice and I would rather you rule on it than find it later.

@JustVugg

Copy link
Copy Markdown
Owner

Merged current dev into your branch: the only conflicts were Makefile prerequisite lines (dev added serve_budget.h in #1712 and the build-flags stamp in Makefile.deepseek-v4 in #1707), resolved as the union with fp8_format.h kept. On the merge, colibri builds without warnings and test_qt_addrow, test_shard_kvb_refuse, test_cuda_lut_gate, test_cuda_fmt_guard, test_makefile_deps and the V4 build-flags test pass. The summary now matches the CUDA comment, thank you. On the NaN block-scale refusal: accepted as policy, a poisoned scale on a format that was refused outright until now should stop loudly. Merging once CI is green; it goes into 1.12.1.

@JustVugg
JustVugg merged commit 1d7defa into JustVugg:dev Sep 23, 2026
29 checks passed
JustVugg pushed a commit to cameron/colibri that referenced this pull request Sep 23, 2026
…ot.h and the new CPU-only test rules next to this branch's sse41_kernels.h and SSE4.1 gates
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants