Qwen3.8-Flash-Next: placement study, and it lands on CT 123's single card - #71
Conversation
…t sweep Now that the EPYC box has 251 GiB of RAM, the first model here that does not fit in VRAM becomes reachable: Qwen3.8-Flash-Next (qwen4exp, 180B total / 6B active), UD-Q4_K_XL at 111.33 GB, across both cards with the PLE n-gram table and a tunable share of the routed experts in system RAM. CT 123 is stopped to free GPU 2. CT 120 grows to 48 cores / 160 GiB / swap 0 and its /models disk to 320G. Everything is reversible: ct120-cutover.sh to-qwen36 detaches GPU 2, restores the limits and brings qwen3.6-35b-a3b back, with both GGUFs and both llama.cpp builds left on disk. pro-v620/qwen38-flash-next/ install.sh idempotent push of env + serve script + unit into CT 120 qwen38fn.env placement config; MODEL_CPU_MOE is the VRAM/RAM dial llamacpp-serve-qwen38fn two loud guards: N-GPU count, and all shards present qwen38fn-download.sh resumable, sha256-verified 112 GB pull ct120-cutover.sh to-qwen38fn / to-qwen36 / status placement-sweep.sh round-robin sweep of --n-cpu-moe x context depth placement-probe.py decode/prefill by prompt class, + template-contract asserts summarize-sweep.py renders the VRAM-freed-versus-decode-lost trade thermal-guard.sh EPYC-correct, fails closed on a missing sensor Findings that correct the sizing note and CLAUDE.md: - qwen4exp is already in b10678, so "exactly one build short / needs b10679+" was wrong and the model was runnable here before this work. The pin still moves to b11018, for three merged qwen4exp follow-ups that post-date b10678 (#27880, #27941, #28023), scoped to this model's LLAMACPP_DIR so qwen3.6 is untouched. - Multi-GPU is the PENALISED path for this architecture: llama.cpp #28699 measured the QSA indexer's pooled rows crossing inter-GPU links every layer at 2x decode cost on a layer split, and that fix is an open draft. Hence ONE_GPU=true exists as a first-class control. - Decode degrades with context depth for the same reason, a cost no bandwidth arithmetic models, so the sweep has a depth axis and numbers are quoted with it. - The reasoning contract is inverted from Qwen3.8-27B: the template raises on "none" and "high", and unset means xhigh, which never answers. Handled server side with --reasoning off and asserted by the probe. - The PLE offload tensor is per_layer_token_embd; the note's ngram_embedding alternative is the safetensors name and matches nothing in a GGUF. - The thermal watchdog's GPU_SERVICE_MAP sent a GPU-2 trip to 123:llama-swap, a no-op now that CT 123 is stopped, which would have left the real load on an overheating card. The cutover sets and restores it. - gpu-ab-bench/thermal-guard.sh and sample-gpus.py still name the B550 PCI addresses, so the guard dies on its first read and fails silently open. Noted, and replaced for this platform rather than patched in place. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…direction All sixteen flags the serve script passes are confirmed present in b11018's --help. Two things the check turned up: - --n-cpu-moe keeps the experts of the FIRST N layers on the CPU, not the last. The comment and README said the opposite. - --reasoning-effort exists as a SERVER flag (minimal|low|medium|high|xhigh|max), so there is a fallback if --reasoning off ever misbehaves with this template. Wired as MODEL_REASONING_EFFORT, empty by default, with a note never to set it to "none" or "high" — the two values this template raises on. Also pins MODEL_EXPECTED_GPUS and EXTRA_ARGS in the env file rather than leaving them to the serve script's defaults, since placement-sweep.sh writes both. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…on Vulkan Corrects the claim made two commits ago. "qwen4exp is already in b10678, so the KB's one-build-short note is wrong" was only half true, and the wrong half is the load-bearing one. The arch STRING is in b10678. But Qwen3.8-Flash-Next uses hyper-connections (hc_count 4), and the Vulkan backend could not run them until qwen4exp: add hc ops (#28901) 2026-09-16 vulkan: support qwen4exp hc ops (#28988) 2026-09-17 04:34 UTC b11013 is the first build containing #28988; b11010 is not (verified with the compare API against the commit sha). So this model could not have run on this box's Vulkan stack at any point before today, and b10678 is NOT a rollback target for it — grepping the arch string on b10678 promises a model that cannot execute. Check that the BACKEND can run an arch, not that the arch is registered. Both provisioning scripts and CT 120's /opt/llamacpp/current now point at b11018. Prior builds stay in /opt/llamacpp, so rollback is still a symlink flip — just not past b11013 for this model. Other items in b10678..b11018 that matter here: - models : fix GDN normalization from max to rsqrt (#28068) is a CORRECTNESS fix for Qwen3.6-35B-A3B itself, which is a Gated-DeltaNet hybrid — not a neighbouring model. Verified after the bump: qwen3.6 serves clean output with no <think> leakage. - memory : avoid allocating V cache for indexer (#28330) frees VRAM on qwen4exp. - Vulkan: sparse Flash Attention (#28105), topk_moe prefill fusion (#28422), type-aligned GET_ROWS (#28253 — the op the PLE lookup uses), small-M matrix optimisations for qwen (#28457), argsort data race and OOB fix (#28705). - For CT 123 when it returns: spec fix for mtmd chunks with DFlash (#28587) and server fix for speculation after an image (#28715) both address the coder's exact DFlash2+mmproj configuration.⚠️ CT 123 is stopped and its container is still on b10678; the script pin moved, so a rebuild is correct, but a plain `pct start 123` brings back the old build. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…onfig On first load llama-server prints: tensor overrides to CPU are used with mmap enabled - consider using --load-mode none for better performance That is this model's whole arrangement: -ot pushes the ~28.7 GB PLE table to the CPU and --n-cpu-moe pushes expert layers there too. Under mmap those pages are FILE-BACKED, so the kernel can evict them and fault them back from the SATA 860 EVO — turning a bandwidth measurement into a disk measurement, which is precisely the failure the sizing note warned about. --load-mode none reads them into anonymous memory, which with swap 0 and MemorySwapMax=0 genuinely stays put. Exposed as MODEL_LOAD_MODE, left empty (auto) so the sweep can measure it rather than assume it. -lm also replaces the --mmap/--mlock/--dio flags deprecated in this same llama.cpp range (#28334), so it is the forward-looking spelling. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…s mandatory
It runs. Output is coherent, non-degenerate (8-gram ratio 1.0) and reproducible
at temperature 0. These look to be the first numbers for this architecture on a
Vulkan/RDNA2 target anywhere — upstream validated CPU and CUDA only, and the
Vulkan hyper-connection ops merged hours before this ran.
Three findings, in order of how much they change.
1. --tensor-split is REQUIRED with --n-cpu-moe, and nothing warns you.
--n-cpu-moe N moves the experts of the FIRST N layers to the CPU, so 0..N-1
are light and N..47 are heavy (~1.56 GB each). llama.cpp's default split
divides 48 layers evenly BY COUNT and therefore hands card 2 all the heavy
ones. At -ncmoe 20:
default split GPU1 13.4 GiB / 0.2 GTT GPU2 30.7 GiB / 9.3 GTT 6.6 t/s
-ts 34,14 GPU1 25.1 GiB / 14 MiB GPU2 21.2 GiB / 14 MiB 11.74 t/s
+78% decode and ~3x prefill from rebalancing alone. The cards were never short
of memory in total — 53 GiB of demand against 60 GiB of capacity. Pure
maldistribution. placement-sweep.sh now derives the split per ncmoe value
(card1 = N + (48-N)/2); a fixed split would be wrong for every other N.
2. Nothing is saturated, so the sizing note's method cannot predict this.
Measured 11.74 t/s against the note's 51. During decode the two cards ALTERNATE
(6-83% each) and the host CPU sits at 33-43%. Every token walks 20 CPU expert
layers, then card 1's, then card 2's, synchronising at each handoff — a latency
cost that is invisible to an "active bytes / bandwidth" model. Every
hybrid-placement figure in large-moe-build-shapes.md is an upper bound that has
now been falsified by ~4x, mine included.
3. --load-mode none is llama.cpp's own startup advice here and it is wrong:
11.18 t/s vs 11.74 on auto, and it pulls 31.6 GiB into GTT. Kept on auto.
Also measured:
- --reasoning off NEUTRALISES the template trap, opposite to what the template
alone implies. With it set, reasoning_effort of none/high/low/medium/xhigh all
return clean content, because the raising values never reach the template. So a
caller carrying CT 123's settings over cannot break this server.
placement-probe.py --contract asserted the REVERSE, which was wrong; fixed to
assert "a caller cannot break this" instead.
- Load: 2m38s cold off the SATA SSD, 40-46s warm. That bounds the two-profile
elasticity idea at ~45s per flip — fine at a role handoff, not per request.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The GPU is the only thermal gap on this board. gpu-blower-control's README already said the BMC "can only ever react to CPU, board and DIMM temps", but nothing recorded how to read those sensors or what normal looks like — so it read as a limitation rather than as coverage we can rely on. Now documented in CLAUDE.md next to the existing "the BMC has no GPU temperature sensor" note, and in gpu-blower-control/README.md: - `ipmitool sdr type Temperature` gives CPU Temp, MB Temp, Card Side Temp, Onboard LAN Temp, and one sensor per memory channel, TEMP_CPU1_DDR4A..H. - An unpopulated channel reports "No Reading", which is the easiest way to see which slots are filled without opening the case. dmidecode cannot do this: ASRock labels every slot "Locator: DIMM 0". - Baseline 2026-09-17 under the qwen4exp sweep, 4 of 8 channels populated (C, D, G, H): DIMMs 45-50 C, CPU 41 C, fans 1200-2400 RPM. RDIMMs throttle near 85 C and the BMC exposes no upper threshold on those sensors, so there is ~35 C of headroom. Expect +5-10 C when A/B/E/F are filled — they are currently acting as airflow gaps. - Therefore no CPU or RAM watchdog is needed, and none should be built. One inference worth keeping: flat DIMM temps are an independent check on whether a workload is really memory-bandwidth-bound. They did not budge while CPU utilisation swung 2% -> 50%, which corroborates the sweep's finding — from a different sensor entirely — that the hybrid qwen4exp placement is latency-bound rather than saturating the ~92 GB/s the sizing note assumes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
I recorded it as "two adjacent pairs, not the usual one-per-quadrant 4-DIMM guidance", which implied it was suboptimal. It is not: the placement was chosen from the board documentation and validated by the owner's own tests, and all four sticks run at Configured Memory Speed 3200 MT/s. Corrected in CLAUDE.md, gpu-blower-control/README.md and the bmc-sensor-coverage memory, and stated positively rather than just removed — so a future reader (or I) does not re-raise it from generic folklore and try to "fix" a working config. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
All three were in tooling I wrote today, and two of them produced wrong output rather than an error, which is the worse kind. 1. set_env_var wrote UNQUOTED values, so EXTRA_ARGS="--device Vulkan0" landed in the env file as `EXTRA_ARGS=--device Vulkan0`. llamacpp-serve-qwen38fn does `set -a; source ...`, which parses that as an assignment followed by a COMMAND — bash tried to run `Vulkan0`, the serve script exited 127, systemd crash-looped the unit, and the ONE_GPU control died on its first config. Rewritten to build the line with json.dumps inside the container, so quoting and escaping are not my problem; the python deliberately contains no single quotes so it survives being wrapped in them. sed needed & and | escaped plus another shell quoting layer, which is what broke it. 2. summarize-sweep.py still read the probe's OLD contract keys. I had changed the probe to record `harmless` instead of `raised`, so the summary rendered "reasoning_effort none rejected as expected: NO" and "empty <think> siphoned to reasoning_content: NO — check --reasoning-format auto" while the raw JSON said every value was harmless with clean content. A false failure in the one artifact a human actually reads. The contract section now explains what it is asserting (a caller cannot break this server, NOT that the template rejects anything) and checks the thing that matters: no <think> leaked into content. An absent reasoning_content is not a fault — it means no thought block was emitted, which is the point of --reasoning off. 3. At n_cpu_moe 48 the derived split is "48,0", so card 2 gets nothing and the row is effectively SINGLE-GPU while being labelled two-GPU. The placement itself is defensible — there are no heavy layers left to spread, and splitting the non-expert layers would only add an inter-GPU hop — but it must not be read as a two-card data point. Now logged loudly and recorded in .single_gpu_rows. Also: placement-sweep.sh now cd's to its own directory, because systemd-run and cron do not inherit a working directory and the helper scripts resolve relative to it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replaces the single 11.7 t/s data point with the completed matrix and the three conclusions that follow from it. - VRAM has almost no leverage past a point: -ncmoe 48 holding 9.2 GiB of VRAM MATCHES -ncmoe 34 holding 31.3 GiB (8.6 vs 8.4 at d0; 8.2 vs 7.9 at depth). 22 GiB buys nothing across that span. +48 GiB total is +51% at d0 but only +23% at depth 8000 — VRAM's value halves at realistic context depth, which matters because the sizing note prices cards on short-prompt arithmetic. - Reserving ~20 GB for a second model costs 11-14% (-ncmoe 20 -> 28, which frees 23.6 GiB — comfortably a 27B guest). That answers the question this work started from. - Cost per GiB is non-monotonic (0.04 / 0.12 / 0.11 / 0.06 t/s per GiB), so do not interpolate: VRAM is nearly free to give away below -ncmoe 20 and above 34. - Depth costs 5-22%, inversely to how much is offloaded — the more on CPU, the smaller a share the QSA indexer is. Consistent with #28699 being the depth cost. Also flags in-place that the -ncmoe 48 row is effectively single-GPU (derived split 48,0; GPU2 measured at 16 MiB) so it is not misread as a two-card point, and names -ncmoe 34 as the honest two-vs-one comparison. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds pro-v620/qwen38-flash-next/membench.sh (STREAM + a random pointer-chase latency probe) and records the 4-channel results. Re-run verbatim at 8 sticks; the value is that both stick counts come from one harness. STREAM on 4 channels, best at 8 threads: Copy 80.3, Scale 51.2, Add 56.2, Triad 56.3 GB/s. This settles the open question the opposite way to my expectation. STREAM reports APPLICATION bytes; Triad moves 24 B/iter in those terms but a normal store first reads the line it overwrites, so real DRAM traffic is 32 B/iter: STREAM Triad 56.3 app -> 75.1 GB/s DRAM traffic 73.3% of theoretical stressapptest 76.4 74.6% STREAM Copy 80.3 (non-temporal stores, no RFO) 78.4% Triad-corrected and stressapptest agree within 2%, so stressapptest was ACCURATE, not an understatement — I had claimed it understated and that real STREAM would reach 77-87. Copy's 80.3 is provably RFO-free (an RFO correction would put it above theoretical), so ~80 GB/s / 78.4% of peak is the real ceiling. The honest per-channel figure is 19.1-20.1 GB/s; the old 23.1 was 90.2% of theoretical and never reachable. Also: 8 threads beat 16 and 32, so four channels saturate early and the acceptance soak's 16 threads cost only 3%. That will invert at eight channels — the re-run must sweep threads. Latency, 4 GiB working set: 141.36 ns/load WITH huge pages, 226.95 without. THP here is `madvise`, not `always`, so the probe has to call madvise(MADV_HUGEPAGE) itself or a random walk pays a page-table walk per access — a 38% error, and my first run made exactly that mistake. Fixed in the harness. 141 ns is not attributable to 3DS without a flat 2Rx4 set on the same harness. On the 3DS question, researched against JEDEC JESD79-4 and it inverts the concern: the core latency chain (CL/tRCD/tRP/tRAS) is unchanged because each die is a standard DDR4 die at the same bin; the _slr/_dlr timing variants 3DS adds are generally more relaxed than same-rank equivalents; and on refresh the delivered part is FAVOURED. At 64 GB x4, 2S2Rx4 3DS implies 8 Gb dies (~4.5% refresh overhead) while the advertised flat 2Rx4 implies monolithic 16 Gb (~14.1%). Stacking avoids the penalty the ordered part would have carried. ⛔ The 8-stick test could not run: A/B/E/F are empty, confirmed by IPMI No Reading on those channels. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The governor finding (-52% from acpi-cpufreq `powersave` pinning 1500 MHz) invalidated every placement number measured so far, so the whole matrix has to be re-measured at full clock and then MTP and --parallel layered on top. That is more hours than a supervised session, so it runs as a chained host-side pipeline. overnight-all.sh lock + thermal-guard precondition + part 1 -> part 2 -> report overnight.sh 1 governor (pick winner, PERSIST it), 2 revalidate overnight-part2.sh 3 MTP, 4 --parallel, 5 graph splits, 6 restore revalidate.sh the -ncmoe curve, threads, and the three placement EXTREMES concurrency-probe.py per-stream vs aggregate throughput, slot count read back morning-report.sh auto-derived recommendation, so the night answers even unwatched stagelib.sh the stage runner, shared by both halves Four defects found and fixed while building it, each of which would have produced confident wrong numbers: * `timeout <shell function>` exits 127 instantly. Because every stage was wrapped to "never stall the night", all six stages reported failure and the pipeline completed in two seconds having measured nothing — the failure-isolation meant to protect the run is precisely what hid its collapse. stagelib.sh enforces the deadline with a watchdog process and its own process group (so pct/python descendants die too), and refuses a stage name that is not a defined function. * No knob reset between cells. MODEL_KV_TYPE=q8_0 and MODEL_MMPROJ_ON_CPU=true were found still set on the live server from an earlier extreme, and cell() never cleared them, so every row would have silently inherited them. reset_env() now puts every knob back to the production default and cells override only what they vary. (Third time a leaked knob has invalidated a cell here.) * The MTP build was written as two cherry-picks onto b11018. PR #28097 already CONTAINS #27836's three commits rebased, and both report mergeable:false against master, so the pick would have failed. One checkout of #28097, plus a no-speculation control on that same binary — comparing two speculative configs measures agreement, not correctness. * Output hashes are not a valid gate at depth. Measured across the governor test: reps at d0 always agree, while a d8000 cell disagrees with ITSELF intermittently at temperature 0 with top_k 1 — hybrid CPU/GPU expert reduction order varies with thread scheduling and ~5.8k tokens of accumulated state is enough to flip a token. Every hash comparison and reproducibility gate is now d0-only; medians at depth are fine. Also: placement-probe.py records draft_n/draft_n_accepted so a speculative config is scored by ACCEPTANCE rather than tok/s, per-prompt-class because acceptance is prompt-dependent. The drafter layout is chosen by measurement (two were downloaded and only one matches what the loader wants), and the concurrency matrix prices a usable slot size — at 65536 total ctx, --parallel 4 leaves only 16k per slot. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A plain `git checkout pr-28097` would have built a binary that cannot run this model.
PR #28097 is the right PR — it already contains #27836's three NextN/MTP commits
rebased, plus the draft-head-only GGUF layout that the drafter downloaded here
actually uses. But its base is **307 commits behind b11018**, and what it lacks
includes `35822afe5 vulkan: support qwen4exp hc ops` (#28988) — the commit that makes
qwen4exp's hyper-connection ops work on Vulkan at all. So the PR as published cannot
offload this architecture to a V620, and any speculation measured on it would have
been meaningless. It also predates #25483 (skip unneeded MoE work in the mul_mm
coopmat1 path) and #28996, both directly relevant to a Vulkan MoE.
The four commits are therefore rebased onto b11018, with the resolution saved under
mtp-patches/ so it is reproducible rather than living only on the builder LXC. Two
commits conflicted in src/models/qwen4exp.cpp because b11018 had moved underneath the
PR: n_ff_exp became a per-layer array, every hc_*_norm / ple_norm_* / hc_head_norm
gamma was reshaped {hc_dim} -> {n_embd, hc} with TENSOR_ALLOW_RESHAPE, the generic
loader now reads LLM_KV_NEXTN_PREDICT_LAYERS before the arch hook, and b11018's PLE
row-count block became a strict superset of the PR's.
Resolution rule throughout: keep b11018's shapes and logic, OR in the PR's flags. That
is unambiguous rather than a judgement call because every create_tensor call outside
the conflicted hunks already used `flags` — those lines merged cleanly, which is what
pins down what the conflicted ones must look like. The PR's require_weight() rewrite of
the PLE block was deliberately not taken: it regresses b11018's no-tensor case and its
`const int64_t ple_rows` cannot compile against b11018's assigning loop.
Verified: 4 commits over b11018 with b11018 an ancestor, and `g++ -fsyntax-only` on the
resolved file using the build's own flags. The build stage now asserts the commit shape
before compiling, and asserts both `draft-mtp` in --help and ggml-vulkan in the output
afterwards — the Vulkan backend being the entire reason for the rebase.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…hiding a GTT spill
Caught mid-run by a new external sampler. At `-ncmoe 15` with the rule this repo
documents
c1 = ncmoe + (48 - ncmoe) / 2 -> 31,17
card 1 sat at 30665 MiB with **39 MiB of VRAM free and 2060 MiB spilled to GTT**, while
card 2 had 3280 MiB free. That is the ~12x decode collapse the startup loud-guard does
not catch, so the cell reads as "a slow placement" rather than "a broken one" — which is
exactly how it has been read until now.
The imbalance measures 1620 MiB against ~1640 MiB per heavy layer: **exactly one layer.**
The documented rule balances by WEIGHT alone, but three other things live on the card
that owns a layer:
* the KV cache. 12 of the 48 layers are full attention (Qwen Sparse Attention every
4th), and at c1=31 card 1 owns 7 of them against card 2's 5.
* the vision projector, +1.11 GiB, which lands on a single device.
* per-device compute buffers.
So the corrected rule is `c1 = ncmoe + (48 - ncmoe) / 2 - 1`, and a new stage tests it at
the placements that were at or over the edge (plus the documented value at 15 as the
control that reproduces the spill, and the incumbent 20 to see whether a config that
already fit also gains). It reports **free VRAM and GTT for both cards** and a
fits/tight/SPILLED verdict, because "used VRAM" looks healthy right up to the cliff.
Its winner — the fastest placement verified NOT to be spilling — is then what the MTP,
--parallel and restore stages use, in preference to ranking the re-validation data, which
was taken before the spill was known and can rank a broken cell first.
placement-sampler.sh logs free VRAM, GTT, busy %, junction/mem temps, watts, busy-core
p90 clock, RAPL package energy, host memory and every populated
DIMM channel to placement-telemetry.jsonl
It is a separate process rather than an edit to revalidate.sh on purpose: bash reads a
script incrementally, so rewriting one mid-run can corrupt its execution. This is the
non-invasive way to add instrumentation to a sweep that is already hours in.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…not guessed The sweep records a cell's throughput but not its VRAM headroom or GTT, and that gap is load-bearing: below roughly 1 GiB of free VRAM RADV spills to GTT, and the startup loud-guard does not catch it. A spilled cell posts a plausible number and reads as "a slow placement" rather than "a broken one". Joining the two immediately produced a result worth having. The first cell of the re-validation, `-ncmoe 15` at the documented split: ncmoe15-t32-r1 13.68 t/s d0 13.57 t/s d8k 39 MiB free / 2060 MiB GTT SPILLED It is **the fastest cell so far while spilling**. That is the opposite of what this repo's existing GTT note would predict, and the difference is what spills: the 12x collapse on record was the KV CACHE landing in GTT (CT 123, 2.0 tok/s), whereas here it is ~2 GiB of expert weight, read by a subset of experts once per token, on a workload already bound by a ~43 ms fixed engine-overhead floor that masks the extra hop. The splitfix stage measures the cost directly against the corrected split rather than leaving it as inference. Also worth recording from the same data: CPU package power is 80 W under this load against 54 W idle (RAPL), the GPUs peak at 119/147 W, and schedutil's busy-core p90 clock ranges 2400-3300 MHz -- it does boost to the full 3300, it just does not sit there.⚠️ The tool reads FREE VRAM, not used. Used looks unremarkable right up to the cliff: the cell above sat at 30665 MiB used, which says nothing, while having 39 MiB free. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…actually fits
"Most versatile" is not only how much VRAM is left for a second model — it is also how
long a context this thing can serve. That was a guess; now it is arithmetic plus a
measurement.
gguf-kv.py reads a GGUF's metadata header with no dependencies and sizes the cache. For
Qwen3.8-Flash-Next:
48 blocks, full_attention_interval 4 -> only **12 blocks hold a KV cache**
head_count_kv 2, key_length 256, value_length 256
=> 2 * 12 * 2 * (256+256) * 2 B = **24.0 KiB/token at f16**, 12.0 at q8_0
ctx f16 q8_0
65536 1.50 GiB 0.75 GiB
131072 3.00 GiB 1.50 GiB
262144 6.00 GiB 3.00 GiB <- the model's native maximum
The other 36 blocks are Gated DeltaNet, whose state is fixed-size and
context-independent — which is the whole reason a 180B model can hold a 256k window on two
32 GB cards at all. Getting this wrong by treating `n_layer` as 48 would overstate the
cache 4x.
Two independent cross-checks that the parse is right: the PLE head vocab sizes sum to
~320M rows at 160 dims, which is 28.8 GB at this quant and matches the ~28.7 GB the env
file claims; and expert_count 512 x ff 640 x embd 2560 x 3 matrices comes to ~1.42 GB of
expert weight per layer against the ~1.56 GB the placement dial assumes.
A new stage 2c then confirms it by VRAM delta rather than trusting the metadata, because
flat VRAM across two context sizes can mean the KV moved to host memory rather than got
cheaper — so GTT is printed on every row. It tests the fastest placement at 131072/f16 and
at the full native 262144/q8_0, plus 262144/f16 on a roomier placement, to price the real
choice: quantise the cache, or give up GPU-resident experts. Each row also probes at 32k
depth, since a long window is worthless if throughput collapses inside it.
Stage order stays as asked: every other test, then MTP, then --parallel last.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…cannot fit at f16 The "-1 layer" correction committed earlier was derived from a single cell. Three cells in, with GTT counted as demand that did not fit, the picture is different: ncmoe split card 1 card 2 total spare of 65536 imbalance 15 31,17 34789 29506 64295 1241 +5283 20 34,14 32046 24741 56787 8749 +7305 28 38,10 26521 18254 44775 20761 +8267 Two corrections to my own earlier claim: * **The imbalance is not one layer and it grows with -ncmoe.** At ~1500 MiB of expert weight per layer — derived from the 15->20 pair, against the ~1.56 GB the env file assumes — moving one layer shifts the imbalance by ~2x that, so the correction is roughly **-2 layers**, and larger further up the curve. The stage now brackets -1/-2/-3 at the hungry end rather than testing one guess, and also tests the correction at -ncmoe 20 and 28 where the curve actually lives. * 🔴 **-ncmoe 15 cannot safely fit at all on f16 KV with the projector on GPU, whatever the split.** Total demand is 64295 MiB against 65536 of capacity: ~620 MiB per card even perfectly balanced, which is below the ~1024 MiB where RADV starts spilling. That is why the first cell posted 13.68 t/s *while* holding 2060 MiB in GTT. Freeing 1.86 GiB — q8_0 KV (lossless here because the server runs --reasoning off) plus --no-mmproj-offload — raises it to ~1573 MiB/card, so the two genuine max-VRAM candidates are now tested as such: 15 and 14 at the corrected split with that shape. Consequences threaded through the rest of the pipeline: * The winner is no longer "fastest cell". A spilled cell can post a competitive number, so split_cell writes its headroom verdict to a `.verdict` sidecar and the picker refuses SPILLED outright, on top of the existing degeneracy gate. * The winner now carries its whole SHAPE (ncmoe, split, KV type, projector placement), not just the placement. The max-VRAM candidate only fits *because* of q8_0 + projector-on-CPU, and the MTP / --parallel / restore stages previously reset both to "production defaults" — which would have silently pushed the winning config back over the edge. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…fter round 1 Measured pace forced a re-plan. Cells are running 12 -> 15 -> 16 min and rising (the CPU-heavy end of the curve is slower), so revalidate's 2 rounds plus its 6-cell confirmation pass comes to ~7.5 h. That is past the stage's own 6 h timeout, so it would have been killed mid-round-2 and lost the confirmation pass anyway, while pushing MTP and --parallel past morning. Two changes. 1. **The warm pass no longer runs all three prompt classes** (`--classes`, new in placement-probe.py). Warming is about weights, the graph and the deep-prefill indexer state, none of which is prompt-class specific, and at depth each class costs a full 8k prefill — six 8k prefills per cell is over half the wall clock. The MEASURE pass keeps all three, because decode speed is prompt-dependent and one class is not this model's throughput. The flag defaults to all classes, so the already-running sweep is unaffected. 2. **trim-revalidate.sh stops the sweep when round 2 begins**, appending its reasoning to RESULTS.md first so the record says why rather than showing an unexplained TIMED OUT. Round 2 is worth less than it looks: it would be a second sample of a curve measured at the DOCUMENTED --tensor-split, and that split is now known to overload card 1 by ~2 layers — -ncmoe 15 and 20 both spill to GTT there. The split-correction stage re-measures the same placements at corrected splits with 2 reps, so that is where the placement conclusion comes from. Round 1 stands as single-sample context, not as the answer. Spending the remaining hours on splitfix, contextsweep and MTP is strictly more information than a second sample of mis-split cells. Round 1 so far, all at the documented split, so all three are context rather than conclusions: -ncmoe 15 -> 13.68 / 13.57 t/s (d0/d8k) but SPILLED, 20 -> 12.69 / 11.53 also spilling marginally, 28 -> 11.29 / 10.56 with real headroom. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…experts Round 1 of the re-validation is in, and the headline is not the column anyone was watching. Decode and prefill do not degrade at the same rate as experts move to the CPU: -ncmoe decode d0 decode d8k prefill d8k total VRAM 15 13.68 13.57 78.9 58089 (spilled) 20 12.69 11.53 63.6 52540 (spilled) 28 11.29 10.56 47.4 40527 34 10.45 9.44 38.8 31267 48 10.18 8.70 29.0 9443 Prefill falls **-63%** across that range against decode's -36%, and it is the binding constraint in absolute terms: an 8k prompt costs **101 s before the first token** at the best placement and **276 s** at -ncmoe 48. Decode at 13.7 tok/s is perfectly usable. What makes this model feel slow is prefill, and no amount of decode tuning touches that. Two consequences: * There is no single best placement. Long-prompt or agentic work wants -ncmoe as low as fits, because prefill dominates. Short-prompt chat can take -ncmoe 48, which costs only -2.6% of decode against -ncmoe 34 while freeing another 21.8 GiB, and leaves the whole model in 9.4 GiB of VRAM -- one card, with 23 GiB spare for a second service. * The d0 prefill column is an artifact and must not be quoted: those prompts are ~30 tokens, so the figure is fixed overhead, not throughput. Only the d8k column means anything, which is why it is the one reported. Also: splitfix is now two-phase. "Does this shape fit" is a MEMORY question answered by mem_info_vram_free + mem_info_gtt_used at load, so spending a 12-minute throughput probe on it was waste. Phase 1 load-checks all nine candidates in ~2 min each (including a tiny completion first, because llama.cpp allocates some buffers lazily and reading headroom straight after /health would miss them); phase 2 probes only the three lowest -ncmoe shapes that survived, plus the documented split at 15 as the control that reproduces the spill. That is ~50 min instead of ~88, and strictly better science: every candidate gets a headroom verdict instead of only the ones there was time to probe. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Measured at -ncmoe 20, round 1: threads decode d0 decode d8k prefill d8k 8 13.13 11.87 63.8 32 12.69 11.53 63.6 delta +3.5% +2.9% +0.3% 8 threads is faster on decode and neutral on prefill, so it is free -- and it hands back 24 physical cores that CT 120 currently holds for nothing. This corroborates the STREAM result already in the README from a completely different direction: 8 threads saturated the four populated memory channels at 80.3 GB/s while 32 measured WORSE at 74.8. The CPU-side expert FFN is bandwidth-bound GEMV, so past the point where the channels are saturated more threads is contention, not throughput. Two consequences wired in: * The --parallel sweep now includes 8 threads (1 and 4 streams) rather than only 16 and 32. It is ADDED, not substituted: with four concurrent streams contending for the same expert FFN the ordering can invert, and 32 stays as the incumbent control. * The restore stage no longer hardcodes --threads 32. It reads back the best single-stream thread count the concurrency stage measured, so the box is left on the measured value rather than the inherited one. Restoring 32 over a better measured number would have been a silent regression, and the kind nobody notices because nothing errors.⚠️ Still n=1 per cell. Spread at full clock is ~1.4%, so +3.5% is above noise but not by a wide margin; the threads-16 cell lands next and a monotonic 8 < 16 < 32 ordering is what would confirm it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous commit said 8 threads was the winner on the strength of 8-vs-32 alone. With the 16-thread cell in, that was the wrong conclusion from an incomplete sweep: threads decode d0 decode d8k vs 32 8 13.13 11.87 +3.5% / +2.9% 16 13.19 12.11 +3.9% / +5.0% 32 12.69 11.53 -- What survives: **32 is clearly contention**, -4 to -5% against either lower value, and the STREAM result still explains why (8 threads saturated the four populated channels at 80.3 GB/s where 32 measured worse at 74.8 -- bandwidth-bound GEMV gains nothing past saturation and then starts to contend). What does not survive: "use 8". 8 and 16 are within ~2% at n=1, with 16 ahead at depth, so the honest statement is a tie between them rather than a winner. The --parallel stage measures all three at 2 reps, which is what should settle it. The restore fallback moves 8 -> 16 accordingly. It only applies if the concurrency stage produced nothing; normally restore reads the measured value back. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This repo's standing guidance is that "q8_0 KV costs zero throughput (marginally faster,
unchanged at 2x context)". That was measured on CT 123's qwen3.8-27b and **does not
transfer to qwen4exp**. Measured tonight at the same placement and thread count:
d0 d8000 VRAM
f16, projector GPU 12.69 11.53 52540 MiB
q8_0, projector CPU 12.17 9.87 50639 MiB
delta -4.1% -14.4% -1901 MiB
The cost scales with context depth, which is mechanistically what this architecture would
predict: qwen4exp decode is dominated by the QSA indexer rescoring the cached context every
token (llama.cpp #28699), so a quantised cache means dequantising on every indexer pass.
At d0 there is almost no context to rescore and the cost is -4%; at 8k it is -14%.
Two reasons to believe it rather than treat it as noise:
* -14.4% is an order of magnitude above the ~1.4% run-to-run spread at full clock.
* The confound points the WRONG way. The q8_0 cell had ~1900 MiB more headroom and less
GTT than the f16 control, so spill pressure would have made it the faster of the two.
It was slower anyway.
Consequence for the recommendation: "free up VRAM with q8_0" is a worse trade here than it
looked. The frugal shape does not even buy a whole extra placement step -- it saves
~1900 MiB while one layer of experts costs ~1500 -- so -ncmoe 14 with q8_0 + projector on
CPU (63895 MiB) lands only just below -ncmoe 15 with the production shape (64295), and
still spills: 820 MiB/card even perfectly balanced, under the ~1024 MiB RADV floor.
So the max-VRAM placement that actually fits is one step back:
--n-cpu-moe 15 --tensor-split 29,19 --cache-type-k/v q8_0 --no-mmproj-offload
62394 MiB total, ~1571 MiB/card balanced, above the floor. That is exactly the candidate
splitfix phase 1 load-checks, so it gets confirmed rather than assumed -- and now it has to
be scored against its ~14% depth cost, not just its headroom.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ad INVERTED Every number in this README was taken with the governor pinned to `powersave`, i.e. all 64 threads at 1500 MHz against a 3308 MHz maximum. Re-measuring at `schedutil` did not just move the magnitudes; it reversed the direction of two conclusions the file stated confidently. Superseded tables are deleted rather than annotated. **Depth cost inverted.** The old table read a clean monotonic "the more that sits on the CPU, the LESS depth hurts" (−22% at -ncmoe 15 down to −5% at 48). At full clock it runs the other way: −0.8% at 15 rising to −14.5% at 48. The old trend was an artifact of every core being parked at minimum frequency. **-ncmoe 48 vs 34 flipped sign.** Old data had 48 (8.6) beating 34 (8.4), which supported "22 GiB of VRAM buys nothing". The claim survives in magnitude — 48 is now −2.6% against 34 for 21.8 GiB freed — but the ordering was wrong. New, and the most decision-relevant thing in the file: **prefill is the binding constraint, not decode.** Across the placement curve prefill falls 63% where decode falls 36%, and in absolute terms an 8k prompt costs 101 s before the first token at the best placement and 276 s at -ncmoe 48. Decode at 13.7 t/s is fine; the wait is not. So there is no single best placement — long prompts want -ncmoe as low as fits, short prompts can take 48 and free 54.8 GiB. Also corrected: **the `--tensor-split` rule of thumb this file gave is ~2 layers off**, and it was hiding a spill. `card1 = N + (48-N)/2` balances by weight alone, ignoring that the card owning a layer also holds its KV (12 of 48 blocks are full attention, unevenly split), the 1.11 GiB projector, and per-device compute buffers. Measured imbalance is +5283/+7305/ +8267 MiB at -ncmoe 15/20/28 — growing with -ncmoe, not one layer. At -ncmoe 15 card 1 sat at 39 MiB free with 2060 MiB in GTT while card 2 had 3280 MiB free, and that cell still posted the sweep's fastest decode, which is exactly how a broken cell passes for a slow one. And two new sections: q8_0 KV costs ~14% at depth here (it does not on CT 123's qwen3.8-27b, which is where the "q8_0 is free" note came from), and --threads is an inverted U peaking at 16 rather than the inherited 32. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Answering the question directly: running this model entirely on the CPU to free both V620s costs too much. config decode d0 decode d8k prefill d8k 8k TTFT VRAM -ncmoe 20, --threads 16 13.19 12.11 63.6 126 s 52540 MiB, both -ncmoe 48 10.18 8.70 29.0 276 s 9443 MiB, ONE card -ngl 0 CPU-only, --threads 16 6.04 5.51 26.4 303 s 2181 MiB -54% against the fastest fitting GPU config, -41% against -ncmoe 48. The percentage is not the real argument though. **-ncmoe 48 already runs the whole model in 9.4 GiB on a single card** — with no heavy layers left the derived split is 48,0, so card 2 holds nothing and is free for another service, with ~23 GiB still spare on card 1. There is therefore never a reason to reach for CPU-only in order to free a GPU: -ncmoe 48 frees one outright and is 69% faster. A mechanistic detail worth keeping: CPU-only costs only -9% of PREFILL against -ncmoe 48 (26.4 vs 29.0 t/s), because at that placement prefill is already CPU-bound on the experts. What the cards buy at -ncmoe 48 is almost entirely the attention and non-expert path during decode, and that is worth +69%. So "the GPU helps prefill" is the wrong intuition here; it helps decode.⚠️ Recorded honestly: "both cards free" is not literally what was measured. reset_env leaves the vision projector resident, so ~2.1 GiB stays on a card. --no-mmproj-offload frees that too, at a cost to image encoding only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…s, which killed splitfix
splitfix died with rc=1 after 0 minutes and the pipeline moved on to contextsweep — which
consumes splitfix's winner, so it ran in the wrong order with a fallback placement. The
cause is a bash rule subtle enough to be worth writing down:
lbl="nc${nc}-c${c1}$([ -n "$kv" ] && echo "-q8")$([ "$mp" = true ] && echo "-projcpu")"
**A variable assignment takes the exit status of its command substitution.** Both
substitutions here exit 1 whenever their test is false, so the assignment returned 1 and
`set -Eeuo pipefail` aborted the function on its very first candidate. The trace makes it
unmistakable once seen:
++ '[' -n '' ']'
++ '[' '' = true ']'
+ lbl=nc15-c31 <- returns 1
What made it look impossible is that the *identical* construct two lines below, passed as a
command argument, is completely safe: there the exit status belongs to the command, not the
substitution. A grep for the pattern found exactly one instance in assignment position and
three in argument or ||-guarded position, which is why only this one was fatal.
Fixed by building the label with plain conditionals and a trailing `:` to keep the function's
exit status clean. Verified by re-running the whole stage in a stubbed harness — it now
completes with all nine candidates load-checked and phase 2 plus the control invoked.
Also here: chain-report.sh, because overnight-all.sh had to be killed to re-order part 2, so
the report generation it would have run at the end needed re-arming. It waits for part 2,
runs morning-report.sh, and stops the telemetry sampler so its file is closed for analysis.
And RESULTS.md is tidied: two tables whose headers were emitted before the stage died are
removed, and the failure line now records that this was a script bug rather than a
measurement, so nobody reads it as "the split correction could not be measured".
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ifact Commit ce74a40 claimed q8_0 KV costs ~14% at depth on qwen4exp, contradicting this repo's standing note. That was wrong. Measured as a clean pair at a placement that FITS (-ncmoe 16, split 30,18, neither cell spilling, same thread count): shape decode d0 decode d8k free c1 / c2 f16 / projector on GPU 14.16 13.07 1305 / 1553 MiB q8_0 / projector on CPU 14.17 13.45 2881 / 1878 MiB +0.1% +2.9% +1.6 GiB Identical at d0 and marginally FASTER at depth, while freeing 1.6 GiB — which is exactly what the original note said. The -14% came from a pair where both cells sat on the documented (spilling) --tensor-split, and it does not reproduce once the placement fits. The methodological lesson is the part worth keeping: **a spilled configuration does not degrade uniformly.** It cost that pair far more at depth than at d0, which is indistinguishable from a depth-scaling penalty by inspection. So: establish a fitting placement first, then vary one knob. I also varied two things at once there (KV type AND projector placement), which means even the sign was unattributable — still true of the new pair, so the honest claim is that *the shape* is free, not that q8_0 specifically is. I defended the original number at the time by arguing the confound pointed the wrong way (the q8_0 cell had more headroom and was still slower). That argument was not wrong so much as insufficient: "more headroom" was inferred from total VRAM rather than measured per-card, and a marginal cell's behaviour is not linear in its headroom. A clean pair beats an argument about a dirty one. Consequence for the recommendation: the q8_0 + projector-on-CPU shape is now the preferred one at -ncmoe 16 — same speed, 1.6 GiB more headroom on a placement that is only "tight" in f16. Both rows are kept because --no-mmproj-offload costs 3-5x on image encoding, so anyone who cares about vision latency should take the f16 row knowingly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…on measurements
Two stages failed tonight for reasons that had nothing to do with the hardware.
**contextsweep, rc=1 in 2 seconds.** When best_split.txt grew from two fields to four
(ncmoe, c1, kv, mmproj) I updated s3_mtp's `read -r NC C1 KV MP` and forgot s2c_context,
which still did `read -r NC C1`. Bash puts the remainder in the last variable, so C1 became
"30 q8_0 true" and `$(( 48 - c1 ))` was an arithmetic syntax error two calls later. Both
readers now consume all four fields, and the unused ones are named `_KV`/`_MP` so the next
person can see they are deliberate.
**mtp, rc=2 after 87 seconds — but the build SUCCEEDED.** 585 targets linked and both
guards passed (`--spec-type draft-mtp` present, ggml-vulkan present). Then:
tar: pr28097-mtp: Cannot stat: No such file or directory
I renamed the build output to mtp-b11018 when the rebase replaced the PR checkout, and left
the tar/push lines on the old name. The 58 MB artifact is intact on CT 201, so the fix adds
MTP_SKIP_BUILD=true, which reuses it — and **re-asserts both guards against the existing
build before trusting it**, since reusing an artifact is only safe if it still carries the
flag and the backend it was built for.
Recovery shape, and the reason it is not just "edit and re-run": part 2 was still executing
its remaining stages, and **bash reads a script incrementally — rewriting one mid-run shifts
byte offsets and can corrupt execution.** So the fixes went to overnight-part2-fixed.sh and
three standalone runners EXTRACT their stage functions from it with sed rather than
duplicating them, which means the fixed copy stays the single source of truth:
contextsweep-standalone.sh the context sweep
mtp-standalone.sh MTP with MTP_SKIP_BUILD=true
restore-standalone.sh the final restore
chain-finish.sh waits for part 2, then runs the three in order + the report
Order matters: restore runs LAST, because contextsweep and mtp both rewrite the model env,
so part 2's own restore is stale by the time they finish.
restore-standalone.sh is a file rather than an inline `bash -c` in the chain because
inlining it needed three levels of nested quoting around an embedded python heredoc — which
is precisely the shape of bug that cost this pipeline two stages tonight.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nder load Measured at the winning placement, the thread ranking is not stable across concurrency: threads par 1 par 2 par 4 8 14.25 -- 29.37 <- best solo, WORST at 4 streams 16 14.09 22.10 30.88 <- within 1% solo, best at 2 and 4 32 13.22 20.46 30.46 8 threads saturates the four populated memory channels for a single stream, which is why it wins there and why STREAM predicted it. But four concurrent streams present more parallel work than 8 threads can cover, so it loses 4.9% to 16 at par 4. **16 is the right all-round setting**: within 1% of best solo, best at both 2 and 4 streams. That exposed a flaw in my own selection logic, which read only the par-1 rows and would therefore have written 8 into best_threads.txt — the worst value for concurrent use — for the restore stage to apply. It now scores each thread count by its MEAN RELATIVE aggregate across every parallel level measured (relative, so the level with the biggest absolute aggregate does not dominate the mean), and only considers thread counts measured at every level, so a value that happens to have been tried only where it looks good cannot win by omission. Worth recording as a general trap: **a knob tuned at one concurrency level can be actively wrong at another**, and single-stream benchmarking is the default that hides it. The same caution applies to the thread count this repo pins elsewhere. Also from this stage, two things that change how its numbers should be read: * Thread count STOPS MATTERING under concurrency — the t16-vs-t32 gap is +6.6% at 1 stream, +8.0% at 2, and +1.4% at 4. Tune --threads for the solo case; if the box runs concurrent it barely matters. * Concurrency scaling beats a fixed-overhead model. Fitting per-token time = F + V to the t32 pair gives F=53.5 ms / V=22.1 ms and predicts 28.2 t/s at 4 streams; measured is 30.46, i.e. 8% better. V is not constant — llama.cpp batches concurrent decode, so the per-stream work itself shrinks with batch size (V falls to ~19.4 ms at 4 streams). The ~72% fixed fraction is still the headline, but it is a floor on the amortisable share, not an exact split. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The --parallel matrix is complete, at the winning placement (-ncmoe 16, split 30,18, q8_0 KV, projector on CPU). Three results. **Scaling.** 1 -> 2 streams is 1.55x aggregate, 1 -> 4 is 2.30x, at 58% of the solo per-stream rate (7.73 vs 14.09). Good for batch work, poor for one interactive user. ✅ **Doubling the total context budget at 4 streams is FREE.** --ctx-size 131072 gives each of four slots 32k instead of 16k for 30.52 vs 30.46 t/s aggregate — identical within noise. So concurrency and context are not in tension here: the KV cache is 12 KiB/token at q8_0 (24 at f16), so the extra 0.75 GiB fits the winning placement's headroom. The corollary is a 配置 to avoid: --parallel 4 at --ctx-size 65536 buys nothing and leaves each caller 16k. 🔴 **--threads inverts with load.** 8 threads wins solo (14.25 vs 14.09 at 16, 13.22 at 32) because it already saturates the four populated memory channels — exactly what STREAM predicted — and then loses 4.9% at four streams, where the parallel work exceeds what 8 threads can cover. 16 is the right default: within 1% of best solo, best at 2 and 4 streams. The gap also shrinks with load (+6.6% / +8.0% / +1.4% at 1 / 2 / 4 streams), so past a certain batch size the setting stops mattering. The general trap is worth stating plainly, because it caught this pipeline's own selection logic: **a knob tuned at one concurrency level can be actively wrong at another, and single-stream benchmarking is the default that hides it.** And concurrency gives a third, independent estimate of the fixed-overhead floor. Fitting per-token time = F + V to the --threads 32 pair gives F = 53.5 ms, V = 22.1 ms, i.e. 71% fixed — corroborating the ~43 ms floor found earlier by two other routes.⚠️ But that fit predicts 28.2 t/s at four streams against a measured 30.46, so V is not constant: llama.cpp batches concurrent decode and the per-stream work itself shrinks with batch size (V ~19.4 ms at four streams). 71% is a floor on the amortisable share, not an exact split.⚠️ Noted in the table: these come from concurrency-probe.py, whose prompt set differs from placement-probe.py, so they are comparable within that table only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Part 2's restore failed rc=1 with `BTH: unbound variable`. When I added the measured
thread-count lookup, the edit used replace(..., 1) and the anchor
`setv MODEL_THREADS 32; setv MODEL_PARALLEL 1; ...` occurs more than once, so the `local BTH`
declaration landed in a different function while s6_restore kept the `${BTH}` reference in
its final note. Declared properly now, with 16 as the fallback.
⚠️ **I initially misdiagnosed this as another instance of the `[ test ] && assign` pattern
and went looking for all eight of them. That was wrong**: a failing `[` inside an `&&` list
is explicitly exempt from set -e (it is not the final command of the list), which is exactly
why splitfix phase 1 ran straight through eight of them tonight. The bug was **set -u**, a
different mechanism, and the audit was a red herring. Recording it because "same shape as
the last bug" is a tempting and, here, misleading heuristic.
`bash -n` and shellcheck both pass on an unbound-variable bug, so neither would ever have
caught this. What does catch it is running the function with everything stubbed under the
real flags — the technique that found the splitfix bug — so that is now a committed tool
rather than a thing I improvise each time:
stage-harness.sh extracts each stage function from the script and runs it under
set -Eeuo pipefail with setv/pct/curl/ip/wait_up/cells stubbed
Run on the remaining stages it reports s2c_context, s3_mtp and s2b_split all clean, so the
rest of the night has no further landmines of this class. ⚠️ It must run on the HOST: BSD
sed on macOS cannot handle the `{` in the extraction address, which made it report false
failures for all three until moved.
What part 2's restore did manage before dying is worth noting, since the failure was on its
LAST line: every setv ran, the model was restarted, and `pct start 121` ran — so only the
summary note was lost. restore-standalone.sh re-runs it at the end of the finish chain from
the fixed copy anyway.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…op half its own tables Dry-ran morning-report.sh against the night's real JSON before it fires, which was worth doing — it had four defects, two of them the kind that produce a confident wrong answer. 🔴 **It would have recommended a SPILLED configuration.** "The answer" took the top row of the placement table, which ranks by throughput with a gate that only checks degeneracy and d0 reproducibility — nothing about VRAM headroom. The top row was `ncmoe15-t32` at 13.68 t/s, measured at the documented split while holding 2060 MiB in GTT. That is precisely the trap this whole run exists to expose, reproduced by my own reporting tool. The headline now comes from best_split.txt, which the split stage writes only after refusing every candidate whose verdict was SPILLED, and the placement table carries a warning that it shows the shape of the curve rather than a ranking of usable configs. 🔴 **The split table showed 1 of 4 rows, then 2 of 4.** `(\w+)` cannot match `q8-projcpu` (hyphen), and after fixing that a mandatory third group still dropped the f16 cells, whose labels have no shape suffix at all (`split-nc16-c30.json`). Now optional, defaulting to "f16 / proj GPU". **The context table would have listed every configuration twice.** Each context cell writes two files — a 3-class shallow probe and a `-deep` 32k one, split because a 32k prefill costs ~13 min per request — and a bare `ctx*.json` glob counted the deep file as another cell, so each config appeared once with no d0 and once with no d32k. They are joined now, and the 32k prefill rate is reported alongside, since that is what decides whether a long window is usable. **And a NameError**: the headline lookup referenced the `split` dict from above where it is built. Moved below it. Lesson worth keeping: a reporting tool needs the same gates as the measurements it reports. Mine ranked throughput and ignored headroom, which is the exact error the measurements were designed to catch — and it would have shipped the wrong recommendation with every number in it technically correct. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
My contextsweep trimmer did `kill -TERM -<pgid>` using the pgid it read from the contextsweep
process, and that pgid was **the chain's own process group** — so it killed chain-finish.sh
too, and MTP, restore and the report never ran.
The instructive part is that the revalidate trimmer did exactly the same thing earlier and
was safe. The difference:
* `stage` runs its function as `set -m; ( "$@" ) &`, which puts it in its OWN process
group, so a group-kill hits only that stage.
* `chain-finish.sh` invoked `./contextsweep-standalone.sh` as a plain command, which
INHERITS the parent's group. Same kill, entirely different blast radius.
⚠️ Rule going in the script: before `kill -TERM -<pgid>`, check the pgid is not the
supervisor's own. `ps -o pgid= -p $$` in the supervisor is the value that must not match.
chain-finish2.sh logs its own pgid at startup so the next trimmer has something to compare
against.
Nothing measured was lost — contextsweep's cell 1 completed and wrote its shallow result
before the kill, and the MTP build was already on disk. chain-finish2.sh re-runs the three
remaining steps (MTP with MTP_SKIP_BUILD=true, then restore, then the report), and MTP is
now running at the winning placement.
Also worth recording from cell 1 before it was trimmed: at `ctx 131072` with f16 KV on
`-ncmoe 16` the config spills (226 MiB free, 183 MiB GTT) and **prefill collapses** — six
requests took 50 minutes, implying an 8k prefill fell from ~63 t/s to roughly 10. That is the
~12x GTT collapse this repo documents, observed in its classic form for the first time
tonight. It is also why f16 is the wrong KV type for a long window here: q8_0 halves the
cache and the --parallel stage already showed 131072 fitting at that shape.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…is untouched Commit b14d846 claimed that at ctx 131072 with f16 KV the spill made prefill "collapse from ~63 t/s to roughly 10". That is wrong. I inferred it from the cell taking 50 minutes of wall clock, without reading prefill_tps_median, which was in the probe's JSON the whole time. Measured, same placement, only the context budget changed: ctx d0 decode d8k decode d8k prefill headroom 65536 14.16 13.07 77.0 1337/1553 MiB free, no spill 131072 12.92 10.34 76.7 226/573 MiB free, 183 MiB GTT -8.8% -20.9% -0.4% Prefill is unaffected. The spill costs decode, and roughly twice as much at depth as at a short prompt — which is the mechanically sensible shape: prefill streams weights in large batches, while decode at depth touches the KV cache every token, so a cache partly in GTT means a host round-trip per token instead of per batch. Two things to carry forward: * **Do not infer a spill's cost from wall-clock time.** A slow cell invites a story, and `prefill_tps_median` is in every probe's output to settle it instead. * This repo's ~12x GTT figure came from a case where the WHOLE KV cache sat in GTT under a much larger over-commit. A marginal 183 MiB spill is a different regime and does not behave the same way, so that number should not be applied to any spill indiscriminately. The practical recommendation does not change — use q8_0 for a long window, because it halves the cache and 131072 then fits rather than spilling, which the --parallel stage independently confirmed — but it now rests on the right mechanism. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… caused it
Both drafters abort the instant the MTP graph is constructed:
ggml.c:2264: GGML_ASSERT(ggml_can_repeat(b, a)) failed
llama_model_qwen4exp::graph::build_hc_mix(...)
llama_model_qwen4exp::graph_mtp::graph_mtp(...)
The build itself is healthy: the no-speculation control on that exact binary gives 13.88 t/s
against 14.17 for b11018-baseline at the same placement, so it loads the target and serves
normally. Only the MTP path is broken.
**And it is broken because of the conflict resolution I chose — though resolving the other
way would not have helped.** b11018 reshaped every hyper-connection gamma from {hc_dim} to
{n_embd, hc} with TENSOR_ALLOW_RESHAPE and updated its TRUNK graph to match. I kept b11018's
shapes, which is right: the trunk works, as the control proves. But PR #28097's MTP graph
predates that reshape and is written against {hc_dim}, so it broadcasts mismatched shapes.
Taking the PR's shapes instead would just move the abort into the trunk, which is the path
that has to work. There is no resolution of those two commits that satisfies both graphs.
The real fix is to port the PR's MTP graph to b11018's hc convention — a code change to the
MTP caller of build_hc_mix, not a merge decision. That belongs upstream or in a patch written
with the hc layout in hand. Guessing at it risks producing wrong OUTPUT rather than a clean
crash, which is worse than no measurement, so it is explicitly not attempted.
⚠️ Worth recording about my own pre-flight: I checked that both drafters carry the metadata
the loader's mtp_only probe needs and concluded the layout-mismatch failure mode was "ruled
out". That validated the LOADER path and said nothing about the GRAPH path, which is where it
actually fails. Metadata satisfying a loader does not establish that the graph built from it
is shape-correct — the pre-flight was necessary, not sufficient, and I overstated it.
Everything is staged for a one-command retry when #28097 rebases: drafters verified, build at
/root/builds/mtp-b11018 (58 MB, both guards asserted), patches in mtp-patches/, runner
./mtp-standalone.sh with MTP_SKIP_BUILD=true.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ured
Adds the recommendation section the night was for, at the top of the measured results, plus
two corrections.
DEFAULT, and it is both the fastest and the roomiest:
--n-cpu-moe 16 --tensor-split 30,18 --threads 16 \
--cache-type-k q8_0 --cache-type-v q8_0 --no-mmproj-offload \
--ctx-size 131072 --parallel 1
-> 14.17 t/s short prompt, 13.45 at 8k, 2.3/1.3 GiB free, no spill
-ncmoe 16 is the FLOOR: nothing below it fits in any shape at any split. q8_0 plus
projector-on-CPU is free (identical at d0, +2.9% at depth), frees 1.6 GiB, and is what lets
131072 of context fit.
The give-a-card-back option is worth knowing: **-ncmoe 48 runs the whole model in 9.4 GiB on
ONE card**, leaving the other entirely free, for -28% of decode. CPU-only is strictly worse
than that at 6.04 t/s (-57%) and frees only 2 GiB more.
Concurrency: 2.30x aggregate at four streams (30.9 t/s), and doubling the context budget is
free -- so --parallel 4 at --ctx-size 65536 is never right.
And the ceiling is now explained rather than described. Decode is ~72% fixed per-token
overhead by three independent routes (the -ncmoe curve fit and a clock experiment both give
~43 ms; the concurrency fit gives 53.5 of 71 ms), and GGML_SCHED_DEBUG names it: ~30 graph
splits per token, which matches the placement almost exactly (16 CPU-expert layers x 2
crossings, plus the card boundary and the PLE lookup) at ~1.1 ms per crossing = 61% of the
fixed term. The lever is fewer CPU<->GPU crossings, not bandwidth or clocks -- both of which,
along with PCIe, core count and disk, were ruled out separately.
Two corrections in this commit:
* The MTP control was quoted as 13.88 t/s. That is the median across BOTH depths, not d0.
Per depth it is 14.42 / 13.43 against b11018-baseline's 14.17 / 13.45 — so the rebased
build is marginally FASTER than baseline, not slower, which strengthens rather than
weakens the claim that only its MTP graph path is broken.
* The spill-costs-decode-not-prefill retraction is now confirmed a second time,
independently: the spilled nc15-c31 cell has prefill 79.9 t/s against 77.0 for the fitting
nc16-c30 — unaffected, marginally higher — while its decode is -11.4% at d0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…free
The recommendation was assembled from separate cells: the split stage measured that
placement at --ctx-size 65536, while restore left the box at 131072 (inherited from the
context sweep). Those are not the same configuration, so the delivered one was re-probed
directly rather than assumed.
live, ctx 131072 splitfix, ctx 65536
decode d0 14.46 14.17
decode 8k 13.37 13.45
prefill 8k 77.2 76.9
VRAM free 2297/1255 MiB 2881/1878 MiB
**--ctx-size 131072 is free**: identical throughput within noise, d0 marginally higher, and
card 2 still at 1255 MiB — above the ~1024 MiB floor where RADV spills. So the 128k window
costs nothing and is kept.
The contract assertions also pass on the live config, which is the part that would silently
break callers rather than show up as a number: bare chat answers with clean content, no
<think> pair leaks into `content`, and no reasoning_effort value a caller might send breaks
the server. That last one matters because this model's template raises on "none" and "high"
— the exact values CT 123's coder uses — so --reasoning off is doing its job of making the
trap unreachable from outside.
`all_reps_agree=False` here is the documented depth non-reproducibility (d8000 disagrees
with itself at temperature 0 on this hybrid placement), not degeneracy: `degenerate=False`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…up on restart I asserted all night, in the restore stage's own output and in every status update, that CT 121's 22:00 backup was "missed tonight, trigger by hand". That was an assumption I never checked, and it is wrong. Verified: CT 121 came up at 06:50:18Z and `Backup hermes 2026-09-18` (fbdbbe08) committed at **06:50:54Z** — four seconds later. Hermes' cron catches up a missed job when the container restarts, so a planned overnight outage needs no intervention, and silence in #backups during one is not automatically a gap. 🔴 **But the same catch-up behaviour caused real collateral damage, and the timing points at this benchmark.** Two jobs fired at that instant. The backup completed in 61 s. The KB Freshness Refresh ran **53 minutes**, hit its 2700 s budget, staged **0 of 448** entries and quarantined five as runaways: ai-tools/openworker, penguin-harness, thinking-orbs, tokentab, turbo-fieldfare That window (06:50-07:43Z) is exactly when the context sweep was restarting CT 120's server every few minutes and driving it flat out. The refresh consumes CT 120's API, so those timeouts are an artifact of a saturated box, not of bad entries — the five quarantines should be reversed and re-run. The restore stage's note is corrected to say what is actually true, and to warn about the part that matters: every cron catches up at the same moment the container returns, so anything still benchmarking will be contended with. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
47 files / 6707 lines -> 31 / 5737. Everything removed was single-use scaffolding whose question is closed, and every FINDING it produced is already in the README. Recovery chains and trimmers, which existed only because the pipeline was mid-flight: chain-finish.sh, chain-finish2.sh, chain-report.sh, chain-pr28699.sh, chain-one-gpu.sh, trim-revalidate.sh, trim-contextsweep.sh, contextsweep-standalone.sh, restore-standalone.sh. Two of these carried the group-kill hazard that took out their own supervisor, so keeping them around as templates would be actively harmful. Investigations that answered their question: cpuset-test.sh and cpuset-variance.sh (the "cpuset lottery" turned out to be the governor), warmup-test.sh (warm-up is built into the probes now), power-test.sh (RAPL numbers recorded; the method is in the host-power-saving memory), structural-tests.sh (whole-layer -ot rejected at -42%), confirm-sweep.sh and tune-sweep.sh (superseded by revalidate.sh + placement-sweep.sh). Two fixes so the survivors stand alone rather than depending on deleted or host-only files: * mtp-standalone.sh pointed at overnight-part2-fixed.sh, which only ever existed on the host as a way to patch a script that was executing. The in-repo overnight-part2.sh is now byte-identical (verified by sha256), so it points there and works from a fresh checkout. Its header also drops the recovery narrative and states the current blocker instead. * overnight-all.sh still pgrep'd for the deleted chain-reval.sh. Removed. Kept deliberately, because they generalise past this model: gguf-kv.py (sizes any GGUF's KV cache from metadata — the figure that gets overstated 4x on a hybrid), stage-harness.sh (catches the set -u class of bug that bash -n and shellcheck both pass), stagelib.sh, attribute-telemetry.py + placement-sampler.sh, build-llamacpp.sh, and the overnight pipeline itself, which is worth re-running once the four extra DIMMs land. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…rect the one-card advice
The placement-dial section gave approximate VRAM ("~59 GB", "~8 GB") and never said what
-ncmoe 48 actually leaves resident. Replaced with measurement and structure.
**What 48 means.** --n-cpu-moe moves ONLY the routed experts, so everything else is resident
at every setting: attention for all 48 layers (12 Qwen Sparse Attention + 36 Gated
DeltaNet), the per-layer shared expert and router, norms, hyper-connections, output, and the
KV cache. -ncmoe 48 is therefore "attention plus the dense path on the GPU, 9.4 GiB, all
routed experts in RAM" — the floor of GPU residency, not an off switch.
**Why the experts and not the layers.** Routed experts are ~61% of the model but only 2.0%
is read per token (topk 10 of 512); attention and the dense path are read in full every
token. Measured rather than argued: moving whole layers to CPU cost -42% against -28% for
moving only the experts at the same VRAM saving.
**VRAM per setting, measured** — read after a load plus one completion so lazily-allocated
buffers are counted, at the corrected split:
ncmoe 15 62.3 G SPILLED at every split (64.2 G demand vs 64 G capacity)
ncmoe 16 59.3 G q8_0/projCPU, tight 14.17 / 13.45 <- the floor for two cards
ncmoe 16 61.1 G f16/projGPU, tight 14.16 / 13.07
ncmoe 20 55.3 G fits 13.03 / 12.06
ncmoe 28 43.6 G fits ~11.3 / ~10.6
one card: ncmoe 34 30.4 G 13.01 / 12.23 <- the floor for one card
ncmoe 36 27.4 G 12.43 / 11.81
ncmoe 48 9.2 G 10.18 / 8.70
🔴 **And a correction to this file's own advice.** It recommended -ncmoe 48 for handing a
card back. That is 28% slower than necessary: -ncmoe 34 on ONE card gives 13.01 t/s, only
-10% off the two-card best, with card 2 completely idle. Pick 48 only when you also want
~23 GiB spare on the remaining card.
Also adds the matched-pair result this rests on: at -ncmoe 34, one card is +20.9% at d0 and
+24.2% at depth against two, with prefill IDENTICAL (+0.5%). The second card is a decode
penalty (~21%, llama.cpp #28699's inter-GPU indexer traffic) and a prefill win (+97%,
because batching amortises the crossing). Net +11% decode, +97% prefill — so it earns its
place on time-to-first-token, not on tok/s.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…per-layer VRAM is STABLE The one-GPU sweep finished. Full single-card curve (--device Vulkan0, q8_0 + projector on CPU, threads 16, ctx 65536): -ncmoe VRAM used free decode d0 / d8k prefill 8k 34 30.4 G 1.7 G 13.01 / 12.23 39.1 36 27.4 G 4.6 G 12.43 / 11.81 37.1 40 21.6 G 10.4 G 11.40 / 10.94 33.7 48 9.2 G 23.3 G 10.18 / 8.70 29.0 ✅ Expert weight is ~1.5 GiB/layer and that figure is STABLE: 1502 MiB/layer over 34->36, 1502 over 36->40, 1581 over 40->48, averaging 1547. That stability is why every VRAM prediction made from the average during the sweep landed within ~120 MiB of measurement across four different shapes, and it is what makes the placement dial predictable.⚠️ Recorded because I nearly wrote the opposite into this file. I computed the deltas as 1502 / 2018 / 1323 MiB/layer — a ±35% spread — and had begun attributing it to UD-Q4_K_XL being Unsloth's *dynamic* quant assigning bit-widths by layer sensitivity. That is a plausible-sounding mechanism for a number that does not exist: the spread came from mixing the sweep's own reading for -ncmoe 40 (free 10679 -> used 22089) with a live reading taken later in the cell (20024). The ~2 GiB gap is transient compute buffer. **Read VRAM from one consistent moment**, or you can invent a non-uniformity and a story to explain it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…wen3.6
The placement study is finished, and its conclusion changes where this model should live.
**One card beats two by 21-24% on decode** for this architecture (llama.cpp #28699 ships the
QSA indexer's pooled rows across the inter-GPU link every token), and Qwen3.8-Flash-Next is
prefill-bound at ~205 s to first token on an 8k prompt — far too slow for Hermes. So:
* **CT 120 -> qwen3.6-35b-a3b on GPU 1** (`ct120-cutover.sh to-qwen36`). Verified 74.8 tok/s
and a clean chat endpoint. This is Hermes' model and it is 5x faster for agent work.
* **CT 123 -> Qwen3.8-Flash-Next on GPU 2 alone**, at the measured best single-card config:
`-ncmoe 34 --tensor-split "" --device Vulkan0 --threads 16 --cache-type-k/v q8_0
--no-mmproj-offload`, ctx 65536. Verified serving at 28964 MiB with 1738 free, no spill.
* **llama-swap removed** from CT 123 (service stopped and disabled; its six models left on
disk untouched at the owner's request — /models grown 180G -> 240G rather than deleting
anything).
* **CT 122 `coder-runner` destroyed.** The loop moved to Multica, so it was unused and only
something to keep patched. Its ssh keypair was removed from CT 121 with it, and the
provisioning script stays as the recipe.
New: `qwen38fn-gpu2.env`, the single-card config as a committed file rather than a hand-edit,
and `install.sh` grows an `ENV_FILE` override so both deployments are first-class:
VMID=120 ./install.sh # two cards, -ncmoe 16
VMID=123 ENV_FILE=qwen38fn-gpu2.env ./install.sh # one card, -ncmoe 34
Two things that had to be fixed to make CT 123 work, both worth knowing:
* **CT 123 was on llama.cpp b10678, which cannot run `qwen4exp` on Vulkan at all** (b11013 is
the floor). Brought to b11018 by copying `/opt/llamacpp/b11018-baseline` from CT 120 —
same Ubuntu 24.04 / glibc 2.39, so the binary moves. Copied the *proven* build rather than
the prebuilt release, which had never served this model.
* **The thermal watchdog's `GPU_SERVICE_MAP` named `123:llama-swap`**, a service that no
longer exists. A map naming a dead service turns a thermal trip into a no-op that leaves
the real load cooking an overheating card. Now `123:llamacpp-qwen38fn`, in both the live
env and `ct120-cutover.sh`.
CLAUDE.md updated throughout: the architecture line, the VMID allocation, the coder-runner
convention (now marked decommissioned, with the isolation rule retained), the common
commands, the build pins, the watchdog map and the token-accounting note.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…DE.md Found while briefing a reviewer on this PR. The overnight sweep established that the obvious split rule overcommits card 1 by ~2 layers and spills anyway, and the study README records the correction -- but it propagated nowhere else. - placement-sweep.sh and revalidate.sh both still derived the UNCORRECTED `N + (48-N)/2`. This is functional, not cosmetic: it is the derivation that silently spilled -ncmoe 15/16/20 into GTT. Corrected to `N + (48-N)/2 - 2`, which reproduces the measured splits exactly (16 -> 30,18 -- the deployed config -- 20 -> 32,16, 28 -> 36,12). `-ncmoe 48` is exempt and clamped to `48,0`: with no heavy layers to rebalance, the -2 would hand card 2 two light layers plus an inter-GPU hop for nothing, and would break the "*,0" single-GPU detection. Keeping the two derivations identical matters -- otherwise a revalidation measures a different placement than the sweep it is checking. - CLAUDE.md asserted the uncorrected rule AND claimed placement-sweep.sh derived it, so fixing the doc alone would have made it wrong a new way. - CLAUDE.md still headlined "~4x below the sizing note, 11.74 t/s", which the overnight curve superseded: the best two-card placement is 14.46 t/s (13.37 at 8k) against a predicted 72, i.e. ~5x. Replaced with the real numbers, the graph-split mechanism, the calibrated latency model, and the two findings that actually drove the deployment -- prefill is the binding constraint (104 s vs 205 s TTFT at 8k) and one card beats two by 21-24% at matched placement. - revalidate.sh and overnight.sh carried the powersave penalty as -52%. That was one noisy pre-warm-up arm; the settled figure is ~30%. - stage-harness.sh was executable with `set -Eeuo pipefail` on line 1 and no shebang (shellcheck SC2148, an error): it would run under the caller's shell, where -E is not POSIX. Verified: shellcheck clean at -S warning on every shell file in the PR, all Python compiles, and no function in the PR ends with a `[ ] && ...` construct -- the `set -e` trap that shape causes applies to a function's last command, not to a loop body or if-branch, which is tested rather than assumed here. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Systematic grep for the pattern found in the last commit: a finding gets corrected in the README and nowhere else. Two of these are functional and one is actively dangerous. 🔴 qwen38fn.env shipped a SPILLING default. It is the install.sh default (CT 120, two cards) and carried MODEL_CPU_MOE=20 with the uncorrected MODEL_TENSOR_SPLIT=34,14 -- the exact combination the README measured at 101 MiB in GTT with 823 MiB free, under the ~1024 MiB RADV floor, for 12.69 t/s. Now -ncmoe 16 + 30,18, the measured best two-card cell (14.46 / 13.37 at 8k), with the rule comment corrected. 🔴 qwen38fn-gpu2.env told a future operator that b10678 is "a working rollback target" because it "CAN load the arch". It cannot execute this model at all: the arch string is registered there, but the Vulkan hyper-connection ops only landed in #28988 (b11013 is the floor). A rollback following that comment would produce a server that cannot load.⚠️ The retracted "q8_0 costs ~14% at depth" figure survived in the VRAM-floor candidate table and inverted its verdict. The clean pair has q8_0 + projector-on-CPU BETTER on both axes -- +2.9% at depth and 1.6 GiB cheaper -- so the table recommended the wrong row, and both env files followed it. The retraction also re-opens -ncmoe 15, whose only recorded reason for rejection was that 14%; flagged as untested rather than excluded, and the contradiction with "below 16 does not fit in any shape" is noted.⚠️ The threads correction landed in one section and three places still carried the superseded single-stream reading: the README's "inverted U with the peak at 16" heading (8 and 16 are within 0.5% at n=1), the summary row claiming 16 "is best across every concurrency level" (8 wins solo by ~1%, then loses 4.9% at four streams), and overnight-part2.sh's own selection comment. All now point at the concurrency section that settles it. Also: CLAUDE.md recorded 34,14 at -ncmoe 20 as "no spill" (it held 101 MiB in GTT -- it read clean only against a 9.3 GiB spill); CLAUDE.md described llama-swap in the present tense as CT 123's service while saying elsewhere it was removed, now marked decommissioned-but-recoverable at the section head; the last "4x" in the README is ~5x; one broken anchor I introduced (an emoji heading leaves a leading space, so the slug takes three hyphens). Verified: shellcheck clean at -S warning, bash -n clean, README anchors all resolve, EnvironmentFile syntax clean (no inline comments after a value), and the serve script does honour MODEL_KV_TYPE / MODEL_MMPROJ_ON_CPU so the new keys are live rather than decorative. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ginal All 14 findings reproduced independently before fixing. The three P1s are real safety defects in the cutover and watchdog paths; the live deployment was never affected. P1 — the forward cutover disabled thermal protection on BOTH cards. It mapped both GPUs to `120:llamacpp`, then disabled that unit and started `120:llamacpp-qwen38fn` instead, so a trip stopped a dead unit while the real load kept running — on the one configuration that drives both cards. Now maps the unit that actually runs, and `assert_map_owns_cards` checks the invariant (every mapped service must own its card) after both directions. P1 — GPU exclusivity did not survive a reboot. `assert_ct123_stopped` was a one-shot check, and CT 123 kept onboot=1 plus its GPU-2 mounts, so the next host boot put two llama-servers on one card. The cutover now clears CT 123's onboot and restores it on the way back; `status` surfaces the hazard. P1 — the committed watchdog default named `123:llama-swap`, removed on 2026-09-18, so a fresh install left GPU 2 mapped to a dead unit. The in-script fallback was worse: B550-era PCI addresses that do not exist here. 🔴 The two-card default was MARGINAL, and this is the reviewer's best catch. Every throughput number came from a sweep pinned at batch/ubatch 1024/256 while the launcher defaults to 4096/1024 and neither env file said so. Loaded end-to-end to settle it: 1024/256 min free 2881 / 1878 MiB no spill prefill 77.2 t/s 4096/1024 min free 2922 / 877 MiB GTT 306 MiB prefill 173-236 t/s Card 2 fell under the ~1024 MiB RADV spill floor and card 1's GTT hit 306 against a ~70 idle floor. Both env files now pin explicit, measured values (two-card 1024/256; one-card 4096/1024, which is verified safe at -ncmoe 34). ✅ New finding: the right --tensor-split is BATCH-SIZE DEPENDENT — the compute buffer scales with batch and lands on the card holding the heavy layers, which is also why the `- 2` correction exists. At 4096/1024 this placement wants ~31,17, untested. Also fixed: the guard treated an unreadable sensor as 0 °C (failing open, against its own stated contract) — a missing sensor is now over-temp; the stage harness printed FAIL and exited 0, and could not tell a broken stage from a failed extraction; the split selector inferred KV/projector shape from filenames and silently dropped unmatched cells, so it now reads a recorded .shape sidecar and logs every skip; the KV A/B ran its two arms at different placements, recreating the confound that produced the retracted 14% claim; `-ncmoe 15` was described as never load-checked when saved results show it SPILLED at 29,19 (3256/89 MiB); three more copies of the split rule (one uncorrected, one using `- 1`); the -54% governor figure; the contract claim that inverted what the probe asserts; `-ncmoe 48 matches 34`; 64295 MiB mislabelled 64.2 GiB (it is 62.79, and the constraint is per-card margin); the patches README not saying MTP aborts in the graph; and the false "all clean" lint claim, whose verification loop iterated an empty file list. Consolidation, per the review's root-cause note: CLAUDE.md's section 108 -> 51 lines with the README as the single detailed record, and the JEDEC 3DS analysis condensed to its conclusions. Gate: `git ls-files '*.sh' | xargs shellcheck -S warning` clean on 0.11.0 (the reviewer's exact command; SC2034 false positives from eval'd stage bodies now carry file-scoped directives with reasons), bash -n clean, py_compile clean. Fixes verified non-vacuously: the guard yields 999 for a missing sensor where it yielded 0, and the harness exits 1 on an injected `return 42` where it exited 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… was wrong The follow-up is right on all five findings, including that my completion claim was not supported by 8c2a79e. Two of these are defects I created while fixing the originals. 🔴 P1 — the thermal guard still failed open, and my test is why I missed it. Validating the CONCATENATION `${jr}${mr}` means one empty reading beside one numeric reading concatenates to a numeric string, passes validation, and the empty one becomes 0 C in the arithmetic. Only BOTH sensors missing tripped — the single case my test covered. Each reading is now validated separately, and losing the hwmon directory mid-run routes the same way (the startup check only proves it existed at startup). Verified by breaking each reading independently, one at a time: junction-missing, mem-missing, both-missing and hwmon-gone all trip; both-ok does not. P2 — `assert_map_owns_cards` checked none of what its name claimed. It never read any container's GPU bindings, accepted is-enabled OR is-active so an enabled-but-dead unit passed, never checked both cards had entries, and returned 0 on every path while logging "every mapped unit exists and is live". That is the same defect class this review set out to trace, committed in the fix for it. It now verifies the target container actually binds that card's render node, reports active / configured-but-not-serving / mismatch as distinct states (the middle one is the legitimate state of a deliberately stopped CT 123 after rollback), checks both cards are mapped, and returns non-zero. Reproduced the reviewer's scenario: CT 120 binding only GPU 1 with a map claiming both now fails with rc=1. Deferred via MAP_BAD so the status dump still prints, then exits non-zero. P2 — the harness still reported success after an INTERNAL command failure. `if ( "$fn" ); then` disables errexit for the whole if-condition and that suppression reaches inside the function, so a failing command mid-stage did not abort it; the stage ran on and returned 0. Counting failures could not help because the failure was never surfaced. Now runs each stage as a standalone subshell with errexit re-armed, reading $? on the next line. Verified with an internal failure whose later commands still succeed: FAIL rc=42, harness exit 1. Also added the missing stage() negative control. P2 — the surviving claims. Six locations in the model README still said "64.2 GiB", "never load-checked", "nobody re-tested" and "does not fit in any shape at any split" — I had fixed the env-file copies and not these. And pro-v620/README.md was never opened: it still described CT 123 as the llama-swap coding server, mapped the watchdog to 123:llama-swap, and carried the entire B550 topology (0000:06:00.0, "PCIe-3 (chipset) slot") in its intro and in copy-pasteable command examples. P3 — "a copy of the release tarball" was asserted, not checked. Same source revision (build 11018, commit c9a5eeeb3) but different builds: GNU 11.4.0 vs 13.3.0, sha256 4ac7aa75 vs 32687325. Worth more than the wording: every sweep number was measured on b11018-baseline while the two-card env ships llama-b11018, so a throughput comparison against the sweep crosses two binaries. The 2026-09-18 load check did use the shipped binary, so the shipped config itself is validated. -ncmoe 15 is closed the way the follow-up argued — by headroom POLICY, not by physics. Adopted its wording plus an explicit >=2048 MiB/card acceptance rule (the README already marked anything below that "tight"), and dropped my "arithmetic rules out the rest", which overstated the case. Gate: shellcheck -S warning clean on 0.11.0 by both the all-files and changed-files commands, bash -n clean, py_compile clean, README anchors resolve. One pre-existing broken anchor in pro-v620/README.md (context-window--concurrency) also breaks on origin/main and is left alone. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…gs, and a retraction that was itself wrong All six findings hold. Three are defects introduced by the previous round's fixes, and one is a correction that corrected something true. P2 — the new ownership assertion BROKE THE ORDINARY ROLLBACK. `pct exec` cannot run in a stopped container, and `to-qwen36` leaves CT 123 stopped by design, so the "CONFIGURED but not serving" state was unreachable for exactly the container that needs it: both queries failed, MAP_BAD was set, and a completely correct rollback exited 1 with a false protection alarm. Container -stopped and unit-inactive are now distinct states, the stopped target reports its bindings as correct with runtime state DEFERRED, and a genuine mismatch still fails. P2 — the >=2048 MiB/card policy rejected the configurations this file ships: the two-card default's minimum is 1878 MiB, the one-card default's 1739, the 131072 row's 1255. A universal rule invented inside the paragraph that needed it is not a policy. Reframed: >=2048 MiB is required for a NEW or UNVALIDATED placement; an existing placement below it is acceptable only with a measured per-card minimum recorded at its shipped batch size, and the three exceptions are named. No live setting changed to suit the prose. P2 — "the shipped config is validated" combined two configurations. The release-binary check ran at 4096/1024; the file ships 1024/256, whose numbers come from the baseline build's sweep. The exact pair was never tested. Stated as the verification gap it is, with the one thing that can be said for it: smaller buffers cannot use more memory than the run that already fit. P2 — the stage() negative control passed without running its body, four ways: wrong arity (`stage <name> <timeout> <fn>`, so $2 was unbound), stderr hidden so the invocation error read as success, `stage()` returning 0 by design after recording a failure so its exit status could never be the assertion, and the harness's `sleep` stub firing the watchdog instantly. It now asserts the recorded note and a post-failure sentinel, restores the real sleep, and runs a succeeding body through the same path to prove it is not vacuous. Verified by three separate mutations: as committed both controls pass; making the failing body succeed fires the vacuity check; breaking stage() itself to record success is also caught.⚠️ Two bugs in my own verification, found while checking the above: the control wrapped stage() in `set +e`, which propagated into its subshell and disabled the very errexit under test; and `local IFS=','` in the assertion stays in effect for the whole function, rewriting `$*` for everything called inside it. The IFS scope is now narrowed to the split itself. P2 — four more current-state claims: gpu2.env still called 15 "the minimum that fits", and the parent README still said the second card runs llama-swap and that CT 123 "cannot" expose token counters (it now runs llama-server with --metrics and does expose them). The model README's "stock map" warning is relabelled historical, keeping the general hazard. P3 — the provenance retraction was itself wrong and is withdrawn. It compared CT 123's b11018-baseline against CT 120's llama-b11018 and concluded "not a copy" from the differing hashes, but CT 120 holds BOTH trees and its baseline is byte-identical to CT 123's (32687325 vs the release's 4ac7aa75). The original copy claim stands. Lesson recorded where it happened: name the tree, not the container. Gate: shellcheck -S warning clean on 0.11.0 by both commands, bash -n clean, py_compile clean, anchors resolve. The pre-existing context-window--concurrency anchor in pro-v620/README.md still breaks on origin/main and is left alone. The mocked cutover transition test is still pending and now has one more case to cover: the deliberately stopped rollback target. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both items promised in the last round. Env consolidation. The review's point was that repeating the correction narrative in env comments had not prevented propagation drift and had made the files unreadable — 276 lines across the two, 199 of them comment, for 40 settings. Now 204 lines and 128 comment, with every setting carrying its value, its trap and a pointer, and the history living only in README.md. All 40 VAR=value lines are byte-identical to the previous commit, verified mechanically rather than by eye. Two real defects surfaced while rewriting: comment blocks had drifted off their settings in both files (MODEL_LOAD_MODE's text sat above MODEL_TENSOR_SPLIT; the never-use-auto-context warning was orphaned after MODEL_THREADS), and the q8_0 <-> --reasoning off XOR warning was missing entirely from qwen38fn-gpu2.env — the LIVE config, the one actually running q8_0. Cutover transition test (cutover-transition-test.sh). Drives the real script against a mocked pct/systemctl and asserts the invariant that matters: at no point may two containers be able to hold GPU 2, running or after a reboot. Six scenarios — refusal while CT 123 runs, forward cutover with an ordering check that the release precedes the attach, a simulated host reboot in the forward state, failure injected during the release, failure injected after it plus recovery by rollback, and the reverse cutover with CT 123 deliberately left stopped (the R3-1 regression). To make it possible the script's paths are now overridable (VMID, CONF, BAK_DIR, WATCHDOG_ENV, LXC_CONF_DIR, DRI_BY_PATH), defaulting to the production values. Nothing else sets them. 🔴 The test is proven non-vacuous. Five mutations, applied one at a time, each caught with the right message: watchdog map pointing at the disabled unit; CT 123 autostart never cleared; the assertion blind to a stopped container; autostart not restored on rollback; GPU 2 never detached on rollback.⚠️ And the mutation run earned its keep immediately — two mutations first read as "not caught" because the test ABORTED instead of reporting. `rel="$(grep ... | head | cut)"` with no match: grep exits 1, pipefail propagates, the assignment inherits that status and set -e kills the test. That is the same bash trap this folder hit once before, this time inside the test written to catch such things. Fixed with `|| true`, and the mutation harness now distinguishes caught / not-caught / aborted. Also noted in passing: `sed -i` in set_watchdog_map is GNU-only, fine on the Proxmox host but not portable; the test shims it so it runs anywhere. Gate: shellcheck -S warning clean on 0.11.0 by both commands, bash -n clean, py_compile clean, transition test PASS, stage harness PASS. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Verifying the round-3 triage at HEAD found finding 2 only half fixed. The definition was correctly narrowed to new/unvalidated placements, but four references elsewhere still named "this folder's >=2048 MiB/card policy" without that qualifier — including the summary table, which is where a reader lands first and would infer exactly the global rule the finding was about, one the shipped configurations violate. Same propagation failure as the rest of this review: the definition moved and its references did not. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review triage — all 6 findings from the
|
| exercised? | |
|---|---|
| release binary | yes — at 4096/1024 (the run that found 877 MiB) |
| 1024/256 | yes — on b11018-baseline (the sweep) |
| release binary at 1024/256 | no — never tested as a pair |
Recorded as a verification gap with the one honest mitigation: smaller buffers cannot use more memory than the run that already fit. No rerun performed or requested.
4. P2 — the stage() negative control passed without calling its body — FIXED
All four mechanisms confirmed: wrong arity (stage <name> <timeout> <fn>, so $2 was unbound), stderr hidden so the invocation error read as success, stage() returning 0 by design so its exit status could never be the assertion, and the sleep stub firing the watchdog instantly.
It asserts the recorded note plus a post-failure sentinel now, restores the real sleep, and runs a succeeding body through the same path. Verified non-vacuously — with the failing body made to succeed:
🔴 FAIL stage() did not record the failure, or ran past it
🔴 1 harness check(s) FAILED
5. P2 — current-state claims in downstream locations — FIXED
All four. The two worth naming: CT 123 "cannot" expose token counters (it runs llama-server --metrics and does — your zero-counter check was the right shape of evidence, proving exposure not usage), and the parent README still saying the second card runs llama-swap. The "stock map" warning is relabelled historical with the general hazard kept.
6. P3 — the provenance correction compared the wrong tree — FIXED, retraction withdrawn
CT120 baseline: 32687325… CT120 release: 4ac7aa75… CT123 baseline: 32687325…
CT 120 holds both trees, its baseline is byte-identical to CT 123's, and the retraction compared CT 123's baseline against CT 120's release. The original copy history stands. Withdrawn, with the lesson recorded where it happened: name the tree, not the container.
Also landed
Env consolidation (your closing note): repeating the correction narrative in env comments hadn't prevented drift and had made the files unreadable — 276 lines across the two, 199 comment, for 40 settings. Now 204 and 128, each setting carrying its value, its trap and a pointer, history only in the README. All 40 VAR=value lines byte-identical, checked mechanically. Rewriting surfaced two more defects: comment blocks had drifted off their settings in both files, and the q8_0 ↔ --reasoning off XOR warning was missing entirely from the live config, the one actually running q8_0.
cutover-transition-test.sh — the one follow-up test you argued for. Drives the real script against a mocked pct/systemctl, asserting that at no point may two containers hold GPU 2, running or after a reboot. Six scenarios including failure injected during and after the release, and the deliberately stopped rollback target. Script paths are overridable now (defaulting to production) to make it possible.
Proven non-vacuous — five mutations, one at a time, each caught with the right message: watchdog map → disabled unit; autostart never cleared; assertion blind to a stopped container; autostart not restored on rollback; GPU 2 never detached on rollback.
rel="$(grep … | head | cut)" with no match exits 1, pipefail propagates, the assignment inherits it, set -e kills the run. The same trap this folder hit before, inside the test written to catch such things. Fixed with || true; the harness now distinguishes caught / not-caught / aborted.
Gate
git ls-files '*.sh' | xargs shellcheck -S warning CLEAN (0.11.0)
git diff --name-only -z origin/main...HEAD … | xargs -0 shellcheck -S warning CLEAN
bash -n (all shell files) CLEAN
py_compile (all python) CLEAN
cutover-transition-test.sh PASS
stage-harness.sh PASS
The context-window--concurrency anchor in pro-v620/README.md is still broken; it also breaks on origin/main, so it predates this PR and fixing it means guessing the intended target. Left alone deliberately.
No live service was touched in this round — everything is mocked and runs in temp dirs.
Open by choice, not oversight: the native 262144 window, re-measuring two cards if llama.cpp #28699's per-device indexer fix lands, -ncmoe 16 + 31,17 + 4096/1024, and the DIMM-count comparison when the sticks arrive.
Two findings from the 45b7f47 review, both fixed. Neither is a regression from the recent commits — finding 1 is an older path the new transition test made visible, which is the test doing its job. 🔴 P1 — a failed `systemctl disable` was swallowed in both directions: pct exec "$VMID" -- systemctl disable --now llamacpp 2>/dev/null || true If that failed the script carried on attaching GPU 2, restarting the container, rewriting the watchdog map and enabling the incoming unit, leaving BOTH model services enabled and active. They then contend for the card and for port 1234, and the map points only at the incoming one — so a thermal trip sheds the wrong load. Reproduced against the committed test with the mocked disable forced to fail, and with container start given reboot semantics: OBSERVED CT 120 units after a FAILED outgoing disable: 120 llamacpp enabled active 120 llamacpp-qwen38fn enabled active all cutover transition checks passed <- the test could not see it Fixed narrowly, via a retire_unit helper used in both directions.⚠️ The `|| true` on the disable itself is KEPT and that is deliberate: the unit may legitimately be absent on a fresh install or in a direction already taken. What is no longer optional is the verification — still enabled, or still active, is now a fatal precondition failure. Tolerate the command failing; never tolerate the unit surviving. Retirement precedes the GPU mutation in both directions (204 before 210, 237 before 241), so an abort leaves no half-attached card. The test gains the invariant it was missing — the OUTGOING unit must be disabled AND inactive after each direction — plus scenario 3c, where the retirement fails and the script must refuse before touching GPU 2. The mock's `start` now reactivates enabled units, which is what makes "still enabled" a falsifiable claim rather than bookkeeping. Non-vacuity proven by restoring the pre-fix one-liner: three distinct failures, including "cut over while the old model server was still enabled".⚠️ Two smaller mutations did NOT fail the test, and that is correct rather than vacuous — removing either the is-enabled or the is-active check leaves the other one catching the scenario. They guard different real states (comes back on reboot vs running right now), so both stay. P2 — the README's validation record claimed more than was measured. Its title said "verification of the shipped two-card default" while what ran was the THEN-default at 4096/1024, and its sweep row silently came from a different binary. Retitled to name the previous default, with a binary column, the three-case exercised/not-exercised matrix moved in from the env, and the reduced-batch mitigation labelled as the inference it is rather than an observed headroom result. "Do not quote 2881 / 1878 MiB as this configuration's headroom" is now stated where someone would read it. Gate: shellcheck -S warning clean on 0.11.0 by both commands, bash -n clean, py_compile clean, transition test PASS (22 checks), stage harness PASS. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review triage —
|
| exercised? | |
|---|---|
| the release binary | yes — at 4096/1024, the run that found 877 MiB |
| 1024/256 | yes — on b11018-baseline, the sweep |
| the release binary AT 1024/256 (what ships) | no |
And the mitigation is labelled as the inference it is: smaller compute buffers cannot use more memory than the run that already fit, so the shipped pair should have strictly more headroom — sound, and still not a measurement. "Do not quote 2881 / 1878 MiB as this configuration's headroom" now appears where someone would read it. No rerun performed or requested.
Gate
git ls-files '*.sh' | xargs shellcheck -S warning CLEAN (0.11.0)
… origin/main...HEAD … | xargs -0 shellcheck -S warning CLEAN
bash -n / py_compile CLEAN
cutover-transition-test.sh PASS (22 checks)
stage-harness.sh PASS
context-window--concurrency in pro-v620/README.md is still broken and still breaks on origin/main — predates this PR, left alone rather than guessing its target.
Nothing ran against the host this round; all cutover execution was against mocks in temporary paths.
Net -49 lines across 12 files. No fact removed, no trap removed — only the narrative scaffolding around them. Across four review rounds I kept correcting things by APPENDING: "an earlier revision said X, that was wrong, actually Y". Thirteen such blocks accumulated in the README, CLAUDE.md, both env files and six scripts. The reviewer pointed out twice that this had not prevented propagation drift, and it hadn't — the volume made the files harder to read while the wrong claims kept surviving somewhere else anyway. The rule applied here: if a claim was wrong, replace it with the right one. A reader needs the current fact and the trap that will bite them; they do not need the history of my mistakes. What was kept, deliberately: every warning that stops someone re-introducing a defect, rephrased as a property of the system rather than as a confession. "The `- 2` is load-bearing: card 1 also carries the output head and a larger KV share, so the obvious rule overcommits it and spills" is a trap. "CORRECTED 2026-09-18, this used to say ..." is a diary entry. Same information, and only one of them is useful at 2am. Biggest removals: the `-ncmoe 15` closure no longer opens with "two wrong things were said about 15 in a row"; the q8_0 section states that the ~14% belonged to a GTT spill rather than narrating its own retraction; the candidate table gives its verdict instead of explaining which retracted figure produced the previous one; CLAUDE.md's provenance note names the three build trees instead of withdrawing a withdrawal; and the parent README states the ROMED8-2T topology rather than quoting the B550 line it replaced. Also trimmed one pre-existing instance of the same pattern in create-lxc-llama-swap-gpu2.sh (the n-predict comment). The copy in coder-runner/ is left alone — that file is outside this PR's diff and the container is decommissioned. Gate: shellcheck -S warning clean on 0.11.0, bash -n clean, py_compile clean, cutover-transition-test.sh PASS, stage-harness.sh PASS, README anchors resolve. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Deploys Qwen3.8-Flash-Next (
qwen4exp, 180B/6B, UD-Q4_K_XL, 111 GB) — the first modelhere that does not fit in VRAM — and settles where it should run with a full placement study.
Where it landed
qwen3.6-35b-a3b— 74.8 tok/s, Hermes' model-ncmoe 34— 13.0 tok/sllama-swapremoved from CT 123 (models left on disk;/modelsgrown 180→240 GB ratherthan deleting anything). CT 122
coder-runnerdestroyed — the loop moved to Multica, soit was unused and only something to keep patched.
The three findings that drove it
🔴 Prefill is the constraint, not decode. Across the placement curve prefill falls 63%
where decode falls 36%, and an 8k prompt costs 104 s before the first token at the best
two-card placement, 205 s on one card. Decode at 13–14 tok/s is fine; the wait is not.
That is why this model is wrong for Hermes and why
qwen3.6went back on CT 120.🔴 One card beats two by 21–24% on decode, with prefill identical. Measured as a matched
pair at
-ncmoe 34. llama.cpp #28699ships the QSA indexer's pooled rows across the inter-GPU link every layer, so a split is a
fixed per-token penalty — decode pays it every token, prefill batches and amortises it. Net,
the second card is worth +11% decode but +97% prefill.
🔴 The documented
--tensor-splitrule is ~2 layers off, and the error grows with-ncmoe. It was silently spilling-ncmoe15/16/20 to GTT. Corrected toN + (48-N)/2 - 2; at-ncmoe 20that turns 823 MiB free + a spill into 4 GiB free.Also fixed on the box
cpu-powersave.servicecost −30% of decode —acpi-cpufreq'spowersavepins allcores to 1500 MHz. Replaced by
cpu-governor.servicerunningschedutil, persistent.qwen4expon Vulkan at all (b11013 isthe floor). Brought to b11018.
123:llama-swap), whichwould have turned a thermal trip into a no-op.
--threadswas 32; 16 is better and the optimum inverts with concurrency.Not measured
MTP speculation is blocked in the graph, not the loader. The rebase of
#28097 onto b11018 builds and serves
(control 14.4 tok/s) but aborts on a shape mismatch in
build_hc_mix: b11018 reshaped thehyper-connection gammas and updated its trunk graph, while the PR's MTP graph predates that.
No merge resolution satisfies both graphs — it needs the MTP graph ported. Drafters,
build and patches are staged for a one-command retry (
mtp-patches/,mtp-standalone.sh).Also open: the full native 262144 window at a roomier placement.
Tooling worth keeping past this model
gguf-kv.pysizes any GGUF's KV cache from metadata — the figure that gets overstated 4× ona hybrid, because
n_layerthere means full-attention layers only (12 of 48 here, so24 KiB/token).
stage-harness.shcatches theset -uclass of bug thatbash -nandshellcheck both pass.
attribute-telemetry.pyjoins sampler telemetry to sweep cells so aspilled cell is labelled rather than guessed at.
Corrections recorded in the README
Several claims I made during the run and then had to retract, kept with the measurements that
refute them: q8_0 KV does not cost 14% at depth (that was a spill artifact — it is free at
a placement that fits); a spill costs decode, not prefill; and
-ncmoe 48is not theright way to hand a card back (
-ncmoe 34on one card is 28% faster). The generalisablelesson is that a spilled configuration does not degrade uniformly and can convincingly
impersonate a depth-scaling penalty.
🤖 Generated with Claude Code