Skip to content

Qwen3.8-Flash-Next: placement study, and it lands on CT 123's single card - #71

Merged
marchah merged 49 commits into
mainfrom
marchah/Qwen-3.8-Flash
Sep 18, 2026
Merged

marchah merged 49 commits into
mainfrom
marchah/Qwen-3.8-Flash

Conversation

@marchah

@marchah marchah commented Sep 18, 2026

Copy link
Copy Markdown
Owner

Deploys Qwen3.8-Flash-Next (qwen4exp, 180B/6B, UD-Q4_K_XL, 111 GB) — the first model
here that does not fit in VRAM — and settles where it should run with a full placement study.

Where it landed

CT 120 GPU 1 back to qwen3.6-35b-a3b — 74.8 tok/s, Hermes' model
CT 123 GPU 2 Qwen3.8-Flash-Next on one card, -ncmoe 34 — 13.0 tok/s

llama-swap removed from CT 123 (models left on disk; /models grown 180→240 GB rather
than deleting anything). CT 122 coder-runner destroyed — the loop moved to Multica, so
it was unused and only something to keep patched.

The three findings that drove it

🔴 Prefill is the constraint, not decode. Across the placement curve prefill falls 63%
where decode falls 36%, and an 8k prompt costs 104 s before the first token at the best
two-card placement, 205 s on one card. Decode at 13–14 tok/s is fine; the wait is not.
That is why this model is wrong for Hermes and why qwen3.6 went back on CT 120.

🔴 One card beats two by 21–24% on decode, with prefill identical. Measured as a matched
pair at -ncmoe 34. llama.cpp #28699
ships the QSA indexer's pooled rows across the inter-GPU link every layer, so a split is a
fixed per-token penalty — decode pays it every token, prefill batches and amortises it. Net,
the second card is worth +11% decode but +97% prefill.

🔴 The documented --tensor-split rule is ~2 layers off, and the error grows with
-ncmoe. It was silently spilling -ncmoe 15/16/20 to GTT. Corrected to
N + (48-N)/2 - 2; at -ncmoe 20 that turns 823 MiB free + a spill into 4 GiB free.

Also fixed on the box

  • cpu-powersave.service cost −30% of decode — acpi-cpufreq's powersave pins all
    cores to 1500 MHz. Replaced by cpu-governor.service running schedutil, persistent.
  • CT 123 was on llama.cpp b10678, which cannot run qwen4exp on Vulkan at all (b11013 is
    the floor). Brought to b11018.
  • The thermal watchdog named a service that no longer exists (123:llama-swap), which
    would have turned a thermal trip into a no-op.
  • --threads was 32; 16 is better and the optimum inverts with concurrency.

Not measured

MTP speculation is blocked in the graph, not the loader. The rebase of
#28097 onto b11018 builds and serves
(control 14.4 tok/s) but aborts on a shape mismatch in build_hc_mix: b11018 reshaped the
hyper-connection gammas and updated its trunk graph, while the PR's MTP graph predates that.
No merge resolution satisfies both graphs — it needs the MTP graph ported. Drafters,
build and patches are staged for a one-command retry (mtp-patches/, mtp-standalone.sh).

Also open: the full native 262144 window at a roomier placement.

Tooling worth keeping past this model

gguf-kv.py sizes any GGUF's KV cache from metadata — the figure that gets overstated 4× on
a hybrid, because n_layer there means full-attention layers only (12 of 48 here, so
24 KiB/token). stage-harness.sh catches the set -u class of bug that bash -n and
shellcheck both pass. attribute-telemetry.py joins sampler telemetry to sweep cells so a
spilled cell is labelled rather than guessed at.

Corrections recorded in the README

Several claims I made during the run and then had to retract, kept with the measurements that
refute them: q8_0 KV does not cost 14% at depth (that was a spill artifact — it is free at
a placement that fits); a spill costs decode, not prefill; and -ncmoe 48 is not the
right way to hand a card back (-ncmoe 34 on one card is 28% faster). The generalisable
lesson is that a spilled configuration does not degrade uniformly and can convincingly
impersonate a depth-scaling penalty.

🤖 Generated with Claude Code

marchah and others added 30 commits September 17, 2026 07:00
…t sweep

Now that the EPYC box has 251 GiB of RAM, the first model here that does not fit
in VRAM becomes reachable: Qwen3.8-Flash-Next (qwen4exp, 180B total / 6B active),
UD-Q4_K_XL at 111.33 GB, across both cards with the PLE n-gram table and a tunable
share of the routed experts in system RAM.

CT 123 is stopped to free GPU 2. CT 120 grows to 48 cores / 160 GiB / swap 0 and
its /models disk to 320G. Everything is reversible: ct120-cutover.sh to-qwen36
detaches GPU 2, restores the limits and brings qwen3.6-35b-a3b back, with both
GGUFs and both llama.cpp builds left on disk.

pro-v620/qwen38-flash-next/
  install.sh                 idempotent push of env + serve script + unit into CT 120
  qwen38fn.env               placement config; MODEL_CPU_MOE is the VRAM/RAM dial
  llamacpp-serve-qwen38fn    two loud guards: N-GPU count, and all shards present
  qwen38fn-download.sh       resumable, sha256-verified 112 GB pull
  ct120-cutover.sh           to-qwen38fn / to-qwen36 / status
  placement-sweep.sh         round-robin sweep of --n-cpu-moe x context depth
  placement-probe.py         decode/prefill by prompt class, + template-contract asserts
  summarize-sweep.py         renders the VRAM-freed-versus-decode-lost trade
  thermal-guard.sh           EPYC-correct, fails closed on a missing sensor

Findings that correct the sizing note and CLAUDE.md:

- qwen4exp is already in b10678, so "exactly one build short / needs b10679+" was
  wrong and the model was runnable here before this work. The pin still moves to
  b11018, for three merged qwen4exp follow-ups that post-date b10678 (#27880,
  #27941, #28023), scoped to this model's LLAMACPP_DIR so qwen3.6 is untouched.
- Multi-GPU is the PENALISED path for this architecture: llama.cpp #28699 measured
  the QSA indexer's pooled rows crossing inter-GPU links every layer at 2x decode
  cost on a layer split, and that fix is an open draft. Hence ONE_GPU=true exists
  as a first-class control.
- Decode degrades with context depth for the same reason, a cost no bandwidth
  arithmetic models, so the sweep has a depth axis and numbers are quoted with it.
- The reasoning contract is inverted from Qwen3.8-27B: the template raises on
  "none" and "high", and unset means xhigh, which never answers. Handled server
  side with --reasoning off and asserted by the probe.
- The PLE offload tensor is per_layer_token_embd; the note's ngram_embedding
  alternative is the safetensors name and matches nothing in a GGUF.
- The thermal watchdog's GPU_SERVICE_MAP sent a GPU-2 trip to 123:llama-swap,
  a no-op now that CT 123 is stopped, which would have left the real load on an
  overheating card. The cutover sets and restores it.
- gpu-ab-bench/thermal-guard.sh and sample-gpus.py still name the B550 PCI
  addresses, so the guard dies on its first read and fails silently open. Noted,
  and replaced for this platform rather than patched in place.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…direction

All sixteen flags the serve script passes are confirmed present in b11018's
--help. Two things the check turned up:

- --n-cpu-moe keeps the experts of the FIRST N layers on the CPU, not the last.
  The comment and README said the opposite.
- --reasoning-effort exists as a SERVER flag (minimal|low|medium|high|xhigh|max),
  so there is a fallback if --reasoning off ever misbehaves with this template.
  Wired as MODEL_REASONING_EFFORT, empty by default, with a note never to set it
  to "none" or "high" — the two values this template raises on.

Also pins MODEL_EXPECTED_GPUS and EXTRA_ARGS in the env file rather than leaving
them to the serve script's defaults, since placement-sweep.sh writes both.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…on Vulkan

Corrects the claim made two commits ago. "qwen4exp is already in b10678, so the
KB's one-build-short note is wrong" was only half true, and the wrong half is the
load-bearing one.

The arch STRING is in b10678. But Qwen3.8-Flash-Next uses hyper-connections
(hc_count 4), and the Vulkan backend could not run them until

  qwen4exp: add hc ops                 (#28901)  2026-09-16
  vulkan: support qwen4exp hc ops      (#28988)  2026-09-17 04:34 UTC

b11013 is the first build containing #28988; b11010 is not (verified with the
compare API against the commit sha). So this model could not have run on this
box's Vulkan stack at any point before today, and b10678 is NOT a rollback
target for it — grepping the arch string on b10678 promises a model that cannot
execute. Check that the BACKEND can run an arch, not that the arch is registered.

Both provisioning scripts and CT 120's /opt/llamacpp/current now point at b11018.
Prior builds stay in /opt/llamacpp, so rollback is still a symlink flip — just not
past b11013 for this model.

Other items in b10678..b11018 that matter here:

- models : fix GDN normalization from max to rsqrt (#28068) is a CORRECTNESS fix
  for Qwen3.6-35B-A3B itself, which is a Gated-DeltaNet hybrid — not a
  neighbouring model. Verified after the bump: qwen3.6 serves clean output with
  no <think> leakage.
- memory : avoid allocating V cache for indexer (#28330) frees VRAM on qwen4exp.
- Vulkan: sparse Flash Attention (#28105), topk_moe prefill fusion (#28422),
  type-aligned GET_ROWS (#28253 — the op the PLE lookup uses), small-M matrix
  optimisations for qwen (#28457), argsort data race and OOB fix (#28705).
- For CT 123 when it returns: spec fix for mtmd chunks with DFlash (#28587) and
  server fix for speculation after an image (#28715) both address the coder's
  exact DFlash2+mmproj configuration.

⚠️ CT 123 is stopped and its container is still on b10678; the script pin moved,
so a rebuild is correct, but a plain `pct start 123` brings back the old build.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…onfig

On first load llama-server prints:

  tensor overrides to CPU are used with mmap enabled - consider using
  --load-mode none for better performance

That is this model's whole arrangement: -ot pushes the ~28.7 GB PLE table to the
CPU and --n-cpu-moe pushes expert layers there too. Under mmap those pages are
FILE-BACKED, so the kernel can evict them and fault them back from the SATA
860 EVO — turning a bandwidth measurement into a disk measurement, which is
precisely the failure the sizing note warned about. --load-mode none reads them
into anonymous memory, which with swap 0 and MemorySwapMax=0 genuinely stays put.

Exposed as MODEL_LOAD_MODE, left empty (auto) so the sweep can measure it rather
than assume it. -lm also replaces the --mmap/--mlock/--dio flags deprecated in
this same llama.cpp range (#28334), so it is the forward-looking spelling.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…s mandatory

It runs. Output is coherent, non-degenerate (8-gram ratio 1.0) and reproducible
at temperature 0. These look to be the first numbers for this architecture on a
Vulkan/RDNA2 target anywhere — upstream validated CPU and CUDA only, and the
Vulkan hyper-connection ops merged hours before this ran.

Three findings, in order of how much they change.

1. --tensor-split is REQUIRED with --n-cpu-moe, and nothing warns you.
   --n-cpu-moe N moves the experts of the FIRST N layers to the CPU, so 0..N-1
   are light and N..47 are heavy (~1.56 GB each). llama.cpp's default split
   divides 48 layers evenly BY COUNT and therefore hands card 2 all the heavy
   ones. At -ncmoe 20:

     default split     GPU1 13.4 GiB / 0.2 GTT   GPU2 30.7 GiB / 9.3 GTT    6.6 t/s
     -ts 34,14         GPU1 25.1 GiB /  14 MiB   GPU2 21.2 GiB /  14 MiB   11.74 t/s

   +78% decode and ~3x prefill from rebalancing alone. The cards were never short
   of memory in total — 53 GiB of demand against 60 GiB of capacity. Pure
   maldistribution. placement-sweep.sh now derives the split per ncmoe value
   (card1 = N + (48-N)/2); a fixed split would be wrong for every other N.

2. Nothing is saturated, so the sizing note's method cannot predict this.
   Measured 11.74 t/s against the note's 51. During decode the two cards ALTERNATE
   (6-83% each) and the host CPU sits at 33-43%. Every token walks 20 CPU expert
   layers, then card 1's, then card 2's, synchronising at each handoff — a latency
   cost that is invisible to an "active bytes / bandwidth" model. Every
   hybrid-placement figure in large-moe-build-shapes.md is an upper bound that has
   now been falsified by ~4x, mine included.

3. --load-mode none is llama.cpp's own startup advice here and it is wrong:
   11.18 t/s vs 11.74 on auto, and it pulls 31.6 GiB into GTT. Kept on auto.

Also measured:

- --reasoning off NEUTRALISES the template trap, opposite to what the template
  alone implies. With it set, reasoning_effort of none/high/low/medium/xhigh all
  return clean content, because the raising values never reach the template. So a
  caller carrying CT 123's settings over cannot break this server.
  placement-probe.py --contract asserted the REVERSE, which was wrong; fixed to
  assert "a caller cannot break this" instead.
- Load: 2m38s cold off the SATA SSD, 40-46s warm. That bounds the two-profile
  elasticity idea at ~45s per flip — fine at a role handoff, not per request.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The GPU is the only thermal gap on this board. gpu-blower-control's README
already said the BMC "can only ever react to CPU, board and DIMM temps", but
nothing recorded how to read those sensors or what normal looks like — so it
read as a limitation rather than as coverage we can rely on.

Now documented in CLAUDE.md next to the existing "the BMC has no GPU temperature
sensor" note, and in gpu-blower-control/README.md:

- `ipmitool sdr type Temperature` gives CPU Temp, MB Temp, Card Side Temp,
  Onboard LAN Temp, and one sensor per memory channel, TEMP_CPU1_DDR4A..H.
- An unpopulated channel reports "No Reading", which is the easiest way to see
  which slots are filled without opening the case. dmidecode cannot do this:
  ASRock labels every slot "Locator: DIMM 0".
- Baseline 2026-09-17 under the qwen4exp sweep, 4 of 8 channels populated
  (C, D, G, H): DIMMs 45-50 C, CPU 41 C, fans 1200-2400 RPM. RDIMMs throttle
  near 85 C and the BMC exposes no upper threshold on those sensors, so there is
  ~35 C of headroom. Expect +5-10 C when A/B/E/F are filled — they are currently
  acting as airflow gaps.
- Therefore no CPU or RAM watchdog is needed, and none should be built.

One inference worth keeping: flat DIMM temps are an independent check on whether
a workload is really memory-bandwidth-bound. They did not budge while CPU
utilisation swung 2% -> 50%, which corroborates the sweep's finding — from a
different sensor entirely — that the hybrid qwen4exp placement is latency-bound
rather than saturating the ~92 GB/s the sizing note assumes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
I recorded it as "two adjacent pairs, not the usual one-per-quadrant 4-DIMM
guidance", which implied it was suboptimal. It is not: the placement was chosen
from the board documentation and validated by the owner's own tests, and all four
sticks run at Configured Memory Speed 3200 MT/s.

Corrected in CLAUDE.md, gpu-blower-control/README.md and the bmc-sensor-coverage
memory, and stated positively rather than just removed — so a future reader (or
I) does not re-raise it from generic folklore and try to "fix" a working config.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
All three were in tooling I wrote today, and two of them produced wrong output
rather than an error, which is the worse kind.

1. set_env_var wrote UNQUOTED values, so EXTRA_ARGS="--device Vulkan0" landed in
   the env file as `EXTRA_ARGS=--device Vulkan0`. llamacpp-serve-qwen38fn does
   `set -a; source ...`, which parses that as an assignment followed by a COMMAND
   — bash tried to run `Vulkan0`, the serve script exited 127, systemd crash-looped
   the unit, and the ONE_GPU control died on its first config. Rewritten to build
   the line with json.dumps inside the container, so quoting and escaping are not
   my problem; the python deliberately contains no single quotes so it survives
   being wrapped in them. sed needed & and | escaped plus another shell quoting
   layer, which is what broke it.

2. summarize-sweep.py still read the probe's OLD contract keys. I had changed the
   probe to record `harmless` instead of `raised`, so the summary rendered
   "reasoning_effort none rejected as expected: NO" and "empty <think> siphoned to
   reasoning_content: NO — check --reasoning-format auto" while the raw JSON said
   every value was harmless with clean content. A false failure in the one artifact
   a human actually reads. The contract section now explains what it is asserting
   (a caller cannot break this server, NOT that the template rejects anything) and
   checks the thing that matters: no <think> leaked into content. An absent
   reasoning_content is not a fault — it means no thought block was emitted, which
   is the point of --reasoning off.

3. At n_cpu_moe 48 the derived split is "48,0", so card 2 gets nothing and the row
   is effectively SINGLE-GPU while being labelled two-GPU. The placement itself is
   defensible — there are no heavy layers left to spread, and splitting the
   non-expert layers would only add an inter-GPU hop — but it must not be read as a
   two-card data point. Now logged loudly and recorded in .single_gpu_rows.

Also: placement-sweep.sh now cd's to its own directory, because systemd-run and
cron do not inherit a working directory and the helper scripts resolve relative
to it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Replaces the single 11.7 t/s data point with the completed matrix and the three
conclusions that follow from it.

- VRAM has almost no leverage past a point: -ncmoe 48 holding 9.2 GiB of VRAM
  MATCHES -ncmoe 34 holding 31.3 GiB (8.6 vs 8.4 at d0; 8.2 vs 7.9 at depth). 22
  GiB buys nothing across that span. +48 GiB total is +51% at d0 but only +23% at
  depth 8000 — VRAM's value halves at realistic context depth, which matters
  because the sizing note prices cards on short-prompt arithmetic.
- Reserving ~20 GB for a second model costs 11-14% (-ncmoe 20 -> 28, which frees
  23.6 GiB — comfortably a 27B guest). That answers the question this work
  started from.
- Cost per GiB is non-monotonic (0.04 / 0.12 / 0.11 / 0.06 t/s per GiB), so do
  not interpolate: VRAM is nearly free to give away below -ncmoe 20 and above 34.
- Depth costs 5-22%, inversely to how much is offloaded — the more on CPU, the
  smaller a share the QSA indexer is. Consistent with #28699 being the depth cost.

Also flags in-place that the -ncmoe 48 row is effectively single-GPU (derived
split 48,0; GPU2 measured at 16 MiB) so it is not misread as a two-card point,
and names -ncmoe 34 as the honest two-vs-one comparison.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Adds pro-v620/qwen38-flash-next/membench.sh (STREAM + a random pointer-chase
latency probe) and records the 4-channel results. Re-run verbatim at 8 sticks;
the value is that both stick counts come from one harness.

STREAM on 4 channels, best at 8 threads: Copy 80.3, Scale 51.2, Add 56.2,
Triad 56.3 GB/s.

This settles the open question the opposite way to my expectation. STREAM
reports APPLICATION bytes; Triad moves 24 B/iter in those terms but a normal
store first reads the line it overwrites, so real DRAM traffic is 32 B/iter:

  STREAM Triad  56.3 app -> 75.1 GB/s DRAM traffic   73.3% of theoretical
  stressapptest 76.4                                 74.6%
  STREAM Copy   80.3 (non-temporal stores, no RFO)   78.4%

Triad-corrected and stressapptest agree within 2%, so stressapptest was
ACCURATE, not an understatement — I had claimed it understated and that real
STREAM would reach 77-87. Copy's 80.3 is provably RFO-free (an RFO correction
would put it above theoretical), so ~80 GB/s / 78.4% of peak is the real
ceiling. The honest per-channel figure is 19.1-20.1 GB/s; the old 23.1 was
90.2% of theoretical and never reachable.

Also: 8 threads beat 16 and 32, so four channels saturate early and the
acceptance soak's 16 threads cost only 3%. That will invert at eight channels —
the re-run must sweep threads.

Latency, 4 GiB working set: 141.36 ns/load WITH huge pages, 226.95 without. THP
here is `madvise`, not `always`, so the probe has to call madvise(MADV_HUGEPAGE)
itself or a random walk pays a page-table walk per access — a 38% error, and my
first run made exactly that mistake. Fixed in the harness. 141 ns is not
attributable to 3DS without a flat 2Rx4 set on the same harness.

On the 3DS question, researched against JEDEC JESD79-4 and it inverts the
concern: the core latency chain (CL/tRCD/tRP/tRAS) is unchanged because each die
is a standard DDR4 die at the same bin; the _slr/_dlr timing variants 3DS adds
are generally more relaxed than same-rank equivalents; and on refresh the
delivered part is FAVOURED. At 64 GB x4, 2S2Rx4 3DS implies 8 Gb dies (~4.5%
refresh overhead) while the advertised flat 2Rx4 implies monolithic 16 Gb
(~14.1%). Stacking avoids the penalty the ordered part would have carried.

⛔ The 8-stick test could not run: A/B/E/F are empty, confirmed by IPMI
No Reading on those channels.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The governor finding (-52% from acpi-cpufreq `powersave` pinning 1500 MHz)
invalidated every placement number measured so far, so the whole matrix has to be
re-measured at full clock and then MTP and --parallel layered on top. That is more
hours than a supervised session, so it runs as a chained host-side pipeline.

  overnight-all.sh   lock + thermal-guard precondition + part 1 -> part 2 -> report
  overnight.sh       1 governor (pick winner, PERSIST it), 2 revalidate
  overnight-part2.sh 3 MTP, 4 --parallel, 5 graph splits, 6 restore
  revalidate.sh      the -ncmoe curve, threads, and the three placement EXTREMES
  concurrency-probe.py  per-stream vs aggregate throughput, slot count read back
  morning-report.sh  auto-derived recommendation, so the night answers even unwatched
  stagelib.sh        the stage runner, shared by both halves

Four defects found and fixed while building it, each of which would have produced
confident wrong numbers:

* `timeout <shell function>` exits 127 instantly. Because every stage was wrapped to
  "never stall the night", all six stages reported failure and the pipeline completed
  in two seconds having measured nothing — the failure-isolation meant to protect the
  run is precisely what hid its collapse. stagelib.sh enforces the deadline with a
  watchdog process and its own process group (so pct/python descendants die too), and
  refuses a stage name that is not a defined function.
* No knob reset between cells. MODEL_KV_TYPE=q8_0 and MODEL_MMPROJ_ON_CPU=true were
  found still set on the live server from an earlier extreme, and cell() never cleared
  them, so every row would have silently inherited them. reset_env() now puts every
  knob back to the production default and cells override only what they vary. (Third
  time a leaked knob has invalidated a cell here.)
* The MTP build was written as two cherry-picks onto b11018. PR #28097 already CONTAINS
  #27836's three commits rebased, and both report mergeable:false against master, so
  the pick would have failed. One checkout of #28097, plus a no-speculation control on
  that same binary — comparing two speculative configs measures agreement, not
  correctness.
* Output hashes are not a valid gate at depth. Measured across the governor test: reps
  at d0 always agree, while a d8000 cell disagrees with ITSELF intermittently at
  temperature 0 with top_k 1 — hybrid CPU/GPU expert reduction order varies with thread
  scheduling and ~5.8k tokens of accumulated state is enough to flip a token. Every
  hash comparison and reproducibility gate is now d0-only; medians at depth are fine.

Also: placement-probe.py records draft_n/draft_n_accepted so a speculative config is
scored by ACCEPTANCE rather than tok/s, per-prompt-class because acceptance is
prompt-dependent. The drafter layout is chosen by measurement (two were downloaded and
only one matches what the loader wants), and the concurrency matrix prices a usable
slot size — at 65536 total ctx, --parallel 4 leaves only 16k per slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A plain `git checkout pr-28097` would have built a binary that cannot run this model.

PR #28097 is the right PR — it already contains #27836's three NextN/MTP commits
rebased, plus the draft-head-only GGUF layout that the drafter downloaded here
actually uses. But its base is **307 commits behind b11018**, and what it lacks
includes `35822afe5 vulkan: support qwen4exp hc ops` (#28988) — the commit that makes
qwen4exp's hyper-connection ops work on Vulkan at all. So the PR as published cannot
offload this architecture to a V620, and any speculation measured on it would have
been meaningless. It also predates #25483 (skip unneeded MoE work in the mul_mm
coopmat1 path) and #28996, both directly relevant to a Vulkan MoE.

The four commits are therefore rebased onto b11018, with the resolution saved under
mtp-patches/ so it is reproducible rather than living only on the builder LXC. Two
commits conflicted in src/models/qwen4exp.cpp because b11018 had moved underneath the
PR: n_ff_exp became a per-layer array, every hc_*_norm / ple_norm_* / hc_head_norm
gamma was reshaped {hc_dim} -> {n_embd, hc} with TENSOR_ALLOW_RESHAPE, the generic
loader now reads LLM_KV_NEXTN_PREDICT_LAYERS before the arch hook, and b11018's PLE
row-count block became a strict superset of the PR's.

Resolution rule throughout: keep b11018's shapes and logic, OR in the PR's flags. That
is unambiguous rather than a judgement call because every create_tensor call outside
the conflicted hunks already used `flags` — those lines merged cleanly, which is what
pins down what the conflicted ones must look like. The PR's require_weight() rewrite of
the PLE block was deliberately not taken: it regresses b11018's no-tensor case and its
`const int64_t ple_rows` cannot compile against b11018's assigning loop.

Verified: 4 commits over b11018 with b11018 an ancestor, and `g++ -fsyntax-only` on the
resolved file using the build's own flags. The build stage now asserts the commit shape
before compiling, and asserts both `draft-mtp` in --help and ggml-vulkan in the output
afterwards — the Vulkan backend being the entire reason for the rebase.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…hiding a GTT spill

Caught mid-run by a new external sampler. At `-ncmoe 15` with the rule this repo
documents

    c1 = ncmoe + (48 - ncmoe) / 2        ->  31,17

card 1 sat at 30665 MiB with **39 MiB of VRAM free and 2060 MiB spilled to GTT**, while
card 2 had 3280 MiB free. That is the ~12x decode collapse the startup loud-guard does
not catch, so the cell reads as "a slow placement" rather than "a broken one" — which is
exactly how it has been read until now.

The imbalance measures 1620 MiB against ~1640 MiB per heavy layer: **exactly one layer.**
The documented rule balances by WEIGHT alone, but three other things live on the card
that owns a layer:

  * the KV cache. 12 of the 48 layers are full attention (Qwen Sparse Attention every
    4th), and at c1=31 card 1 owns 7 of them against card 2's 5.
  * the vision projector, +1.11 GiB, which lands on a single device.
  * per-device compute buffers.

So the corrected rule is `c1 = ncmoe + (48 - ncmoe) / 2 - 1`, and a new stage tests it at
the placements that were at or over the edge (plus the documented value at 15 as the
control that reproduces the spill, and the incumbent 20 to see whether a config that
already fit also gains). It reports **free VRAM and GTT for both cards** and a
fits/tight/SPILLED verdict, because "used VRAM" looks healthy right up to the cliff.

Its winner — the fastest placement verified NOT to be spilling — is then what the MTP,
--parallel and restore stages use, in preference to ranking the re-validation data, which
was taken before the spill was known and can rank a broken cell first.

  placement-sampler.sh  logs free VRAM, GTT, busy %, junction/mem temps, watts, busy-core
                        p90 clock, RAPL package energy, host memory and every populated
                        DIMM channel to placement-telemetry.jsonl

It is a separate process rather than an edit to revalidate.sh on purpose: bash reads a
script incrementally, so rewriting one mid-run can corrupt its execution. This is the
non-invasive way to add instrumentation to a sweep that is already hours in.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…not guessed

The sweep records a cell's throughput but not its VRAM headroom or GTT, and that gap is
load-bearing: below roughly 1 GiB of free VRAM RADV spills to GTT, and the startup
loud-guard does not catch it. A spilled cell posts a plausible number and reads as "a
slow placement" rather than "a broken one".

Joining the two immediately produced a result worth having. The first cell of the
re-validation, `-ncmoe 15` at the documented split:

  ncmoe15-t32-r1   13.68 t/s d0   13.57 t/s d8k   39 MiB free / 2060 MiB GTT   SPILLED

It is **the fastest cell so far while spilling**. That is the opposite of what this repo's
existing GTT note would predict, and the difference is what spills: the 12x collapse on
record was the KV CACHE landing in GTT (CT 123, 2.0 tok/s), whereas here it is ~2 GiB of
expert weight, read by a subset of experts once per token, on a workload already bound by
a ~43 ms fixed engine-overhead floor that masks the extra hop. The splitfix stage measures
the cost directly against the corrected split rather than leaving it as inference.

Also worth recording from the same data: CPU package power is 80 W under this load against
54 W idle (RAPL), the GPUs peak at 119/147 W, and schedutil's busy-core p90 clock ranges
2400-3300 MHz -- it does boost to the full 3300, it just does not sit there.

⚠️ The tool reads FREE VRAM, not used. Used looks unremarkable right up to the cliff: the
cell above sat at 30665 MiB used, which says nothing, while having 39 MiB free.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…actually fits

"Most versatile" is not only how much VRAM is left for a second model — it is also how
long a context this thing can serve. That was a guess; now it is arithmetic plus a
measurement.

gguf-kv.py reads a GGUF's metadata header with no dependencies and sizes the cache. For
Qwen3.8-Flash-Next:

  48 blocks, full_attention_interval 4  ->  only **12 blocks hold a KV cache**
  head_count_kv 2, key_length 256, value_length 256
    => 2 * 12 * 2 * (256+256) * 2 B  =  **24.0 KiB/token at f16**, 12.0 at q8_0

  ctx      f16        q8_0
  65536    1.50 GiB   0.75 GiB
  131072   3.00 GiB   1.50 GiB
  262144   6.00 GiB   3.00 GiB    <- the model's native maximum

The other 36 blocks are Gated DeltaNet, whose state is fixed-size and
context-independent — which is the whole reason a 180B model can hold a 256k window on two
32 GB cards at all. Getting this wrong by treating `n_layer` as 48 would overstate the
cache 4x.

Two independent cross-checks that the parse is right: the PLE head vocab sizes sum to
~320M rows at 160 dims, which is 28.8 GB at this quant and matches the ~28.7 GB the env
file claims; and expert_count 512 x ff 640 x embd 2560 x 3 matrices comes to ~1.42 GB of
expert weight per layer against the ~1.56 GB the placement dial assumes.

A new stage 2c then confirms it by VRAM delta rather than trusting the metadata, because
flat VRAM across two context sizes can mean the KV moved to host memory rather than got
cheaper — so GTT is printed on every row. It tests the fastest placement at 131072/f16 and
at the full native 262144/q8_0, plus 262144/f16 on a roomier placement, to price the real
choice: quantise the cache, or give up GPU-resident experts. Each row also probes at 32k
depth, since a long window is worthless if throughput collapses inside it.

Stage order stays as asked: every other test, then MTP, then --parallel last.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…cannot fit at f16

The "-1 layer" correction committed earlier was derived from a single cell. Three cells in,
with GTT counted as demand that did not fit, the picture is different:

  ncmoe  split    card 1   card 2    total   spare of 65536   imbalance
  15     31,17     34789    29506    64295             1241      +5283
  20     34,14     32046    24741    56787             8749      +7305
  28     38,10     26521    18254    44775            20761      +8267

Two corrections to my own earlier claim:

* **The imbalance is not one layer and it grows with -ncmoe.** At ~1500 MiB of expert
  weight per layer — derived from the 15->20 pair, against the ~1.56 GB the env file
  assumes — moving one layer shifts the imbalance by ~2x that, so the correction is roughly
  **-2 layers**, and larger further up the curve. The stage now brackets -1/-2/-3 at the
  hungry end rather than testing one guess, and also tests the correction at -ncmoe 20 and
  28 where the curve actually lives.

* 🔴 **-ncmoe 15 cannot safely fit at all on f16 KV with the projector on GPU, whatever the
  split.** Total demand is 64295 MiB against 65536 of capacity: ~620 MiB per card even
  perfectly balanced, which is below the ~1024 MiB where RADV starts spilling. That is why
  the first cell posted 13.68 t/s *while* holding 2060 MiB in GTT. Freeing 1.86 GiB — q8_0
  KV (lossless here because the server runs --reasoning off) plus --no-mmproj-offload —
  raises it to ~1573 MiB/card, so the two genuine max-VRAM candidates are now tested as
  such: 15 and 14 at the corrected split with that shape.

Consequences threaded through the rest of the pipeline:

* The winner is no longer "fastest cell". A spilled cell can post a competitive number, so
  split_cell writes its headroom verdict to a `.verdict` sidecar and the picker refuses
  SPILLED outright, on top of the existing degeneracy gate.
* The winner now carries its whole SHAPE (ncmoe, split, KV type, projector placement), not
  just the placement. The max-VRAM candidate only fits *because* of q8_0 + projector-on-CPU,
  and the MTP / --parallel / restore stages previously reset both to "production defaults" —
  which would have silently pushed the winning config back over the edge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…fter round 1

Measured pace forced a re-plan. Cells are running 12 -> 15 -> 16 min and rising (the
CPU-heavy end of the curve is slower), so revalidate's 2 rounds plus its 6-cell
confirmation pass comes to ~7.5 h. That is past the stage's own 6 h timeout, so it would
have been killed mid-round-2 and lost the confirmation pass anyway, while pushing MTP and
--parallel past morning.

Two changes.

1. **The warm pass no longer runs all three prompt classes** (`--classes`, new in
   placement-probe.py). Warming is about weights, the graph and the deep-prefill indexer
   state, none of which is prompt-class specific, and at depth each class costs a full 8k
   prefill — six 8k prefills per cell is over half the wall clock. The MEASURE pass keeps
   all three, because decode speed is prompt-dependent and one class is not this model's
   throughput. The flag defaults to all classes, so the already-running sweep is unaffected.

2. **trim-revalidate.sh stops the sweep when round 2 begins**, appending its reasoning to
   RESULTS.md first so the record says why rather than showing an unexplained TIMED OUT.

Round 2 is worth less than it looks: it would be a second sample of a curve measured at the
DOCUMENTED --tensor-split, and that split is now known to overload card 1 by ~2 layers —
-ncmoe 15 and 20 both spill to GTT there. The split-correction stage re-measures the same
placements at corrected splits with 2 reps, so that is where the placement conclusion comes
from. Round 1 stands as single-sample context, not as the answer. Spending the remaining
hours on splitfix, contextsweep and MTP is strictly more information than a second sample
of mis-split cells.

Round 1 so far, all at the documented split, so all three are context rather than
conclusions: -ncmoe 15 -> 13.68 / 13.57 t/s (d0/d8k) but SPILLED, 20 -> 12.69 / 11.53 also
spilling marginally, 28 -> 11.29 / 10.56 with real headroom.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…experts

Round 1 of the re-validation is in, and the headline is not the column anyone was
watching. Decode and prefill do not degrade at the same rate as experts move to the CPU:

  -ncmoe   decode d0   decode d8k   prefill d8k   total VRAM
  15           13.68        13.57          78.9        58089   (spilled)
  20           12.69        11.53          63.6        52540   (spilled)
  28           11.29        10.56          47.4        40527
  34           10.45         9.44          38.8        31267
  48           10.18         8.70          29.0         9443

Prefill falls **-63%** across that range against decode's -36%, and it is the binding
constraint in absolute terms: an 8k prompt costs **101 s before the first token** at the
best placement and **276 s** at -ncmoe 48. Decode at 13.7 tok/s is perfectly usable. What
makes this model feel slow is prefill, and no amount of decode tuning touches that.

Two consequences:

* There is no single best placement. Long-prompt or agentic work wants -ncmoe as low as
  fits, because prefill dominates. Short-prompt chat can take -ncmoe 48, which costs only
  -2.6% of decode against -ncmoe 34 while freeing another 21.8 GiB, and leaves the whole
  model in 9.4 GiB of VRAM -- one card, with 23 GiB spare for a second service.
* The d0 prefill column is an artifact and must not be quoted: those prompts are ~30
  tokens, so the figure is fixed overhead, not throughput. Only the d8k column means
  anything, which is why it is the one reported.

Also: splitfix is now two-phase. "Does this shape fit" is a MEMORY question answered by
mem_info_vram_free + mem_info_gtt_used at load, so spending a 12-minute throughput probe on
it was waste. Phase 1 load-checks all nine candidates in ~2 min each (including a tiny
completion first, because llama.cpp allocates some buffers lazily and reading headroom
straight after /health would miss them); phase 2 probes only the three lowest -ncmoe shapes
that survived, plus the documented split at 15 as the control that reproduces the spill.
That is ~50 min instead of ~88, and strictly better science: every candidate gets a
headroom verdict instead of only the ones there was time to probe.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Measured at -ncmoe 20, round 1:

  threads   decode d0   decode d8k   prefill d8k
  8             13.13        11.87          63.8
  32            12.69        11.53          63.6
  delta         +3.5%        +2.9%          +0.3%

8 threads is faster on decode and neutral on prefill, so it is free -- and it hands back 24
physical cores that CT 120 currently holds for nothing. This corroborates the STREAM result
already in the README from a completely different direction: 8 threads saturated the four
populated memory channels at 80.3 GB/s while 32 measured WORSE at 74.8. The CPU-side expert
FFN is bandwidth-bound GEMV, so past the point where the channels are saturated more
threads is contention, not throughput.

Two consequences wired in:

* The --parallel sweep now includes 8 threads (1 and 4 streams) rather than only 16 and 32.
  It is ADDED, not substituted: with four concurrent streams contending for the same
  expert FFN the ordering can invert, and 32 stays as the incumbent control.
* The restore stage no longer hardcodes --threads 32. It reads back the best single-stream
  thread count the concurrency stage measured, so the box is left on the measured value
  rather than the inherited one. Restoring 32 over a better measured number would have been
  a silent regression, and the kind nobody notices because nothing errors.

⚠️ Still n=1 per cell. Spread at full clock is ~1.4%, so +3.5% is above noise but not by a
wide margin; the threads-16 cell lands next and a monotonic 8 < 16 < 32 ordering is what
would confirm it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous commit said 8 threads was the winner on the strength of 8-vs-32 alone. With
the 16-thread cell in, that was the wrong conclusion from an incomplete sweep:

  threads   decode d0   decode d8k   vs 32
  8             13.13        11.87   +3.5% / +2.9%
  16            13.19        12.11   +3.9% / +5.0%
  32            12.69        11.53   --

What survives: **32 is clearly contention**, -4 to -5% against either lower value, and the
STREAM result still explains why (8 threads saturated the four populated channels at
80.3 GB/s where 32 measured worse at 74.8 -- bandwidth-bound GEMV gains nothing past
saturation and then starts to contend).

What does not survive: "use 8". 8 and 16 are within ~2% at n=1, with 16 ahead at depth, so
the honest statement is a tie between them rather than a winner. The --parallel stage
measures all three at 2 reps, which is what should settle it.

The restore fallback moves 8 -> 16 accordingly. It only applies if the concurrency stage
produced nothing; normally restore reads the measured value back.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This repo's standing guidance is that "q8_0 KV costs zero throughput (marginally faster,
unchanged at 2x context)". That was measured on CT 123's qwen3.8-27b and **does not
transfer to qwen4exp**. Measured tonight at the same placement and thread count:

                        d0      d8000    VRAM
  f16, projector GPU    12.69   11.53    52540 MiB
  q8_0, projector CPU   12.17    9.87    50639 MiB
  delta                 -4.1%   -14.4%   -1901 MiB

The cost scales with context depth, which is mechanistically what this architecture would
predict: qwen4exp decode is dominated by the QSA indexer rescoring the cached context every
token (llama.cpp #28699), so a quantised cache means dequantising on every indexer pass.
At d0 there is almost no context to rescore and the cost is -4%; at 8k it is -14%.

Two reasons to believe it rather than treat it as noise:

* -14.4% is an order of magnitude above the ~1.4% run-to-run spread at full clock.
* The confound points the WRONG way. The q8_0 cell had ~1900 MiB more headroom and less
  GTT than the f16 control, so spill pressure would have made it the faster of the two.
  It was slower anyway.

Consequence for the recommendation: "free up VRAM with q8_0" is a worse trade here than it
looked. The frugal shape does not even buy a whole extra placement step -- it saves
~1900 MiB while one layer of experts costs ~1500 -- so -ncmoe 14 with q8_0 + projector on
CPU (63895 MiB) lands only just below -ncmoe 15 with the production shape (64295), and
still spills: 820 MiB/card even perfectly balanced, under the ~1024 MiB RADV floor.

So the max-VRAM placement that actually fits is one step back:

    --n-cpu-moe 15 --tensor-split 29,19 --cache-type-k/v q8_0 --no-mmproj-offload

62394 MiB total, ~1571 MiB/card balanced, above the floor. That is exactly the candidate
splitfix phase 1 load-checks, so it gets confirmed rather than assumed -- and now it has to
be scored against its ~14% depth cost, not just its headroom.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ad INVERTED

Every number in this README was taken with the governor pinned to `powersave`, i.e. all 64
threads at 1500 MHz against a 3308 MHz maximum. Re-measuring at `schedutil` did not just
move the magnitudes; it reversed the direction of two conclusions the file stated
confidently. Superseded tables are deleted rather than annotated.

**Depth cost inverted.** The old table read a clean monotonic "the more that sits on the
CPU, the LESS depth hurts" (−22% at -ncmoe 15 down to −5% at 48). At full clock it runs the
other way: −0.8% at 15 rising to −14.5% at 48. The old trend was an artifact of every core
being parked at minimum frequency.

**-ncmoe 48 vs 34 flipped sign.** Old data had 48 (8.6) beating 34 (8.4), which supported
"22 GiB of VRAM buys nothing". The claim survives in magnitude — 48 is now −2.6% against 34
for 21.8 GiB freed — but the ordering was wrong.

New, and the most decision-relevant thing in the file: **prefill is the binding constraint,
not decode.** Across the placement curve prefill falls 63% where decode falls 36%, and in
absolute terms an 8k prompt costs 101 s before the first token at the best placement and
276 s at -ncmoe 48. Decode at 13.7 t/s is fine; the wait is not. So there is no single best
placement — long prompts want -ncmoe as low as fits, short prompts can take 48 and free
54.8 GiB.

Also corrected: **the `--tensor-split` rule of thumb this file gave is ~2 layers off**, and
it was hiding a spill. `card1 = N + (48-N)/2` balances by weight alone, ignoring that the
card owning a layer also holds its KV (12 of 48 blocks are full attention, unevenly split),
the 1.11 GiB projector, and per-device compute buffers. Measured imbalance is +5283/+7305/
+8267 MiB at -ncmoe 15/20/28 — growing with -ncmoe, not one layer. At -ncmoe 15 card 1 sat
at 39 MiB free with 2060 MiB in GTT while card 2 had 3280 MiB free, and that cell still
posted the sweep's fastest decode, which is exactly how a broken cell passes for a slow one.

And two new sections: q8_0 KV costs ~14% at depth here (it does not on CT 123's
qwen3.8-27b, which is where the "q8_0 is free" note came from), and --threads is an
inverted U peaking at 16 rather than the inherited 32.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Answering the question directly: running this model entirely on the CPU to free both V620s
costs too much.

  config                        decode d0   decode d8k   prefill d8k   8k TTFT   VRAM
  -ncmoe 20, --threads 16           13.19        12.11          63.6     126 s   52540 MiB, both
  -ncmoe 48                         10.18         8.70          29.0     276 s    9443 MiB, ONE card
  -ngl 0 CPU-only, --threads 16      6.04         5.51          26.4     303 s    2181 MiB

-54% against the fastest fitting GPU config, -41% against -ncmoe 48.

The percentage is not the real argument though. **-ncmoe 48 already runs the whole model in
9.4 GiB on a single card** — with no heavy layers left the derived split is 48,0, so card 2
holds nothing and is free for another service, with ~23 GiB still spare on card 1. There is
therefore never a reason to reach for CPU-only in order to free a GPU: -ncmoe 48 frees one
outright and is 69% faster.

A mechanistic detail worth keeping: CPU-only costs only -9% of PREFILL against -ncmoe 48
(26.4 vs 29.0 t/s), because at that placement prefill is already CPU-bound on the experts.
What the cards buy at -ncmoe 48 is almost entirely the attention and non-expert path during
decode, and that is worth +69%. So "the GPU helps prefill" is the wrong intuition here; it
helps decode.

⚠️ Recorded honestly: "both cards free" is not literally what was measured. reset_env leaves
the vision projector resident, so ~2.1 GiB stays on a card. --no-mmproj-offload frees that
too, at a cost to image encoding only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…s, which killed splitfix

splitfix died with rc=1 after 0 minutes and the pipeline moved on to contextsweep — which
consumes splitfix's winner, so it ran in the wrong order with a fallback placement. The
cause is a bash rule subtle enough to be worth writing down:

    lbl="nc${nc}-c${c1}$([ -n "$kv" ] && echo "-q8")$([ "$mp" = true ] && echo "-projcpu")"

**A variable assignment takes the exit status of its command substitution.** Both
substitutions here exit 1 whenever their test is false, so the assignment returned 1 and
`set -Eeuo pipefail` aborted the function on its very first candidate. The trace makes it
unmistakable once seen:

    ++ '[' -n '' ']'
    ++ '[' '' = true ']'
    + lbl=nc15-c31          <- returns 1

What made it look impossible is that the *identical* construct two lines below, passed as a
command argument, is completely safe: there the exit status belongs to the command, not the
substitution. A grep for the pattern found exactly one instance in assignment position and
three in argument or ||-guarded position, which is why only this one was fatal.

Fixed by building the label with plain conditionals and a trailing `:` to keep the function's
exit status clean. Verified by re-running the whole stage in a stubbed harness — it now
completes with all nine candidates load-checked and phase 2 plus the control invoked.

Also here: chain-report.sh, because overnight-all.sh had to be killed to re-order part 2, so
the report generation it would have run at the end needed re-arming. It waits for part 2,
runs morning-report.sh, and stops the telemetry sampler so its file is closed for analysis.

And RESULTS.md is tidied: two tables whose headers were emitted before the stage died are
removed, and the failure line now records that this was a script bug rather than a
measurement, so nobody reads it as "the split correction could not be measured".

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ifact

Commit ce74a40 claimed q8_0 KV costs ~14% at depth on qwen4exp, contradicting this repo's
standing note. That was wrong. Measured as a clean pair at a placement that FITS
(-ncmoe 16, split 30,18, neither cell spilling, same thread count):

  shape                        decode d0   decode d8k   free c1 / c2
  f16 / projector on GPU           14.16        13.07   1305 / 1553 MiB
  q8_0 / projector on CPU          14.17        13.45   2881 / 1878 MiB
                                   +0.1%        +2.9%   +1.6 GiB

Identical at d0 and marginally FASTER at depth, while freeing 1.6 GiB — which is exactly
what the original note said. The -14% came from a pair where both cells sat on the
documented (spilling) --tensor-split, and it does not reproduce once the placement fits.

The methodological lesson is the part worth keeping: **a spilled configuration does not
degrade uniformly.** It cost that pair far more at depth than at d0, which is
indistinguishable from a depth-scaling penalty by inspection. So: establish a fitting
placement first, then vary one knob. I also varied two things at once there (KV type AND
projector placement), which means even the sign was unattributable — still true of the new
pair, so the honest claim is that *the shape* is free, not that q8_0 specifically is.

I defended the original number at the time by arguing the confound pointed the wrong way
(the q8_0 cell had more headroom and was still slower). That argument was not wrong so much
as insufficient: "more headroom" was inferred from total VRAM rather than measured per-card,
and a marginal cell's behaviour is not linear in its headroom. A clean pair beats an
argument about a dirty one.

Consequence for the recommendation: the q8_0 + projector-on-CPU shape is now the preferred
one at -ncmoe 16 — same speed, 1.6 GiB more headroom on a placement that is only "tight" in
f16. Both rows are kept because --no-mmproj-offload costs 3-5x on image encoding, so anyone
who cares about vision latency should take the f16 row knowingly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…on measurements

Two stages failed tonight for reasons that had nothing to do with the hardware.

**contextsweep, rc=1 in 2 seconds.** When best_split.txt grew from two fields to four
(ncmoe, c1, kv, mmproj) I updated s3_mtp's `read -r NC C1 KV MP` and forgot s2c_context,
which still did `read -r NC C1`. Bash puts the remainder in the last variable, so C1 became
"30 q8_0 true" and `$(( 48 - c1 ))` was an arithmetic syntax error two calls later. Both
readers now consume all four fields, and the unused ones are named `_KV`/`_MP` so the next
person can see they are deliberate.

**mtp, rc=2 after 87 seconds — but the build SUCCEEDED.** 585 targets linked and both
guards passed (`--spec-type draft-mtp` present, ggml-vulkan present). Then:

    tar: pr28097-mtp: Cannot stat: No such file or directory

I renamed the build output to mtp-b11018 when the rebase replaced the PR checkout, and left
the tar/push lines on the old name. The 58 MB artifact is intact on CT 201, so the fix adds
MTP_SKIP_BUILD=true, which reuses it — and **re-asserts both guards against the existing
build before trusting it**, since reusing an artifact is only safe if it still carries the
flag and the backend it was built for.

Recovery shape, and the reason it is not just "edit and re-run": part 2 was still executing
its remaining stages, and **bash reads a script incrementally — rewriting one mid-run shifts
byte offsets and can corrupt execution.** So the fixes went to overnight-part2-fixed.sh and
three standalone runners EXTRACT their stage functions from it with sed rather than
duplicating them, which means the fixed copy stays the single source of truth:

  contextsweep-standalone.sh   the context sweep
  mtp-standalone.sh            MTP with MTP_SKIP_BUILD=true
  restore-standalone.sh        the final restore
  chain-finish.sh              waits for part 2, then runs the three in order + the report

Order matters: restore runs LAST, because contextsweep and mtp both rewrite the model env,
so part 2's own restore is stale by the time they finish.

restore-standalone.sh is a file rather than an inline `bash -c` in the chain because
inlining it needed three levels of nested quoting around an embedded python heredoc — which
is precisely the shape of bug that cost this pipeline two stages tonight.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…nder load

Measured at the winning placement, the thread ranking is not stable across concurrency:

  threads   par 1    par 2    par 4
  8         14.25      --     29.37    <- best solo, WORST at 4 streams
  16        14.09    22.10    30.88    <- within 1% solo, best at 2 and 4
  32        13.22    20.46    30.46

8 threads saturates the four populated memory channels for a single stream, which is why it
wins there and why STREAM predicted it. But four concurrent streams present more parallel
work than 8 threads can cover, so it loses 4.9% to 16 at par 4. **16 is the right all-round
setting**: within 1% of best solo, best at both 2 and 4 streams.

That exposed a flaw in my own selection logic, which read only the par-1 rows and would
therefore have written 8 into best_threads.txt — the worst value for concurrent use — for
the restore stage to apply. It now scores each thread count by its MEAN RELATIVE aggregate
across every parallel level measured (relative, so the level with the biggest absolute
aggregate does not dominate the mean), and only considers thread counts measured at every
level, so a value that happens to have been tried only where it looks good cannot win by
omission.

Worth recording as a general trap: **a knob tuned at one concurrency level can be actively
wrong at another**, and single-stream benchmarking is the default that hides it. The same
caution applies to the thread count this repo pins elsewhere.

Also from this stage, two things that change how its numbers should be read:

* Thread count STOPS MATTERING under concurrency — the t16-vs-t32 gap is +6.6% at 1 stream,
  +8.0% at 2, and +1.4% at 4. Tune --threads for the solo case; if the box runs concurrent
  it barely matters.
* Concurrency scaling beats a fixed-overhead model. Fitting per-token time = F + V to the
  t32 pair gives F=53.5 ms / V=22.1 ms and predicts 28.2 t/s at 4 streams; measured is
  30.46, i.e. 8% better. V is not constant — llama.cpp batches concurrent decode, so the
  per-stream work itself shrinks with batch size (V falls to ~19.4 ms at 4 streams). The
  ~72% fixed fraction is still the headline, but it is a floor on the amortisable share,
  not an exact split.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The --parallel matrix is complete, at the winning placement (-ncmoe 16, split 30,18, q8_0
KV, projector on CPU). Three results.

**Scaling.** 1 -> 2 streams is 1.55x aggregate, 1 -> 4 is 2.30x, at 58% of the solo
per-stream rate (7.73 vs 14.09). Good for batch work, poor for one interactive user.

✅ **Doubling the total context budget at 4 streams is FREE.** --ctx-size 131072 gives each
of four slots 32k instead of 16k for 30.52 vs 30.46 t/s aggregate — identical within noise.
So concurrency and context are not in tension here: the KV cache is 12 KiB/token at q8_0
(24 at f16), so the extra 0.75 GiB fits the winning placement's headroom. The corollary is a
配置 to avoid: --parallel 4 at --ctx-size 65536 buys nothing and leaves each caller 16k.

🔴 **--threads inverts with load.** 8 threads wins solo (14.25 vs 14.09 at 16, 13.22 at 32)
because it already saturates the four populated memory channels — exactly what STREAM
predicted — and then loses 4.9% at four streams, where the parallel work exceeds what 8
threads can cover. 16 is the right default: within 1% of best solo, best at 2 and 4 streams.
The gap also shrinks with load (+6.6% / +8.0% / +1.4% at 1 / 2 / 4 streams), so past a
certain batch size the setting stops mattering.

The general trap is worth stating plainly, because it caught this pipeline's own selection
logic: **a knob tuned at one concurrency level can be actively wrong at another, and
single-stream benchmarking is the default that hides it.**

And concurrency gives a third, independent estimate of the fixed-overhead floor. Fitting
per-token time = F + V to the --threads 32 pair gives F = 53.5 ms, V = 22.1 ms, i.e. 71%
fixed — corroborating the ~43 ms floor found earlier by two other routes. ⚠️ But that fit
predicts 28.2 t/s at four streams against a measured 30.46, so V is not constant: llama.cpp
batches concurrent decode and the per-stream work itself shrinks with batch size (V ~19.4 ms
at four streams). 71% is a floor on the amortisable share, not an exact split.

⚠️ Noted in the table: these come from concurrency-probe.py, whose prompt set differs from
placement-probe.py, so they are comparable within that table only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Part 2's restore failed rc=1 with `BTH: unbound variable`. When I added the measured
thread-count lookup, the edit used replace(..., 1) and the anchor
`setv MODEL_THREADS 32; setv MODEL_PARALLEL 1; ...` occurs more than once, so the `local BTH`
declaration landed in a different function while s6_restore kept the `${BTH}` reference in
its final note. Declared properly now, with 16 as the fallback.

⚠️ **I initially misdiagnosed this as another instance of the `[ test ] && assign` pattern
and went looking for all eight of them. That was wrong**: a failing `[` inside an `&&` list
is explicitly exempt from set -e (it is not the final command of the list), which is exactly
why splitfix phase 1 ran straight through eight of them tonight. The bug was **set -u**, a
different mechanism, and the audit was a red herring. Recording it because "same shape as
the last bug" is a tempting and, here, misleading heuristic.

`bash -n` and shellcheck both pass on an unbound-variable bug, so neither would ever have
caught this. What does catch it is running the function with everything stubbed under the
real flags — the technique that found the splitfix bug — so that is now a committed tool
rather than a thing I improvise each time:

  stage-harness.sh   extracts each stage function from the script and runs it under
                     set -Eeuo pipefail with setv/pct/curl/ip/wait_up/cells stubbed

Run on the remaining stages it reports s2c_context, s3_mtp and s2b_split all clean, so the
rest of the night has no further landmines of this class. ⚠️ It must run on the HOST: BSD
sed on macOS cannot handle the `{` in the extraction address, which made it report false
failures for all three until moved.

What part 2's restore did manage before dying is worth noting, since the failure was on its
LAST line: every setv ran, the model was restarted, and `pct start 121` ran — so only the
summary note was lost. restore-standalone.sh re-runs it at the end of the finish chain from
the fixed copy anyway.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…op half its own tables

Dry-ran morning-report.sh against the night's real JSON before it fires, which was worth
doing — it had four defects, two of them the kind that produce a confident wrong answer.

🔴 **It would have recommended a SPILLED configuration.** "The answer" took the top row of
the placement table, which ranks by throughput with a gate that only checks degeneracy and
d0 reproducibility — nothing about VRAM headroom. The top row was `ncmoe15-t32` at 13.68
t/s, measured at the documented split while holding 2060 MiB in GTT. That is precisely the
trap this whole run exists to expose, reproduced by my own reporting tool. The headline now
comes from best_split.txt, which the split stage writes only after refusing every candidate
whose verdict was SPILLED, and the placement table carries a warning that it shows the shape
of the curve rather than a ranking of usable configs.

🔴 **The split table showed 1 of 4 rows, then 2 of 4.** `(\w+)` cannot match `q8-projcpu`
(hyphen), and after fixing that a mandatory third group still dropped the f16 cells, whose
labels have no shape suffix at all (`split-nc16-c30.json`). Now optional, defaulting to
"f16 / proj GPU".

**The context table would have listed every configuration twice.** Each context cell writes
two files — a 3-class shallow probe and a `-deep` 32k one, split because a 32k prefill costs
~13 min per request — and a bare `ctx*.json` glob counted the deep file as another cell, so
each config appeared once with no d0 and once with no d32k. They are joined now, and the
32k prefill rate is reported alongside, since that is what decides whether a long window is
usable.

**And a NameError**: the headline lookup referenced the `split` dict from above where it is
built. Moved below it.

Lesson worth keeping: a reporting tool needs the same gates as the measurements it reports.
Mine ranked throughput and ignored headroom, which is the exact error the measurements were
designed to catch — and it would have shipped the wrong recommendation with every number in
it technically correct.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
marchah and others added 17 commits September 18, 2026 03:45
My contextsweep trimmer did `kill -TERM -<pgid>` using the pgid it read from the contextsweep
process, and that pgid was **the chain's own process group** — so it killed chain-finish.sh
too, and MTP, restore and the report never ran.

The instructive part is that the revalidate trimmer did exactly the same thing earlier and
was safe. The difference:

  * `stage` runs its function as `set -m; ( "$@" ) &`, which puts it in its OWN process
    group, so a group-kill hits only that stage.
  * `chain-finish.sh` invoked `./contextsweep-standalone.sh` as a plain command, which
    INHERITS the parent's group. Same kill, entirely different blast radius.

⚠️ Rule going in the script: before `kill -TERM -<pgid>`, check the pgid is not the
supervisor's own. `ps -o pgid= -p $$` in the supervisor is the value that must not match.
chain-finish2.sh logs its own pgid at startup so the next trimmer has something to compare
against.

Nothing measured was lost — contextsweep's cell 1 completed and wrote its shallow result
before the kill, and the MTP build was already on disk. chain-finish2.sh re-runs the three
remaining steps (MTP with MTP_SKIP_BUILD=true, then restore, then the report), and MTP is
now running at the winning placement.

Also worth recording from cell 1 before it was trimmed: at `ctx 131072` with f16 KV on
`-ncmoe 16` the config spills (226 MiB free, 183 MiB GTT) and **prefill collapses** — six
requests took 50 minutes, implying an 8k prefill fell from ~63 t/s to roughly 10. That is the
~12x GTT collapse this repo documents, observed in its classic form for the first time
tonight. It is also why f16 is the wrong KV type for a long window here: q8_0 halves the
cache and the --parallel stage already showed 131072 fitting at that shape.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…is untouched

Commit b14d846 claimed that at ctx 131072 with f16 KV the spill made prefill "collapse from
~63 t/s to roughly 10". That is wrong. I inferred it from the cell taking 50 minutes of wall
clock, without reading prefill_tps_median, which was in the probe's JSON the whole time.

Measured, same placement, only the context budget changed:

  ctx       d0 decode   d8k decode   d8k prefill   headroom
  65536         14.16        13.07          77.0   1337/1553 MiB free, no spill
  131072        12.92        10.34          76.7   226/573 MiB free, 183 MiB GTT
                -8.8%       -20.9%         -0.4%

Prefill is unaffected. The spill costs decode, and roughly twice as much at depth as at a
short prompt — which is the mechanically sensible shape: prefill streams weights in large
batches, while decode at depth touches the KV cache every token, so a cache partly in GTT
means a host round-trip per token instead of per batch.

Two things to carry forward:

* **Do not infer a spill's cost from wall-clock time.** A slow cell invites a story, and
  `prefill_tps_median` is in every probe's output to settle it instead.
* This repo's ~12x GTT figure came from a case where the WHOLE KV cache sat in GTT under a
  much larger over-commit. A marginal 183 MiB spill is a different regime and does not
  behave the same way, so that number should not be applied to any spill indiscriminately.

The practical recommendation does not change — use q8_0 for a long window, because it halves
the cache and 131072 then fits rather than spilling, which the --parallel stage independently
confirmed — but it now rests on the right mechanism.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… caused it

Both drafters abort the instant the MTP graph is constructed:

    ggml.c:2264: GGML_ASSERT(ggml_can_repeat(b, a)) failed
      llama_model_qwen4exp::graph::build_hc_mix(...)
      llama_model_qwen4exp::graph_mtp::graph_mtp(...)

The build itself is healthy: the no-speculation control on that exact binary gives 13.88 t/s
against 14.17 for b11018-baseline at the same placement, so it loads the target and serves
normally. Only the MTP path is broken.

**And it is broken because of the conflict resolution I chose — though resolving the other
way would not have helped.** b11018 reshaped every hyper-connection gamma from {hc_dim} to
{n_embd, hc} with TENSOR_ALLOW_RESHAPE and updated its TRUNK graph to match. I kept b11018's
shapes, which is right: the trunk works, as the control proves. But PR #28097's MTP graph
predates that reshape and is written against {hc_dim}, so it broadcasts mismatched shapes.
Taking the PR's shapes instead would just move the abort into the trunk, which is the path
that has to work. There is no resolution of those two commits that satisfies both graphs.

The real fix is to port the PR's MTP graph to b11018's hc convention — a code change to the
MTP caller of build_hc_mix, not a merge decision. That belongs upstream or in a patch written
with the hc layout in hand. Guessing at it risks producing wrong OUTPUT rather than a clean
crash, which is worse than no measurement, so it is explicitly not attempted.

⚠️ Worth recording about my own pre-flight: I checked that both drafters carry the metadata
the loader's mtp_only probe needs and concluded the layout-mismatch failure mode was "ruled
out". That validated the LOADER path and said nothing about the GRAPH path, which is where it
actually fails. Metadata satisfying a loader does not establish that the graph built from it
is shape-correct — the pre-flight was necessary, not sufficient, and I overstated it.

Everything is staged for a one-command retry when #28097 rebases: drafters verified, build at
/root/builds/mtp-b11018 (58 MB, both guards asserted), patches in mtp-patches/, runner
./mtp-standalone.sh with MTP_SKIP_BUILD=true.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ured

Adds the recommendation section the night was for, at the top of the measured results, plus
two corrections.

DEFAULT, and it is both the fastest and the roomiest:

  --n-cpu-moe 16 --tensor-split 30,18 --threads 16 \
    --cache-type-k q8_0 --cache-type-v q8_0 --no-mmproj-offload \
    --ctx-size 131072 --parallel 1
  -> 14.17 t/s short prompt, 13.45 at 8k, 2.3/1.3 GiB free, no spill

-ncmoe 16 is the FLOOR: nothing below it fits in any shape at any split. q8_0 plus
projector-on-CPU is free (identical at d0, +2.9% at depth), frees 1.6 GiB, and is what lets
131072 of context fit.

The give-a-card-back option is worth knowing: **-ncmoe 48 runs the whole model in 9.4 GiB on
ONE card**, leaving the other entirely free, for -28% of decode. CPU-only is strictly worse
than that at 6.04 t/s (-57%) and frees only 2 GiB more.

Concurrency: 2.30x aggregate at four streams (30.9 t/s), and doubling the context budget is
free -- so --parallel 4 at --ctx-size 65536 is never right.

And the ceiling is now explained rather than described. Decode is ~72% fixed per-token
overhead by three independent routes (the -ncmoe curve fit and a clock experiment both give
~43 ms; the concurrency fit gives 53.5 of 71 ms), and GGML_SCHED_DEBUG names it: ~30 graph
splits per token, which matches the placement almost exactly (16 CPU-expert layers x 2
crossings, plus the card boundary and the PLE lookup) at ~1.1 ms per crossing = 61% of the
fixed term. The lever is fewer CPU<->GPU crossings, not bandwidth or clocks -- both of which,
along with PCIe, core count and disk, were ruled out separately.

Two corrections in this commit:

* The MTP control was quoted as 13.88 t/s. That is the median across BOTH depths, not d0.
  Per depth it is 14.42 / 13.43 against b11018-baseline's 14.17 / 13.45 — so the rebased
  build is marginally FASTER than baseline, not slower, which strengthens rather than
  weakens the claim that only its MTP graph path is broken.
* The spill-costs-decode-not-prefill retraction is now confirmed a second time,
  independently: the spilled nc15-c31 cell has prefill 79.9 t/s against 77.0 for the fitting
  nc16-c30 — unaffected, marginally higher — while its decode is -11.4% at d0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…free

The recommendation was assembled from separate cells: the split stage measured that
placement at --ctx-size 65536, while restore left the box at 131072 (inherited from the
context sweep). Those are not the same configuration, so the delivered one was re-probed
directly rather than assumed.

                       live, ctx 131072    splitfix, ctx 65536
  decode d0                     14.46                   14.17
  decode 8k                     13.37                   13.45
  prefill 8k                     77.2                    76.9
  VRAM free               2297/1255 MiB           2881/1878 MiB

**--ctx-size 131072 is free**: identical throughput within noise, d0 marginally higher, and
card 2 still at 1255 MiB — above the ~1024 MiB floor where RADV spills. So the 128k window
costs nothing and is kept.

The contract assertions also pass on the live config, which is the part that would silently
break callers rather than show up as a number: bare chat answers with clean content, no
<think> pair leaks into `content`, and no reasoning_effort value a caller might send breaks
the server. That last one matters because this model's template raises on "none" and "high"
— the exact values CT 123's coder uses — so --reasoning off is doing its job of making the
trap unreachable from outside.

`all_reps_agree=False` here is the documented depth non-reproducibility (d8000 disagrees
with itself at temperature 0 on this hybrid placement), not degeneracy: `degenerate=False`.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…up on restart

I asserted all night, in the restore stage's own output and in every status update, that
CT 121's 22:00 backup was "missed tonight, trigger by hand". That was an assumption I never
checked, and it is wrong.

Verified: CT 121 came up at 06:50:18Z and `Backup hermes 2026-09-18` (fbdbbe08) committed at
**06:50:54Z** — four seconds later. Hermes' cron catches up a missed job when the container
restarts, so a planned overnight outage needs no intervention, and silence in #backups during
one is not automatically a gap.

🔴 **But the same catch-up behaviour caused real collateral damage, and the timing points at
this benchmark.** Two jobs fired at that instant. The backup completed in 61 s. The KB
Freshness Refresh ran **53 minutes**, hit its 2700 s budget, staged **0 of 448** entries and
quarantined five as runaways:

  ai-tools/openworker, penguin-harness, thinking-orbs, tokentab, turbo-fieldfare

That window (06:50-07:43Z) is exactly when the context sweep was restarting CT 120's server
every few minutes and driving it flat out. The refresh consumes CT 120's API, so those
timeouts are an artifact of a saturated box, not of bad entries — the five quarantines should
be reversed and re-run.

The restore stage's note is corrected to say what is actually true, and to warn about the
part that matters: every cron catches up at the same moment the container returns, so
anything still benchmarking will be contended with.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
47 files / 6707 lines -> 31 / 5737. Everything removed was single-use scaffolding whose
question is closed, and every FINDING it produced is already in the README.

Recovery chains and trimmers, which existed only because the pipeline was mid-flight:
chain-finish.sh, chain-finish2.sh, chain-report.sh, chain-pr28699.sh, chain-one-gpu.sh,
trim-revalidate.sh, trim-contextsweep.sh, contextsweep-standalone.sh,
restore-standalone.sh. Two of these carried the group-kill hazard that took out their own
supervisor, so keeping them around as templates would be actively harmful.

Investigations that answered their question: cpuset-test.sh and cpuset-variance.sh (the
"cpuset lottery" turned out to be the governor), warmup-test.sh (warm-up is built into the
probes now), power-test.sh (RAPL numbers recorded; the method is in the host-power-saving
memory), structural-tests.sh (whole-layer -ot rejected at -42%), confirm-sweep.sh and
tune-sweep.sh (superseded by revalidate.sh + placement-sweep.sh).

Two fixes so the survivors stand alone rather than depending on deleted or host-only files:

* mtp-standalone.sh pointed at overnight-part2-fixed.sh, which only ever existed on the
  host as a way to patch a script that was executing. The in-repo overnight-part2.sh is now
  byte-identical (verified by sha256), so it points there and works from a fresh checkout.
  Its header also drops the recovery narrative and states the current blocker instead.
* overnight-all.sh still pgrep'd for the deleted chain-reval.sh. Removed.

Kept deliberately, because they generalise past this model: gguf-kv.py (sizes any GGUF's KV
cache from metadata — the figure that gets overstated 4x on a hybrid), stage-harness.sh
(catches the set -u class of bug that bash -n and shellcheck both pass), stagelib.sh,
attribute-telemetry.py + placement-sampler.sh, build-llamacpp.sh, and the overnight pipeline
itself, which is worth re-running once the four extra DIMMs land.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…rect the one-card advice

The placement-dial section gave approximate VRAM ("~59 GB", "~8 GB") and never said what
-ncmoe 48 actually leaves resident. Replaced with measurement and structure.

**What 48 means.** --n-cpu-moe moves ONLY the routed experts, so everything else is resident
at every setting: attention for all 48 layers (12 Qwen Sparse Attention + 36 Gated
DeltaNet), the per-layer shared expert and router, norms, hyper-connections, output, and the
KV cache. -ncmoe 48 is therefore "attention plus the dense path on the GPU, 9.4 GiB, all
routed experts in RAM" — the floor of GPU residency, not an off switch.

**Why the experts and not the layers.** Routed experts are ~61% of the model but only 2.0%
is read per token (topk 10 of 512); attention and the dense path are read in full every
token. Measured rather than argued: moving whole layers to CPU cost -42% against -28% for
moving only the experts at the same VRAM saving.

**VRAM per setting, measured** — read after a load plus one completion so lazily-allocated
buffers are counted, at the corrected split:

  ncmoe 15  62.3 G  SPILLED at every split (64.2 G demand vs 64 G capacity)
  ncmoe 16  59.3 G  q8_0/projCPU, tight     14.17 / 13.45   <- the floor for two cards
  ncmoe 16  61.1 G  f16/projGPU,  tight     14.16 / 13.07
  ncmoe 20  55.3 G  fits                    13.03 / 12.06
  ncmoe 28  43.6 G  fits                    ~11.3 / ~10.6
  one card: ncmoe 34  30.4 G   13.01 / 12.23   <- the floor for one card
            ncmoe 36  27.4 G   12.43 / 11.81
            ncmoe 48   9.2 G   10.18 / 8.70

🔴 **And a correction to this file's own advice.** It recommended -ncmoe 48 for handing a
card back. That is 28% slower than necessary: -ncmoe 34 on ONE card gives 13.01 t/s, only
-10% off the two-card best, with card 2 completely idle. Pick 48 only when you also want
~23 GiB spare on the remaining card.

Also adds the matched-pair result this rests on: at -ncmoe 34, one card is +20.9% at d0 and
+24.2% at depth against two, with prefill IDENTICAL (+0.5%). The second card is a decode
penalty (~21%, llama.cpp #28699's inter-GPU indexer traffic) and a prefill win (+97%,
because batching amortises the crossing). Net +11% decode, +97% prefill — so it earns its
place on time-to-first-token, not on tok/s.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…per-layer VRAM is STABLE

The one-GPU sweep finished. Full single-card curve (--device Vulkan0, q8_0 + projector on
CPU, threads 16, ctx 65536):

  -ncmoe   VRAM used   free    decode d0 / d8k   prefill 8k
  34       30.4 G      1.7 G   13.01 / 12.23     39.1
  36       27.4 G      4.6 G   12.43 / 11.81     37.1
  40       21.6 G     10.4 G   11.40 / 10.94     33.7
  48        9.2 G     23.3 G   10.18 /  8.70     29.0

✅ Expert weight is ~1.5 GiB/layer and that figure is STABLE: 1502 MiB/layer over 34->36,
1502 over 36->40, 1581 over 40->48, averaging 1547. That stability is why every VRAM
prediction made from the average during the sweep landed within ~120 MiB of measurement
across four different shapes, and it is what makes the placement dial predictable.

⚠️ Recorded because I nearly wrote the opposite into this file. I computed the deltas as
1502 / 2018 / 1323 MiB/layer — a ±35% spread — and had begun attributing it to UD-Q4_K_XL
being Unsloth's *dynamic* quant assigning bit-widths by layer sensitivity. That is a
plausible-sounding mechanism for a number that does not exist: the spread came from mixing
the sweep's own reading for -ncmoe 40 (free 10679 -> used 22089) with a live reading taken
later in the cell (20024). The ~2 GiB gap is transient compute buffer. **Read VRAM from one
consistent moment**, or you can invent a non-uniformity and a story to explain it.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…wen3.6

The placement study is finished, and its conclusion changes where this model should live.
**One card beats two by 21-24% on decode** for this architecture (llama.cpp #28699 ships the
QSA indexer's pooled rows across the inter-GPU link every token), and Qwen3.8-Flash-Next is
prefill-bound at ~205 s to first token on an 8k prompt — far too slow for Hermes. So:

* **CT 120 -> qwen3.6-35b-a3b on GPU 1** (`ct120-cutover.sh to-qwen36`). Verified 74.8 tok/s
  and a clean chat endpoint. This is Hermes' model and it is 5x faster for agent work.
* **CT 123 -> Qwen3.8-Flash-Next on GPU 2 alone**, at the measured best single-card config:
  `-ncmoe 34 --tensor-split "" --device Vulkan0 --threads 16 --cache-type-k/v q8_0
  --no-mmproj-offload`, ctx 65536. Verified serving at 28964 MiB with 1738 free, no spill.
* **llama-swap removed** from CT 123 (service stopped and disabled; its six models left on
  disk untouched at the owner's request — /models grown 180G -> 240G rather than deleting
  anything).
* **CT 122 `coder-runner` destroyed.** The loop moved to Multica, so it was unused and only
  something to keep patched. Its ssh keypair was removed from CT 121 with it, and the
  provisioning script stays as the recipe.

New: `qwen38fn-gpu2.env`, the single-card config as a committed file rather than a hand-edit,
and `install.sh` grows an `ENV_FILE` override so both deployments are first-class:

    VMID=120                            ./install.sh   # two cards, -ncmoe 16
    VMID=123 ENV_FILE=qwen38fn-gpu2.env ./install.sh   # one card,  -ncmoe 34

Two things that had to be fixed to make CT 123 work, both worth knowing:

* **CT 123 was on llama.cpp b10678, which cannot run `qwen4exp` on Vulkan at all** (b11013 is
  the floor). Brought to b11018 by copying `/opt/llamacpp/b11018-baseline` from CT 120 —
  same Ubuntu 24.04 / glibc 2.39, so the binary moves. Copied the *proven* build rather than
  the prebuilt release, which had never served this model.
* **The thermal watchdog's `GPU_SERVICE_MAP` named `123:llama-swap`**, a service that no
  longer exists. A map naming a dead service turns a thermal trip into a no-op that leaves
  the real load cooking an overheating card. Now `123:llamacpp-qwen38fn`, in both the live
  env and `ct120-cutover.sh`.

CLAUDE.md updated throughout: the architecture line, the VMID allocation, the coder-runner
convention (now marked decommissioned, with the isolation rule retained), the common
commands, the build pins, the watchdog map and the token-accounting note.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…DE.md

Found while briefing a reviewer on this PR. The overnight sweep established
that the obvious split rule overcommits card 1 by ~2 layers and spills
anyway, and the study README records the correction -- but it propagated
nowhere else.

- placement-sweep.sh and revalidate.sh both still derived the UNCORRECTED
  `N + (48-N)/2`. This is functional, not cosmetic: it is the derivation
  that silently spilled -ncmoe 15/16/20 into GTT. Corrected to
  `N + (48-N)/2 - 2`, which reproduces the measured splits exactly
  (16 -> 30,18 -- the deployed config -- 20 -> 32,16, 28 -> 36,12).
  `-ncmoe 48` is exempt and clamped to `48,0`: with no heavy layers to
  rebalance, the -2 would hand card 2 two light layers plus an inter-GPU
  hop for nothing, and would break the "*,0" single-GPU detection.
  Keeping the two derivations identical matters -- otherwise a revalidation
  measures a different placement than the sweep it is checking.

- CLAUDE.md asserted the uncorrected rule AND claimed placement-sweep.sh
  derived it, so fixing the doc alone would have made it wrong a new way.

- CLAUDE.md still headlined "~4x below the sizing note, 11.74 t/s", which
  the overnight curve superseded: the best two-card placement is 14.46 t/s
  (13.37 at 8k) against a predicted 72, i.e. ~5x. Replaced with the real
  numbers, the graph-split mechanism, the calibrated latency model, and the
  two findings that actually drove the deployment -- prefill is the binding
  constraint (104 s vs 205 s TTFT at 8k) and one card beats two by 21-24%
  at matched placement.

- revalidate.sh and overnight.sh carried the powersave penalty as -52%.
  That was one noisy pre-warm-up arm; the settled figure is ~30%.

- stage-harness.sh was executable with `set -Eeuo pipefail` on line 1 and
  no shebang (shellcheck SC2148, an error): it would run under the caller's
  shell, where -E is not POSIX.

Verified: shellcheck clean at -S warning on every shell file in the PR,
all Python compiles, and no function in the PR ends with a `[ ] && ...`
construct -- the `set -e` trap that shape causes applies to a function's
last command, not to a loop body or if-branch, which is tested rather than
assumed here.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Systematic grep for the pattern found in the last commit: a finding gets
corrected in the README and nowhere else. Two of these are functional and
one is actively dangerous.

🔴 qwen38fn.env shipped a SPILLING default. It is the install.sh default
   (CT 120, two cards) and carried MODEL_CPU_MOE=20 with the uncorrected
   MODEL_TENSOR_SPLIT=34,14 -- the exact combination the README measured at
   101 MiB in GTT with 823 MiB free, under the ~1024 MiB RADV floor, for
   12.69 t/s. Now -ncmoe 16 + 30,18, the measured best two-card cell
   (14.46 / 13.37 at 8k), with the rule comment corrected.

🔴 qwen38fn-gpu2.env told a future operator that b10678 is "a working
   rollback target" because it "CAN load the arch". It cannot execute this
   model at all: the arch string is registered there, but the Vulkan
   hyper-connection ops only landed in #28988 (b11013 is the floor). A
   rollback following that comment would produce a server that cannot load.

⚠️ The retracted "q8_0 costs ~14% at depth" figure survived in the
   VRAM-floor candidate table and inverted its verdict. The clean pair has
   q8_0 + projector-on-CPU BETTER on both axes -- +2.9% at depth and
   1.6 GiB cheaper -- so the table recommended the wrong row, and both env
   files followed it. The retraction also re-opens -ncmoe 15, whose only
   recorded reason for rejection was that 14%; flagged as untested rather
   than excluded, and the contradiction with "below 16 does not fit in any
   shape" is noted.

⚠️ The threads correction landed in one section and three places still
   carried the superseded single-stream reading: the README's "inverted U
   with the peak at 16" heading (8 and 16 are within 0.5% at n=1), the
   summary row claiming 16 "is best across every concurrency level" (8 wins
   solo by ~1%, then loses 4.9% at four streams), and overnight-part2.sh's
   own selection comment. All now point at the concurrency section that
   settles it.

Also: CLAUDE.md recorded 34,14 at -ncmoe 20 as "no spill" (it held 101 MiB
in GTT -- it read clean only against a 9.3 GiB spill); CLAUDE.md described
llama-swap in the present tense as CT 123's service while saying elsewhere
it was removed, now marked decommissioned-but-recoverable at the section
head; the last "4x" in the README is ~5x; one broken anchor I introduced
(an emoji heading leaves a leading space, so the slug takes three hyphens).

Verified: shellcheck clean at -S warning, bash -n clean, README anchors all
resolve, EnvironmentFile syntax clean (no inline comments after a value),
and the serve script does honour MODEL_KV_TYPE / MODEL_MMPROJ_ON_CPU so the
new keys are live rather than decorative.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ginal

All 14 findings reproduced independently before fixing. The three P1s are
real safety defects in the cutover and watchdog paths; the live deployment
was never affected.

P1 — the forward cutover disabled thermal protection on BOTH cards. It
mapped both GPUs to `120:llamacpp`, then disabled that unit and started
`120:llamacpp-qwen38fn` instead, so a trip stopped a dead unit while the
real load kept running — on the one configuration that drives both cards.
Now maps the unit that actually runs, and `assert_map_owns_cards` checks the
invariant (every mapped service must own its card) after both directions.

P1 — GPU exclusivity did not survive a reboot. `assert_ct123_stopped` was a
one-shot check, and CT 123 kept onboot=1 plus its GPU-2 mounts, so the next
host boot put two llama-servers on one card. The cutover now clears CT 123's
onboot and restores it on the way back; `status` surfaces the hazard.

P1 — the committed watchdog default named `123:llama-swap`, removed on
2026-09-18, so a fresh install left GPU 2 mapped to a dead unit. The
in-script fallback was worse: B550-era PCI addresses that do not exist here.

🔴 The two-card default was MARGINAL, and this is the reviewer's best catch.
Every throughput number came from a sweep pinned at batch/ubatch 1024/256
while the launcher defaults to 4096/1024 and neither env file said so. Loaded
end-to-end to settle it:

  1024/256   min free 2881 / 1878 MiB   no spill        prefill 77.2 t/s
  4096/1024  min free 2922 /  877 MiB   GTT 306 MiB     prefill 173-236 t/s

Card 2 fell under the ~1024 MiB RADV spill floor and card 1's GTT hit 306
against a ~70 idle floor. Both env files now pin explicit, measured values
(two-card 1024/256; one-card 4096/1024, which is verified safe at -ncmoe 34).
✅ New finding: the right --tensor-split is BATCH-SIZE DEPENDENT — the compute
buffer scales with batch and lands on the card holding the heavy layers,
which is also why the `- 2` correction exists. At 4096/1024 this placement
wants ~31,17, untested.

Also fixed: the guard treated an unreadable sensor as 0 °C (failing open,
against its own stated contract) — a missing sensor is now over-temp; the
stage harness printed FAIL and exited 0, and could not tell a broken stage
from a failed extraction; the split selector inferred KV/projector shape from
filenames and silently dropped unmatched cells, so it now reads a recorded
.shape sidecar and logs every skip; the KV A/B ran its two arms at different
placements, recreating the confound that produced the retracted 14% claim;
`-ncmoe 15` was described as never load-checked when saved results show it
SPILLED at 29,19 (3256/89 MiB); three more copies of the split rule (one
uncorrected, one using `- 1`); the -54% governor figure; the contract claim
that inverted what the probe asserts; `-ncmoe 48 matches 34`; 64295 MiB
mislabelled 64.2 GiB (it is 62.79, and the constraint is per-card margin);
the patches README not saying MTP aborts in the graph; and the false "all
clean" lint claim, whose verification loop iterated an empty file list.

Consolidation, per the review's root-cause note: CLAUDE.md's section 108 ->
51 lines with the README as the single detailed record, and the JEDEC 3DS
analysis condensed to its conclusions.

Gate: `git ls-files '*.sh' | xargs shellcheck -S warning` clean on 0.11.0
(the reviewer's exact command; SC2034 false positives from eval'd stage
bodies now carry file-scoped directives with reasons), bash -n clean,
py_compile clean. Fixes verified non-vacuously: the guard yields 999 for a
missing sensor where it yielded 0, and the harness exits 1 on an injected
`return 42` where it exited 0.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… was wrong

The follow-up is right on all five findings, including that my completion
claim was not supported by 8c2a79e. Two of these are defects I created while
fixing the originals.

🔴 P1 — the thermal guard still failed open, and my test is why I missed it.
Validating the CONCATENATION `${jr}${mr}` means one empty reading beside one
numeric reading concatenates to a numeric string, passes validation, and the
empty one becomes 0 C in the arithmetic. Only BOTH sensors missing tripped —
the single case my test covered. Each reading is now validated separately,
and losing the hwmon directory mid-run routes the same way (the startup check
only proves it existed at startup). Verified by breaking each reading
independently, one at a time: junction-missing, mem-missing, both-missing and
hwmon-gone all trip; both-ok does not.

P2 — `assert_map_owns_cards` checked none of what its name claimed. It never
read any container's GPU bindings, accepted is-enabled OR is-active so an
enabled-but-dead unit passed, never checked both cards had entries, and
returned 0 on every path while logging "every mapped unit exists and is
live". That is the same defect class this review set out to trace, committed
in the fix for it. It now verifies the target container actually binds that
card's render node, reports active / configured-but-not-serving / mismatch as
distinct states (the middle one is the legitimate state of a deliberately
stopped CT 123 after rollback), checks both cards are mapped, and returns
non-zero. Reproduced the reviewer's scenario: CT 120 binding only GPU 1 with
a map claiming both now fails with rc=1. Deferred via MAP_BAD so the status
dump still prints, then exits non-zero.

P2 — the harness still reported success after an INTERNAL command failure.
`if ( "$fn" ); then` disables errexit for the whole if-condition and that
suppression reaches inside the function, so a failing command mid-stage did
not abort it; the stage ran on and returned 0. Counting failures could not
help because the failure was never surfaced. Now runs each stage as a
standalone subshell with errexit re-armed, reading $? on the next line.
Verified with an internal failure whose later commands still succeed:
FAIL rc=42, harness exit 1. Also added the missing stage() negative control.

P2 — the surviving claims. Six locations in the model README still said
"64.2 GiB", "never load-checked", "nobody re-tested" and "does not fit in any
shape at any split" — I had fixed the env-file copies and not these. And
pro-v620/README.md was never opened: it still described CT 123 as the
llama-swap coding server, mapped the watchdog to 123:llama-swap, and carried
the entire B550 topology (0000:06:00.0, "PCIe-3 (chipset) slot") in its
intro and in copy-pasteable command examples.

P3 — "a copy of the release tarball" was asserted, not checked. Same source
revision (build 11018, commit c9a5eeeb3) but different builds: GNU 11.4.0 vs
13.3.0, sha256 4ac7aa75 vs 32687325. Worth more than the wording: every sweep
number was measured on b11018-baseline while the two-card env ships
llama-b11018, so a throughput comparison against the sweep crosses two
binaries. The 2026-09-18 load check did use the shipped binary, so the
shipped config itself is validated.

-ncmoe 15 is closed the way the follow-up argued — by headroom POLICY, not by
physics. Adopted its wording plus an explicit >=2048 MiB/card acceptance rule
(the README already marked anything below that "tight"), and dropped my
"arithmetic rules out the rest", which overstated the case.

Gate: shellcheck -S warning clean on 0.11.0 by both the all-files and
changed-files commands, bash -n clean, py_compile clean, README anchors
resolve. One pre-existing broken anchor in pro-v620/README.md
(context-window--concurrency) also breaks on origin/main and is left alone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…gs, and

a retraction that was itself wrong

All six findings hold. Three are defects introduced by the previous round's
fixes, and one is a correction that corrected something true.

P2 — the new ownership assertion BROKE THE ORDINARY ROLLBACK. `pct exec`
cannot run in a stopped container, and `to-qwen36` leaves CT 123 stopped by
design, so the "CONFIGURED but not serving" state was unreachable for exactly
the container that needs it: both queries failed, MAP_BAD was set, and a
completely correct rollback exited 1 with a false protection alarm. Container
-stopped and unit-inactive are now distinct states, the stopped target
reports its bindings as correct with runtime state DEFERRED, and a genuine
mismatch still fails.

P2 — the >=2048 MiB/card policy rejected the configurations this file ships:
the two-card default's minimum is 1878 MiB, the one-card default's 1739, the
131072 row's 1255. A universal rule invented inside the paragraph that needed
it is not a policy. Reframed: >=2048 MiB is required for a NEW or UNVALIDATED
placement; an existing placement below it is acceptable only with a measured
per-card minimum recorded at its shipped batch size, and the three exceptions
are named. No live setting changed to suit the prose.

P2 — "the shipped config is validated" combined two configurations. The
release-binary check ran at 4096/1024; the file ships 1024/256, whose numbers
come from the baseline build's sweep. The exact pair was never tested. Stated
as the verification gap it is, with the one thing that can be said for it:
smaller buffers cannot use more memory than the run that already fit.

P2 — the stage() negative control passed without running its body, four ways:
wrong arity (`stage <name> <timeout> <fn>`, so $2 was unbound), stderr hidden
so the invocation error read as success, `stage()` returning 0 by design after
recording a failure so its exit status could never be the assertion, and the
harness's `sleep` stub firing the watchdog instantly. It now asserts the
recorded note and a post-failure sentinel, restores the real sleep, and runs a
succeeding body through the same path to prove it is not vacuous. Verified by
three separate mutations: as committed both controls pass; making the failing
body succeed fires the vacuity check; breaking stage() itself to record
success is also caught.

⚠️ Two bugs in my own verification, found while checking the above: the
control wrapped stage() in `set +e`, which propagated into its subshell and
disabled the very errexit under test; and `local IFS=','` in the assertion
stays in effect for the whole function, rewriting `$*` for everything called
inside it. The IFS scope is now narrowed to the split itself.

P2 — four more current-state claims: gpu2.env still called 15 "the minimum
that fits", and the parent README still said the second card runs llama-swap
and that CT 123 "cannot" expose token counters (it now runs llama-server with
--metrics and does expose them). The model README's "stock map" warning is
relabelled historical, keeping the general hazard.

P3 — the provenance retraction was itself wrong and is withdrawn. It compared
CT 123's b11018-baseline against CT 120's llama-b11018 and concluded "not a
copy" from the differing hashes, but CT 120 holds BOTH trees and its baseline
is byte-identical to CT 123's (32687325 vs the release's 4ac7aa75). The
original copy claim stands. Lesson recorded where it happened: name the tree,
not the container.

Gate: shellcheck -S warning clean on 0.11.0 by both commands, bash -n clean,
py_compile clean, anchors resolve. The pre-existing context-window--concurrency
anchor in pro-v620/README.md still breaks on origin/main and is left alone.
The mocked cutover transition test is still pending and now has one more case
to cover: the deliberately stopped rollback target.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Both items promised in the last round.

Env consolidation. The review's point was that repeating the correction
narrative in env comments had not prevented propagation drift and had made
the files unreadable — 276 lines across the two, 199 of them comment, for 40
settings. Now 204 lines and 128 comment, with every setting carrying its
value, its trap and a pointer, and the history living only in README.md.
All 40 VAR=value lines are byte-identical to the previous commit, verified
mechanically rather than by eye.

Two real defects surfaced while rewriting: comment blocks had drifted off
their settings in both files (MODEL_LOAD_MODE's text sat above
MODEL_TENSOR_SPLIT; the never-use-auto-context warning was orphaned after
MODEL_THREADS), and the q8_0 <-> --reasoning off XOR warning was missing
entirely from qwen38fn-gpu2.env — the LIVE config, the one actually running
q8_0.

Cutover transition test (cutover-transition-test.sh). Drives the real script
against a mocked pct/systemctl and asserts the invariant that matters: at no
point may two containers be able to hold GPU 2, running or after a reboot.
Six scenarios — refusal while CT 123 runs, forward cutover with an ordering
check that the release precedes the attach, a simulated host reboot in the
forward state, failure injected during the release, failure injected after it
plus recovery by rollback, and the reverse cutover with CT 123 deliberately
left stopped (the R3-1 regression).

To make it possible the script's paths are now overridable (VMID, CONF,
BAK_DIR, WATCHDOG_ENV, LXC_CONF_DIR, DRI_BY_PATH), defaulting to the
production values. Nothing else sets them.

🔴 The test is proven non-vacuous. Five mutations, applied one at a time,
each caught with the right message: watchdog map pointing at the disabled
unit; CT 123 autostart never cleared; the assertion blind to a stopped
container; autostart not restored on rollback; GPU 2 never detached on
rollback.

⚠️ And the mutation run earned its keep immediately — two mutations first
read as "not caught" because the test ABORTED instead of reporting.
`rel="$(grep ... | head | cut)"` with no match: grep exits 1, pipefail
propagates, the assignment inherits that status and set -e kills the test.
That is the same bash trap this folder hit once before, this time inside the
test written to catch such things. Fixed with `|| true`, and the mutation
harness now distinguishes caught / not-caught / aborted.

Also noted in passing: `sed -i` in set_watchdog_map is GNU-only, fine on the
Proxmox host but not portable; the test shims it so it runs anywhere.

Gate: shellcheck -S warning clean on 0.11.0 by both commands, bash -n clean,
py_compile clean, transition test PASS, stage harness PASS.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Verifying the round-3 triage at HEAD found finding 2 only half fixed. The
definition was correctly narrowed to new/unvalidated placements, but four
references elsewhere still named "this folder's >=2048 MiB/card policy"
without that qualifier — including the summary table, which is where a reader
lands first and would infer exactly the global rule the finding was about,
one the shipped configurations violate.

Same propagation failure as the rest of this review: the definition moved and
its references did not.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@marchah

marchah commented Sep 18, 2026

Copy link
Copy Markdown
Owner Author

Review triage — all 6 findings from the 0da0a5e review

All six reproduced, all six fixed. Nothing discarded, nothing escalated. Fixes in 1bb2c1f, 45b7f47 and 9343fd3.

Three of these were defects introduced by the previous round's fixes, and one was a correction that corrected something true — noted per finding rather than glossed.


1. P2 — normal rollback falsely failed the ownership assertion — FIXED

Reproduced: pct exec cannot reach a stopped container, and to-qwen36 leaves CT 123 stopped by design, so CONFIGURED but not serving was unreachable for exactly the container that needed it. Both queries failed → MAP_BAD=1 → a correct rollback exited 1 with a false protection alarm. In that flow the fix was worse than the bug it replaced.

Container-stopped and unit-inactive are distinct states now. Verified with a mocked host (CT 120 running, CT 123 stopped, bindings correct):

0000:03:00.0 -> 120:llamacpp — owns the card, unit ACTIVE
0000:83:00.0 -> 123:llamacpp-qwen38fn — binding and map correct; CT 123 is stopped,
    so runtime service state is DEFERRED (not verifiable from outside a stopped CT)
rc=0

A genuine binding mismatch still returns 1, and cutover-transition-test.sh scenario 4 now covers this terminal state.

2. P2 — the ≥2048 MiB/card policy rejected the shipped configs — FIXED (and it was only half-fixed until 9343fd3)

Confirmed: a global rule excludes the two-card default (1878), the one-card default (1739) and the 131072 row (1255). A universal rule invented inside the paragraph that needs it isn't a policy.

Reframed: ≥2048 MiB for a new or unvalidated placement; an existing placement below it is acceptable only with a measured per-card minimum recorded at its shipped batch size, with the three exceptions named.

⚠️ Verifying this at HEAD caught it incomplete — the definition was scoped but four references still said "this folder's ≥2048 MiB/card policy", including the summary table, where a reader lands first. Same propagation failure as the rest of this review. Fixed in 9343fd3; every mention now carries its scope.

3. P2 — "the shipped config is validated" spanned a cross-product — FIXED

Correct. Stated as a table in qwen38fn.env:

exercised?
release binary yes — at 4096/1024 (the run that found 877 MiB)
1024/256 yes — on b11018-baseline (the sweep)
release binary at 1024/256 no — never tested as a pair

Recorded as a verification gap with the one honest mitigation: smaller buffers cannot use more memory than the run that already fit. No rerun performed or requested.

4. P2 — the stage() negative control passed without calling its body — FIXED

All four mechanisms confirmed: wrong arity (stage <name> <timeout> <fn>, so $2 was unbound), stderr hidden so the invocation error read as success, stage() returning 0 by design so its exit status could never be the assertion, and the sleep stub firing the watchdog instantly.

It asserts the recorded note plus a post-failure sentinel now, restores the real sleep, and runs a succeeding body through the same path. Verified non-vacuously — with the failing body made to succeed:

🔴 FAIL stage() did not record the failure, or ran past it
🔴 1 harness check(s) FAILED

5. P2 — current-state claims in downstream locations — FIXED

All four. The two worth naming: CT 123 "cannot" expose token counters (it runs llama-server --metrics and does — your zero-counter check was the right shape of evidence, proving exposure not usage), and the parent README still saying the second card runs llama-swap. The "stock map" warning is relabelled historical with the general hazard kept.

6. P3 — the provenance correction compared the wrong tree — FIXED, retraction withdrawn

CT120 baseline: 32687325…   CT120 release: 4ac7aa75…   CT123 baseline: 32687325…

CT 120 holds both trees, its baseline is byte-identical to CT 123's, and the retraction compared CT 123's baseline against CT 120's release. The original copy history stands. Withdrawn, with the lesson recorded where it happened: name the tree, not the container.


Also landed

Env consolidation (your closing note): repeating the correction narrative in env comments hadn't prevented drift and had made the files unreadable — 276 lines across the two, 199 comment, for 40 settings. Now 204 and 128, each setting carrying its value, its trap and a pointer, history only in the README. All 40 VAR=value lines byte-identical, checked mechanically. Rewriting surfaced two more defects: comment blocks had drifted off their settings in both files, and the q8_0 ↔ --reasoning off XOR warning was missing entirely from the live config, the one actually running q8_0.

cutover-transition-test.sh — the one follow-up test you argued for. Drives the real script against a mocked pct/systemctl, asserting that at no point may two containers hold GPU 2, running or after a reboot. Six scenarios including failure injected during and after the release, and the deliberately stopped rollback target. Script paths are overridable now (defaulting to production) to make it possible.

Proven non-vacuous — five mutations, one at a time, each caught with the right message: watchdog map → disabled unit; autostart never cleared; assertion blind to a stopped container; autostart not restored on rollback; GPU 2 never detached on rollback.

⚠️ The mutation run immediately caught a bug in the test itself: two mutations read as "not caught" because the test aborted rather than reported — rel="$(grep … | head | cut)" with no match exits 1, pipefail propagates, the assignment inherits it, set -e kills the run. The same trap this folder hit before, inside the test written to catch such things. Fixed with || true; the harness now distinguishes caught / not-caught / aborted.

Gate

git ls-files '*.sh' | xargs shellcheck -S warning          CLEAN (0.11.0)
git diff --name-only -z origin/main...HEAD … | xargs -0 shellcheck -S warning   CLEAN
bash -n (all shell files)                                   CLEAN
py_compile (all python)                                     CLEAN
cutover-transition-test.sh                                  PASS
stage-harness.sh                                            PASS

The context-window--concurrency anchor in pro-v620/README.md is still broken; it also breaks on origin/main, so it predates this PR and fixing it means guessing the intended target. Left alone deliberately.

No live service was touched in this round — everything is mocked and runs in temp dirs.

Open by choice, not oversight: the native 262144 window, re-measuring two cards if llama.cpp #28699's per-device indexer fix lands, -ncmoe 16 + 31,17 + 4096/1024, and the DIMM-count comparison when the sticks arrive.

Two findings from the 45b7f47 review, both fixed. Neither is a regression
from the recent commits — finding 1 is an older path the new transition test
made visible, which is the test doing its job.

🔴 P1 — a failed `systemctl disable` was swallowed in both directions:

    pct exec "$VMID" -- systemctl disable --now llamacpp 2>/dev/null || true

If that failed the script carried on attaching GPU 2, restarting the
container, rewriting the watchdog map and enabling the incoming unit, leaving
BOTH model services enabled and active. They then contend for the card and
for port 1234, and the map points only at the incoming one — so a thermal
trip sheds the wrong load. Reproduced against the committed test with the
mocked disable forced to fail, and with container start given reboot
semantics:

    OBSERVED CT 120 units after a FAILED outgoing disable:
      120 llamacpp            enabled active
      120 llamacpp-qwen38fn   enabled active
    all cutover transition checks passed        <- the test could not see it

Fixed narrowly, via a retire_unit helper used in both directions. ⚠️ The
`|| true` on the disable itself is KEPT and that is deliberate: the unit may
legitimately be absent on a fresh install or in a direction already taken.
What is no longer optional is the verification — still enabled, or still
active, is now a fatal precondition failure. Tolerate the command failing;
never tolerate the unit surviving. Retirement precedes the GPU mutation in
both directions (204 before 210, 237 before 241), so an abort leaves no
half-attached card.

The test gains the invariant it was missing — the OUTGOING unit must be
disabled AND inactive after each direction — plus scenario 3c, where the
retirement fails and the script must refuse before touching GPU 2. The mock's
`start` now reactivates enabled units, which is what makes "still enabled" a
falsifiable claim rather than bookkeeping.

Non-vacuity proven by restoring the pre-fix one-liner: three distinct
failures, including "cut over while the old model server was still enabled".
⚠️ Two smaller mutations did NOT fail the test, and that is correct rather
than vacuous — removing either the is-enabled or the is-active check leaves
the other one catching the scenario. They guard different real states (comes
back on reboot vs running right now), so both stay.

P2 — the README's validation record claimed more than was measured. Its title
said "verification of the shipped two-card default" while what ran was the
THEN-default at 4096/1024, and its sweep row silently came from a different
binary. Retitled to name the previous default, with a binary column, the
three-case exercised/not-exercised matrix moved in from the env, and the
reduced-batch mitigation labelled as the inference it is rather than an
observed headroom result. "Do not quote 2881 / 1878 MiB as this
configuration's headroom" is now stated where someone would read it.

Gate: shellcheck -S warning clean on 0.11.0 by both commands, bash -n clean,
py_compile clean, transition test PASS (22 checks), stage harness PASS.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@marchah

marchah commented Sep 18, 2026

Copy link
Copy Markdown
Owner Author

Review triage — 45b7f47 follow-up: 2 findings, 2 fixed

Both reproduced before fixing. Neither is a regression from the recent commits — finding 1 is an older path that the new transition test made visible, which is the test earning its keep. Fixed in 6cef25d.


1. P1 — a failed service retirement survived the cutover — FIXED

Reproduced against the committed test, with the mocked disable forced to fail and container start given reboot semantics (production script unmodified, verified by diff):

OBSERVED CT 120 units after a FAILED outgoing disable:
    120 llamacpp            enabled active
    120 llamacpp-qwen38fn   enabled active
✅ all cutover transition checks passed          <- exit 0; the test could not see it

Both units enabled and active. They contend for the card and for port 1234, and the watchdog maps only to the incoming unit — so a thermal trip sheds the wrong load. Exactly as you described.

Fix: a retire_unit helper used in both directions.

⚠️ The || true on the disable itself is kept, deliberately — the unit may legitimately be absent on a fresh install or in a direction already taken. What is no longer optional is the verification: still enabled, or still active, is now a fatal precondition failure. Tolerate the command failing; never tolerate the unit surviving. That distinction is applied in both directions, as you asked, rather than a blanket removal.

Retirement precedes the GPU mutation in both paths (line 204 before 210, 237 before 241), so an abort can't leave a half-attached card.

Test extended with the invariant it was missing — the outgoing unit must be disabled AND inactive after each direction — plus a new scenario 3c where retirement fails and the script must refuse before touching GPU 2. The mock's start now reactivates enabled units, which is what makes "still enabled" falsifiable rather than bookkeeping.

Non-vacuity, by restoring the pre-fix one-liner:

🔴 cut over while the old model server was still enabled — two servers, one card
🔴 GPU 2 was attached despite the failed retirement
🔴 the incoming unit was enabled alongside a surviving outgoing unit
🔴 3 transition check(s) FAILED

⚠️ Two smaller mutations did not fail the test — removing either the is-enabled or the is-active check leaves the other catching the scenario. That's redundancy, not vacuity: they guard different real states (comes back on reboot vs running right now), and a partial disable --now can produce either. Both stay.

2. P2 — the validation record claimed more than was measured — FIXED

You're right that the caveat never reached the canonical record, and that the section title overclaimed. It said "verification of the shipped two-card default" while what ran was the then-default at 4096/1024, and the sweep row silently came from a different binary.

Retitled to name the previous default, with a binary column added and the exercised/not-exercised matrix moved in from the env:

exercised?
the release binary yes — at 4096/1024, the run that found 877 MiB
1024/256 yes — on b11018-baseline, the sweep
the release binary AT 1024/256 (what ships) no

And the mitigation is labelled as the inference it is: smaller compute buffers cannot use more memory than the run that already fit, so the shipped pair should have strictly more headroom — sound, and still not a measurement. "Do not quote 2881 / 1878 MiB as this configuration's headroom" now appears where someone would read it. No rerun performed or requested.


Gate

git ls-files '*.sh' | xargs shellcheck -S warning                CLEAN (0.11.0)
… origin/main...HEAD … | xargs -0 shellcheck -S warning          CLEAN
bash -n / py_compile                                             CLEAN
cutover-transition-test.sh                                       PASS (22 checks)
stage-harness.sh                                                 PASS

context-window--concurrency in pro-v620/README.md is still broken and still breaks on origin/main — predates this PR, left alone rather than guessing its target.

Nothing ran against the host this round; all cutover execution was against mocks in temporary paths.

Net -49 lines across 12 files. No fact removed, no trap removed — only the
narrative scaffolding around them.

Across four review rounds I kept correcting things by APPENDING: "an earlier
revision said X, that was wrong, actually Y". Thirteen such blocks
accumulated in the README, CLAUDE.md, both env files and six scripts. The
reviewer pointed out twice that this had not prevented propagation drift, and
it hadn't — the volume made the files harder to read while the wrong claims
kept surviving somewhere else anyway.

The rule applied here: if a claim was wrong, replace it with the right one. A
reader needs the current fact and the trap that will bite them; they do not
need the history of my mistakes.

What was kept, deliberately: every warning that stops someone re-introducing
a defect, rephrased as a property of the system rather than as a confession.
"The `- 2` is load-bearing: card 1 also carries the output head and a larger
KV share, so the obvious rule overcommits it and spills" is a trap. "CORRECTED
2026-09-18, this used to say ..." is a diary entry. Same information, and only
one of them is useful at 2am.

Biggest removals: the `-ncmoe 15` closure no longer opens with "two wrong
things were said about 15 in a row"; the q8_0 section states that the ~14%
belonged to a GTT spill rather than narrating its own retraction; the
candidate table gives its verdict instead of explaining which retracted figure
produced the previous one; CLAUDE.md's provenance note names the three build
trees instead of withdrawing a withdrawal; and the parent README states the
ROMED8-2T topology rather than quoting the B550 line it replaced.

Also trimmed one pre-existing instance of the same pattern in
create-lxc-llama-swap-gpu2.sh (the n-predict comment). The copy in
coder-runner/ is left alone — that file is outside this PR's diff and the
container is decommissioned.

Gate: shellcheck -S warning clean on 0.11.0, bash -n clean, py_compile clean,
cutover-transition-test.sh PASS, stage-harness.sh PASS, README anchors
resolve.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@marchah
marchah merged commit 2fabf5f into main Sep 18, 2026
@marchah
marchah deleted the marchah/Qwen-3.8-Flash branch September 18, 2026 20:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant