Skip to content

docs(benchmarks): #1947 resume plan after the CPU swap and disk expansion - #3014

Merged
Xore merged 9 commits into
mainfrom
docs/1947-resume-plan
Sep 5, 2026
Merged

Xore merged 9 commits into
mainfrom
docs/1947-resume-plan

Conversation

@Xore

@Xore Xore commented Sep 5, 2026

Copy link
Copy Markdown
Owner

Adds docs/benchmarks/plans/2026-09-05-1947-resume-plan.md.

The #1947 work area has been destroyed once (#2971) and mis-restored once, and three of its phase scripts were lost because they were never committed (#2985). A resume plan that lives only in /mnt-1/benchmarks has the same failure mode, so this one goes in the repo. New folder docs/benchmarks/plans/ for it.

Written from live inspection of the homeserver and every open benchmark issue on 2026-09-05.

What it records

The hardware delta that prompted it. Xeon Silver 4110 (8c/16t) → Xeon Gold 5220R (24c/48t); /var 153 G free at 92 % → 6.5 T free. RAM is unchanged at 6×16 GiB with two slots still empty, and is now the only remaining hardware limit — it bounds the 123 B / 218 B RAM-offload rows and nothing else.

Completion ledger. Phase 1 complete (38 models). Phase 2 at 36/52 both tiers, 16 unstarted, 0 Tier-A-only and running — the relaunched sweep is scoring Tier B, unlike the run rejected on 2026-09-04, and it resolved 21 of the 22 TIER_A_ONLY markers. Phases 2.5/3/5 blocked on #2985. Phase 4 sitting in draft PR #2641. 152 Tier A / 144 Tier B result files.

The five unpullable roster entries are itemised with the call for each — one of them (XORTRON.LARGE at 429) is a transient rate limit and just needs a retry, and it is #2245's deliberate 123 B offload row.

What the new hardware actually buys. sweep_extra.sh's unconditional ollama rm after each model existed because /var was at 92 %; at 6.5 T free it is now actively harmful, since every re-measure re-pulls tens of GB. And the CPU swap's own A/B has never been measured — STATE-2026-08-31 recorded a five-model pre-swap timing table precisely so it could be, and no post-swap comparison exists anywhere. The plan notes the two post-swap data points we have are suggestive but not comparable, and that a load average of 20 from CI would swamp the effect either way.

An ordered resume path P0–P8, each step naming its owning issue, plus the eight rules already paid for in wasted hours: the a99e765 pin (14 cases, max 69) is load-bearing; never let two scoring vintages share a table; 0 ≠ empty; cold slot or the numbers are contaminated; uniform ±0 is a symptom not a result; the injection axis reads as coverage-verified-resistance-not-measured; verify the benchmark leg and not the pull; restore operational scripts from the repo and never from scripts-copy/.

Closes with a runbook, a key-paths table, and four decisions that need an operator (RAM, #2641's disposition, the unpullable entries, and whether CI can be quiesced for the CPU A/B).

Not in scope here

Docs only — no script changes. The sweep_extra.sh free-space gate described in §2.1 is deliberately left for its own PR, and must not be applied to the copy driving the in-flight sweep.

Refs #1947, #2985, #2971, #2279, #2245, #2969, #1805, #1804

…sion

The sweep's work area has been destroyed once (#2971) and mis-restored once,
and three of its phase scripts were lost because they were never committed
(#2985). A resume plan that lives only in /mnt-1/benchmarks has the same
failure mode, so it goes in the repo.

Records, from live inspection on 2026-09-05:

- the hardware delta that prompted this — Xeon Silver 4110 -> Gold 5220R
  (8c/16t -> 24c/48t) and /var 153 G free -> 6.5 T free. RAM is unchanged at
  6x16 GiB with two slots empty, and is now the only remaining hardware limit.
- the completion ledger: phase 1 complete, phase 2 at 36/52 both tiers with
  16 unstarted (5 of them unpullable, itemised with the call for each),
  phases 2.5/3/5 blocked on #2985, phase 4 sitting in draft PR #2641.
- what the new hardware actually buys: `sweep_extra.sh`'s unconditional
  `ollama rm` is now harmful rather than necessary, and the CPU swap's own
  A/B has never been measured against the pre-swap timing table it was
  recorded for.
- an ordered resume path P0-P8 with the owning issue for each step, the
  eight rules already paid for in wasted hours (pin a99e765 / 14 cases /
  max 69, no mixed scoring vintages, 0 != empty, cold slot, uniform +-0 is a
  symptom, injection axis reads as coverage-not-resistance, verify the
  benchmark leg and not the pull, restore scripts from the repo), a runbook,
  and four decisions that need an operator.

Refs #1947, #2985, #2971, #2279, #2245, #2969, #1805, #1804
@github-actions

github-actions Bot commented Sep 5, 2026

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

Xore added 8 commits September 5, 2026 15:36
The homeserver is the GPU host for the #1947 model benchmark as well as the
CI box, and CI is that benchmark's largest source of measurement noise:
identical work on the same model spread 62% (78.2 vs 126.7 min/run) purely
from runner contention, with load average sitting at 15-27 under seven
instances. #1947's premise is that a delta smaller than the error bar is not
a result, so a run taken against an unknown CI-driven load is not a
measurement -- and the resume plan's CPU-swap A/B (step P5) cannot be
answered at all without quiescing it.

supermicro-ci-3..7 are `systemctl disable --now` on the host, not
deregistered: registration, _work dir and tool cache survive, so restoring
them is one command and needs no token. supermicro + supermicro-ci-2 stay.
The separate honeypot-home deploy runner is untouched -- different label,
no redundancy behind it.

Seven was never a measured optimum, only the ceiling #2572 allowed once one
instance stopped serialising the matrix; two instances on the Gold 5220R
(24c/48t) have more cores behind them than four had on the Silver 4110.

Documents the count and the restore command in docs/CI-CD.md, and records in
the resume plan that the pre-swap baseline was itself taken under seven
instances -- so the post-swap side is the quieter half of that comparison and
uptime has to be recorded per run either way.

Refs #1947, #2572
… resume plan

The sweep's clone at /mnt-1/benchmarks/APIARY sits on a detached HEAD at
a99e765, and every completed run writes its transcript dir into
docs/benchmarks/runs/ as an untracked file. As of today that is 127 run dirs,
24 MB, all from 2026-09-04/05, on no branch, committed nowhere, in the one
directory the rebuild already wiped once.

Per #1944 the transcripts are the deliverable -- "no transcripts" is reason 4
of the six that made this epic necessary, and #2266's offline re-scoring reads
them. Losing them turns every phase-2 row back into a number nobody can check.

They must NOT be committed in that clone: a commit moves HEAD off a99e765 and
resume_phases.sh hard-aborts on exactly that string, so the pin is the guard.
Mirrored to the workstation instead, and the plan now carries the rsync, the
reason the obvious fix is wrong, and the note that they get committed properly
at write-up time on a branch cut from main.

Also names PR #2641's branch in the ledger so the phase-4 results and the
running phase-2 sweep are not read as the same work.

Refs #1947, #1944, #2266, #2971
…keep-alias pulled weights

Checked #2245 against the live host. Four of its premises no longer hold:

- llama-quantize IS present (/usr/lib/ollama/llama-quantize in ghidra-ollama-1,
  with --allow-requantize), so the requant-from-GGUF path needs nothing new --
  but there is no llama.cpp checkout anywhere on the box, so
  convert_hf_to_gguf.py does not exist and the rex86-eval container
  model-quant-benchmark/README.md assumes is not running. The issue's "harness
  we already have, this is applied not built" describes the repo, not the host.
- vram_samples.tsv, the empirical spill list the requant plan rests on, was
  wiped with the work area, is in no mirror, and is NOT re-derivable: the result
  JSONs record no VRAM, no served size and no CPU/GPU split.
- All ten source models named in requant_plan.txt have been deleted by
  sweep_extra.sh's `ollama rm`. One survives under a #2738 alias; at least three
  of the rest cannot be re-pulled at all.
- The scorer is undecided: step 3 says corpus_eval.py via llama-server, step 4
  demands comparison against rows scored by record_baseline.py on a99e765. Two
  scorers, incomparable numbers -- reason 6 of the six behind #1947.

keep_and_sample.sh runs beside the sweep and changes nothing about it. It
samples `ollama ps` back into vram_samples.tsv (served size, CPU/GPU split and
served context exist only while a model is loaded -- sample or lose), and gives
every pulled tag a keep/<slug>:src alias. `ollama cp` makes a second manifest
over the SAME blobs -- verified, cp then rm of the copy left the blob count
unchanged at 39 -- so the sweep's rm drops only its own manifest and the weights
survive at zero extra disk.

Deliberately not an edit to sweep_extra.sh: bash reads a running script lazily
by byte offset, and the alias names cannot collide with a roster entry, so the
sweep's own "already local -> will not delete" test is unaffected either way.

First datum back: GLM-4.6-REAP-218B:i1-IQ1_S is 44 GB on disk, 57 GB served at
ctx 32768, 65%/35% CPU/GPU -- confirming the KV arithmetic and showing served
size cannot be inferred from disk size.

Refs #2245, #2985, #2971, #1947
…uences

Answered 2026-09-05:

- #2245 is scored by record_baseline.py on the a99e765 pin, Ollama-served, so
  self-quant rows land in the same matrix as their as-published twins.
  corpus_eval.py + llama-server is not used for this issue.
- #2245 runs the full clean f16 ladder, not the cheap requant probe.
- PR #2641 stays DRAFT and gets regenerated from the completed run, so #1805
  and the gemma-4-26B-A4B ghidra promotion stay parked until the matrix is
  whole. Its branch is now a synthesis input -- do not delete it.
- The two free DIMM slots get filled before phase 3.

Two of those create prerequisites, which reorders the plan: phase 3 is gated on
a RAM install and an HF token, so phase 5 (slots) moves ahead of it and the
short jobs fill the gaps.

The f16 ladder's toolchain turns out not to be a blocker at all:
ghcr.io/ggml-org/llama.cpp:full pulls clean on the homeserver and carries
convert_hf_to_gguf.py, llama-quantize and llama-server in one image, so
"stand up llama.cpp" is a docker pull rather than a build, with no host package
installs. Verified live. What remains is the HF token.

The RAM finding is bigger than headroom. dmidecode shows DIMME1 and DIMMF1
empty, i.e. channels E and F entirely unpopulated -- the box runs 4-channel on
a 6-channel Gold 5220R. CPU-offloaded inference is memory-bandwidth bound, not
thread bound (llama-server sits at ~16 of 48 cores while 65% of the 218B is on
the CPU), so filling those two slots is a bandwidth change on exactly the rows
this decision cares about. Part number and the ECC/registered caveat recorded.

Refs #1947, #2245, #1805, #2641
# Conflicts:
#	tests/docs/test_1609_installer_shared_framework.py
…the probe

Phase 2 ran without the cold protocol's worker-stop step: hp-llm-worker spent
~11 hours issuing a competing /api/chat every ~4 minutes against the same
OLLAMA_MAX_LOADED_MODELS=1 slot (#2582 recurring, #3023). Phase 2 carries 12
escalations against phase 4's uniformly +-0 cold cells, which looked like the
contamination that demoted #1805-c to a survey.

Measured it instead of re-running two days of GPU on the suspicion:

  DeepHat-V1-7B:Q4_K_M    4.7 GB resident   A  contended [63,63]     cold [63,63]
  ravenx-cyberagent-35b   21 GB spills      B  contended [64,63,64]  cold [64,64,64]

The resident control is load-bearing -- no CPU-offload nondeterminism is
possible there, so contention would have to show up in it, and it reproduced
exactly. The escalated cell resolved to the majority the N=3 protocol had
already chosen. Verdict: keep phase 2, annotate it, do not re-run. Escalations
cost GPU time, not accuracy.

Checking all nine escalated cells surfaced a separate defect: eight resolve
2-of-3, but GLM-4.6-REAP-218B Tier B is [55, 57, 54] -- three distinct values,
so N=3 cannot produce a majority and that cell has no defensible score. It must
not be published as any of the three. sweep_extra.sh escalates once and has no
handling for "still no majority"; that wants an UNRESOLVED marker and an N>=5
path.

Commits both drivers, per #2985's rule that a script driving real benchmark runs
must not live only on the host:

- coldprobe.sh asserts its preconditions rather than assuming them (sweep
  stopped, no run in flight, hp-llm-worker down, tierb-cache present, head ==
  a99e765) and writes to a separate directory so no roster row is polluted.
- probe_at_gap.sh takes the next natural gap: it stops only the sweep's driver,
  never an in-flight record_baseline, because do_run's own `rm -f` cleanup is
  skipped when the parent dies -- which is how the 2026-08-31 abort left a
  valid-looking partial result carrying 11 of 14 cases.

Also records the #3031 DNS incident: a container resolver predating #2974's host
fix still had 1.1.1.1/8.8.8.8 first, port 53 to them is blocked here, so every
lookup burned ~8s against ollama's 30s pull deadline -- 12 models lost in four
minutes, followed by a false EXTRA_COMPLETE.

Refs #1947, #3023, #3031, #2582, #2641, #2985, #2974
'the ~16 MB of scripts/roster' read as a citation of a tracked path
'scripts/roster', which does not exist. It was prose, not a path -- rephrased
rather than allowlisted, since allowlisting a non-path would weaken the #2458
check for a real one later.

Refs #1947, #2458
@Xore
Xore merged commit 7a617e1 into main Sep 5, 2026
126 checks passed
@Xore
Xore deleted the docs/1947-resume-plan branch September 5, 2026 18:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant