docs(benchmarks): #1947 resume plan after the CPU swap and disk expansion - #3014
Merged
Merged
Conversation
…sion The sweep's work area has been destroyed once (#2971) and mis-restored once, and three of its phase scripts were lost because they were never committed (#2985). A resume plan that lives only in /mnt-1/benchmarks has the same failure mode, so it goes in the repo. Records, from live inspection on 2026-09-05: - the hardware delta that prompted this — Xeon Silver 4110 -> Gold 5220R (8c/16t -> 24c/48t) and /var 153 G free -> 6.5 T free. RAM is unchanged at 6x16 GiB with two slots empty, and is now the only remaining hardware limit. - the completion ledger: phase 1 complete, phase 2 at 36/52 both tiers with 16 unstarted (5 of them unpullable, itemised with the call for each), phases 2.5/3/5 blocked on #2985, phase 4 sitting in draft PR #2641. - what the new hardware actually buys: `sweep_extra.sh`'s unconditional `ollama rm` is now harmful rather than necessary, and the CPU swap's own A/B has never been measured against the pre-swap timing table it was recorded for. - an ordered resume path P0-P8 with the owning issue for each step, the eight rules already paid for in wasted hours (pin a99e765 / 14 cases / max 69, no mixed scoring vintages, 0 != empty, cold slot, uniform +-0 is a symptom, injection axis reads as coverage-not-resistance, verify the benchmark leg and not the pull, restore scripts from the repo), a runbook, and four decisions that need an operator. Refs #1947, #2985, #2971, #2279, #2245, #2969, #1805, #1804
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
The homeserver is the GPU host for the #1947 model benchmark as well as the CI box, and CI is that benchmark's largest source of measurement noise: identical work on the same model spread 62% (78.2 vs 126.7 min/run) purely from runner contention, with load average sitting at 15-27 under seven instances. #1947's premise is that a delta smaller than the error bar is not a result, so a run taken against an unknown CI-driven load is not a measurement -- and the resume plan's CPU-swap A/B (step P5) cannot be answered at all without quiescing it. supermicro-ci-3..7 are `systemctl disable --now` on the host, not deregistered: registration, _work dir and tool cache survive, so restoring them is one command and needs no token. supermicro + supermicro-ci-2 stay. The separate honeypot-home deploy runner is untouched -- different label, no redundancy behind it. Seven was never a measured optimum, only the ceiling #2572 allowed once one instance stopped serialising the matrix; two instances on the Gold 5220R (24c/48t) have more cores behind them than four had on the Silver 4110. Documents the count and the restore command in docs/CI-CD.md, and records in the resume plan that the pre-swap baseline was itself taken under seven instances -- so the post-swap side is the quieter half of that comparison and uptime has to be recorded per run either way. Refs #1947, #2572
… resume plan The sweep's clone at /mnt-1/benchmarks/APIARY sits on a detached HEAD at a99e765, and every completed run writes its transcript dir into docs/benchmarks/runs/ as an untracked file. As of today that is 127 run dirs, 24 MB, all from 2026-09-04/05, on no branch, committed nowhere, in the one directory the rebuild already wiped once. Per #1944 the transcripts are the deliverable -- "no transcripts" is reason 4 of the six that made this epic necessary, and #2266's offline re-scoring reads them. Losing them turns every phase-2 row back into a number nobody can check. They must NOT be committed in that clone: a commit moves HEAD off a99e765 and resume_phases.sh hard-aborts on exactly that string, so the pin is the guard. Mirrored to the workstation instead, and the plan now carries the rsync, the reason the obvious fix is wrong, and the note that they get committed properly at write-up time on a branch cut from main. Also names PR #2641's branch in the ledger so the phase-4 results and the running phase-2 sweep are not read as the same work. Refs #1947, #1944, #2266, #2971
…keep-alias pulled weights Checked #2245 against the live host. Four of its premises no longer hold: - llama-quantize IS present (/usr/lib/ollama/llama-quantize in ghidra-ollama-1, with --allow-requantize), so the requant-from-GGUF path needs nothing new -- but there is no llama.cpp checkout anywhere on the box, so convert_hf_to_gguf.py does not exist and the rex86-eval container model-quant-benchmark/README.md assumes is not running. The issue's "harness we already have, this is applied not built" describes the repo, not the host. - vram_samples.tsv, the empirical spill list the requant plan rests on, was wiped with the work area, is in no mirror, and is NOT re-derivable: the result JSONs record no VRAM, no served size and no CPU/GPU split. - All ten source models named in requant_plan.txt have been deleted by sweep_extra.sh's `ollama rm`. One survives under a #2738 alias; at least three of the rest cannot be re-pulled at all. - The scorer is undecided: step 3 says corpus_eval.py via llama-server, step 4 demands comparison against rows scored by record_baseline.py on a99e765. Two scorers, incomparable numbers -- reason 6 of the six behind #1947. keep_and_sample.sh runs beside the sweep and changes nothing about it. It samples `ollama ps` back into vram_samples.tsv (served size, CPU/GPU split and served context exist only while a model is loaded -- sample or lose), and gives every pulled tag a keep/<slug>:src alias. `ollama cp` makes a second manifest over the SAME blobs -- verified, cp then rm of the copy left the blob count unchanged at 39 -- so the sweep's rm drops only its own manifest and the weights survive at zero extra disk. Deliberately not an edit to sweep_extra.sh: bash reads a running script lazily by byte offset, and the alias names cannot collide with a roster entry, so the sweep's own "already local -> will not delete" test is unaffected either way. First datum back: GLM-4.6-REAP-218B:i1-IQ1_S is 44 GB on disk, 57 GB served at ctx 32768, 65%/35% CPU/GPU -- confirming the KV arithmetic and showing served size cannot be inferred from disk size. Refs #2245, #2985, #2971, #1947
…uences Answered 2026-09-05: - #2245 is scored by record_baseline.py on the a99e765 pin, Ollama-served, so self-quant rows land in the same matrix as their as-published twins. corpus_eval.py + llama-server is not used for this issue. - #2245 runs the full clean f16 ladder, not the cheap requant probe. - PR #2641 stays DRAFT and gets regenerated from the completed run, so #1805 and the gemma-4-26B-A4B ghidra promotion stay parked until the matrix is whole. Its branch is now a synthesis input -- do not delete it. - The two free DIMM slots get filled before phase 3. Two of those create prerequisites, which reorders the plan: phase 3 is gated on a RAM install and an HF token, so phase 5 (slots) moves ahead of it and the short jobs fill the gaps. The f16 ladder's toolchain turns out not to be a blocker at all: ghcr.io/ggml-org/llama.cpp:full pulls clean on the homeserver and carries convert_hf_to_gguf.py, llama-quantize and llama-server in one image, so "stand up llama.cpp" is a docker pull rather than a build, with no host package installs. Verified live. What remains is the HF token. The RAM finding is bigger than headroom. dmidecode shows DIMME1 and DIMMF1 empty, i.e. channels E and F entirely unpopulated -- the box runs 4-channel on a 6-channel Gold 5220R. CPU-offloaded inference is memory-bandwidth bound, not thread bound (llama-server sits at ~16 of 48 cores while 65% of the 218B is on the CPU), so filling those two slots is a bandwidth change on exactly the rows this decision cares about. Part number and the ECC/registered caveat recorded. Refs #1947, #2245, #1805, #2641
# Conflicts: # tests/docs/test_1609_installer_shared_framework.py
…the probe Phase 2 ran without the cold protocol's worker-stop step: hp-llm-worker spent ~11 hours issuing a competing /api/chat every ~4 minutes against the same OLLAMA_MAX_LOADED_MODELS=1 slot (#2582 recurring, #3023). Phase 2 carries 12 escalations against phase 4's uniformly +-0 cold cells, which looked like the contamination that demoted #1805-c to a survey. Measured it instead of re-running two days of GPU on the suspicion: DeepHat-V1-7B:Q4_K_M 4.7 GB resident A contended [63,63] cold [63,63] ravenx-cyberagent-35b 21 GB spills B contended [64,63,64] cold [64,64,64] The resident control is load-bearing -- no CPU-offload nondeterminism is possible there, so contention would have to show up in it, and it reproduced exactly. The escalated cell resolved to the majority the N=3 protocol had already chosen. Verdict: keep phase 2, annotate it, do not re-run. Escalations cost GPU time, not accuracy. Checking all nine escalated cells surfaced a separate defect: eight resolve 2-of-3, but GLM-4.6-REAP-218B Tier B is [55, 57, 54] -- three distinct values, so N=3 cannot produce a majority and that cell has no defensible score. It must not be published as any of the three. sweep_extra.sh escalates once and has no handling for "still no majority"; that wants an UNRESOLVED marker and an N>=5 path. Commits both drivers, per #2985's rule that a script driving real benchmark runs must not live only on the host: - coldprobe.sh asserts its preconditions rather than assuming them (sweep stopped, no run in flight, hp-llm-worker down, tierb-cache present, head == a99e765) and writes to a separate directory so no roster row is polluted. - probe_at_gap.sh takes the next natural gap: it stops only the sweep's driver, never an in-flight record_baseline, because do_run's own `rm -f` cleanup is skipped when the parent dies -- which is how the 2026-08-31 abort left a valid-looking partial result carrying 11 of 14 cases. Also records the #3031 DNS incident: a container resolver predating #2974's host fix still had 1.1.1.1/8.8.8.8 first, port 53 to them is blocked here, so every lookup burned ~8s against ollama's 30s pull deadline -- 12 models lost in four minutes, followed by a false EXTRA_COMPLETE. Refs #1947, #3023, #3031, #2582, #2641, #2985, #2974
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
docs/benchmarks/plans/2026-09-05-1947-resume-plan.md.The #1947 work area has been destroyed once (#2971) and mis-restored once, and three of its phase scripts were lost because they were never committed (#2985). A resume plan that lives only in
/mnt-1/benchmarkshas the same failure mode, so this one goes in the repo. New folderdocs/benchmarks/plans/for it.Written from live inspection of the homeserver and every open benchmark issue on 2026-09-05.
What it records
The hardware delta that prompted it. Xeon Silver 4110 (8c/16t) → Xeon Gold 5220R (24c/48t);
/var153 G free at 92 % → 6.5 T free. RAM is unchanged at 6×16 GiB with two slots still empty, and is now the only remaining hardware limit — it bounds the 123 B / 218 B RAM-offload rows and nothing else.Completion ledger. Phase 1 complete (38 models). Phase 2 at 36/52 both tiers, 16 unstarted, 0 Tier-A-only and running — the relaunched sweep is scoring Tier B, unlike the run rejected on 2026-09-04, and it resolved 21 of the 22
TIER_A_ONLYmarkers. Phases 2.5/3/5 blocked on #2985. Phase 4 sitting in draft PR #2641. 152 Tier A / 144 Tier B result files.The five unpullable roster entries are itemised with the call for each — one of them (
XORTRON.LARGEat429) is a transient rate limit and just needs a retry, and it is #2245's deliberate 123 B offload row.What the new hardware actually buys.
sweep_extra.sh's unconditionalollama rmafter each model existed because/varwas at 92 %; at 6.5 T free it is now actively harmful, since every re-measure re-pulls tens of GB. And the CPU swap's own A/B has never been measured — STATE-2026-08-31 recorded a five-model pre-swap timing table precisely so it could be, and no post-swap comparison exists anywhere. The plan notes the two post-swap data points we have are suggestive but not comparable, and that a load average of 20 from CI would swamp the effect either way.An ordered resume path P0–P8, each step naming its owning issue, plus the eight rules already paid for in wasted hours: the
a99e765pin (14 cases, max 69) is load-bearing; never let two scoring vintages share a table;0≠ empty; cold slot or the numbers are contaminated; uniform±0is a symptom not a result; the injection axis reads as coverage-verified-resistance-not-measured; verify the benchmark leg and not the pull; restore operational scripts from the repo and never fromscripts-copy/.Closes with a runbook, a key-paths table, and four decisions that need an operator (RAM, #2641's disposition, the unpullable entries, and whether CI can be quiesced for the CPU A/B).
Not in scope here
Docs only — no script changes. The
sweep_extra.shfree-space gate described in §2.1 is deliberately left for its own PR, and must not be applied to the copy driving the in-flight sweep.Refs #1947, #2985, #2971, #2279, #2245, #2969, #1805, #1804