From ff2bb0eaef562063f43c95c7b967a67911ff0f27 Mon Sep 17 00:00:00 2001 From: Xore Date: Sat, 5 Sep 2026 15:26:53 +0200 Subject: [PATCH 1/7] docs(benchmarks): #1947 resume plan after the CPU swap and disk expansion MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The sweep's work area has been destroyed once (#2971) and mis-restored once, and three of its phase scripts were lost because they were never committed (#2985). A resume plan that lives only in /mnt-1/benchmarks has the same failure mode, so it goes in the repo. Records, from live inspection on 2026-09-05: - the hardware delta that prompted this — Xeon Silver 4110 -> Gold 5220R (8c/16t -> 24c/48t) and /var 153 G free -> 6.5 T free. RAM is unchanged at 6x16 GiB with two slots empty, and is now the only remaining hardware limit. - the completion ledger: phase 1 complete, phase 2 at 36/52 both tiers with 16 unstarted (5 of them unpullable, itemised with the call for each), phases 2.5/3/5 blocked on #2985, phase 4 sitting in draft PR #2641. - what the new hardware actually buys: `sweep_extra.sh`'s unconditional `ollama rm` is now harmful rather than necessary, and the CPU swap's own A/B has never been measured against the pre-swap timing table it was recorded for. - an ordered resume path P0-P8 with the owning issue for each step, the eight rules already paid for in wasted hours (pin a99e765 / 14 cases / max 69, no mixed scoring vintages, 0 != empty, cold slot, uniform +-0 is a symptom, injection axis reads as coverage-not-resistance, verify the benchmark leg and not the pull, restore scripts from the repo), a runbook, and four decisions that need an operator. Refs #1947, #2985, #2971, #2279, #2245, #2969, #1805, #1804 --- .../plans/2026-09-05-1947-resume-plan.md | 365 ++++++++++++++++++ ...> test_1609_installer_shared_framework.py} | 0 2 files changed, 365 insertions(+) create mode 100644 docs/benchmarks/plans/2026-09-05-1947-resume-plan.md rename tests/docs/{test_1609_installer_framework_parity.py => test_1609_installer_shared_framework.py} (100%) diff --git a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md new file mode 100644 index 000000000..6de26b60c --- /dev/null +++ b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md @@ -0,0 +1,365 @@ +# #1947 benchmark — resume plan after the CPU swap and disk expansion + +**Written** 2026-09-05, from live inspection of `homeserver` and every open +benchmark issue. Supersedes nothing; it sits *beside* +`/mnt-1/benchmarks/STATE-2026-09-05-fix-stage.md`, which remains the authority on +what was repaired on 2026-09-04/05. + +**Why this file is in the repo and not in the work area.** `/mnt-1/benchmarks/` +has now been destroyed once (#2971) and partially mis-restored once. Anything +whose loss costs GPU-hours belongs in git. Same reason #2985 exists. + +--- + +## 1. Where we actually are + +### 1.1 The hardware changed — measured, not assumed + +| | before (STATE-2026-08-31) | now (verified 2026-09-05) | +|---|---|---| +| CPU | Xeon **Silver 4110**, 8c/16t, 2.10 GHz, 11 MiB L3 | Xeon **Gold 5220R**, 24c/48t, 2.20 GHz | +| `/var` (Ollama blob store) | ~800 G, **92 % full, 153 G free** | **7.0 T, 6.5 T free** (`/dev/sdc1`) | +| `/mnt-1` (work area) | 1.8 T, 1.7 T free | 1.8 T, 1.7 T free — unchanged | +| RAM | 91 G, 6×16 GiB, **2 slots empty** | 93 G, 6×16 GiB, **2 slots still empty** | +| GPU | RTX 4000 Ada, 20475 MiB | unchanged, driver 610.57.04 | + +Two of the three constraints that shaped the sweep are gone. **The RAM upgrade +still has not happened** — that is the one remaining hardware limit, and it is +the binding one for the 123 B / 218 B RAM-offload rows. + +### 1.2 A sweep is running right now — and it is producing valid data + +Phase 2 was relaunched `2026-09-04T22:14:34Z` on the repaired setup and is alive: + +``` +xore timeout 10800 python3 analysis/ghidra/benchmarks/corpus/record_baseline.py … +``` + +Unlike the 2026-09-04 run that was rejected, **Tier B is scoring** — e.g. +`done B …Gemma4-31B-QAT… run2 score=63`. The `tierb-cache` regeneration held. + +Current cell: `GLM-4.6-REAP-218B-A32B-Derestricted:i1-IQ1_S`, Tier B run2, +started 13:22 (≈50 min/run, fully CPU-offloaded, 44 GB resident). + +### 1.3 Completion ledger + +| phase | what | state | +|---|---|---| +| 1 | main roster, 38 models, both tiers | **complete** (`FULLRUN_COMPLETE` in `fullrun.log`) | +| 2 | extra roster, 52 entries | **36 both tiers / 16 not started / 0 Tier-A-only** — running | +| 2.5 | harmony re-run of empty-answer models (#2279) | **blocked** — `gptoss_rerun.sh` missing (#2985) | +| 3 | self-requantisation sweep (#2245) | **blocked** — `requant_sweep.sh` missing (#2985) | +| 4 | ghidra-slot cold-protocol cohort (#1805-c) | measured; **PR #2641 still OPEN and DRAFT** | +| 5 | sessions + Rev·Deck slots across the roster | **blocked** — `slots_sweep.sh` missing (#2985) | + +Result files in `1947full/`: **152 Tier A, 144 Tier B**. +`TIER_A_ONLY_*.status` markers: 22 written, **21 now resolved by the relaunch**; +one still open — `qwen3.5_4b`, an older gap with an undetermined cause. + +### 1.4 The 16 phase-2 entries not started + +Five have already failed to pull, and they need a **roster decision, not a +retry**: + +| tag | failure | call | +|---|---|---| +| `mradermacher/XORTRON.CriminalComputing.LARGE.2026.3-i1:i1-IQ2_XXS` | `429` | **transient — retry.** This is #2245's deliberate 123 B RAM-offload row | +| `ahmedandaloes/CyberStrike-OffSec-35B:Q3_K_M` | `400` gated repo | request access or mark `UNMEASURED` | +| `protoLabsAI/ThinkingCap-Qwen3.6-27B-abliterated-MTP:Q4_K_M` | manifest realm error | mark `UNMEASURED` unless a mirror exists | +| `baronllm-llama3.1:q6_k` | local tag, file missing | needs local import; BaronLLM was gated (#1947 body) | +| `mradermacher/CyberPal2.0-20B:Q4_K_M` | pulls, emits **0 tokens** | **`UNMEASURABLE`** per #2696/#2982 — never a score | + +The remaining eleven are unattempted and mostly 27–35 B class. Three of them are +the **`gemma-4-26B-A4B-it-ultra-uncensored-heretic` Q3/Q4/Q5 quant ladder** — +that is the promote candidate from PR #2641, so those are the highest-value +cells left in phase 2. + +--- + +## 2. What the new hardware actually buys, and what it does not + +### 2.1 Disk: stop deleting weights + +`sweep_extra.sh:153` runs `ollama rm "$TAG"` after each model, and logs +`removed … (free now 6698G)`. That step existed because `/var` was at 92 %. It +is now **actively harmful**: every re-measure (phase 2.5, phase 3, phase 5, the +Ollama 0.33.3 requalification, any re-score) re-pulls tens of GB it already had. + +The whole roster at Q4 is on the order of 700 GB–1 TB against **6.5 T free**. + +**Change:** gate the removal behind a free-space floor rather than doing it +unconditionally — keep the weights while `/var` has, say, >1 T free. This is a +change to a committed script (`analysis/ghidra/benchmarks/corpus/sweep_extra.sh`), +so it goes through a PR, and **it must not be applied to the copy driving the +in-flight sweep** — mid-run edits are exactly the split-vintage defect #1947 +rule 6 exists to prevent. Land it, then deploy it at the phase-2/2.5 boundary. + +### 2.2 CPU: the swap's own A/B has never been measured + +The swap was made *to be tested* — STATE-2026-08-31 records a pre-swap +five-model timing table for precisely that purpose, and **no post-swap +comparison has been recorded anywhere.** The two data points we have are +suggestive but not comparable (different models and quants): + +| | pre-swap (Silver 4110) | post-swap (Gold 5220R) | +|---|---|---| +| 27 B Q4, ~19 GB, marginal fit | 18.6 – 27.6 min/run | — | +| 27 B IQ4_XS, ~16 GB, resident | — | 6.5 min/run | +| 31 B Q4, marginal fit | — | 22 – 26 min/run | +| 32 B Q8_0, ~34 GB, heavy spill | 78 – 127 min/run | — | +| 218 B IQ1_S, ~44 GB, full spill | — | ~50 min/run | + +**Anything that fits in VRAM is GPU-bound and should barely move. Judge the swap +on the spill rows.** A 218 B model at 50 min/run where a 32 B at Q8 previously +cost 78–127 is a strong hint, but it is a hint, not a measurement. + +The honest test is cheap: re-run the same five tags from the STATE table under +the same protocol. See §3, step **P5**. + +**Caveat that invalidated the pre-swap numbers and will invalidate these:** +load average on the homeserver right now is **15.4 / 20.6 / 21.3** from CI +runners, Elasticsearch and node. The pre-swap Trendyol rows spread 62 % on +contention alone. Quiesce CI on both sides of the comparison, or record +`uptime` with every run and compare like for like. Without that, the CPU swap +will appear to have done whatever CI happened to be doing. + +### 2.3 RAM: still 6×16 GiB with 2 slots empty + +This is the one thing that did not change, and it still bounds the top of the +roster. `GLM-4.6-REAP-218B:IQ1_S` is running at 44 GB resident inside 93 GB +while the live stack holds ~25 GB. `XORTRON.LARGE` at IQ2_XXS is the next one. +Both are thin. Two 16 GiB (or 32 GiB) DIMMs would take the box to 125 G/157 G +and make the offload rows comfortable rather than marginal — **but nothing in +this plan is blocked on it.** Both rows are measurable today; they are just slow. + +--- + +## 3. The plan, in order + +Each step names the issue that owns it. Nothing here changes the epic's rules — +the point is to finish the measurement, not to relax the bar. + +### P0 — Let phase 2 finish. Do not touch the running sweep. `#1947` + +11 unattempted pullable entries remain, mostly 27–35 B. At post-swap observed +rates (~6–26 min/run for resident/marginal, 4 runs per model) that is roughly +**15–25 h of GPU plus pull time**, and the pull time is now the larger half. + +Verification that it is still healthy, not just alive: + +```bash +ssh homeserver 'tail -20 /mnt-1/benchmarks/extra.log' +# must show BOTH "done A … score=" and "done B … score=" lines. +# Tier-B-only-failing is the exact 2026-09-04 defect; if you see GIVEUP on every +# tier B, stop the sweep — do not let it write MODEL_DONE over a hole. +ssh homeserver 'ls /mnt-1/benchmarks/1947full/*.status' +``` + +Do **not** measure progress by watching the model store grow. That measures a +pull, not a benchmark — the mistake that cost 8 h on 2026-09-04. + +### P1 — Rewrite and commit the three missing scripts. `#2985` — the real blocker + +`gptoss_rerun.sh`, `requant_sweep.sh` and `slots_sweep.sh` were never committed +and did not survive the rebuild. **Phases 2.5, 3 and 5 cannot run until they +exist.** This is the critical path for everything after phase 2, and it needs no +GPU, so it can be done entirely while P0 runs. + +- Rewrite from each phase's own issue: #2279 (harmony re-run), #2245 (requant + ladder), #1947 phase 5 (slots). +- Commit under `analysis/ghidra/benchmarks/corpus/` beside `sweep_extra.sh`, + with the same "operational copy lives at `/mnt-1/benchmarks/…`" header. +- While there, audit and commit the rest of the uncommitted operational set: + `preseed.sh`, `recreate-local-tags.sh`, `shutdown_watcher.sh`, + `resume_phases.sh`, `chain.sh`, `chain2b.sh`, `chain3.sh`, + `check-roster-name-lengths.sh`, `recover-oversized-models.sh`. +- Fold the `sweep_extra.sh` free-space gate from §2.1 into the same PR or a + sibling one. + +### P2 — Phase 2.5: harmony re-run. `#2279` + +Only step 3 of that issue is left; the harness fix (#2679) is already merged +**and is exactly the pinned `a99e765`**, so no code change and no re-pin is +needed. Four rows: `gpt-oss:20b`, +`GPT-OSS-Cybersecurity-20B-Merged-heretic` (both i1 and non-i1), and +`CyberPal2.0-20B` (which is `GptOssForCausalLM` despite the name — tag-substring +family detection misses it). + +Three of those four are known **zero-token** (#2696, #2982). Their pre-fix +`11/69` and `3/69` cells are written up as **unmeasured, never as scores**. + +### P3 — Phase 3: self-requantisation sweep. `#2245` + +Now genuinely cheap: the models are already pulled (once §2.1 lands) and +`/mnt-1` has 1.7 T for `llama-quantize` churn with `/var` no longer at 89 %. + +Note the framing #2245 already fixed: requantising an already-quantised GGUF is +lossy-on-lossy, so a Q4→Q3 row is a **fit-and-cost probe, not a clean quality +datapoint**. A spilling model gets two rows — as-published-with-offload, and +self-quantised-to-fit. The measured impossibility results +(`GLM-4.6-Derestricted-v3` at 0.36 bpw, `XORTRON.LARGE` at 1.04 bpw to fit +16 GB) stand as recorded rejections and are not re-opened. + +### P4 — Phase 5: sessions + Rev·Deck slots across the roster. `#1947` + +Parts 1–2 covered these slots for 33 models under the **contended** regime, +before the cold-slot protocol existed (#2641/#2642/#2646). Phase 5 is what makes +the sessions/revdeck columns comparable to the ghidra column. Run it cold: +sequential, `ollama stop ` between runs, `hp-llm-worker` and +`ghidra-revdeck-1` stopped. + +### P5 — Measure the CPU swap. `#1947`, new comment + +Re-run the five tags from STATE-2026-08-31's timing table under the same +protocol, with CI quiesced and `uptime` recorded per run: + +``` +GLM-4.7-Flash Q4 (fits) pre: 2.0 – 2.3 min/run +Ornith-1.5-9B Q4 (fits) pre: 2.6 – 2.8 +Titus Q4 (fits) pre: 3.7 – 4.8 +Qwen3.8-27B Q4 ~19 GB (marginal) pre: 18.6 – 27.6 +Trendyol 32B Q8_0 ~34 GB(spills) pre: 78.2 – 126.7 +``` + +Expected shape of the answer: the first three flat, the last two materially +faster. If the first three move, the measurement is contaminated by load, not +by the CPU. + +**Trendyol resume gotcha, still live:** `sweep_extra.sh` calls a model finished +when `tierA__run1.json` and `tierB__run1.json` both exist. Trendyol +has both, so a plain re-run **skips it** and leaves Tier B at N=1. Delete its +Tier B run1 first, or run the cell by hand. The quarantined partial is +`aborted-partial-tierB_Trendyol_run2.json.quarantine` — score 48 from 11 of 14 +cases; its real Tier B is 63. Never let that file back into a results glob. + +### P6 — Requalify Ollama 0.33.3. `#2969` + +Fully specified, needs the GPU free, and is **short**. It is the natural +occupant of any gap between phases. Blocks the runtime pin bump; the compose and +`origin/main` currently agree on tag *and* digest, so there is no drift to +untangle first. + +### P7 — Synthesis, and the PRs that are still open + +- **PR #2641 is OPEN and still DRAFT.** #1947's own protocol comment says it + "must be regenerated from the completed run rather than merged as it stands." + Decide explicitly: regenerate, or merge part 4 as a standalone and write the + full matrix separately. +- **#2641 carries no `Closes #1805`** — #1805 will not auto-close on merge. + The #2922 pattern; hand-close it. +- **#1804 Tier C** (`LLM4Decompile-Ref`) is still unmeasured, and blocked on + #1805's own Tier C row. +- **#1804 candidate 2** has numbers from the phase-1 data but they are + keyword-group scores, not the pooled-claims scoring. A transcript read is the + honest next step before it becomes a claim. +- **#2983** — `GHIDRA_VERSION=11.3.2` has to be exported by hand for + `ghidra_cache.py` because the headless service publishes no version. Fix it or + the next cache regeneration silently keys on an operator's memory. +- **#2986** (ml-worker Tier 2 accuracy) is a separate track with its own metrics + and does not gate any of the above. + +### P8 — Durability, so this is not restored from a mirror a third time + +The work area currently has **one** off-host copy: `~/apiary-bench-snapshots/` +on this workstation. `hermes` (192.168.42.253) rejects every key in the fleet — +its copy is a frozen artifact, not a live mirror, and the docs still point at it. + +- Fold the ~16 MB of scripts/roster into `scripts/backup-essentials.sh` so it + inherits the three-location fan-out. It qualifies under that script's own + "small config, no ES data or payloads" rule. +- Or fix `hermes`' keys and say so in the runbook. Either way, stop pointing at + a host nothing can reach. +- Note the workstation USB (`/run/media/xore/00586654-…`) is a udisks + auto-mount: check `mountpoint -q` before writing or an unattended run fills + the root filesystem instead. + +--- + +## 4. Rules that must not be broken while resuming + +These are the ones already paid for in wasted hours. + +1. **The pin `a99e765` is load-bearing.** 14 corpus cases, max **69**, and the + pre-recalibration injection gate. Merged `main` defaults to 17 cases, max + **79**. `resume_phases.sh` hard-aborts if the head moves. Do not pull it + forward mid-sweep. +2. **Never let two scoring vintages share a table.** Phase 1/2 (a99e765) and + PR #2641's part 4 (22f01c2) are both /69 but different harness commits — + state that explicitly in the writeup rather than assuming they compose. +3. **`0` ≠ empty.** A zero-token model is `UNMEASURABLE`, a model that never ran + is `UNMEASURED`. Neither is a score. `mark_unmeasured()` exists precisely so + "never ran" cannot be read as "ran badly" (#2696, #2982). +4. **Cold slot or the numbers are contaminated.** A warm slot does not + reproduce: `qwen3:14b` measured 60/61 warm with 1 of 14 answers identical, + 60/60 cold with 14 of 14. `hp-llm-worker` holding `qwen3:14b` resident at + CONTEXT 8192 is what made the incumbent's row wrong (#2642, #2644, #2646). +5. **Uniform `±0` is a symptom, not a result.** #2642 is open on why repeats are + byte-identical. Until it is settled, do not quote ±0.58 as this harness's + noise floor, and put repeats on the axes that are *not* fixed — seed, quant, + prompt phrasing — rather than re-deriving identical bytes. +6. **The injection axis reads as "coverage verified, resistance not measured"** + until the v3 positive control has fired (#2694, PR #2697 merged). The four + historical Tier A injection failures were re-read from transcripts and are + **false positives** of the forbidden-term matcher. "Score does not buy past + that gate" must not be applied to those four rows. +7. **Verify the benchmark leg, not the pull.** `du -sh` on the model store is + not evidence that anything is being measured. +8. **Restore operational scripts from the repo, never from `scripts-copy/`.** + The stale 4117-byte `sweep_extra.sh` cost a day; `origin/main`'s is 7213 + bytes and carries #2728's `PRESEED` workaround, pull-stderr logging, + `mark_unmeasured()` and #2738's roster pre-flight. + +--- + +## 5. Runbook + +```bash +# state of the running sweep +ssh homeserver 'tail -f /mnt-1/benchmarks/extra.log' +ssh homeserver 'cd /mnt-1/benchmarks && ls 1947full/tierA_*.json | wc -l && ls 1947full/tierB_*.json | wc -l' +ssh homeserver 'cat /mnt-1/benchmarks/1947full/failures.txt' +ssh homeserver 'uptime; nvidia-smi --query-gpu=memory.used,utilization.gpu --format=csv' + +# resume phase 2 from cold (only if nothing is running) +ssh homeserver 'cd /mnt-1/benchmarks && setsid nohup bash sweep_extra.sh >> extra.log 2>&1 Date: Sat, 5 Sep 2026 15:36:59 +0200 Subject: [PATCH 2/7] ops(ci): cut honeypot-ci runner instances from seven to two The homeserver is the GPU host for the #1947 model benchmark as well as the CI box, and CI is that benchmark's largest source of measurement noise: identical work on the same model spread 62% (78.2 vs 126.7 min/run) purely from runner contention, with load average sitting at 15-27 under seven instances. #1947's premise is that a delta smaller than the error bar is not a result, so a run taken against an unknown CI-driven load is not a measurement -- and the resume plan's CPU-swap A/B (step P5) cannot be answered at all without quiescing it. supermicro-ci-3..7 are `systemctl disable --now` on the host, not deregistered: registration, _work dir and tool cache survive, so restoring them is one command and needs no token. supermicro + supermicro-ci-2 stay. The separate honeypot-home deploy runner is untouched -- different label, no redundancy behind it. Seven was never a measured optimum, only the ceiling #2572 allowed once one instance stopped serialising the matrix; two instances on the Gold 5220R (24c/48t) have more cores behind them than four had on the Silver 4110. Documents the count and the restore command in docs/CI-CD.md, and records in the resume plan that the pre-swap baseline was itself taken under seven instances -- so the post-swap side is the quieter half of that comparison and uptime has to be recorded per run either way. Refs #1947, #2572 --- docs/CI-CD.md | 35 +++++++++++++++++++ .../plans/2026-09-05-1947-resume-plan.md | 21 ++++++++--- 2 files changed, 51 insertions(+), 5 deletions(-) diff --git a/docs/CI-CD.md b/docs/CI-CD.md index 75514ab93..c555494dd 100644 --- a/docs/CI-CD.md +++ b/docs/CI-CD.md @@ -809,6 +809,41 @@ runner, that cache survives between job runs on this same machine, so the second and every later run skips the download entirely. This is most of where the actual speed win comes from, not raw CPU. +#### How many instances, and why it is 2 during a benchmark window + +The box ran **seven** `honeypot-ci` instances (`supermicro`, +`supermicro-ci-2` .. `-7`) plus the separate `honeypot-home` deploy runner. +Reduced to **two** (`supermicro` + `supermicro-ci-2`) on 2026-09-05. + +This is a shared box, not a CI box. The homeserver is also the GPU host for +the #1947 model benchmark, and CI is that benchmark's largest source of +measurement noise: identical work on the same model spread **62%** (78.2 vs +126.7 min/run) purely from runner contention, and load average sat at +**15-27** with seven instances live. #1947's whole premise is that scores +separated by less than the error bar are not results, so a benchmark run +taken against an unknown, CI-driven load is not a measurement -- see +`docs/benchmarks/plans/2026-09-05-1947-resume-plan.md` step P5, which cannot +be answered at all without this. + +Seven was never a measured optimum; it was the ceiling #2572 allowed once +one instance stopped serialising the matrix. The CPU swap (Xeon Silver 4110 +8c/16t -> Xeon Gold 5220R 24c/48t) also means two instances today have more +cores behind them than four did before. + +To restore the extra instances once the benchmark window closes: + +```bash +ssh homeserver 'for n in 3 4 5 6 7; do + sudo systemctl enable --now actions.runner.Xore-APIARY.supermicro-ci-$n.service +done' +``` + +They are only `systemctl disable --now`, not deregistered -- the registration, +`_work` dir and tool cache all survive, so re-enabling costs nothing and needs +no token. **Stopping a busy instance fails its in-flight job**, so cancel the +run first or wait for idle; the 2026-09-05 reduction cancelled a run and cost +16 jobs a re-run. + ### Buildx layer cache: `type=local` on the homeserver (#2822) `containers.yml` used to export every image's layer cache to `type=gha`, diff --git a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md index 6de26b60c..5f7a0a402 100644 --- a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md +++ b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md @@ -117,11 +117,22 @@ The honest test is cheap: re-run the same five tags from the STATE table under the same protocol. See §3, step **P5**. **Caveat that invalidated the pre-swap numbers and will invalidate these:** -load average on the homeserver right now is **15.4 / 20.6 / 21.3** from CI -runners, Elasticsearch and node. The pre-swap Trendyol rows spread 62 % on -contention alone. Quiesce CI on both sides of the comparison, or record -`uptime` with every run and compare like for like. Without that, the CPU swap -will appear to have done whatever CI happened to be doing. +the pre-swap Trendyol rows spread **62 %** on contention alone (78.2 vs +126.7 min/run, identical work), and load average on the homeserver was +**15–27** with seven `honeypot-ci` runner instances live. Quiesce CI on both +sides of the comparison, or record `uptime` with every run and compare like +for like. Without that, the CPU swap will appear to have done whatever CI +happened to be doing. + +**Acted on 2026-09-05:** the seven `honeypot-ci` runner instances were +reduced to **two** (`supermicro` + `supermicro-ci-2`); `supermicro-ci-3`..`-7` +are `systemctl disable --now`, not deregistered, so re-enabling them is one +command and needs no token. The separate `honeypot-home` deploy runner is +untouched — it carries a different label with no redundancy behind it. See +`docs/CI-CD.md`, "How many instances, and why it is 2 during a benchmark +window". Note the pre-swap baseline in the table above was itself taken under +seven instances, so the post-swap side is the *quieter* side of the +comparison — record `uptime` per run either way and say so in the writeup. ### 2.3 RAM: still 6×16 GiB with 2 slots empty From a43f0d1126c3f40268e924189aafb82c7e576418 Mon Sep 17 00:00:00 2001 From: Xore Date: Sat, 5 Sep 2026 15:53:31 +0200 Subject: [PATCH 3/7] docs(benchmarks): record the phase-2 transcript durability gap in the resume plan The sweep's clone at /mnt-1/benchmarks/APIARY sits on a detached HEAD at a99e765, and every completed run writes its transcript dir into docs/benchmarks/runs/ as an untracked file. As of today that is 127 run dirs, 24 MB, all from 2026-09-04/05, on no branch, committed nowhere, in the one directory the rebuild already wiped once. Per #1944 the transcripts are the deliverable -- "no transcripts" is reason 4 of the six that made this epic necessary, and #2266's offline re-scoring reads them. Losing them turns every phase-2 row back into a number nobody can check. They must NOT be committed in that clone: a commit moves HEAD off a99e765 and resume_phases.sh hard-aborts on exactly that string, so the pin is the guard. Mirrored to the workstation instead, and the plan now carries the rsync, the reason the obvious fix is wrong, and the note that they get committed properly at write-up time on a branch cut from main. Also names PR #2641's branch in the ledger so the phase-4 results and the running phase-2 sweep are not read as the same work. Refs #1947, #1944, #2266, #2971 --- .../plans/2026-09-05-1947-resume-plan.md | 31 +++++++++++++++++-- 1 file changed, 28 insertions(+), 3 deletions(-) diff --git a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md index 5f7a0a402..80ae0d13b 100644 --- a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md +++ b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md @@ -49,7 +49,7 @@ started 13:22 (≈50 min/run, fully CPU-offloaded, 44 GB resident). | 2 | extra roster, 52 entries | **36 both tiers / 16 not started / 0 Tier-A-only** — running | | 2.5 | harmony re-run of empty-answer models (#2279) | **blocked** — `gptoss_rerun.sh` missing (#2985) | | 3 | self-requantisation sweep (#2245) | **blocked** — `requant_sweep.sh` missing (#2985) | -| 4 | ghidra-slot cold-protocol cohort (#1805-c) | measured; **PR #2641 still OPEN and DRAFT** | +| 4 | ghidra-slot cold-protocol cohort (#1805-c) | measured; **PR #2641 still OPEN and DRAFT** on `worktree-bench-1805c-results` | | 5 | sessions + Rev·Deck slots across the roster | **blocked** — `slots_sweep.sh` missing (#2985) | Result files in `1947full/`: **152 Tier A, 144 Tier B**. @@ -271,7 +271,31 @@ untangle first. ### P8 — Durability, so this is not restored from a mirror a third time -The work area currently has **one** off-host copy: `~/apiary-bench-snapshots/` +**The phase-2 transcripts are the exposed part, not the scripts.** The sweep's +clone at `/mnt-1/benchmarks/APIARY` sits on a **detached HEAD** at `a99e765`, +and every run it completes writes a transcript dir into +`docs/benchmarks/runs/` as an *untracked* file. As of 2026-09-05 that is +**127 run dirs, 24 MB, all from 2026-09-04/05**, on no branch, committed +nowhere, in the one directory that has already been wiped once. + +Per #1944 the transcripts *are* the deliverable — "no transcripts" is reason 4 +of the six that made this epic necessary, and offline re-scoring (#2266) reads +them. Losing them turns every phase-2 row back into an unrecheckable number. + +**Do not commit them in that clone.** A commit moves HEAD off `a99e765`, and +`resume_phases.sh` hard-aborts on exactly that string — the pin is the guard. +Mirror them instead: + +```bash +rsync -a homeserver:/mnt-1/benchmarks/APIARY/docs/benchmarks/runs/ \ + ~/apiary-bench-snapshots/homeserver-workarea/transcripts-phase2/ +``` + +Done once on 2026-09-05. It needs redoing as the sweep proceeds, or better, +folding into the same automation as the rest of P8. They get committed properly +at write-up time, on a branch cut from `main`, the way parts 1–4 were. + +Beyond that, the work area has **one** off-host copy: `~/apiary-bench-snapshots/` on this workstation. `hermes` (192.168.42.253) rejects every key in the fleet — its copy is a frozen artifact, not a live mirror, and the docs still point at it. @@ -354,7 +378,8 @@ Key paths on `homeserver`: | path | what | |---|---| | `/mnt-1/benchmarks/` | work area (**not backed up by `backup-essentials.sh`**) | -| `/mnt-1/benchmarks/APIARY` | pinned checkout, `a99e765` — must not move | +| `/mnt-1/benchmarks/APIARY` | pinned checkout, **detached HEAD** at `a99e765` — must not move, and must not be committed into | +| `…/APIARY/docs/benchmarks/runs/` | phase-2 transcripts, written **untracked**; mirror them, do not commit them here (P8) | | `/mnt-1/benchmarks/1947full/` | all result JSON, per-run logs, status markers | | `/mnt-1/benchmarks/models_extra_all.txt` | phase-2 roster, carrying the #2746/#2695/#2696 annotations | | `/mnt-1/benchmarks/tierb-cache/` | Ghidra decompilation cache; **its absence fails every Tier B run** | From 98c758345d6881ccaf5ddf9a5b3ad94700923781 Mon Sep 17 00:00:00 2001 From: Xore Date: Sat, 5 Sep 2026 16:06:47 +0200 Subject: [PATCH 4/7] =?UTF-8?q?benchmarks:=20stop=20losing=20#2245's=20inp?= =?UTF-8?q?uts=20=E2=80=94=20re-create=20the=20VRAM=20sampler,=20keep-alia?= =?UTF-8?q?s=20pulled=20weights?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Checked #2245 against the live host. Four of its premises no longer hold: - llama-quantize IS present (/usr/lib/ollama/llama-quantize in ghidra-ollama-1, with --allow-requantize), so the requant-from-GGUF path needs nothing new -- but there is no llama.cpp checkout anywhere on the box, so convert_hf_to_gguf.py does not exist and the rex86-eval container model-quant-benchmark/README.md assumes is not running. The issue's "harness we already have, this is applied not built" describes the repo, not the host. - vram_samples.tsv, the empirical spill list the requant plan rests on, was wiped with the work area, is in no mirror, and is NOT re-derivable: the result JSONs record no VRAM, no served size and no CPU/GPU split. - All ten source models named in requant_plan.txt have been deleted by sweep_extra.sh's `ollama rm`. One survives under a #2738 alias; at least three of the rest cannot be re-pulled at all. - The scorer is undecided: step 3 says corpus_eval.py via llama-server, step 4 demands comparison against rows scored by record_baseline.py on a99e765. Two scorers, incomparable numbers -- reason 6 of the six behind #1947. keep_and_sample.sh runs beside the sweep and changes nothing about it. It samples `ollama ps` back into vram_samples.tsv (served size, CPU/GPU split and served context exist only while a model is loaded -- sample or lose), and gives every pulled tag a keep/:src alias. `ollama cp` makes a second manifest over the SAME blobs -- verified, cp then rm of the copy left the blob count unchanged at 39 -- so the sweep's rm drops only its own manifest and the weights survive at zero extra disk. Deliberately not an edit to sweep_extra.sh: bash reads a running script lazily by byte offset, and the alias names cannot collide with a roster entry, so the sweep's own "already local -> will not delete" test is unaffected either way. First datum back: GLM-4.6-REAP-218B:i1-IQ1_S is 44 GB on disk, 57 GB served at ctx 32768, 65%/35% CPU/GPU -- confirming the KV arithmetic and showing served size cannot be inferred from disk size. Refs #2245, #2985, #2971, #1947 --- .../benchmarks/corpus/keep_and_sample.sh | 69 +++++++++++++++++++ .../plans/2026-09-05-1947-resume-plan.md | 57 ++++++++++++--- 2 files changed, 118 insertions(+), 8 deletions(-) create mode 100755 analysis/ghidra/benchmarks/corpus/keep_and_sample.sh diff --git a/analysis/ghidra/benchmarks/corpus/keep_and_sample.sh b/analysis/ghidra/benchmarks/corpus/keep_and_sample.sh new file mode 100755 index 000000000..e758849f0 --- /dev/null +++ b/analysis/ghidra/benchmarks/corpus/keep_and_sample.sh @@ -0,0 +1,69 @@ +#!/usr/bin/env bash +# keep_and_sample.sh -- runs alongside sweep_extra.sh, changes nothing about it. +# +# Two jobs, both read-mostly: +# +# 1. RE-CREATE THE LOST SAMPLER (#2245). /mnt-1/benchmarks/vram_samples.tsv was +# the empirical record of served size and CPU/GPU split per model -- the input +# #2245's category-1 "which models actually spill" list was supposed to rest +# on. It was wiped with the rest of the work area (#2971) and is in no mirror, +# and the result JSONs record no VRAM at all, so it cannot be re-derived after +# the fact. Sample `ollama ps` while each model is loaded or it is lost again. +# +# 2. STOP LOSING THE REQUANT INPUTS. sweep_extra.sh:153 does `ollama rm "$TAG"` +# on every tag it pulled -- correct when /var was at 92%, wrong now that it +# has 6.5T free, and actively destructive to #2245: all ten sources named in +# requant_plan.txt have already been deleted this way. `ollama cp` makes a +# second manifest over the SAME blobs (verified: cp+rm of the copy left the +# blob count unchanged), so a keep-alias costs no disk and survives the +# sweep's rm of the roster tag. +# +# Deliberately does NOT edit sweep_extra.sh: bash reads a running script lazily +# by byte offset, and the keep-aliases use names no roster entry can match, so +# the sweep's own "already local -> will not delete" test at line 119 is +# unaffected either way. +# +# Stop with: pkill -f keep_and_sample.sh (leaves every alias in place) +set -u +OLLAMA=ghidra-ollama-1 +SAMPLES=/mnt-1/benchmarks/vram_samples.tsv +KEPT=/mnt-1/benchmarks/kept-aliases.tsv +INTERVAL=60 + +oll() { docker exec "$OLLAMA" ollama "$@" 2>/dev/null; } + +[ -f "$SAMPLES" ] || printf 'ts\tname\tsize\tprocessor\tcontext\n' > "$SAMPLES" +[ -f "$KEPT" ] || printf 'ts\troster_tag\tkeep_alias\n' > "$KEPT" + +echo "$(date -u +%FT%TZ) KEEPSAMPLE_START interval=${INTERVAL}s" + +while true; do + ts=$(date -u +%FT%TZ) + + # --- 1. sample whatever is resident right now ------------------------------- + # `ollama ps` columns, whitespace-separated: + # $1 NAME $2 ID $3+$4 SIZE ("57 GB") $5+$6 PROCESSOR ("65%/35% CPU/GPU") + # $7 CONTEXT $8.. UNTIL + # SIZE is the *served* footprint (weights + KV at the served context), which + # is the number #2245 needs -- not the on-disk GGUF size. PROCESSOR is the + # spill: anything not "100% GPU" is a category-1 requantization candidate. + oll ps | tail -n +2 | awk -v ts="$ts" 'NF>=7 { + printf "%s\t%s\t%s %s\t%s %s\t%s\n", ts, $1, $3, $4, $5, $6, $7 + }' >> "$SAMPLES" + + # --- 2. keep-alias anything new the sweep pulled ---------------------------- + oll list | tail -n +2 | awk '{print $1}' | while IFS= read -r tag; do + [ -z "$tag" ] && continue + case "$tag" in keep/*) continue;; esac + slug=$(printf '%s' "$tag" | tr ':/' '__' | tr '[:upper:]' '[:lower:]') + alias="keep/${slug}:src" + if ! oll list | awk '{print $1}' | grep -qixF "$alias"; then + if oll cp "$tag" "$alias" >/dev/null 2>&1; then + printf '%s\t%s\t%s\n' "$ts" "$tag" "$alias" >> "$KEPT" + echo "$ts KEPT $tag -> $alias" + fi + fi + done + + sleep "$INTERVAL" +done diff --git a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md index 80ae0d13b..b5b7ef786 100644 --- a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md +++ b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md @@ -201,15 +201,56 @@ Three of those four are known **zero-token** (#2696, #2982). Their pre-fix ### P3 — Phase 3: self-requantisation sweep. `#2245` -Now genuinely cheap: the models are already pulled (once §2.1 lands) and -`/mnt-1` has 1.7 T for `llama-quantize` churn with `/var` no longer at 89 %. - -Note the framing #2245 already fixed: requantising an already-quantised GGUF is +Checked against the live host on 2026-09-05; **four of this issue's premises no +longer hold** and the full write-up is on #2245. Summary: + +**Tooling.** `llama-quantize` **is** present — `/usr/lib/ollama/llama-quantize` +inside `ghidra-ollama-1`, with `--allow-requantize` — so the requant-from-GGUF +path needs nothing new. But there is **no llama.cpp checkout anywhere on the +host**, so `convert_hf_to_gguf.py` does not exist, and the `rex86-eval` +container that `model-quant-benchmark/README.md` assumes (llama.cpp at +`/work/llama.cpp`) is not running. The issue's "harness we already have — this +is applied, not built" describes the repo, not the box. The clean f16 ladder — +which is the hypothesis the issue is actually about — needs that toolchain stood +up first, best as a container (`ghcr.io/ggml-org/llama.cpp:full` carries both +pieces) rather than host installs. + +**Inputs.** `requant_plan.txt` (10 rows) survived. `vram_samples.tsv` — the +empirical spill list the plan was supposed to rest on — did not, is in no +mirror, and is **not re-derivable**: the result JSONs record no VRAM, no served +size and no CPU/GPU split. And **all ten source models have been deleted** by +`sweep_extra.sh`'s `ollama rm`; one survives under the #2738 alias +`ravenx-cyberagent-35b:Q4_K_M`, and at least three of the rest cannot be +re-pulled at all (two `PULL_FAILED`, one never pulled). + +**Mitigated on 2026-09-05** by `/mnt-1/benchmarks/keep_and_sample.sh`, running +beside the sweep and changing nothing about it: it samples `ollama ps` back into +`vram_samples.tsv`, and gives every pulled tag a `keep/:src` alias. +`ollama cp` makes a second manifest over the *same* blobs — verified, `cp` then +`rm` of the copy left the blob count unchanged — so the sweep's `rm` drops only +its own manifest and the weights survive at zero extra disk. First datum back: +`GLM-4.6-REAP-218B:i1-IQ1_S` is **44 GB on disk → 57 GB served at ctx 32768, +65%/35% CPU/GPU**, which confirms the issue's KV arithmetic and shows served +size cannot be inferred from disk size. + +**The undecided question.** #2245 step 3 says score with +`engine-benchmark/corpus_eval.py` via `llama-server`; step 4 says winners must +beat the #1795/#1947 rows, which were scored by `record_baseline.py` at tiers +A/B on `a99e765`. Those are two different scorers and their numbers do not +compare — rule 2 above, and reason 6 of the six behind #1947. +`requant_plan.txt`'s third column is already a *new ollama tag*, so settle it +that way: **requantise → `ollama create` → `record_baseline.py` both tiers on +`a99e765`**, and each self-quant row lands beside its as-published twin in the +same matrix. + +Unchanged and still correct: requantising an already-quantised GGUF is lossy-on-lossy, so a Q4→Q3 row is a **fit-and-cost probe, not a clean quality -datapoint**. A spilling model gets two rows — as-published-with-offload, and -self-quantised-to-fit. The measured impossibility results -(`GLM-4.6-Derestricted-v3` at 0.36 bpw, `XORTRON.LARGE` at 1.04 bpw to fit -16 GB) stand as recorded rejections and are not re-opened. +datapoint**; a spilling model gets two rows, as-published-with-offload and +self-quantised-to-fit, and that pairing is the comparison. The measured +impossibility results (`GLM-4.6-Derestricted-v3` at 0.36 bpw, +`XORTRON.LARGE` at 1.04 bpw to fit 16 GB) stand as recorded rejections. XORTRON's +`IQ2_XXS` pull failed today on a transient `429` and is worth retrying — it is +the deliberate RAM-offload row. ### P4 — Phase 5: sessions + Rev·Deck slots across the roster. `#1947` From 31302666a1e11036591233ff57103082b893604f Mon Sep 17 00:00:00 2001 From: Xore Date: Sat, 5 Sep 2026 16:16:16 +0200 Subject: [PATCH 5/7] docs(benchmarks): record the four operator decisions and their consequences Answered 2026-09-05: - #2245 is scored by record_baseline.py on the a99e765 pin, Ollama-served, so self-quant rows land in the same matrix as their as-published twins. corpus_eval.py + llama-server is not used for this issue. - #2245 runs the full clean f16 ladder, not the cheap requant probe. - PR #2641 stays DRAFT and gets regenerated from the completed run, so #1805 and the gemma-4-26B-A4B ghidra promotion stay parked until the matrix is whole. Its branch is now a synthesis input -- do not delete it. - The two free DIMM slots get filled before phase 3. Two of those create prerequisites, which reorders the plan: phase 3 is gated on a RAM install and an HF token, so phase 5 (slots) moves ahead of it and the short jobs fill the gaps. The f16 ladder's toolchain turns out not to be a blocker at all: ghcr.io/ggml-org/llama.cpp:full pulls clean on the homeserver and carries convert_hf_to_gguf.py, llama-quantize and llama-server in one image, so "stand up llama.cpp" is a docker pull rather than a build, with no host package installs. Verified live. What remains is the HF token. The RAM finding is bigger than headroom. dmidecode shows DIMME1 and DIMMF1 empty, i.e. channels E and F entirely unpopulated -- the box runs 4-channel on a 6-channel Gold 5220R. CPU-offloaded inference is memory-bandwidth bound, not thread bound (llama-server sits at ~16 of 48 cores while 65% of the 218B is on the CPU), so filling those two slots is a bandwidth change on exactly the rows this decision cares about. Part number and the ECC/registered caveat recorded. Refs #1947, #2245, #1805, #2641 --- .../plans/2026-09-05-1947-resume-plan.md | 123 +++++++++++++++--- 1 file changed, 105 insertions(+), 18 deletions(-) diff --git a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md index b5b7ef786..b53139f70 100644 --- a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md +++ b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md @@ -145,6 +145,62 @@ this plan is blocked on it.** Both rows are measurable today; they are just slow --- +## 2.4 Decisions taken 2026-09-05 + +Four open questions were put to the operator and answered. They change the +ordering, so they sit above the plan rather than in a footnote. + +| question | decision | consequence | +|---|---|---| +| #2245 scorer | **`record_baseline.py` on the `a99e765` pin, Ollama-served** | self-quant rows land in the same matrix as their as-published twins. `corpus_eval.py` + `llama-server` is not used for this issue | +| #2245 scope | **full clean f16 ladder** | not the cheap requant probe. Needs a conversion toolchain and HF snapshot downloads — see below | +| PR #2641 | **regenerate from the completed run** | stays DRAFT. #1805 stays open, and the `gemma-4-26B-A4B` ghidra promotion stays parked until the matrix is whole | +| homeserver RAM | **fill the two free slots before phase 3** | phase 3 is now gated on hardware; phase 5 moves ahead of it | + +**Consequence for the order:** phase 3 is no longer the next GPU job. It is +gated on a RAM install *and* an HF token, so **phase 5 (slots) runs before +phase 3**, and the short jobs (#2279, #2969) fill the gaps. P0→P2 are unchanged. + +### The f16 ladder's toolchain is solved + +`ghcr.io/ggml-org/llama.cpp:full` pulled clean on the homeserver and carries +both halves in one image — verified: + +``` +/app/convert_hf_to_gguf.py (+ convert_lora_to_gguf.py, gguf-py/) +/app/llama-quantize (--allow-requantize, --imatrix, --dry-run) +/app/llama-server, llama-bench, llama-cli, … +``` + +So "stand up a llama.cpp toolchain" is a `docker pull`, not a build. No host +package installs, no torch on the box. What remains is an **HF token** — the +f16 path re-downloads snapshots, and `~/.cache/huggingface/` does not exist on +the homeserver at all. That is an operator action and the one hard blocker on +this decision. + +### RAM: two channels are empty, not just two slots + +`dmidecode` says this is worth more than the capacity: + +``` +DIMMA1 A2 B1 C1 D1 D2 populated 6 x 16 GiB +DIMME1 F1 EMPTY +Samsung M393A2K40CB2-CVF, 16 GiB DDR4 RDIMM (Registered/Buffered), +rated 2933 MT/s, running at 2666 (the Gold 5220R's ceiling), single rank +``` + +Channels **E and F are entirely unpopulated**, so the box is running +**4-channel on a 6-channel CPU**. CPU-offloaded inference is memory-bandwidth +bound, not thread bound — `llama-server` on the 218 B row sits at ~16 cores of +48 while 65 % of the model is on the CPU. Filling `DIMME1` + `DIMMF1` takes it +to 6-channel, which is a bandwidth change, not just a headroom change, and it +lands on exactly the rows this decision cares about. + +**What to buy:** 2 × Samsung `M393A2K40CB2-CVF` (or any 16 GiB DDR4-2933/2666 +1Rx4 **registered ECC** RDIMM — unbuffered/UDIMM will not post). Result: +**128 GiB across all six channels**. 2 × 32 GiB in those slots also works +(160 GiB) but mixes ranks; matching the existing part is the safe order. + ## 3. The plan, in order Each step names the issue that owns it. Nothing here changes the epic's rules — @@ -199,10 +255,30 @@ family detection misses it). Three of those four are known **zero-token** (#2696, #2982). Their pre-fix `11/69` and `3/69` cells are written up as **unmeasured, never as scores**. -### P3 — Phase 3: self-requantisation sweep. `#2245` +### P3 — Phase 3: the quantization ladder. `#2245` — **gated, runs after P4** + +**Decided 2026-09-05:** full clean f16 ladder, scored by `record_baseline.py` on +the `a99e765` pin. Gated on the RAM install and an HF token (§2.4), so phase 5 +(P4) takes the GPU first. + +Shape, once ungated, per base: + +1. `docker run ghcr.io/ggml-org/llama.cpp:full` → `convert_hf_to_gguf.py` on the + HF snapshot → f16 GGUF on `/mnt-1`; delete the snapshot. +2. `llama-quantize` down each level from the largest that fits ~16–17 GB of + weights at ctx 32768 through Q2-class. Skip levels already present — + resume-safe per model per level. +3. `ollama create` each level as its own tag, then `record_baseline.py --tier A` + and `--tier B`. Same scorer, same rubric, same vintage as every other row. +4. Pair each level against the as-published row for the same base, and against + the production slot. `vram_samples.tsv` supplies the residency column. + +f16 is quantize input only and gets deleted afterwards unless it is itself a +requested comparison point — a multi-hundred-GB file per model adds up even on +6.4 T. Checked against the live host on 2026-09-05; **four of this issue's premises no -longer hold** and the full write-up is on #2245. Summary: +longer held** and the full write-up is on #2245. Summary: **Tooling.** `llama-quantize` **is** present — `/usr/lib/ollama/llama-quantize` inside `ghidra-ollama-1`, with `--allow-requantize` — so the requant-from-GGUF @@ -252,7 +328,7 @@ impossibility results (`GLM-4.6-Derestricted-v3` at 0.36 bpw, `IQ2_XXS` pull failed today on a transient `429` and is worth retrying — it is the deliberate RAM-offload row. -### P4 — Phase 5: sessions + Rev·Deck slots across the roster. `#1947` +### P4 — Phase 5: sessions + Rev·Deck slots across the roster. `#1947` — **moved ahead of P3** Parts 1–2 covered these slots for 33 models under the **contended** regime, before the cold-slot protocol existed (#2641/#2642/#2646). Phase 5 is what makes @@ -293,10 +369,13 @@ untangle first. ### P7 — Synthesis, and the PRs that are still open -- **PR #2641 is OPEN and still DRAFT.** #1947's own protocol comment says it - "must be regenerated from the completed run rather than merged as it stands." - Decide explicitly: regenerate, or merge part 4 as a standalone and write the - full matrix separately. +- **PR #2641 stays DRAFT and gets regenerated from the completed run** + (decided 2026-09-05, per #1947's own protocol comment). So it is a + *synthesis-time* deliverable, not a merge that can happen now — and #1805 and + the `gemma-4-26B-A4B` ghidra promotion stay parked until the matrix is whole. + Its 62 committed transcript dirs and 2 matrices on + `worktree-bench-1805c-results` are the inputs to that regeneration; do not + delete the branch. - **#2641 carries no `Closes #1805`** — #1805 will not auto-close on merge. The #2922 pattern; hand-close it. - **#1804 Tier C** (`LLM4Decompile-Ref`) is still unmeasured, and blocked on @@ -429,14 +508,22 @@ Key paths on `homeserver`: --- -## 6. Open questions for Xore - -1. **RAM.** Two DIMM slots are free. Adding them is the only thing that would - make the 123 B / 218 B offload rows comfortable. Worth doing before phase 3, - or measure them slow and move on? -2. **PR #2641** — regenerate from the completed run as #1947's protocol comment - demands, or merge part 4 standalone and write the full matrix separately? -3. **The five unpullable roster entries** (§1.4) — chase access for the gated - ones, or record them as `UNMEASURED` with the reason and close the roster? -4. **CI quiescing.** The CPU-swap A/B (P5) is worthless against a load average - of 20. Is pausing the self-hosted runners for that window acceptable? +## 6. What is still needed from the operator + +The four questions of 2026-09-05 are answered and recorded in §2.4. What they +left behind: + +1. **An HF token on the homeserver.** Hard blocker on the decided f16 ladder + (P3) — `~/.cache/huggingface/` does not exist at all — and on BaronLLM, which + was never a size question, only an authentication one. A credential, so it + goes on the host directly and never into an issue or this file. +2. **2 × 16 GiB DDR4 registered ECC RDIMM** into `DIMME1` + `DIMMF1` — §2.4 has + the part number and why it is a memory-bandwidth change rather than just + capacity. Phase 3 waits on it. +3. **The five unpullable roster entries** (§1.4). Default if nobody says + otherwise: retry `XORTRON.LARGE` — its failure was a plain transient `429` — + and record the other four as `UNMEASURED` with the reason attached rather + than chasing gated-repo access. Say so if you want access chased instead. + +Settled, needing nothing further: CI contention — the seven `honeypot-ci` +runner instances were cut to two on 2026-09-05 (§2.2, `docs/CI-CD.md`). From b0c3f6824b078fc35bb89a6bf5dd0a5c5117c325 Mon Sep 17 00:00:00 2001 From: Xore Date: Sat, 5 Sep 2026 18:22:33 +0200 Subject: [PATCH 6/7] benchmarks: measure whether contention corrupted phase 2, and commit the probe Phase 2 ran without the cold protocol's worker-stop step: hp-llm-worker spent ~11 hours issuing a competing /api/chat every ~4 minutes against the same OLLAMA_MAX_LOADED_MODELS=1 slot (#2582 recurring, #3023). Phase 2 carries 12 escalations against phase 4's uniformly +-0 cold cells, which looked like the contamination that demoted #1805-c to a survey. Measured it instead of re-running two days of GPU on the suspicion: DeepHat-V1-7B:Q4_K_M 4.7 GB resident A contended [63,63] cold [63,63] ravenx-cyberagent-35b 21 GB spills B contended [64,63,64] cold [64,64,64] The resident control is load-bearing -- no CPU-offload nondeterminism is possible there, so contention would have to show up in it, and it reproduced exactly. The escalated cell resolved to the majority the N=3 protocol had already chosen. Verdict: keep phase 2, annotate it, do not re-run. Escalations cost GPU time, not accuracy. Checking all nine escalated cells surfaced a separate defect: eight resolve 2-of-3, but GLM-4.6-REAP-218B Tier B is [55, 57, 54] -- three distinct values, so N=3 cannot produce a majority and that cell has no defensible score. It must not be published as any of the three. sweep_extra.sh escalates once and has no handling for "still no majority"; that wants an UNRESOLVED marker and an N>=5 path. Commits both drivers, per #2985's rule that a script driving real benchmark runs must not live only on the host: - coldprobe.sh asserts its preconditions rather than assuming them (sweep stopped, no run in flight, hp-llm-worker down, tierb-cache present, head == a99e765) and writes to a separate directory so no roster row is polluted. - probe_at_gap.sh takes the next natural gap: it stops only the sweep's driver, never an in-flight record_baseline, because do_run's own `rm -f` cleanup is skipped when the parent dies -- which is how the 2026-08-31 abort left a valid-looking partial result carrying 11 of 14 cases. Also records the #3031 DNS incident: a container resolver predating #2974's host fix still had 1.1.1.1/8.8.8.8 first, port 53 to them is blocked here, so every lookup burned ~8s against ollama's 30s pull deadline -- 12 models lost in four minutes, followed by a false EXTRA_COMPLETE. Refs #1947, #3023, #3031, #2582, #2641, #2985, #2974 --- .../ghidra/benchmarks/corpus/coldprobe.sh | 87 +++++++++++++++++++ .../ghidra/benchmarks/corpus/probe_at_gap.sh | 54 ++++++++++++ .../plans/2026-09-05-1947-resume-plan.md | 56 ++++++++++++ 3 files changed, 197 insertions(+) create mode 100755 analysis/ghidra/benchmarks/corpus/coldprobe.sh create mode 100755 analysis/ghidra/benchmarks/corpus/probe_at_gap.sh diff --git a/analysis/ghidra/benchmarks/corpus/coldprobe.sh b/analysis/ghidra/benchmarks/corpus/coldprobe.sh new file mode 100755 index 000000000..0a4114e7b --- /dev/null +++ b/analysis/ghidra/benchmarks/corpus/coldprobe.sh @@ -0,0 +1,87 @@ +#!/usr/bin/env bash +# coldprobe.sh -- #3023: did llm-worker contention move phase-2 scores, or only +# wall clock? +# +# Phase 2 ran without the cold-slot protocol's worker-stop step, so 36 of its 52 +# models were measured while hp-llm-worker issued a competing /api/chat every +# ~4 minutes against the same OLLAMA_MAX_LOADED_MODELS=1 slot. Phase 2 shows 12 +# escalations; phase 4's cold cells were +-0 across three runs on four +# architectures. That is suggestive, not proof -- so measure it instead of +# arguing about it. +# +# Two models, both ALREADY measured and both still local (no pull, no delete): +# +# DeepHat-V1-7B:Q4_K_M 4.7 GB, fully GPU-resident stored A=[63,63] +# -> the control. A resident model has no CPU-offload nondeterminism, so if +# its cold score differs from its contended one, contention moved scores +# and phase 2 needs re-running. +# +# ravenx-cyberagent-35b:Q4_K_M 21 GB, spills to CPU stored B=[64,63,64] +# -> the escalated cell. If cold gives three identical values the escalation +# was contention; if it still disagrees the nondeterminism is intrinsic to +# spilling (float reduction order across threads), which would exonerate +# phase 2 rather than condemn it. +# +# Results go to a SEPARATE directory so no roster row is polluted, and the +# stored phase-2 files are never touched. +# +# Preconditions this asserts rather than assumes: hp-llm-worker down, +# sweep_extra.sh not running, no record_baseline in flight. +set -u +REPO=/mnt-1/benchmarks/APIARY +OUT=/mnt-1/benchmarks/coldprobe +CACHE=/mnt-1/benchmarks/tierb-cache +mkdir -p "$OUT/logs" + +die() { echo "ABORT: $*" >&2; exit 1; } + +pgrep -f "sweep_extra.sh" >/dev/null && die "sweep_extra.sh is running -- stop it between models first" +pgrep -f "record_baseline.py" >/dev/null && die "a record_baseline run is in flight" +docker ps --format '{{.Names}}' | grep -qx hp-llm-worker && die "hp-llm-worker is up -- this probe measures its absence" +[ -d "$CACHE" ] || die "tierb-cache missing" +head=$(git -C "$REPO" rev-parse --short HEAD) +[ "$head" = "a99e765" ] || die "repo head is $head, not a99e765 -- wrong scoring vintage" + +cd "$REPO" || die "no repo" + +run() { # tier tag slug n + local tier="$1" tag="$2" slug="$3" n="$4" + local out="$OUT/cold_tier${tier}_${slug}_run${n}.json" + [ -f "$out" ] && { echo "skip $tier $slug run$n (present)"; return 0; } + local extra=""; [ "$tier" = "B" ] && extra="--ghidra-cache $CACHE" + docker exec ghidra-ollama-1 ollama stop "$tag" >/dev/null 2>&1 # cold every run + sleep 5 + echo "$(date -u +%H:%M:%S) start $tier $slug run$n (load $(cut -d' ' -f1 /proc/loadavg))" + timeout 10800 python3 analysis/ghidra/benchmarks/corpus/record_baseline.py \ + --tier "$tier" $extra --model "$tag" \ + --operator coldprobe-3023 --provenance synthetic \ + --output "$out" > "$OUT/logs/cold_tier${tier}_${slug}_run${n}.log" 2>&1 + local rc=$? + if [ $rc -eq 0 ] && [ -f "$out" ]; then + echo "$(date -u +%H:%M:%S) done $tier $slug run$n score=$(python3 -c "import json;print(json.load(open('$out'))['total_score'])")" + else + echo "$(date -u +%H:%M:%S) FAIL $tier $slug run$n rc=$rc" + rm -f "$out" + fi +} + +echo "=== $(date -u +%FT%TZ) COLDPROBE_START head=$head ===" + +# control: fully resident, stored Tier A = [63, 63] +run A 'hf.co/mradermacher/DeepHat-V1-7B-GGUF:Q4_K_M' 'deephat_7b' 1 +run A 'hf.co/mradermacher/DeepHat-V1-7B-GGUF:Q4_K_M' 'deephat_7b' 2 + +# the escalated cell: spills to CPU, stored Tier B = [64, 63, 64] +run B 'ravenx-cyberagent-35b:Q4_K_M' 'ravenx35b' 1 +run B 'ravenx-cyberagent-35b:Q4_K_M' 'ravenx35b' 2 +run B 'ravenx-cyberagent-35b:Q4_K_M' 'ravenx35b' 3 + +echo +echo "=== verdict inputs ===" +printf 'deephat_7b TierA contended=[63,63] cold=[' +for n in 1 2; do f="$OUT/cold_tierA_deephat_7b_run${n}.json"; [ -f "$f" ] && printf '%s,' "$(python3 -c "import json;print(json.load(open('$f'))['total_score'])")"; done +printf ']\n' +printf 'ravenx35b TierB contended=[64,63,64] cold=[' +for n in 1 2 3; do f="$OUT/cold_tierB_ravenx35b_run${n}.json"; [ -f "$f" ] && printf '%s,' "$(python3 -c "import json;print(json.load(open('$f'))['total_score'])")"; done +printf ']\n' +echo "=== $(date -u +%FT%TZ) COLDPROBE_COMPLETE ===" diff --git a/analysis/ghidra/benchmarks/corpus/probe_at_gap.sh b/analysis/ghidra/benchmarks/corpus/probe_at_gap.sh new file mode 100755 index 000000000..43d63a4c2 --- /dev/null +++ b/analysis/ghidra/benchmarks/corpus/probe_at_gap.sh @@ -0,0 +1,54 @@ +#!/usr/bin/env bash +# probe_at_gap.sh -- run coldprobe.sh in the next natural gap of the phase-2 +# sweep, then put the sweep back. +# +# The gap matters: the probe must measure a genuinely uncontended slot, so it +# cannot run alongside sweep_extra.sh. But sweep_extra must not be killed +# mid-run either -- do_run's own `rm -f "$out"` cleanup is skipped when the +# parent dies, which is how the 2026-08-31 abort left a valid-looking partial +# result carrying 11 of 14 cases (see STATE-2026-08-31-cpu-swap.md). +# +# So: wait for the next MODEL_DONE, stop the sweep's driver only, let any +# in-flight run finish on its own, probe, then relaunch. sweep_extra skips +# every model that already has both tier run1 files, so the relaunch resumes +# exactly where it stopped. +set -u +BASE=/mnt-1/benchmarks +LOG=$BASE/probe_at_gap.log + +say() { echo "$(date -u +%FT%TZ) $*" | tee -a "$LOG"; } + +baseline=$(grep -c MODEL_DONE "$BASE/extra.log") +say "ARMED baseline_model_done=$baseline" + +# 1. wait for the sweep to finish its current model (up to 6h) +for _ in $(seq 1 720); do + [ "$(grep -c MODEL_DONE "$BASE/extra.log")" -gt "$baseline" ] && break + sleep 30 +done +[ "$(grep -c MODEL_DONE "$BASE/extra.log")" -gt "$baseline" ] || { say "TIMEOUT waiting for MODEL_DONE -- doing nothing"; exit 1; } +say "GAP model finished: $(grep MODEL_DONE "$BASE/extra.log" | tail -1)" + +# 2. stop the driver only, never an in-flight record_baseline +pkill -f "$BASE/sweep_extra.sh" 2>/dev/null +pkill -f "bash sweep_extra.sh" 2>/dev/null +sleep 3 +say "driver stopped (sweep_extra procs now: $(pgrep -cf sweep_extra.sh))" + +# 3. let any run that was already started finish by itself (up to 3h) +for _ in $(seq 1 360); do + pgrep -f record_baseline.py >/dev/null || break + sleep 30 +done +pgrep -f record_baseline.py >/dev/null && { say "ABORT: a record_baseline is still in flight after 3h"; exit 1; } +say "slot is idle -- starting cold probe" + +# 4. probe +bash "$BASE/coldprobe.sh" 2>&1 | tee -a "$LOG" +say "COLDPROBE finished rc=$?" + +# 5. put the sweep back +cd "$BASE" || exit 1 +setsid nohup bash "$BASE/sweep_extra.sh" >> "$BASE/extra.log" 2>&1 < /dev/null & +sleep 5 +say "sweep relaunched (procs: $(pgrep -cf sweep_extra.sh))" diff --git a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md index b53139f70..a4fd233f5 100644 --- a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md +++ b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md @@ -201,6 +201,62 @@ lands on exactly the rows this decision cares about. **128 GiB across all six channels**. 2 × 32 GiB in those slots also works (160 GiB) but mixes ranks; matching the existing part is the safe order. +## 2.5 Measured on 2026-09-05 — three findings that change what to trust + +### Phase 2 ran contended, and it did not matter. Measured, not argued. + +`sweep_extra.sh` does `ollama stop ` and nothing else; the cold protocol +#2641 established also requires the live workers stopped. Phase 2 never did +that, and `hp-llm-worker` spent ~11 hours issuing a competing `/api/chat` every +~4 minutes against the same `OLLAMA_MAX_LOADED_MODELS=1` slot (#2582 recurring, +tracked as **#3023**). Phase 2 carries 12 escalations; phase 4's cold cells were +all `±0`. + +Rather than re-run two days of GPU on a suspicion, two already-measured models +were re-run cold (`coldprobe.sh`, artefacts in `/mnt-1/benchmarks/coldprobe/`): + +| model | tier | contended | cold | | +|---|---|---|---|---| +| `DeepHat-V1-7B:Q4_K_M` — 4.7 GB, fully resident | A | `[63, 63]` | **`[63, 63]`** | identical | +| `ravenx-cyberagent-35b:Q4_K_M` — 21 GB, spills | B | `[64, 63, 64]` | **`[64, 64, 64]`** | same majority | + +The resident control is the load-bearing row — no CPU-offload nondeterminism is +possible there, so contention would have to show up in it. It reproduced +exactly. **Verdict: keep phase 2 and annotate it; do not re-run.** Escalations +cost GPU time, not accuracy. + +Annotate the matrix with: phase 2 ran with `hp-llm-worker` live; escalated cells +were resolved by N=3; **phase 2's wall-clock column is not comparable to phase +4's** and must not be used as a speed measure. + +### One cell has no defensible score, and the escalation protocol has a gap + +Of the nine escalated cells, eight resolve 2-of-3. The ninth does not resolve at +all: `GLM-4.6-REAP-218B:i1-IQ1_S` Tier B is **`[55, 57, 54]`** — three distinct +values, so `N=3` cannot yield a majority. It must not be published as 55, 57 or +54. It is also the most CPU-offloaded row measured (57 GB served, 65%/35% +CPU/GPU), so spill nondeterminism is the likelier cause than contention. + +`sweep_extra.sh` escalates once and has **no handling for "still no majority"**. +That needs an `UNRESOLVED` marker and an `N>=5` path, independent of #3023. + +### A stale container resolver silently ate the roster tail + +Between 15:42Z and 15:47Z the sweep lost **12 consecutive models** to +`lookup hf.co on 127.0.0.11:53: i/o timeout` and then wrote `EXTRA_COMPLETE` +anyway — the third appearance of the false-completion class. #2974 fixed the +*host* resolver order; `ghidra-ollama-1` had been up since before that fix and +Docker snapshots the host list at container start, so it still had public +resolvers first — and port 53 egress to them is blocked here (verified: +`1.1.1.1` and `8.8.8.8` time out, `192.168.42.253` responds). Every lookup burned +both timeouts: **~8 s per `getent`**, against ollama's 30 s pull deadline. + +Restarting the container fixed it — `8 s → 0.067 s`. Tracked as **#3031**, which +also carries the durable fix (pin `dns:` in the compose so resolver order does +not depend on when a container happened to start) and the failure-rate guard +`EXTRA_COMPLETE` needs. **Every long-running container predating #2974 has the +same latent fault.** + ## 3. The plan, in order Each step names the issue that owns it. Nothing here changes the epic's rules — From 1b8b643e5a94a293210c2c5891ee92a91c0b794f Mon Sep 17 00:00:00 2001 From: Xore Date: Sat, 5 Sep 2026 18:25:38 +0200 Subject: [PATCH 7/7] docs(benchmarks): fix a doc-path lint false positive in the resume plan 'the ~16 MB of scripts/roster' read as a citation of a tracked path 'scripts/roster', which does not exist. It was prose, not a path -- rephrased rather than allowlisted, since allowlisting a non-path would weaken the #2458 check for a real one later. Refs #1947, #2458 --- docs/benchmarks/plans/2026-09-05-1947-resume-plan.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md index a4fd233f5..3ff243c43 100644 --- a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md +++ b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md @@ -475,7 +475,7 @@ Beyond that, the work area has **one** off-host copy: `~/apiary-bench-snapshots/ on this workstation. `hermes` (192.168.42.253) rejects every key in the fleet — its copy is a frozen artifact, not a live mirror, and the docs still point at it. -- Fold the ~16 MB of scripts/roster into `scripts/backup-essentials.sh` so it +- Fold the ~16 MB of scripts and roster files into `scripts/backup-essentials.sh` so it inherits the three-location fan-out. It qualifies under that script's own "small config, no ES data or payloads" rule. - Or fix `hermes`' keys and say so in the runbook. Either way, stop pointing at