feat(benchmarks): rewrite requant_sweep.sh — the clean f16 quantization ladder (#2245 phase 3) - #3042
Merged
Merged
Conversation
…on ladder (#2245 phase 3) One of the three scripts lost in the 2026-09-03/04 rebuild because it was never committed (#2985). Rewritten from #2245 rather than restored, since nothing of the original survives. ## What it does Per base, resume-safe at every step: HF snapshot -> convert_hf_to_gguf.py -> f16 master (snapshot deleted immediately) -> llama-quantize down each requested level -> ollama create a tag per level -> score. Re-running after an interruption picks up where it stopped rather than re-downloading or re-converting, which matters when an f16 of a 35B base is 70 GB. The operator chose the full clean f16 ladder over the cheap --allow-requantize path, so this converts from original weights every time. #2245's own comment is why: requantizing an already-quantized GGUF is lossy-on-lossy and gives "a fit-and-cost probe, not a clean quality datapoint". ## It does not reimplement scoring sweep_extra.sh already owns the cold-slot protocol, the N=2->3->5 escalation and the UNRESOLVED marker (#3036). Duplicating that would produce a second code path whose numbers cannot be compared with the as-published rows -- rule 6 of #1947's six, and the reason #1805-c had to be demoted to a survey. So BASE/LIST/REPO/ PRESEED/MAXTRY in sweep_extra.sh are now env-overridable, and requant_sweep.sh builds a tag list and hands it over. That also makes the ladder safe by construction: locally created tags are never `ollama pull`ed, so sweep_extra's own "already local -> will not delete" branch protects hours of conversion work from its own cleanup step. ## Toolchain ghcr.io/ggml-org/llama.cpp:full, verified on the homeserver -- it carries convert_hf_to_gguf.py, gguf-py, llama-quantize and llama-gguf-split in one image. There is no llama.cpp checkout and no rex86-eval container on that host, so model-quant-benchmark/README.md's documented layout does not exist there. Containers run --user $(id -u) deliberately: a root-owned file in a work tree is how #3024 bricked every CI runner. ## The plan file is grounded, not guessed f16_ladder_plan.txt names ORIGINAL safetensors repos, not GGUF repos -- there is nothing for convert_hf_to_gguf.py to convert in a GGUF repo. Each base_model was read off the corresponding GGUF repo's card and verified live to exist, be ungated, and carry config.json plus safetensors (Ornith-35B 70 GB, gemma-26B-A4B 52 GB). Every level name was checked against llama-quantize's own type list. Levels bracket the measured envelope rather than following habit: 20475 MiB of VRAM minus 3-5 GB of KV at CONTEXT 32768 leaves a ~16-17 GB weights budget, so each ladder crosses the residency line -- one level above, one at, one or two below. The question is whether parameters bought with bits pay off, and that is unanswerable if the ladder never crosses. Each row's top level matches a quant already measured as-published, giving every ladder a same-base same-scorer control. Refs #2245, #2985, #1947, #3036
The first live run printed five `xargs: unmatched single quote` errors before reaching the first base. Cause: the line was trimmed with `echo ... | xargs` *before* the comment check, so xargs saw the apostrophes in the plan file's own prose comments and errored on each one. Cosmetic in this run -- the `case` still skipped the comment lines -- but it would silently mangle any plan value containing a quote, and it hid the real start of the run behind noise. Trim with bash parameter expansion instead, and do the comment/blank check before any processing. Verified against the real plan file: both rows parse to the expected base/levels/prefix. Refs #2245, #2985
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
…nverting
The gemma-4-26B-A4B ladder failed to convert. The error surfaced 20 frames deep
inside transformers with no mention of the model or the field:
tokenization_utils_base.py, _set_model_specific_special_tokens:
self.SPECIAL_TOKENS_ATTRIBUTES + list(special_tokens.keys())
AttributeError: 'list' object has no attribute 'keys'
Cause, read out of the snapshot rather than guessed:
llmfan46/gemma-4-26B-A4B-it-ultra-uncensored-heretic ships
"extra_special_tokens": ["<|video|>"]
a bare list where transformers requires a dict. The same file's sibling field
is correctly formed --
"model_specific_special_tokens": {"audio_token": "<|audio|>",
"boi_token": "<|image>", ...}
-- so the intended shape is unambiguous and the list is simply malformed
upstream.
Deleting the field would also have "fixed" the crash, and would have been
wrong: <|video|> is genuinely in tokenizer.json's vocab, so dropping it silently
discards a token the weights know about. Re-key it instead, following the
sibling field's own <|x|> -> x_token convention, giving
{"video_token": "<|video|>"}.
The rewrite is logged rather than applied quietly. It is a deviation from the
published artifact, and #1947 rule 5 requires deviations to be recorded -- the
same discipline as the DeepHat TEMPLATE override in #2695.
Generic, not a per-model special case: any repo shipping a list here gets the
same treatment, and a repo that already has a dict is untouched. Verified
against the real file: list ['<|video|>'] -> dict {'video_token': '<|video|>'},
after which the conversion runs and is writing a 50.5 GB f16 across 658 tensors.
Refs #2245, #2985
…etion check The previous generation of chain scripts (chain2b.sh, chain3.sh) polled for the EXTRA_COMPLETE marker, were never committed, and did not survive the rebuild -- taking phases 2.5/3/5 with them (#2985). This replaces that pattern rather than restoring it. It does not wait on EXTRA_COMPLETE, because that marker is written unconditionally at the end of the roster loop no matter how many entries failed, and has now produced a false completion three separate times: on the pull path, on the Tier B path (#2971), and when 12 models died to container DNS in four minutes and the sweep declared itself finished anyway (#3031). Instead it waits on what is actually true or false -- whether every roster entry has both tier files or is explicitly marked UNMEASURED -- and refuses to start while sweep_extra.sh or record_baseline.py is alive. It also declines a suspicious roster: over MAX_UNMEASURED_PCT (default 25%) unmeasured is an infrastructure failure to investigate, not a green light to spend GPU hours on the next phase. That is the guard #3031 asks for, applied at the consumer rather than waiting for the producer to be fixed. Guard verified live: with the sweep running at 38+2/52 it aborts with the counts rather than starting. Refs #1947, #2245, #2985, #3031
This was referenced Sep 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
One of the three scripts lost in the 2026-09-03/04 rebuild because it was never
committed (#2985). Rewritten from #2245, not restored — nothing of the
original survives.
Phase 3 is no longer gated on hardware
It was previously blocked on a RAM upgrade. That gate is dropped: the
measurements say the box already serves models far past the card.
GLM-4.6-REAP-218Bran at 57 GB served, 65%/35% CPU/GPU and scored;
XORTRON.LARGE(123 B) isrunning now at 44 GB, 56%/44%. RAM buys wall-clock, not capability. So the
only real blocker was that this script did not exist.
What it does
Per base, resume-safe at every step:
Re-running after an interruption picks up where it stopped rather than
re-downloading or re-converting — which matters when an f16 of a 35 B base is
70 GB.
The operator chose the full clean f16 ladder over the cheap
--allow-requantizepath. #2245's own comment is why: requantizing analready-quantized GGUF is lossy-on-lossy and yields "a fit-and-cost probe, not a
clean quality datapoint".
It deliberately does not reimplement scoring
sweep_extra.shalready owns the cold-slot protocol, the N=2→3→5 escalation andthe
UNRESOLVEDmarker (#3036). A second copy would produce numbers that cannotbe compared with the as-published rows — rule 6 of #1947's six, and precisely why
#1805-c had to be demoted to a survey. So
BASE/LIST/REPO/PRESEED/MAXTRYare now env-overridable there, and this script builds a tag list and hands it over.
That also makes the ladder safe by construction: locally created tags are
never
ollama pulled, sosweep_extra's ownalready local -> will not deletebranch protects hours of conversion work from its own cleanup step.
Toolchain: one container, no host installs
ghcr.io/ggml-org/llama.cpp:full— verified on the homeserver, carriesconvert_hf_to_gguf.py,gguf-py,llama-quantizeandllama-gguf-split.There is no llama.cpp checkout and no
rex86-evalcontainer on that host, so thelayout
model-quant-benchmark/README.mddocuments does not exist there andshould not be reached for.
Containers run
--user $(id -u):$(id -g)deliberately — a root-owned file in awork tree is exactly how #3024 bricked every CI runner today.
The plan file is grounded, not guessed
f16_ladder_plan.txtnames original safetensors repos, not GGUF repos — thereis nothing for
convert_hf_to_gguf.pyto convert in a GGUF repo. Eachbase_modelwas read off the corresponding GGUF repo's card and verified live:llmfan46/Ornith-1.0-35B-uncensored-hereticllmfan46/gemma-4-26B-A4B-it-ultra-uncensored-hereticEvery level name was checked against
llama-quantize's own type list rather thanassumed.
Levels bracket the measured envelope. 20475 MiB of VRAM minus 3–5 GB of KV at
CONTEXT 32768leaves a ~16–17 GB weights budget, so each ladder crosses theresidency line — one level above it, one at it, one or two below. Whether
parameters bought with bits pay off is unanswerable if the ladder never crosses.
Each row's top level matches a quant already measured as-published, so every
ladder carries a same-base, same-scorer control instead of floating free.
Two arms per #2245 step 1: Ornith-35B dense (published Q4_K_M is 22 GB and
spills — 6.7 min/run against gemma's 2.1 for identical work), and gemma-4-26B-A4B
MoE (PR #2641's promote candidate, fastest in its cohort).
Preconditions asserted, not assumed
Pin is
a99e765or it aborts; plan file, pinned repo, docker, a running ollama,an HF token and the toolchain image are all checked before anything downloads.
bash -nclean on both scripts.Refs #2245, #2985, #1947, #3036