Skip to content

feat(benchmarks): rewrite requant_sweep.sh — the clean f16 quantization ladder (#2245 phase 3) - #3042

Merged
Xore merged 4 commits into
mainfrom
fix/2985-requant-sweep
Sep 5, 2026
Merged

Xore merged 4 commits into
mainfrom
fix/2985-requant-sweep

Conversation

@Xore

@Xore Xore commented Sep 5, 2026

Copy link
Copy Markdown
Owner

One of the three scripts lost in the 2026-09-03/04 rebuild because it was never
committed (#2985). Rewritten from #2245, not restored — nothing of the
original survives.

Phase 3 is no longer gated on hardware

It was previously blocked on a RAM upgrade. That gate is dropped: the
measurements say the box already serves models far past the card. GLM-4.6-REAP-218B
ran at 57 GB served, 65%/35% CPU/GPU and scored; XORTRON.LARGE (123 B) is
running now at 44 GB, 56%/44%. RAM buys wall-clock, not capability. So the
only real blocker was that this script did not exist.

What it does

Per base, resume-safe at every step:

HF snapshot -> convert_hf_to_gguf.py -> f16 master (snapshot deleted at once)
            -> llama-quantize per level -> ollama create per level -> score

Re-running after an interruption picks up where it stopped rather than
re-downloading or re-converting — which matters when an f16 of a 35 B base is
70 GB.

The operator chose the full clean f16 ladder over the cheap
--allow-requantize path. #2245's own comment is why: requantizing an
already-quantized GGUF is lossy-on-lossy and yields "a fit-and-cost probe, not a
clean quality datapoint".

It deliberately does not reimplement scoring

sweep_extra.sh already owns the cold-slot protocol, the N=2→3→5 escalation and
the UNRESOLVED marker (#3036). A second copy would produce numbers that cannot
be compared with the as-published rows — rule 6 of #1947's six, and precisely why
#1805-c had to be demoted to a survey. So BASE/LIST/REPO/PRESEED/MAXTRY
are now env-overridable there, and this script builds a tag list and hands it over.

That also makes the ladder safe by construction: locally created tags are
never ollama pulled, so sweep_extra's own already local -> will not delete
branch protects hours of conversion work from its own cleanup step.

Toolchain: one container, no host installs

ghcr.io/ggml-org/llama.cpp:full — verified on the homeserver, carries
convert_hf_to_gguf.py, gguf-py, llama-quantize and llama-gguf-split.
There is no llama.cpp checkout and no rex86-eval container on that host, so the
layout model-quant-benchmark/README.md documents does not exist there and
should not be reached for.

Containers run --user $(id -u):$(id -g) deliberately — a root-owned file in a
work tree is exactly how #3024 bricked every CI runner today.

The plan file is grounded, not guessed

f16_ladder_plan.txt names original safetensors repos, not GGUF repos — there
is nothing for convert_hf_to_gguf.py to convert in a GGUF repo. Each
base_model was read off the corresponding GGUF repo's card and verified live:

base shards size gated
llmfan46/Ornith-1.0-35B-uncensored-heretic 2 70 GB no
llmfan46/gemma-4-26B-A4B-it-ultra-uncensored-heretic 2 52 GB no

Every level name was checked against llama-quantize's own type list rather than
assumed.

Levels bracket the measured envelope. 20475 MiB of VRAM minus 3–5 GB of KV at
CONTEXT 32768 leaves a ~16–17 GB weights budget, so each ladder crosses the
residency line — one level above it, one at it, one or two below. Whether
parameters bought with bits pay off is unanswerable if the ladder never crosses.
Each row's top level matches a quant already measured as-published, so every
ladder carries a same-base, same-scorer control instead of floating free.

Two arms per #2245 step 1: Ornith-35B dense (published Q4_K_M is 22 GB and
spills — 6.7 min/run against gemma's 2.1 for identical work), and gemma-4-26B-A4B
MoE (PR #2641's promote candidate, fastest in its cohort).

Preconditions asserted, not assumed

Pin is a99e765 or it aborts; plan file, pinned repo, docker, a running ollama,
an HF token and the toolchain image are all checked before anything downloads.
bash -n clean on both scripts.

Refs #2245, #2985, #1947, #3036

Xore added 2 commits September 5, 2026 20:36
…on ladder (#2245 phase 3)

One of the three scripts lost in the 2026-09-03/04 rebuild because it was never
committed (#2985). Rewritten from #2245 rather than restored, since nothing of
the original survives.

## What it does

Per base, resume-safe at every step: HF snapshot -> convert_hf_to_gguf.py -> f16
master (snapshot deleted immediately) -> llama-quantize down each requested level
-> ollama create a tag per level -> score. Re-running after an interruption picks
up where it stopped rather than re-downloading or re-converting, which matters
when an f16 of a 35B base is 70 GB.

The operator chose the full clean f16 ladder over the cheap
--allow-requantize path, so this converts from original weights every time.
#2245's own comment is why: requantizing an already-quantized GGUF is
lossy-on-lossy and gives "a fit-and-cost probe, not a clean quality datapoint".

## It does not reimplement scoring

sweep_extra.sh already owns the cold-slot protocol, the N=2->3->5 escalation and
the UNRESOLVED marker (#3036). Duplicating that would produce a second code path
whose numbers cannot be compared with the as-published rows -- rule 6 of #1947's
six, and the reason #1805-c had to be demoted to a survey. So BASE/LIST/REPO/
PRESEED/MAXTRY in sweep_extra.sh are now env-overridable, and requant_sweep.sh
builds a tag list and hands it over.

That also makes the ladder safe by construction: locally created tags are never
`ollama pull`ed, so sweep_extra's own "already local -> will not delete" branch
protects hours of conversion work from its own cleanup step.

## Toolchain

ghcr.io/ggml-org/llama.cpp:full, verified on the homeserver -- it carries
convert_hf_to_gguf.py, gguf-py, llama-quantize and llama-gguf-split in one
image. There is no llama.cpp checkout and no rex86-eval container on that host,
so model-quant-benchmark/README.md's documented layout does not exist there.
Containers run --user $(id -u) deliberately: a root-owned file in a work tree is
how #3024 bricked every CI runner.

## The plan file is grounded, not guessed

f16_ladder_plan.txt names ORIGINAL safetensors repos, not GGUF repos -- there is
nothing for convert_hf_to_gguf.py to convert in a GGUF repo. Each base_model was
read off the corresponding GGUF repo's card and verified live to exist, be
ungated, and carry config.json plus safetensors (Ornith-35B 70 GB, gemma-26B-A4B
52 GB). Every level name was checked against llama-quantize's own type list.

Levels bracket the measured envelope rather than following habit: 20475 MiB of
VRAM minus 3-5 GB of KV at CONTEXT 32768 leaves a ~16-17 GB weights budget, so
each ladder crosses the residency line -- one level above, one at, one or two
below. The question is whether parameters bought with bits pay off, and that is
unanswerable if the ladder never crosses. Each row's top level matches a quant
already measured as-published, giving every ladder a same-base same-scorer
control.

Refs #2245, #2985, #1947, #3036
The first live run printed five `xargs: unmatched single quote` errors before
reaching the first base. Cause: the line was trimmed with `echo ... | xargs`
*before* the comment check, so xargs saw the apostrophes in the plan file's own
prose comments and errored on each one.

Cosmetic in this run -- the `case` still skipped the comment lines -- but it
would silently mangle any plan value containing a quote, and it hid the real
start of the run behind noise.

Trim with bash parameter expansion instead, and do the comment/blank check
before any processing. Verified against the real plan file: both rows parse to
the expected base/levels/prefix.

Refs #2245, #2985
@github-actions

github-actions Bot commented Sep 5, 2026

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

Xore added 2 commits September 5, 2026 22:30
…nverting

The gemma-4-26B-A4B ladder failed to convert. The error surfaced 20 frames deep
inside transformers with no mention of the model or the field:

    tokenization_utils_base.py, _set_model_specific_special_tokens:
      self.SPECIAL_TOKENS_ATTRIBUTES + list(special_tokens.keys())
    AttributeError: 'list' object has no attribute 'keys'

Cause, read out of the snapshot rather than guessed:
llmfan46/gemma-4-26B-A4B-it-ultra-uncensored-heretic ships

    "extra_special_tokens": ["<|video|>"]

a bare list where transformers requires a dict. The same file's sibling field
is correctly formed --

    "model_specific_special_tokens": {"audio_token": "<|audio|>",
                                      "boi_token": "<|image>", ...}

-- so the intended shape is unambiguous and the list is simply malformed
upstream.

Deleting the field would also have "fixed" the crash, and would have been
wrong: <|video|> is genuinely in tokenizer.json's vocab, so dropping it silently
discards a token the weights know about. Re-key it instead, following the
sibling field's own <|x|> -> x_token convention, giving
{"video_token": "<|video|>"}.

The rewrite is logged rather than applied quietly. It is a deviation from the
published artifact, and #1947 rule 5 requires deviations to be recorded -- the
same discipline as the DeepHat TEMPLATE override in #2695.

Generic, not a per-model special case: any repo shipping a list here gets the
same treatment, and a repo that already has a dict is untouched. Verified
against the real file: list ['<|video|>'] -> dict {'video_token': '<|video|>'},
after which the conversion runs and is writing a 50.5 GB f16 across 658 tensors.

Refs #2245, #2985
…etion check

The previous generation of chain scripts (chain2b.sh, chain3.sh) polled for the
EXTRA_COMPLETE marker, were never committed, and did not survive the rebuild --
taking phases 2.5/3/5 with them (#2985). This replaces that pattern rather than
restoring it.

It does not wait on EXTRA_COMPLETE, because that marker is written
unconditionally at the end of the roster loop no matter how many entries failed,
and has now produced a false completion three separate times: on the pull path,
on the Tier B path (#2971), and when 12 models died to container DNS in four
minutes and the sweep declared itself finished anyway (#3031).

Instead it waits on what is actually true or false -- whether every roster entry
has both tier files or is explicitly marked UNMEASURED -- and refuses to start
while sweep_extra.sh or record_baseline.py is alive.

It also declines a suspicious roster: over MAX_UNMEASURED_PCT (default 25%)
unmeasured is an infrastructure failure to investigate, not a green light to
spend GPU hours on the next phase. That is the guard #3031 asks for, applied at
the consumer rather than waiting for the producer to be fixed.

Guard verified live: with the sweep running at 38+2/52 it aborts with the counts
rather than starting.

Refs #1947, #2245, #2985, #3031
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant