Part of #3079 (round 7). Plan §7, experiments R1–R3. Calibration-set building needs no GPU; llama-imatrix and the scoring queue behind the cold re-run. Extends #2245 — its plain K-quant ladder (phase 3, scoring today) is this issue's control.
What "dynamic requant" means here
Unsloth's Dynamic GGUFs (now 3.0) are not a new format: they are an importance matrix from a curated calibration set + per-tensor bit selection (attention / output / embedding tensors kept higher, expert FFN tensors pushed lower), and the tooling is upstream llama.cpp — llama-imatrix, then llama-quantize --imatrix … --tensor-type <regex>=<type>. Unsloth publishes its imatrix_unsloth.dat per base, which is a usable starting point where the base matches. ollama create --quantize cannot do any of this (q4_K_M / q4_K_S / q8_0 only, no imatrix), which is why quantisation happens in the llama.cpp:full image and Ollama only serves.
The self-quant control already says something: gemma-4-26B-A4B at Q3 / Q4 / Q5 scored 62 / 62 / 62 — plain K-quants do not separate. What is untested is whether an imatrix + mixed-precision quant at the same file size keeps more of the 35 B-class score while fitting the 16–17 GB weights budget at ctx 32768.
R1 — calibration set (no GPU)
/mnt-1/training/calib/: decompiler output (from corpus v1 S3 and S2 — not the 17 test programs), sanitised sessions, REx86 text, and a general-text portion so the imatrix does not overfit to one register; ~2–10 M tokens; sha256 recorded. Decontaminated by round7-3's tool. Compare, per base, against Unsloth's published imatrix where one exists — same quant recipe, two imatrices, is a clean question.
R2 — ladders
| base |
class |
levels to build |
note |
Ornith-1.0-35B (heretic, qwen3_5_moe) |
35 B MoE, top Tier B row, spills at Q4_K_M (22 GB) |
UD-style Q4_K_XL (control vs plain), Q3_K_XL, IQ3_XXS, IQ2_XXS |
the f16 master and plain Q3/IQ3_M rows exist from phase 3 |
Qwen3.8-27B / Qwen3.6-27B heretic / abliterated |
27 B dense, 17–19 GB marginal |
Q4_K_XL, IQ4_XS, Q3_K_XL |
the top dense family in the matrix |
Qwen3.6-35B-A3B abliterated |
35 B MoE, joint top at 64 / 66 from q3_k |
Q3_K_XL, IQ3_XXS |
|
gemma-4-31B QAT / heretic |
31 B dense |
IQ4_XS, Q3_K_XL |
|
XORTRON.CriminalComputing.LARGE 123 B |
RAM-offload class |
IQ2_XXS (the #2245 row that failed on a 429), IQ1_M |
offload measured, not gated; DIMMs make it comfortable, not possible |
GLM-4.6-REAP-218B-A32B |
RAM-offload class, [55, 57, 54] UNRESOLVED at i1-IQ1_S |
UD-IQ1_M, IQ2_XXS with imatrix |
the question is whether a better 1–2-bit recipe resolves the cell |
GLM-4.6-Derestricted-v3 (357 B, 0.36 bpw to fit) stays a measured rejection. Each level: file size, vram_samples.tsv residency and CPU/GPU split, min/run, and both tiers on the round-7 pin. Levels are built from the f16 masters (requant_sweep.sh already does snapshot → f16; extend it with IMATRIX= and TENSOR_TYPES= inputs rather than writing a second driver).
R3 — the trained models' own ladders
Every round7-4/5/6 artefact at Q8_0 / Q6_K / Q4_K_M / imatrix IQ4_XS, so training and quantisation are never confounded — owned here as the recipe, executed by export_to_ollama.sh.
Deliverables
Depends on round7-1 (image pins), round7-3 (decontamination tool). Feeds round7-8.
Part of #3079 (round 7). Plan §7, experiments R1–R3. Calibration-set building needs no GPU;
llama-imatrixand the scoring queue behind the cold re-run. Extends #2245 — its plain K-quant ladder (phase 3, scoring today) is this issue's control.What "dynamic requant" means here
Unsloth's Dynamic GGUFs (now 3.0) are not a new format: they are an importance matrix from a curated calibration set + per-tensor bit selection (attention / output / embedding tensors kept higher, expert FFN tensors pushed lower), and the tooling is upstream llama.cpp —
llama-imatrix, thenllama-quantize --imatrix … --tensor-type <regex>=<type>. Unsloth publishes itsimatrix_unsloth.datper base, which is a usable starting point where the base matches.ollama create --quantizecannot do any of this (q4_K_M / q4_K_S / q8_0 only, no imatrix), which is why quantisation happens in thellama.cpp:fullimage and Ollama only serves.The self-quant control already says something: gemma-4-26B-A4B at Q3 / Q4 / Q5 scored 62 / 62 / 62 — plain K-quants do not separate. What is untested is whether an imatrix + mixed-precision quant at the same file size keeps more of the 35 B-class score while fitting the 16–17 GB weights budget at ctx 32768.
R1 — calibration set (no GPU)
/mnt-1/training/calib/: decompiler output (from corpus v1 S3 and S2 — not the 17 test programs), sanitised sessions, REx86 text, and a general-text portion so the imatrix does not overfit to one register; ~2–10 M tokens; sha256 recorded. Decontaminated by round7-3's tool. Compare, per base, against Unsloth's published imatrix where one exists — same quant recipe, two imatrices, is a clean question.R2 — ladders
Ornith-1.0-35B(heretic,qwen3_5_moe)Qwen3.8-27B/Qwen3.6-27Bheretic / abliteratedQwen3.6-35B-A3Babliteratedq3_kgemma-4-31BQAT / hereticXORTRON.CriminalComputing.LARGE123 B429), IQ1_MGLM-4.6-REAP-218B-A32B[55, 57, 54]UNRESOLVED at i1-IQ1_SGLM-4.6-Derestricted-v3(357 B, 0.36 bpw to fit) stays a measured rejection. Each level: file size,vram_samples.tsvresidency and CPU/GPU split, min/run, and both tiers on the round-7 pin. Levels are built from the f16 masters (requant_sweep.shalready does snapshot → f16; extend it withIMATRIX=andTENSOR_TYPES=inputs rather than writing a second driver).R3 — the trained models' own ladders
Every round7-4/5/6 artefact at Q8_0 / Q6_K / Q4_K_M / imatrix IQ4_XS, so training and quantisation are never confounded — owned here as the recipe, executed by
export_to_ollama.sh.Deliverables
requant_sweep.shextended (imatrix + per-tensor overrides, resume-safe per level), committed beside the phase-3 version with the same operational-copy headerDepends on round7-1 (image pins), round7-3 (decontamination tool). Feeds round7-8.