Skip to content

round7-7: dynamic requant ladders — imatrix + per-tensor bit selection for the 27–35B class and the 100B+ offload class, against #2245's plain ladder #3086

Description

@Xore

Part of #3079 (round 7). Plan §7, experiments R1–R3. Calibration-set building needs no GPU; llama-imatrix and the scoring queue behind the cold re-run. Extends #2245 — its plain K-quant ladder (phase 3, scoring today) is this issue's control.

What "dynamic requant" means here

Unsloth's Dynamic GGUFs (now 3.0) are not a new format: they are an importance matrix from a curated calibration set + per-tensor bit selection (attention / output / embedding tensors kept higher, expert FFN tensors pushed lower), and the tooling is upstream llama.cpp — llama-imatrix, then llama-quantize --imatrix … --tensor-type <regex>=<type>. Unsloth publishes its imatrix_unsloth.dat per base, which is a usable starting point where the base matches. ollama create --quantize cannot do any of this (q4_K_M / q4_K_S / q8_0 only, no imatrix), which is why quantisation happens in the llama.cpp:full image and Ollama only serves.

The self-quant control already says something: gemma-4-26B-A4B at Q3 / Q4 / Q5 scored 62 / 62 / 62 — plain K-quants do not separate. What is untested is whether an imatrix + mixed-precision quant at the same file size keeps more of the 35 B-class score while fitting the 16–17 GB weights budget at ctx 32768.

R1 — calibration set (no GPU)

/mnt-1/training/calib/: decompiler output (from corpus v1 S3 and S2 — not the 17 test programs), sanitised sessions, REx86 text, and a general-text portion so the imatrix does not overfit to one register; ~2–10 M tokens; sha256 recorded. Decontaminated by round7-3's tool. Compare, per base, against Unsloth's published imatrix where one exists — same quant recipe, two imatrices, is a clean question.

R2 — ladders

base class levels to build note
Ornith-1.0-35B (heretic, qwen3_5_moe) 35 B MoE, top Tier B row, spills at Q4_K_M (22 GB) UD-style Q4_K_XL (control vs plain), Q3_K_XL, IQ3_XXS, IQ2_XXS the f16 master and plain Q3/IQ3_M rows exist from phase 3
Qwen3.8-27B / Qwen3.6-27B heretic / abliterated 27 B dense, 17–19 GB marginal Q4_K_XL, IQ4_XS, Q3_K_XL the top dense family in the matrix
Qwen3.6-35B-A3B abliterated 35 B MoE, joint top at 64 / 66 from q3_k Q3_K_XL, IQ3_XXS
gemma-4-31B QAT / heretic 31 B dense IQ4_XS, Q3_K_XL
XORTRON.CriminalComputing.LARGE 123 B RAM-offload class IQ2_XXS (the #2245 row that failed on a 429), IQ1_M offload measured, not gated; DIMMs make it comfortable, not possible
GLM-4.6-REAP-218B-A32B RAM-offload class, [55, 57, 54] UNRESOLVED at i1-IQ1_S UD-IQ1_M, IQ2_XXS with imatrix the question is whether a better 1–2-bit recipe resolves the cell

GLM-4.6-Derestricted-v3 (357 B, 0.36 bpw to fit) stays a measured rejection. Each level: file size, vram_samples.tsv residency and CPU/GPU split, min/run, and both tiers on the round-7 pin. Levels are built from the f16 masters (requant_sweep.sh already does snapshot → f16; extend it with IMATRIX= and TENSOR_TYPES= inputs rather than writing a second driver).

R3 — the trained models' own ladders

Every round7-4/5/6 artefact at Q8_0 / Q6_K / Q4_K_M / imatrix IQ4_XS, so training and quantisation are never confounded — owned here as the recipe, executed by export_to_ollama.sh.

Deliverables

Depends on round7-1 (image pins), round7-3 (decontamination tool). Feeds round7-8.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

analysisPayload analysis pipelineenhancementNew feature or requestllmLLM analysis workermlML worker and GPU scoring

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions