add GLM-5.3-Flash (GLM5-Next) support - #27773
Conversation
106ece6 to
9370c82
Compare
|
Hey @timkhronos great work on the PR! A few requests if possible:
Tagging @ngxson for visibility as well. I re-checked and if (1) + (2) is applied, the quants we uploaded work fine (+ the small shard-1 rewrite) and KLD / PPL are correct under this PR. Seems like a simple alias isn't possible actually :( It breaks the quants made with this PR |
|
Hmmm https://github.com/timkhronos/llama.cpp/pull/9/changes would alias the tensors but it looks a bit problematic hmmm |
|
Throwing up some performance numbers here from the lower end of consumer hardware (128GB DDR5 + 24GB VRAM (4090)). avar6 has some freshly converted imatrix quants from this PR up as of now if anyone else wants to give them a go: https://huggingface.co/avar6/GLM-5.3-Flash-BF16-gguf For the IQ3_S, I am getting roughly 300t/s prefill at 256K context and 2048 b/ub size. Generation speed starts off at around 9t/s and drops down considerably by mid window (~128K) to around 6t/s. This seems to track with the 'pooled indexer keys' issue. The model is fully coherent and seems to be working fine. I don't have PPL/KL numbers at the moment as I still need to generate a logit dump. I have noticed an interesting memory quirk, which I haven't seen before. This is the only model I have ever seen have inconsistent checkpoint sizes. As the context fills the checkpoints grow alarmingly fast in size. At ~90K they are already up to nearly 1.6GB. I don't know if this is an expected behavior for this model arch, or if this is a something which needs to be looked into. Also, something of note for you @danielhanchen which I found last night while looking over the three PRs for this arch. The vision towers between this PR and yours differ as well. This PR reuses the name clip.vision.projector_type = "glm4v" while you built a new one clip.vision.projector_type = "glm5next". Likely not much of an issue given how easy it is to regenerate mmproj files, but it will need to be delt with as well. |
|
Yes I'll re-do the vision! This is fine! @timkhronos I confirmed timkhronos#9 works fine and does not break your GGUFs. We will however need to do a cheap shard-1 update so that should be fine |
…ed up long context decode, fla, and slight MTP improvements.
|
@timkhronos I saw you changed the tensor naming - but my solution I provided was to allow everyone's quants to work - now your own ones you uploaded don't work haha. We still need to provide the shard rewrite for the naming (glm5-next) which we're fine with, but now the DeepSeek convention means you yourself have to reupload all shards or do a tensor rename inplace with a script - was this your intention? |
|
@danielhanchen Hey! I ended up going with the the indexer_compressor naming scheme, as it is closer to what's already there, and I was meaning to ask Avar to reconvert anyways, as his ggufs were made when we were missing quantization protection for some crucial tensors so they are not ideal. Your vision projectors will need reconverting though most likely, and your main model ggufs might be missing the |
|
@timkhronos Hey! I made some shard rewrites to https://huggingface.co/unsloth/GLM-5.3-Flash-GGUF/tree/main/Shard_Rewrite for in preparation! |
|
Vision works on 5c4bd50 with UD-Q4_K_XL, Note for anyone hitting "Failed to load CLIP model": the unsloth |
Native ggml-org#27773 metadata (per-layer head_count_kv, kda head dim/gate, indexer kpool/select_tail/index_share_mtp, NoPE MLA, mHC, MoE, swiglu) while keeping the fork direct-quant infrastructure. Adds writer methods, kpool keys, MODEL_TENSORS[GLM5_NEXT], HC/kpool mappings and GLM5V schema keys. Assisted-by: opencode
Add mandatory layer_norm_eps 1e-6, reshape native KDA q/k/v conv1d to the 4D (1, d_inner, 1, d_conv) graph layout in the ordinary path and the same output shape in the direct-quant canonical manifest and direct records via a pure LocalTensor reshape (byte-identical). Sync head_dim cleanup and A_log flatten to n_head. Assisted-by: opencode
…indexer seq_trim Snapshot before replacing the ggml-org#27752-based glm5next port with upstream PR ggml-org#27773. Kept for reference: the last-layer output_norm/output duplicates that keep the trunk graph on the pipeline when the head is pinned to --device-draft, the PIPEDEC_BODY/HEAD graph split, the mtp_only/trunk_only probes for a draft-only GGUF, load_mtp gating for --model-draft, and llama_memory_hybrid_idx::seq_trim. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
Upstream is converging on PR ggml-org#27773 (arch glm5-next) instead; the fork's port of ggml-org#27752 plus the ggml-org#27754 vision graft goes so that PR can be merged as-is. Drops src/models/glm5next.cpp, the converter, the mtmd graft and the glm5next additions to llama-graph / hybrid-idx / model / tests. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
Upstream's GLM-5.3-Flash port replaces the fork's ggml-org#27752-based one. The one real conflict was two refactors of the same mHC helpers: the fork had moved them into llm_graph_context_dsv4_mla for the SPD/DSpark graphs, the PR into a graph_base<Base> template so glm5-next can stack them on the delta-net base. Resolved by templating the fork's class (llm_graph_context_dsv4_mla_t<Base>, alias llm_graph_context_dsv4_mla) and aliasing the PR's llama_model_deepseek4::graph_base to it, so both the fork's fused-kernel / stream-view helpers and the PR's glm5-next graph compile unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
… trim Re-port of the fork's GLM-5.3-Flash hooks onto the ggml-org#27773 graph: - LLM_GRAPH_TYPE_DECODER_PIPEDEC_BODY ends the trunk at the post-norm hidden state with no out_ids input; graph_pipedec_head is the lm_head over the gathered lane rows on the draft GPU. - With --device-draft pinning the output head, output_norm / output are duplicated onto the last transformer layer so the trunk and body graphs keep ending on the pipeline (the RPC-ends-on-local-backend rule). - A draft-only GGUF (blk.<n_layer>.* + token_embd/output/output_norm, no trunk) loads with the trunk tensors optional, so the 7.4 GiB Q8_0 NextN block can be requantized apart from the trunk and pinned to the GTX 1080. - glm5-next joins the stage-2 / tree eligibility lists. - llama_memory_hybrid_idx::seq_trim marks the pooled-key cache stale. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PtrnaBuYDHGDvy73TRkFbG
The ggml-org#27773 port dropped the rollback-plane writes the RS machinery requires: glm5_conv1d wrote only the main conv slot (old graph wrote one snapshot per rollback slot), and the SSM state went through the raw build_delta_net dispatcher plus a direct main-plane write instead of build_recurrent_attn (which emits the K per-token snapshots). After the first partial accept (speculative/DFlash), rollback rounds read never-written planes and the main plane got clobbered with rollback-computed states: drafts degraded (walk top-k -> absurd ids), acceptance collapsed 0.5 -> 0.09 with per-pos decay, output derailed into garbage at ~2-4 t/s. With n_rs_seq == 0 the bug is invisible, which is why target-only stayed clean. Fix: snapshot loop in glm5_conv1d (identical to the slot-0 main write when n_rs_seq == 0) and build_recurrent_attn for the SSM state. Verified: DFlash n_max=7 clean, acceptance 0.46-0.47, 9.2-11.5 t/s; 100-token run stable, matching pre-migration baseline (0.52).
|
I input an 8MP image with |
|
I'm running into an issue involving a soft_max crash on commit 8134115, I had the model do some of its own investigation and it came up with the following, apologies if it's already fixed in the last few commits, or if bug reports investigated by AI are unwelcome, I'm not knowledgeable enough to chase it down myself:
The symptom is a crash at 262k context consistently across all sessions, my current configuration is 3x NVIDIA GeForce RTX 5090 + 2x NVIDIA RTX PRO 6000 Blackwell Workstation, and I was originally chasing it down as a blackwell bug since a few results came up for softmax crashes on blackwell, but my setup doesn't match any of the conditions. Brief log excerpt: 5.16.078.713 I slot print_timing: id 0 | prompt processing, n_tokens = 260621, progress = 0.99, t = 258.59 s / 1007.87 tokens per second Backtrace (from the same crash, according to and written by AI): Launch flags: --ctx-size 420000 --n-gpu-layers 999 --tensor-split 65,21,21,21,70 --flash-attn on --parallel 1 --ubatch-size 2048 --batch-size 2048 Worth noting in case it affects prefill that this configuration gives an error/warning at launch that it couldn't use pipeline parallel, and that it's falling back. Performance seems to be about the same regardless on my setup, so I've been ignoring it in favor of a higher context window for both Deepseek V4 flash 0731 and this model. Tested across different batch/ubatch sizes, no effect on crashes. |
|
Cross-ref: k-pool softmax gridDim.y fix at timkhronos#11 (for the SOFT_MAX crash at n_kv >= 262144). Also see timkhronos#9 (tensor naming), #27752, #27754, #27917. |
Overview
Add support for GLM -5.3-flash a 320B hybrid model, supporting both text and vision.
Additional information
Architecture
GLM 5.3 flash mixes 34 KDA linear layers with 11 DSA laters, with mHC and Deepseek style Moe. Most of the parts are already in llama.cpp so I reused whatever I could:
What I implemented new:
llama_memory_hybrid_dsa: recurrent state + DSA cache, cloned from hybrid ISWA.Rebased onto llama_memory_hybrid_idx instead of the earlier ISWA clone.the encoder is the same family as glmv4 with per head qk-norm, clamped Swiglu and no post conv norm. It reuses glm4v projector with a swiglu_limit key and an optional image token budget.Added as glm5v as GLM 5.3 Flash requires a different pre processing method than what glm4v uses.Tests
Limitations
Quantized GGUFs converted with this PR are available here.
Requirements