llama.cpp fork with TurboQuant WHT-rotated KV cache & weight compression + Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% throughput).
-
Updated
Aug 6, 2026 - C++
llama.cpp fork with TurboQuant WHT-rotated KV cache & weight compression + Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% throughput).
Qwen3.8-Flash-Next as Cogni-Brain on NVIDIA DGX Spark (GB10): HashK GPU PLE + SGLang NEXTN, 36.8 tok/s code, 100/100 tool-eval, 262K context.
Extract and verify MTP/NextN GGUF speculative draft models
Minimal pinned SGLang recipe for Qwen3.8-Flash-Next on one DGX Spark with exact-FP8 SSD-backed PLE, native 262K context, and NEXTN.
Run 180B-parameter Qwen3.8-Flash-Next on a DGX Spark with full 262K context using SSD-backed memory for ultimate desktop AI performance.
To associate your repository with the nextn topic, visit your repo's landing page and select "manage topics."