Triton kernels for 4-bit MoE inference: grouped NF4/MXFP4 GEMM, INT4 GEMV, FP8 paged attention, and CPU/NVMe expert streaming.
-
Updated
Sep 11, 2026 - Python
Triton kernels for 4-bit MoE inference: grouped NF4/MXFP4 GEMM, INT4 GEMV, FP8 paged attention, and CPU/NVMe expert streaming.
Apple Neural Engine (ANE/NPU) vs GPU for local LLM inference on Apple Silicon / macOS. Core ML (CoreML), Core AI (CoreAI), MLX and Metal benchmarks: prefill, throughput, latency, memory, thermals, INT4/INT8, W4A16/A8W4 quantization, grouped scales and FP16 arithmetic. Reproducible component tests, compatibility findings, English/Chinese articles.
DeepSeek-R1 7B INT4 at 69.3 tok/s on a $300 RTX 3060. Faster than llama.cpp, vLLM, and NVIDIA TensorRT-LLM. Is one developer + Ai really better than the entire industry?
GLM-5.2 744B at 4-bit on Modal 4x H200 via vLLM, plus a static streaming chat UI.
Qwen3.8-Flash-Next (NVIDIA NVFP4 weights) served as W4A16 on 2x DGX Spark GB10 with vLLM, TP=2. Pinned engine build, 4 start-time patches, measured results, and the levers that were tried and rejected.
Argus-built, evidence-first Qwen2.5-0.5B-Instruct-AWQ W4A16 RTL with a verified 24-layer cascade and authenticated Hybrid RTL runtime.
Argus-built, evidence-first Qwen2.5-0.5B-Instruct-AWQ W4A16 RTL with a verified 24-layer cascade and authenticated Hybrid RTL runtime.
Tests worst-environment scale selection for multi-environment W4A16 quantization.
To associate your repository with the w4a16 topic, visit your repo's landing page and select "manage topics."