Skip to content
#

w4a16

Here are 8 public repositories matching this topic...

Language: All
Filter by language

Apple Neural Engine (ANE/NPU) vs GPU for local LLM inference on Apple Silicon / macOS. Core ML (CoreML), Core AI (CoreAI), MLX and Metal benchmarks: prefill, throughput, latency, memory, thermals, INT4/INT8, W4A16/A8W4 quantization, grouped scales and FP16 arithmetic. Reproducible component tests, compatibility findings, English/Chinese articles.

  • Updated Sep 12, 2026
  • Python

Add this topic to your repo

To associate your repository with the w4a16 topic, visit your repo's landing page and select "manage topics."

Learn more