Windows-native BitNet and ternary LLM inference with CPU GGUF, GPU runtime, terminal and browser chat, and release zips.
-
Updated
Mar 20, 2026 - Python
Windows-native BitNet and ternary LLM inference with CPU GGUF, GPU runtime, terminal and browser chat, and release zips.
Run BitNet 1.58-bit and ternary LLMs on Windows with CPU and GPU inference, chat tools, and release-ready builds
Reproducible artifact for the preprint 'On-Chain LLM Inference Under Instruction Budgets' (Aerni, Fluck, Becker 2026) — DOI 10.5281/zenodo.20607598
Add a description, image, and links to the ternary-llm topic page so that developers can more easily learn about it.
To associate your repository with the ternary-llm topic, visit your repo's landing page and select "manage topics."