Exact Mixture-of-Experts inference in C, on the computer you already own.
Trochilus runs Mixture-of-Experts language models with no dependencies: libc and the
operating system's threads, nothing else. One binary picks its kernels at runtime for whatever CPU
it lands on, reads the experts from disk when they do not fit in RAM, and puts part of the work on
an NVIDIA GPU when one is there — without a CUDA toolkit. And it gives the same tokens as
transformers, checked in every build.
Speed may change; the result may not. Threads, SIMD tier, batch size, expert budget, speculative decoding, Windows or Linux, CPU or GPU: each one changes how fast the answer comes, never a bit of it. Every one of them is a test in the gate.
- Runs OLMoE-1B-7B and any GGUF v3 file of its family, straight from Hugging Face: F32, F16, Q8_0, Q4_K and Q6_K weights (so the Q4_K_M files people download).
- CPU kernels for scalar, AVX2, AVX-512 and AVX-512 with VNNI and VBMI, chosen at runtime, each bit-identical to the scalar definition, with a test that the tier you think ran is the one that ran.
- The decode's attention on an NVIDIA GPU, loaded at runtime (hand-written PTX through the driver) with the CPU's bits: the logits of 600 positions and the tokens after a 4000-token prompt are the CPU's byte for byte. Without a GPU nothing changes.
- Experts from disk under a RAM budget: a slot store with an O(1) LRU, unbuffered reads, the prompt read layer by layer, the next layer's experts read while this one computes.
- Speculative decoding (
--spec): drafts taken from the text already in the context, several tokens verified in one pass that loads each weight row from memory once for all of them (a pass of three tokens costs 1.7 times a pass of one). run,chat(the model's own chat template, byte for byteapply_chat_template),generate, andserve, which keeps the model loaded between commands.- A byte-level BPE tokenizer built from the GGUF metadata, token for token Hugging Face
tokenizers, at 22.0 MB/s.
OLMoE-1B-7B, the same GGUF for both engines, both in the same Linux container on one laptop (Ryzen 9 7940HX, 16 cores, 31 GB), a free machine, the night of 2026-09-26 to 27. Tokens per second, the median of 10 runs: two series of 5 for each engine, alternated with the other's.
| llama.cpp | Trochilus | Trochilus / llama.cpp | |
|---|---|---|---|
| Q4_K, prompt of 512 tokens, 4 threads | 220.6 | 220.0 | 1.00× |
| Q4_K, generation after that prompt, 4 threads | 50.9 | 52.4 | 1.03× |
| Q8_0, prompt of 512 tokens, 16 threads | 370.7 | 492.9 | 1.33× |
| Q8_0, prompt of 2048 tokens, 16 threads | 341.3 | 485.2 | 1.42× |
| Q8_0, generation at 512 tokens of context, 8 threads | 33.1 | 33.1 | 1.00× |
| Q8_0, generation at 2048 tokens of context, 8 threads | 27.2 | 26.3 | 0.96× |
The difference is in what is computed: llama.cpp rounds the activations to 8 bits before each matrix
product and keeps its attention cache in 16 bits; Trochilus computes what transformers computes, the
activations in float (for Q4_K, exact integers) and the cache in 32 bits. Every number, the conditions,
and the optimizations that did not pay are in docs/MEASUREMENTS.md.
Under a RAM budget the model is not read twice: with a partial expert budget a 2048-token prompt runs 1.89x faster reading 6 273 MiB instead of 22 880, and 1.22x more with the disk reading the next layer while the cores compute this one.
Every tier must give the scalar definition's bits, so the speed has to come from how the work is arranged, not from rounding. Four ideas do it, each checked in the gate:
Sixteen rows in sixteen lanes. The scalar definition adds a row's products in sixteen interleaved chains and joins them with a fixed tree. Rather than spread one row across a register's sixteen lanes, which must then be joined at the end of every row, Trochilus gives each lane its own weight row: every chain sums in the scalar order, the input is broadcast to sixteen rows at once, and nothing is decoded, scaled or joined inside the loop. The prompt's matrix-product tile runs at 164.7 GFLOP/s on one core, 98.6% of what the CPU can do in float without fused multiply-adds (which would round once where the scalar definition rounds twice), and every output is the scalar definition's, bit for bit.
Q4_K in exact integers, each input converted once. A Q4_K matrix product takes each block of 256 inputs as 32-bit fixed point and sums exact integers, then rounds once: the input's 32 bits are the only approximation, where a float dot product rounds at every step. The order of the sums no longer matters, so any thread count or SIMD width gives the same bits, as fast as the float kernels this replaced. Each input row is converted once and shared: q, k and v read the same converted rows, gate and up read a token's row through a map instead of eight gathered copies, and the SwiGLU converts the down projection's input while it is still in the core's cache.
A byte that can be deduced is not read. In OLMoE's GGUF files the router matrices are 32-bit floats whose low 16 bits are all zero: the model was converted from bf16. The load checks every value, keeps only the top halves, and each row widens them back with a shift: the 32-bit row's bits, from half the bytes, every token. A matrix with one value that does not fit stays in 32 bits.
A correctly rounded exp, proved on every float. Our expf is a 64-entry table and a polynomial,
with the eight hard cases computed at 200 bits; the gate checks it on all 4 278 190 082 float
arguments that are not NaN, against the exact value, in every SIMD tier. (glibc's rounds 0.004% of
them otherwise; ggml's 3.36%, by up to 2 units in the last place.)
Exactness is an invariant, not a hope. Windows and Linux give the same logits byte for byte
on the real model; tiny models built for the purpose are compared with transformers in every gate,
and so is the real model cut to two layers. Speculative decoding cannot change the answer: every
verified row is bit for bit the computation of a single-token pass.
Nothing to configure. At startup the engine measures the cores, the instructions, the free RAM and the disk, and places the work itself; while it runs it times how many threads each kind of pass wants (the decode, and each size of verify pass on its own). It refuses to load if it would leave the machine with less than 2 GB or 10% of its RAM.
git clone https://github.com/namespaceMarcello/trochilus.git
cd trochilus
make # build/trochilus: no dependencies
build/trochilus cpu # what the engine sees of this machine
build/trochilus run -m OLMoE-1B-7B-0125-Instruct-Q4_K_M.gguf -f prompt.txt -n 200
build/trochilus chat -m OLMoE-1B-7B-0125-Instruct-Q4_K_M.ggufC11 with intrinsics and a Makefile: gcc, clang or MinGW-w64, on Linux and Windows. make check is
the gate (lint, a 0-warning build, the tests, ASan, TSan, the oracles). The Python tools under
tools/ are for conversion and for the oracles only; the engine never needs them.
Pre-alpha, under active development, measured on one machine.
| Milestone | Content | State |
|---|---|---|
| M0 | the exact engine: GGUF, CPU kernels for every tier, the OLMoE graph, tokenizer, chat | done |
| M1 | experts from disk under a RAM budget, with no options to set | the store, the budget and the overlapped reads done; the first prompt's cost next |
| M2 | smaller weights: Q4_K, Q6_K, then Q2_K and IQ2 | Q4_K and Q6_K done, Q4_K_M runs |
| M3 | CUDA: the attention, the dense weights, the experts in VRAM | the decode's attention done |
| M4 | a model of hundreds of gigabytes on the same laptop | — |
| M5 | KV checkpoints on disk, an HTTP server, a draft model | speculation from the context, serve |
| M6 | Vulkan and Metal, so the GPU is not only NVIDIA | — |
Not there yet: one model family and one chat template; the GPU does only the decode's attention, on NVIDIA only; no 2-bit formats; no NEON kernels (ARM takes the scalar path); greedy decoding only, and no HTTP API; a model larger than RAM has not been run yet.
Trochilus is written by Claude Code (Anthropic's Opus) as the coding agent, with Marcello Costagliola
leading — the direction, the questions, the reviews and every decision; every commit says so. One
piece of the engine at a time: first read how colibri,
ds4 and llama.cpp solve it,
measure theirs against ours, then build and measure again. A comparison alternates its two sides run by
run and carries an A/A control, the prediction is written before the run, and every mistake becomes an
automatic check (a test, a lint rule, a mutation that must turn red). The engineering log is in docs/:
| Document | Content |
|---|---|
ARCHITECTURE.md |
principles, layers, the correctness ladder, milestones |
STATUS.md |
where the project stands, the decisions, the next step |
MEASUREMENTS.md |
every measurement, including the optimizations that were rejected |
LESSONS.md |
every mistake and discovery, with the check that now prevents it |
ORIGINS.md |
where each idea and file comes from; every piece against the three references |
COMMANDS.md |
benchmarks, measurements, mutations, reports |
Trochilus is written from scratch, on what four other engines taught. They are pinned at a commit in
ref/ and read as primary sources, and every idea taken from one of them is recorded in
docs/ORIGINS.md with the file and function it came from.
| Project | Some of what we learned from it |
|---|---|
| colibri — Apache-2.0 | threads on physical cores; weights read with pread; the pretokenizer regex replayed over codepoints; oracles on tiny generated models; drafts from the prompt's own text; one expert in one slot |
| ds4 — MIT | a persistent thread pool instead of OpenMP; the vocabulary from GGUF metadata; (token, expert) pairs sorted by a counting sort; one weight row against several tokens in registers; the GGUF type table |
| llama.cpp / ggml — MIT | prompts in passes of 512 tokens; the K-quant block layouts; pinning a thread to a processor; undoing a rejected draft; the engine we race |
| ik_llama.cpp — MIT | its K-quant CPU kernels, read before we wrote ours |
transformers and Hugging Face tokenizers define what a correct result is here, and the
first model is OLMoE-1B-7B, from AI2. Only two files carry code from elsewhere — the GGUF type
table in src/format/gguf.c and tools/make_tiny_olmoe.py — and each names its origin in its header.
Apache-2.0 — see LICENSE, and NOTICE for third-party material.
