A course that teaches inference engineering from the ground up, in Rust. You start with what a model is and how numbers are stored. You finish with a from-scratch engine that runs a real pretrained language model (SmolLM2) and serves it over HTTP with streaming, continuous batching, a paged KV cache, quantized weights and speculative decoding. Every step is explained in plain language, with complete code and measurements.
Who it is for: programmers who know basic Rust and want to understand how models are actually run in production, and why engines like vLLM, llama.cpp and TensorRT-LLM are built the way they are. You do not need a machine learning background. Rust basics are not taught, but every Rust feature that matters for inference (borrowing to avoid copies, lifetimes on memory-mapped weights, unsafe SIMD, Send/Sync across threads, async serving) is explained where it first appears.
What "from scratch" means here: the numeric kernels, tensor layout, weight-file parser, tokenizer, transformer, KV cache, sampler, quantization, batching scheduler and paged cache are all written in this repository in plain Rust. Third-party crates are used only for things that are not the point of the lesson (JSON parsing, memory-mapping a file, a thread pool, an async runtime and an HTTP server), and each one is introduced with a reason.
Each chapter is a folder with a lesson (README.md) and a Rust crate you can run and test. Read them in order: each chapter uses what the previous ones built.
| # | Chapter | You will build |
|---|---|---|
| 1 | What inference is | A toy model and a harness that measures latency percentiles, batching throughput and the cost of copying weights |
| 2 | Numbers inside a model | f16 and bf16 from scratch, with exhaustive tests; precision and accumulation experiments |
| 3 | Tensors, strides and views | An owned Tensor and borrowed zero-copy TensorViews: transpose, slice, broadcast without copying |
| 4 | Memory is the bottleneck | Bandwidth and latency probes for every cache level; the roofline model |
| # | Chapter | You will build |
|---|---|---|
| 5 | Matrix multiplication | Matmul from naive to cache-blocked, measured in GFLOP/s |
| 6 | SIMD | Dot products with AVX2/FMA and NEON intrinsics, runtime CPU detection |
| 7 | Threads | Parallel matvec and matmul with scoped threads and rayon; false sharing |
| 8 | Neural network operators | Stable softmax, RMSNorm, LayerNorm, SiLU, GELU, tested against f64 references |
| # | Chapter | You will build |
|---|---|---|
| 9 | Weights on disk | A safetensors parser and writer, zero-copy loading with mmap and lifetimes |
| 10 | A first model, end to end | Train, export, load and serve a small classifier; batching and shared models |
| # | Chapter | You will build |
|---|---|---|
| 11 | Tokenization | Byte-level BPE training and encoding; streaming UTF-8 decoding |
| 12 | Attention | Scaled dot-product attention, causal masking, multi-head, GQA and RoPE |
| 13 | The transformer | A complete Llama-style model; parameter and FLOP counting |
| 14 | The KV cache | Prefill and decode with a KV cache, proven equal to full recomputation |
| 15 | Sampling | Temperature, top-k, top-p, min-p, penalties, and a seeded RNG |
| 16 | Running a real model | SmolLM2-135M in bf16 with its real tokenizer, checked against PyTorch |
| # | Chapter | You will build |
|---|---|---|
| 17 | Measuring and profiling | Sampling and per-matrix profiles of the real model, A/B tests on a noisy machine, a tiled prefill kernel |
| 18 | Quantization I: int8 | Symmetric int8 weights and activations, integer dot products, perplexity checks |
| 19 | Quantization II: 4-bit | Block-wise 4-bit weights with packed nibbles; quality versus size |
| 20 | FlashAttention | Online softmax, tiled attention and split-KV decoding |
| # | Chapter | You will build |
|---|---|---|
| 21 | The engine thread | An engine on its own thread, streaming tokens over channels, cancellation by drop |
| 22 | An HTTP inference server | An OpenAI-style API with server-sent events, backpressure and graceful shutdown |
| 23 | Continuous batching | An iteration-level batch scheduler; throughput versus batch size on SmolLM2 |
| 24 | Paged KV cache | A block allocator, block tables, paged attention and prefix caching |
| 25 | Scheduling and SLOs | Load generators, TTFT/TPOT percentiles, chunked prefill and admission control |
| # | Chapter | You will build |
|---|---|---|
| 26 | Speculative decoding | Draft-and-verify with greedy and rejection sampling; measured on real models |
| 27 | Structured output | Grammar-constrained decoding that always produces valid JSON |
| 28 | More than one device | Tensor, pipeline and expert parallelism, simulated with threads and all-reduce |
| 29 | GPUs | How GPUs run inference; CUDA kernels written in Rust with NVIDIA's cuda-oxide |
| 30 | Inference in production | Capacity planning, layer-by-layer numerical debugging, metrics and reliability |
A glossary collects every term the course defines.
Every lesson follows the same order, so you always know where to look:
- In one sentence, where it fits, what you need, what you will build.
- The intuition: a physical analogy, and where the analogy stops being true.
- Vocabulary: every new term, defined before it is used.
- The concepts in depth.
- The code: the chapter's source, excerpted and explained line by line. The complete files are in the chapter's
src/folder. - Run it: the exact command and the output it produced on the reference machine, followed by what the numbers mean.
- The Rust behind it: the Rust features the code relies on, and why they matter for inference.
- Mistakes you will make.
- How the professionals do it: how production engines handle the same problem.
- Exercises, check yourself questions, a recap, and answers.
Install Rust with rustup. The course is tested with Rust 1.94.
git clone <this repository>
cd rust_inference_engineering
cargo test # every chapter's tests
cargo run --release -p ch01-what-is-inference # run one chapterAlways pass --release when you time anything. Debug builds are 10-50x slower on numeric code.
Chapters 16 onward run SmolLM2-135M-Instruct, a 135-million-parameter chat model released by Hugging Face under the Apache 2.0 license. Download it once (about 270 MB):
./tools/download_model.sh # SmolLM2-135M-Instruct into models/
./tools/download_model.sh 360M # the larger sibling, used as a target in chapter 26Tests that need the model skip themselves when it is absent, so cargo test always works. To keep models somewhere else, set INFER_MODELS to the directory that contains the smollm2-* folders.
Chapter 16 checks our Rust model against the official PyTorch implementation. The small reference outputs are committed in chapters/16-real-model/fixtures/, so you do not need Python. To regenerate them, see tools/README.md.
Every number in the lessons was measured, not estimated, on this machine:
- Intel Xeon cloud VM (Sapphire Rapids generation, 2.1 GHz), 4 vCPUs, 15 GB RAM, Linux
- Supports AVX2, FMA and AVX-512. L2 cache 2 MB per core
- Rust 1.94.1, release builds, no
-C target-cpu=nativeunless a lesson says otherwise
A shared cloud VM is noisy: expect run-to-run variation of 5-10%, sometimes more. Your machine will give different numbers. The lessons are about the shape of the results (what gets faster, by roughly how much, and why), and each lesson explains how to read your own numbers.
Some code cannot run on this machine: GPU code in chapter 29 and multi-machine communication in chapter 28. Those chapters say clearly which parts were verified and how, and never present unverified output as real.
Every chapter must pass four checks before it goes into the course:
cargo build --all-targets
cargo test
cargo clippy --all-targets -- -D warnings
cargo fmt --checkClippy runs with the pedantic group enabled. A few pedantic lints are switched off for the whole workspace in Cargo.toml, because numeric code triggers them constantly and deliberately. The most common are the cast_* lints (usize as f32 and similar): we accept those casts and explain the ranges involved where it matters. Everywhere else, when we do something Clippy objects to, we write #[expect(lint, reason = "...")] next to it. expect rather than allow, so the build fails if the exception stops being needed.
A fifth check keeps the lessons honest:
cargo xtask checkEvery code excerpt in a lesson is marked with the file it came from, and this command verifies that each excerpt still appears in that file, and that every relative link in the Markdown resolves. If the code changes and the lesson does not, the check fails. CI runs all five checks on every push (.github/workflows/ci.yml).
The SIMD code has separate paths for x86-64 (AVX2) and ARM (NEON, for Apple Silicon and ARM servers). The ARM paths were tested under emulation on the reference machine:
cargo test --target aarch64-unknown-linux-gnu -p ch06-simd # needs qemu-user and an aarch64 linkerCode and text are dual-licensed under MIT or Apache-2.0, at your option. SmolLM2 weights are not included in this repository and are licensed separately by Hugging Face (Apache 2.0).