Twenty chapters and two capstones. One while loop that becomes vLLM.
Read the diagram → break the simulator → take the quiz → run the code.
Everyone can call an inference API. Almost nobody can explain why vllm serve is
40× faster than the twenty lines of PyTorch they wrote last week. Closing that
gap is what this course is for.
It starts with a while loop that calls a model. It ends with a paged KV cache,
a preempting scheduler, speculative decoding, tensor parallelism and a streaming
OpenAI-compatible server. Each of those arrives only after you have felt the
specific pain that made it necessary.
# Chapter 1. This is the entire engine.
for _ in range(max_tokens):
logits = model(ids) # ← this line is 99% of the cost
ids.append(sample(logits[-1]))
if ids[-1] == eos: breakNineteen chapters later that loop is unrecognisable, and you will be able to point at any line of vLLM and say what it is for.
No GPU. No downloads. No PyTorch. 6,700 lines of NumPy that run on a laptop in about six seconds.
| You ship on top of an inference engine and want to stop guessing which knob to turn. | Chapters S09–S11 and S19 are the ones your on-call pager cares about. |
| You are interviewing for inference / systems roles. | "Why is decode memory-bound?" and "what does PagedAttention actually page?" are S05 and S06. |
| You read the vLLM source once and bounced off it. | Every chapter names where its idea lives in vLLM, SGLang and llama.cpp. |
| You learn by breaking things. | Twenty-two simulators. Push the KV pool until it thrashes. |
Prerequisites: Python, and roughly what a transformer is. Not calculus, not CUDA.
| 📊 Diagrams of the mechanism | Actual data movement rather than boxes and logos, drawn so you can see where the cost accumulates. |
| 🎛 Simulators you can break | Real implementations running in the page. Push the KV pool until it thrashes. Set top-p above the model's confidence and watch it divide by zero. |
| ✅ 138 quiz questions | Six graded per chapter, each with an explanation that names the section to reread. |
| 🐍 Code that runs | 22 self-contained NumPy files. All 22 pass in about six seconds, in CI, on every push. |
| 🌏 English and 中文 | Fully translated — prose, quizzes, simulator tooltips and every label inside every SVG. |
| Layer | Chapters | You learn |
|---|---|---|
| The Machine | S00 | The roofline · arithmetic intensity · HBM vs SRAM · bf16/fp8 · tensor-core tiling · systolic arrays |
| The Model | S01–S04 | Generation loop · BPE tokenizer · RMSNorm/RoPE/GQA/SwiGLU · sampling |
| Memory & KV Cache | S05–S08 | KV cache · PagedAttention · prefix caching · quantization |
| Batching & Scheduling | S09–S12 | Continuous batching · scheduler & preemption · chunked prefill · FlashAttention |
| Decoding Acceleration | S13–S16 | CUDA graphs · speculative decoding · grammar-constrained output · MoE |
| Distributed Serving | S17–S20 | Tensor/pipeline parallelism · prefill/decode disaggregation · the API · the complete engine |
| Capstone I | S21 | picoLM, in Python — a real C engine ported module by module · GGUF & mmap · K-quants · grammar masking |
| Capstone II | S22 | mini-picoLM, in TypeScript — the same engine rebuilt for the browser tab, runs with Node · ArrayBuffer arena · fused quant matmul |
Each layer answers a question the previous one created:
- The machine has two speeds. …which tells you what "fast" even means here. Now write something slow. So:
- The model runs. …which leaves the loop correct and unusably slow. So:
- The cache fixes that. …which makes memory the scarce resource, not FLOPs. So:
- Paging and quantization fix that. …which saturates one GPU with many requests. So:
- Batching and scheduling fix that. …which is as fast as one machine gets. So:
- You go distributed.
Read them in order the first time. The dependencies are real.
git clone https://github.com/xinbetween/learn-llm-inference-from-scratch
cd learn-llm-inference-from-scratch
# --- the code (one dependency) ---
python3 -m venv .venv && source .venv/bin/activate
pip install numpy
python code/s01_generation_loop.py # start here
python code/run_all.py # all 22 chapters, ~6 seconds
# --- the site ---
npm install
npm run dev # http://localhost:3000Two files do more than print:
python code/s19_server.py --serve # a real OpenAI-compatible SSE server
python code/s20_complete_engine.py # full engine + leave-one-out ablationEvery file proves the chapter's thesis rather than describing it. This is the actual CI output, not a summary of it:
✓ s00 batch-1 decode asks 1 FLOP/byte of a machine that wants 295 — 0.3% of peak
✓ s05 cached and uncached generation are token-for-token identical
✓ s06 paged attention == contiguous attention to 1.7e-16, at 2× utilisation
✓ s08 grouping recovers 13 dB on outlier weights; unfused dequant moves MORE bytes than fp16
✓ s09 2.9× faster drain, zero padding — and better latency at the same time
✓ s11 chunked prefill produces a bit-identical KV cache; 25× better tail ITL at 32k
✓ s12 online softmax is exact; paged and split-K variants agree to 1e-15
✓ s14 speculative decoding's empirical distribution matches the target within 0.002
✓ s15 200/200 schema-valid outputs vs 0/200 unconstrained; 62% of decode steps skipped
✓ s17 Megatron sharding is exact; TP8 loses to TP1 on Ethernet at batch 512
22/22 passed in 5.9s
run_all.py only checks exit codes, so the house rule does the real work: every
claim a chapter makes is checkable by the file next to it. A chapter without
assertions is a chapter that isn't finished.
Some results are uncomfortable and shipped anyway — the S20 ablation reports speculative decoding as a net loss at batch 64. That is the answer, not a bug: speculation spends spare compute, and at that concurrency there is none spare. A course that only shows techniques winning is a brochure.
code/ ALL chapter code lives here
s00..s21_*.py one runnable file per chapter + run_all.py
picolm/ S21 capstone — a real Python package
minipicolm/ S22 capstone — the TypeScript engine, imported
by the browser simulator via the @engine/* alias
src/
app/[locale]/ locale-prefixed routes (/en/…, /zh/…)
content/ structure, not text
chapters.ts chapter metadata: num, layer, loc, file
bodies/s00..s22.ts code snippets, shared by every locale
pages.ts non-chapter page structure
components/
diagrams/ one inline SVG per chapter
sim/ one interactive simulator per chapter
i18n/
config.ts LOCALES, href(), Translated<T>
locales/en/ EVERY English string: ui, sim, diagrams,
locales/zh/ chapters, pages, bodies/s00..s22 — copy a
directory to add a language
lib/ seo.ts, jsonld.ts, site.ts, analytics.ts
scripts/brand/ icon + OG card sources, and the script that rasterises them
The site is a Next.js static export — no server, no database. It builds to a folder of HTML you can host anywhere.
Every push to main builds and publishes to GitHub Pages via
.github/workflows/deploy.yml; no build output
is committed. It serves from a custom domain at the origin root, so the deploy
sets NEXT_PUBLIC_SITE_URL and leaves NEXT_PUBLIC_BASE_PATH unset, and
public/CNAME carries the domain into every artifact.
To host it somewhere else, copy .env.example and set NEXT_PUBLIC_SITE_URL:
canonical URLs, hreflang, the sitemap and the Open Graph tags are absolute, and
they all derive from it. Analytics (GA4 and Google Tag Manager) is opt-in through
the same file — unset, the build loads no third-party scripts at all.
A new language is a mechanical task with a compiler-generated checklist. There is no runtime fallback, on purpose: a half-translated page is worse than an honest one, and TypeScript is a better reviewer than a human reading diffs.
To add a language:
Every translatable string lives under src/i18n/locales/<code>/, one directory per
language. English is the reference shape, and every other locale is type-checked
against it.
-
Add its code to
LOCALESinsrc/i18n/config.tsand fill inLOCALE_META. -
Copy
src/i18n/locales/en/tosrc/i18n/locales/<code>/. That directory is the whole job:File What it holds ui.tsSite chrome, nav, quiz UI sim.tsSimulator labels, readouts, tooltips and notes diagrams.tsEvery label baked into the chapter SVGs chapters.tsLayer blurbs, chapter titles, taglines, theses pages.tsHome, layers, timeline, compare, glossary, setup, references bodies/sNN.tsChapter prose and quizzes, one file per chapter -
Name the new bundle in the six assemblers that pair locales together —
src/i18n/{ui,sim,diagrams}.tsandsrc/content/{chapters,pages,bodies/index}.ts— each of which holds a single{ en, zh, … }literal. -
Run
npx tsc --noEmit. Every remaining gap is a compile error, becauseTranslated<T>is exhaustive overLOCALES. -
npm run buildand check both locales.
Code snippets are never translated: they live once in src/content/bodies/sNN.ts
and are shared by every language. Only a snippet's caption and filename label are
translatable, and those sit with the prose.
Conventions. Code identifiers (d_model, top-p, cu_seqlens), notation
(softmax(S), q = round(w / scale) + zero), paper titles and model names stay in
English — they are what the reader will type or search for. Everything else,
including every diagram label and every simulator tooltip, is translated. Prose
uses a small inline markup subset (**bold**, `code`, [label](/s05/)) so
translators never touch JSX.
Issues and PRs welcome. See CONTRIBUTING.md. Particularly useful:
- Corrections. If a claim is wrong, open an issue with the measurement. This is the most valuable contribution there is.
- Translations. See above — the compiler will tell you exactly what's missing.
- Quiz questions. Six per chapter; more good ones are always welcome.
- A chapter this course is missing. Multimodal inputs, LoRA serving, embeddings and encoder-decoder models are all absent and all interesting.
Built on the work of the people who actually solved these problems — the vLLM, SGLang, FlashAttention, Orca, Sarathi-Serve, GPTQ/AWQ and EAGLE authors, among many others. Full list on the references page.
Special thanks to the minimal engines that make these ideas readable: picoLM, quant.cpp, baseRT and llama.cpp.
This is an educational reimplementation and is not affiliated with any of them. Production engines are the real thing; this teaches you how to read them.
If this helped, a ⭐ makes it findable for the next person.