Skip to content

Latest commit

 

History

338 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Little Gemma

A small, from-scratch C program on CUDA that loads a Gemma 4 model from GGUF and runs it — written to teach how a modern LLM actually executes, in the spirit of Karpathy's llama2.c, but covering a current model end to end: parse GGUF → BPE tokenize → run the transformer → generate text. Every stage validated bit-for-bit against llama.cpp.

text ──► tokenizer ──► token ids ──► forward ──► logits ──► argmax ──► next token
                          ▲                                                │
                          ╰──────────────── append, repeat ◄───────────────╯

~9,700 lines of C/CUDA for the shipped int8-CUDA build, no vendored dependencies — 4,908 CPU · 8,282 f32-CUDA · 9,746 int8-CUDA, counted per binary as the include closure the compiler actually sees (the three backends are mutually exclusive, so no single program is their sum). That 9,746 includes multi-turn socket serving, image and audio understanding, a GPU vision encoder, tensor-core flash-attention prefill, a ring-buffered f16 KV cache, and byte-identical speculative decoding (MTP).

Performance vs llama.cpp

Same day, same GGUFs, same machines; little-gemma measured on its serving path, llama.cpp with llama-bench (best of -fa 0/1, decode at matched context depth). Full tables, methodology, and history: docs/benchmarks.md.

Generation with MTP on (tokens/s) — the shipped configuration, and the number that matters in use; output is byte-identical to plain greedy decoding, always:

device model plain +MTP chat +MTP structured
Jetson Orin NX E4B QAT 20.7 29.9 48.6
Jetson Orin NX 12B QAT 9.8 14.5 20.4
RTX A5000 E4B QAT 134 168.8 282
RTX A5000 12B QAT 70.7 101.5 143.7

The gain is content-dependent — Orin E4B by turn type: prose 29.9, image description 31.5, code 40.7, fully predictable output 48.6 (57.8 at block-4) — and ahead of llama-server's own draft-mtp at every point measured (prose 24.5, code 34.1, image 28.8).

Best-vs-best MTP (experimental; Orin NX QAT, one prose turn, greedy, tokens/s, each config at its optimal block depth). little-gemma runs on its serve path, llama.cpp via --spec-type draft-mtp with the same head — a like-for-like engine comparison. Block depth is LG_MTP_N (a runtime knob) / llama's --spec-draft-n-max, shown at its optimum, which climbs with model size as acceptance headroom grows. Output stays byte-identical to plain greedy at every depth:

model plain little-gemma full head llama.cpp full head little-gemma reduced-vocab
E2B QAT 34.5 49.2 (N=2) 49.0 (n=2) 54.1 (N=3)
E4B QAT 20.7 31.9 (N=4) 30.4 (n=3) 34.0 (N=4)
12B QAT 9.8 16.4 (N=4) 1 17.8 (N=4)

little-gemma's full head is level-to-ahead of llama's on the identical head; its reduced-vocab head — the tied draft head trimmed from 262,144 rows to 16,384, which llama.cpp has no equivalent for — is fastest everywhere, +7–10% over our own full head. Content-type breakdowns and the trim study are in docs/mtp-vocab-trim.md.

Decode (tokens/s, batch 1, speculation off) — ahead on the Jetson, the device this project targets, and at parity on desktop:

device model little-gemma llama.cpp ratio
Jetson Orin NX E4B QAT q4_0 20.7 18.7 1.11×
Jetson Orin NX 12B QAT q4_0 9.8 9.05 1.08×
Jetson Orin NX E2B QAT q4_0 34.5 37.4 0.92×
RTX A5000 E4B / 12B / E2B 134 / 70.7 / 213.9 136.6 / 70.7 / 209.3 0.98× / 1.00× / 1.02×

Prefill (929-token prompts) — 0.8× llama.cpp, consistently:

device model little-gemma llama.cpp ratio
Jetson Orin NX E4B / 12B / E2B 474 / 193 / 834 553 / 232 / 1,020 0.82–0.86×
RTX A5000 E4B / 12B / E2B 4,335 / 2,067 / 7,222 5,254 / 2,365 / 8,785 0.82–0.87×

The RTX A5000 rows across all three tables were measured on a separate Windows workstation. They will be re-based to the RTX PRO 4500 (Blackwell) dev box once identical E4B/E2B GGUFs are in place there; until then they are left un-updated.

The pattern is the project's thesis: decode speed is mostly everything around the matmul — launch overhead, syncs, norms, the PLE path, how the KV walk is split across the GPU — which a few thousand readable lines can do leanly. Prefill runs through llama.cpp's home turf (arch-tuned tensor-core GEMMs); the 2026-07 campaign closed it from ~0.2× to 0.8× and measured the rest to its structural floor. On media turns, time-to-first-token inverts in our favor (1.5–2.4×) — GPU encoder plus arrival-overlapped prefill; see docs/benchmarks.md.

Build

cmake -S . -B build
cmake --build build --config Release

CPU build (run) needs only a C compiler; OpenMP is auto-detected. If the CUDA toolkit is found, CMake also builds run-cuda (readable f32 matmul) and run-cuda-i8 (int8 + tensor cores — the fast one):

cmake --build build --config Release --target run-cuda-i8

All three implement the same model.h; only the compute kernels differ.

Run

run-cuda-i8 -m gemma-4-E4B-it-qat-UD-Q4_K_XL.gguf -p "The capital of France is"

The default E4B and 12B are unsloth's QAT q4_0 builds (with their matched MTP heads, mtp-gemma-4-{E4B,12B}-it.gguf) — QAT-trained for q4_0, and faster than Q4_K_M on both stacks.

Serve conversations over a Unix-domain socket (multi-turn KV cache, raw token stream out — details in docs/serving.md):

run-cuda-i8 -m model.gguf -s /tmp/lg.sock          # server (Ctrl-C to stop)
echo "What is the capital of France?" | nc -N -U /tmp/lg.sock
run -c /tmp/lg.sock                                # or the bundled client

Options:

  • -mm mmproj.gguf — image/audio input over the socket, via mmcat (docs/multimodal.md).
  • -mtp assistant.gguf — speculative decoding, byte-identical output (docs/mtp.md).
  • -sys file — prefill a system turn once at server start.
  • -think N — cap the reasoning channel: 0 off (structural — prompt control of thinking is inert on Gemma 4), N up to N tokens, omitted unlimited (docs/serving.md).
  • -temp/-topk/-topp/-seed — sample instead of greedy.

On Windows the same code serves %TEMP%\lg.sock and the build ships its own socket clients.

Documentation

License

MIT — see LICENSE.

Footnotes

  1. llama.cpp's draft-mtp degenerates into a control-token loop on this 12B QAT GGUF in the build tested (its tokenizer marks some control tokens as normal type); little-gemma decodes the same model + head cleanly.

About

No description, website, or topics provided.

Resources

Stars

18 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages