A from-scratch, single-binary local model runner in Rust — import any model, quantize it to fit, and run it on all the hardware you have: CPU, one GPU, or several. No Ollama, no cloud, no CUDA toolchain.
Mummu imports models from HuggingFace (or disk), (auto-)quantizes them to fit, and runs them natively in Rust on Burn across every device you have. One binary — a runtime probe inventories your GPUs (Vulkan / DX12 / Metal via wgpu) and your CPU (burn-flex) and places the model to use them to the fullest: several GPUs together, with CPU offload when VRAM is short, no feature-split builds and no CUDA toolchain. Models are reimplemented from scratch, generic over the Burn backend, and parity-tested byte-for-byte against a reference so the reimplementations can be trusted.
It exists because two local-first apps — laurelane (a private budgeting cockpit) and Nanna (an always-on local AI presence) — were building the same runner twice. Mummu is that runner, extracted and generalized: each app consumes it as a dependency and keeps only its own domain glue. Laurelane proved the blueprint (Qwen2.5 / LFM2.5 / all-MiniLM ported to Burn, byte-identical parity vs Candle, validated on an RTX 4070 Ti SUPER 16 GB); Mummu is where it lives, hardens, and grows.
- One binary, every device — compile both
Wgpu(Vulkan/DX12/Metal, no CUDA toolchain) andburn-flex(CPU); a runtime probe enumerates all adapters + the CPU and places the model across them — a single GPU, several GPUs together, or GPU + CPU hybrid. No feature-split builds, no per-vendor path. - Models from scratch, generic over
B: Backend— a growing zoo (Qwen2/2.5, Qwen3 dense with per-head q/k norm + decoupled head_dim, LFM2/2.5 hybrid conv+attention, OLMoE sparse mixture-of-experts, all-MiniLM embedder) built on shared blocks (RmsNorm · GQA · RoPE · SwiGLU · top-k-routed expert bank · tied lm-head · depthwise causal conv), with a clean trait to add more. - Trustworthy reimplementations — every port must pass a parity gate: single-forward top-k logits and a short greedy sequence match a reference (Candle, or a local Ollama of the same model) exactly.
- Fast — per-layer KV cache (+ conv-state cache for hybrids), on-GPU argmax (sync only the winning index), sampling, token streaming, cooperative cancellation; kernel
fusion+autotune; an f16 path (f32 attention-score island for numeric safety) that halves VRAM at full speed. - A full model-import suite — pull a model from HuggingFace (by repo id) or from disk and load it: safetensors, PyTorch state dicts, and GGUF (llama.cpp, dequantized) weights;
config.json-driven hyperparameters; tokenizer + chat-template import (HFtokenizers/ SentencePiece / BPE); per-architecture weight-name remapping with a checked load (fail loudly on a key mismatch, never silently zero-init); resumable, shard-aware downloads into a per-user cache; and a declarative model registry so adding a model is a manifest entry, not new code. - Quantize to fit, fill the hardware — a planner probes every GPU + the CPU (VRAM / RAM), then imports or quantizes on the fly (GGUF K-quants, GPTQ / AWQ, or Burn's own int8/int4) and chooses precision + layer placement so the largest model that fits runs and every device is used — sharded across GPUs, spilling cold layers to CPU when needed. Plus a model-management API (download progress, disk usage, remove) apps surface in their settings UI.
- Local embeddings — a from-scratch MiniLM-class sentence embedder (CPU) for fully-offline semantic search.
[dependencies]
mummu = "0.4"Load a checkpoint and decode. The library is async end to end — a decode that waits on the device yields its worker instead of parking it:
use std::path::Path;
use mummu::models::CausalLm;
use mummu::models::qwen2;
async fn answer(prompt: &[u32]) -> Result<Vec<u32>, Box<dyn std::error::Error>> {
// The runtime probe picks the GPU when there is one, the CPU otherwise.
let device = if mummu::backend::use_gpu() {
mummu::backend::gpu_device()
} else {
mummu::backend::cpu_device()
};
// A directory holding config.json / tokenizer.json / model.safetensors.
let dir = Path::new("path/to/qwen2.5-1.5b-instruct");
let model = qwen2::load_from_dir(dir, &device)?;
// `prompt` is token ids from the checkpoint's own tokenizer.
Ok(model.greedy_generate(prompt, 32, &device).await?)
}mummu::cache::ModelSlot holds one checkpoint across calls and serializes device access, which is
what keeps a 16 GB card from trying to hold two multi-GB loads at once. GGUF and PyTorch
checkpoints load through mummu::gguf and mummu::safetensors; mummu::registry turns a model
into a manifest entry rather than new code.
The two schedulers are usable on their own and depend on nothing:
mummu-schedule divides work across heterogeneous devices to minimize
makespan, mummu-mix places per-tensor precision under a byte budget.
For a server rather than a library, mummu-serve exposes /api/chat, /api/models and a web chat
UI; it is built from this repo rather than installed from crates.io.
- Workspace + backends — six crates:
mummu(the library),mummu-mixandmummu-schedule(the two dependency-free schedulers),mummu-serve(HTTP server + chat UI),mummu-app(Tauri desktop shell) andmummu-bench(criterion). One binary compiles bothWgpu(withfusion+autotune) andburn-flex(CPU), with a cached runtime GPU probe and a device inventory that records per-adapter/per-APISHADER_F16, max buffer size, and true VRAM capacity (NVML, loaded at runtime rather than linked, so its absence degrades toNoneand never to a link error; wgpu exposes no portable query), plus the host CPU's cores and total RAM — the planner's (and settings UIs') device set. - Shared blocks, generic over
B: Backend— cache-aware GQA attention (optional per-head q/k RMSNorm), manual RoPE, SwiGLU, and LFM2's double-gated causal short-conv with rolling decode state; unit tests prove prefill+decode ≡ full-forward for both cache kinds. - Checked safetensors + PyTorch import — bf16→backend-float cast adapter, per-architecture key
remaps, and a fail-loud load (never silently zero-init);
config.json-driven hyperparameters.pytorch_model.binstate dicts load through the same checked path (safetensors preferred when both exist) — proven byte-identical on MiniLM's real Hub checkpoint in both formats. tokenizer_config.jsonimport —mummu::tok_config::TokenizerConfigparses the conventions HF keeps besidetokenizer.json: the BOS/EOS/PAD/UNK special-token slots (id-resolved fromadded_tokens_decoder), the whole added-token map,model_max_length, and the raw Jinjachat_template(from the JSON key, or a standalone siblingchat_template.jinjawhen that key is absent — the layout recenttransformerswrites). Total and bounded (malformed input is a loudImportError::Parse, never a panic); it doesn't render Jinja (prompt wrapping stays the byte-verifiedchatrenderers) but gives apps a model's declared ids + template, and detects the template's tool-call convention (Hermes vs LFM) so the right render style is picked from the checkpoint. Two consistency validators catch repackaging bugs:check_ids_against(every added-token id must match the real tokenizer) andcheck_eos_agrees(config.json'seos_token_idmust match the resolved EOS). Cross-checked on real weights: all 26 of Qwen3-0.6B's added-token ids agree byte-for-byte withtokenizer.json, and itsconfig.jsonEOS 151645 agrees with the resolved<|im_end|>. The safetensors loaders enforce all of this at load:load_from_dir(Qwen2/Qwen3/LFM2.5) runs the gate right afterconfig.jsonparses and before any weights are read — a siblingtokenizer_config.jsonwhose EOS disagrees withconfig.json, whose chat-template speaks a different tool-call convention than the family's renderer, or — when atokenizer.jsonsits beside it — whose declared added-token ids don't match that real tokenizer, is a loudImportError::Inconsistentinstead of a model that silently mis-stops, mis-templates, or mis-tokenizes. Both sibling files are optional (a GGUF-derived dir has neither → no behavior change). On a successful safetensors load the parsedTokenizerConfigis surfaced on the returnedLoaded{Qwen2,Qwen3,Lfm2}struct (tokenizer_config), so a consumer reads config-driven EOS/BOS/PAD ids straight off the model; a GGUF load surfacesNone(self-contained).- Fallback chat renderer for un-ported models (optional feature
jinja-template) —mummu::template::ImportedTemplaterenders a checkpoint's own importedchat_templateJinja, so a model whose family has no hardcoded renderer is still promptable from the authority on its prompt format: its own template. The selection rule ships as a value —Renderer::for_checkpoint(family, dir)takes a byte-verified family renderer when one exists and falls back to the template otherwise, never second-guessing the family renderer. Bounded and fail-loud (no template →Absent, bad Jinja →Jinja, a runaway render →TooLargeat 8 MiB), with the config's BOS/EOS/PAD/UNK reaching the render context and assistant tool calls passed structurally so the template writes its own call markers instead of inheriting Hermes'. Verified against the from-scratch path on real checkpoints (tests/imported_render.rs): byte-identical toChatMl::qwen3()on plain (142 B), tools (748 B) and full FC history (324 B), and toChatMl::lfm2()on plain (157 B) and tools (379 B) — the LFM leg also proving the standalonechat_template.jinjafallback and thebos_tokeninjection. The feature is off by default: the zoo's from-scratch renderers cover it byte-for-byte, and a default build carries no Jinja engine. - Import validation — a two-stage error taxonomy:
ImportErrorfor the file→module stage (missing file, parse, load, and anIncompleteper-tensor missing/errored diff) andSanityErrorfor the runtime liveness a checked load can't see — NaN/Inf logits, a vocab-width mismatch, or a degenerate/dead forward.CausalLm::sanity_checkis the post-installgate an app calls to catch a silently-broken import before trusting the model. - Precision is a property of the device, not of a type —
backend::float_dtype(&device)reads the element type back from the device burn 0.22 keeps it on, andbackend::gpu_device_f16()returns a GPU device configured to compute in f16. Device dtype settings lock on first use, sogpu_device_f16is idempotent within a process and returns an error rather than an f32 device under an f16 name — the one failure mode that turns a benchmark into fiction. One process can hold an f16 accelerator beside an f32 host, which the oldGpu/GpuF16type aliases could not express at all. - Unsupported attention shapes are refused, not approximated — every loader parses the
rope_scaling/rope_parametersobject (both spellings, plus the pre-4.38typekey) and thesliding_windowfields, fromconfig.jsonand the GGUF header, and fails the load naming the mode when it is not plain rotary + full causal attention. A YaRN-scaled or windowed checkpoint would otherwise load clean and degrade only far out in the context, where short prompts never look. Presence is not enablement: Qwen2.5 ships an inertsliding_window: 32768behinduse_sliding_window: falseand keeps loading, as does a window spanning the whole trained context. - Five architectures ported and running on real weights — Qwen2/2.5, Qwen3 dense, the LFM2/2.5
hybrid, OLMoE (sparse MoE), and the all-MiniLM sentence embedder; Qwen2.5-1.5B, Qwen3-0.6B, and
LFM2.5-1.2B/230M load and greedy-decode correctly on the reference GPU (wgpu/Vulkan), and
Qwen2.5-0.5B / LFM2.5-230M / OLMoE-1B-7B do the same on the CPU backend. Qwen3 reuses the shared
blocks whole (its per-head q/k RMSNorm, absent qkv bias, and decoupled
head_dimwere all already supported), loads from safetensors and a single Q4_K_M GGUF, and handles Qwen3's<think>reasoning mode. - Mixture-of-experts —
nn::SparseMoeis a softmax top-k router over a fused expert bank ([experts, out, in]— the row-major twin of GGUF'sffn_*_exps, so 64 experts load as three tensors per layer, not 192);models::olmoedrives it config-first offolmoe.*GGUF metadata. OLMoE-1B-7B (64 experts, top-8 routing, 1B active / 7B total) runs from the one official 4.21 GB Q4_K_M file — 16 layers × 64 experts loaded and "2 + 2 equals 4." greedy-decoded at 1.15 s/token on the CPU backend (~28 GB f32 resident; a 16 GB card waits on keep-quantized VRAM, tracked in P9). Attention learned OLMoE's whole-projection q/k RMSNorm placement, inferred from the loaded norm's own width, so every existing checkpoint loads byte-unchanged. It also loads from the HF safetensors release, where the 64 experts are stored as separate tensors across three shards: the importer fusesexperts.{0..63}.{gate,up,down}_projinto the same[64, 1024, 2048]banks, in numeric (not lexicographic) member order, validating group completeness before reading a payload byte. Proven on the real 13.84 GB bf16 checkpoint — 16 layers checked-loaded in 136.1 s, and layer 5 / expert 37'sgate_projread straight from the raw shard bytes is bit-identical to slot 37 of the fused bank across all 2 097 152 values. The fuse streams to a temp file rather than RAM, so a checkpoint this size costs the model's footprint, not the model plus a second copy of itself. The GGUF path streams the same way: its dequant plans the whole output before reading a payload byte, then writes header + f32 tensors straight to a temp file thatburn-storemmaps back, so loading the 1B-7B costs 26.5 GB of measured private commit — the model alone — where the old in-RAM dequant paid for the payload twice on top of it. - All three models are parity-verified — the two-leg P7 gate passes for Qwen2.5-1.5B on the
reference GPU: single-forward top-5 logits match a Candle f32 reference (max |Δlogit| 2.7e-5,
tests/parity_qwen2.rs+ the committedtools/candle-probefixture) and a 24-token greedy sequence matchesollama qwen2.5:1.5b-instruct-fp16byte-for-byte. LFM2.5-1.2B passes both legs against a same-weights llama.cpp reference (tests/parity_lfm2.rs+ thetests/llama_refharness: a localllama-serveron LiquidAI's official BF16 GGUF, raw/completion, prompts as token-id arrays): top-5 first-forward ids match exactly in order and a 24-token greedy sequence is byte-identical. LFM2.5-230M passes the same two legs through the same tier-parameterized gate (top-5 ids exact in order, greedy byte-identical, max |Δlogprob| 3.2e-2) — one config-driven hybrid loader covers both tiers. Qwen3 is parity-verified through the same llama.cpp harness (tests/parity_gguf.rs,qwen3leg): on Qwen3-0.6B Q4_K_M, top-5 first-forward ids match exactly in order and a 24-token greedy sequence —<think>reasoning tokens included — is byte-identical tollama-serveron the same file. OLMoE-1B-7B passes the same llama.cpp gate on its own Q4_K_M (olmoeleg): top-5 ids exact in order, 24-token greedy byte-identical, max |Δlogprob| 3.7e-1 — so the MoE router and expert bank are verified against a reference, not just plausible. The MiniLM embedder matches its Candle reference at cosine 0.99999994 (max |Δcomponent| 1.2e-7,tests/real_minilm.rs). The f16 path is parity-verified too (tests/parity_f16.rs, its own binary becauseGpuF16locks Burn's per-device dtype policy): the same llama.cpp comparison with our side loaded ontoGpuF16passes for Qwen2.5-1.5B and Qwen3-0.6B — top-5 ids exact in order, 24-token greedy byte-identical, max |Δlogprob| 2.5e-1 / 3.9e-1, below the f32 legs' own 2.7e-1 / 4.0e-1. So half precision is a verified path, not merely a live one. - Sampling, streaming, cancellation — temperature / top-k / top-p sampling (deterministic per seed),
per-token streaming through a
ControlFlowcallback, and cooperative between-token cancellation; greedy decoding keeps the argmax on-device. - Function calling (both zoo conventions) — advertise
ToolSpecs throughrender_with_toolsin the convention the model was trained on: Hermes for Qwen2.5/Qwen3 (the exact# Tools/<tool_call>JSON template, results as merged<tool_response>turns) and LFM for LFM2.5 (bare tool JSON on aList of tools:system line, Pythonic calls in<|tool_call_start|>tokens, results as realtoolturns, past-turn</think>stripping); both parsers are bounded with a loud error taxonomy. Proven end-to-end on the real GPU: Qwen2.5-1.5B emitted a parseable Hermes call, LFM2.5-1.2B emitted exactly<|tool_call_start|>[get_weather(city="Paris")]<|tool_call_end|>, and Qwen3-0.6B — from aChatMl::qwen3()prompt selected by its own imported template's convention — emitted a<think>block plus a parseable Hermes call (tests/real_toolcall.rs,tests/real_toolcall_lfm.rs,tests/real_toolcall_qwen3.rs). - Template byte gate — the hardcoded renderers are proven byte-identical to
transformers.apply_chat_templaterendering the checkpoint's own importedchat_template(via thehf-chat-templatedev-dependency): plain, multi-turn, the full Hermes# Toolsblock, and function-call history match byte-for-byte on Qwen3-0.6B, Qwen2.5-1.5B, and LFM2.5-1.2B (LFM's legs cover both tool conventions, itstoolrole turns, its history think-stripping, and the standalonechat_template.jinjaimport path),tests/template_gate.rs.ChatMl::qwen3()carries Qwen3's two template deltas exactly —<think>reasoning stripped from assistant turns at/before the last user query (kept and re-normalized for later turns mid tool loop), and no default system preamble with tools — all three byte-equal against the imported template. The one remaining family divergence is pinned to its exact delta so any other drift fails loudly: Qwen2.5's no-system branding preamble vsqwen2()'s neutral one. Prompt JSON deliberately serializes with Pythonjson.dumpsspacing and insertion-order keys (serde_jsonpreserve_order) — the exact bytes the reference stack renders and models emit back. - f16 inference, validated — Qwen2.5-1.5B runs coherently on
GpuF16(weights + KV in f16, the q·kᵀ attention scores + softmax computed in an f32 island to stop f16 overflow): ~3.6 GiB runner VRAM vs ~7.9 GiB f32, and ~2.3× the decode throughput (27.1 vs 61.8 ms/token, measured 2026-08-09); the parity gate re-passes unchanged on f32, where the island casts are no-ops (bench/BASELINE.md). The same island covers the Qwen3 arch — Qwen3-0.6B decodes coherently in f16 (its qk-norm + decoupled head_dim ride the same f32 scores). - Warm-up API — a freshly-started process decodes its first ~32 tokens at roughly a third of its
steady rate (per-process kernel compilation + pipeline creation; CubeCL persists autotune across
processes but the wgpu runtime caches no compiled kernels).
CausalLm::warm_up(probe_ids, steps, device)pays that cost off the user's critical path — one prefill plusstepsgreedy decode steps on a throwaway cache, bounded and synchronized. Measured on the reference GPU: after a 4.2 s warm-up a cold process's first 32-token burst runs at 41.9 tok/s vs the next burst's 41.0 (un-warmed, that ratio is 0.33×) —mummu-bench/tests/warmup_api_f16.rs, curve in bench/BASELINE.md. - In-process mixed precision is defined behavior — every runtime tensor-creation site pins its
dtype to the backend type (
backend::{float_dtype, int_dtype}), so an f32 (Gpu) and an f16 (GpuF16) model can share one process and one device regardless of Burn's first-touch-locked per-device dtype policy — proven on the real GPU in the historically failing order (tests/real_mixed_dtype.rs: f16 locks the policy first, the f32 model still forwards f32 logits, agrees on the greedy top token, and decodes coherently). - SPIR-V kernels on Vulkan — CubeCL compiles direct SPIR-V (burn's
vulkanfeature) instead of WGSL/naga on Vulkan adapters, worth +30% decode throughput on the reference GPU with parity byte-identical; other APIs (DX12/Metal) transparently keep WGSL in the same binary. - Benchmarked — Qwen2.5-1.5B on the reference GPU, criterion: f32 TTFT 98.2 ms, decode
16.2 tok/s, prefill@2048 597 ms (11.5 GiB whole-card peak ≈ 8.0 GiB runner); f16 TTFT 24.9 ms,
decode 36.8 tok/s, prefill@2048 241 ms (~3.6 GiB runner) — recorded with budgets in
bench/BASELINE.md, enforced by opt-in regression gates
(
mummu-bench/tests/budget{,_f16,_cpu,_moe}.rs, one dtype alias per process). - Precision selection —
mummu::plan::pick_precisionanswers "which dtype fits this card?" from numbers the crate already has: a model'sconfig.jsonshape on one side,backend::inventory()'s per-adapter VRAM andSHADER_F16on the other. It returns the highest precision that fits (f32 before f16) with the projected and usable byte counts behind the decision,Nonewhen no float tier fits — the honest "this needs quantization or several devices", never a silently-worse tier — and never plans f16 on an adapter that doesn't advertise it. Its overhead and headroom constants are calibrated against bench/BASELINE.md, and its tests pin the decisions to real hardware: the 15.7 GiB reference card takes Qwen2.5-1.5B in f32, an 8 GiB card in f16, a 64k context forces a 12 GiB card down to f16, and OLMoE-1B-7B on 16 GiB reports no fit. - Autotune-cache control — CubeCL benchmarks each kernel once and persists the winner to disk, so
later processes start warm; but the cache has no invalidation, so a pick made while the machine was
busy is believed forever (measured 2026-08-09: a tune taken during a contended moment cost 21–27% of
f16 decode in every subsequent process, silently).
mummu::tuneis the repair a settings UI needs:autotune_cache_dir()reports where the picks live — read out of the very config CubeCL discovers, not a hardcoded copy of the rule —autotune_cache_report()measures it, andclear_autotune_cache()removes it so the next launch re-tunes. Bounded and fail-loud (an implausibly deep or wide tree is an error, never a wide delete). Proven on the real GPU (tests/real_autotune_cache.rs). - Model management —
ModelManagergives settings UIs the whole lifecycle over a declarative model catalog (registry::ModelSpec): install with per-chunk download progress,is_installed, per-model disk usage, and traversal-safe removal; model switching ridesModelSlot. - GGUF import, end to end —
mummu::ggufparses the llama.cpp container (typed, bounded metadata; fully validated tensor table) and dequantizes every storage dtype (F32/F16/BF16, the legacy Q4_0/Q4_1/Q5_0/Q5_1/Q8_0 blocks, and the Q2_K–Q6_K superblocks) to f32;qwen2::load_from_ggufturns the one file into a running model — hyperparameters from the GGUF metadata, weights bridged through the same checked-load pipeline as safetensors (tied and untied lm-heads). Proven against the model's true weights (tests/real_gguf.rs): F32 norms bit-exact vs the bf16 safetensors of the same checkpoint, Q4_K rows at cosine 0.9975, and the real Qwen2.5-1.5B Q4_K_M file greedy-decodes "2+2 equals 4." on the GPU with first-token top-1 identical to the bf16 build (logit cosine 0.977). The tokenizer comes from the GGUF too (tokenizer_from_gguf: per-family pre-regex → ByteLevel → BPE, byte-identical ids vs the checkpoint'stokenizer.json) — one .gguf file is the whole model. Works for the LFM2.5 hybrid as well (lfm2::load_from_gguf: layer kinds from the per-layer kv-head array, conv kernels un-squeezed bit-exactly): the official LiquidAI Q4_K_M greedy-decodes "2 + 2 equals 4." with top-1 identical to bf16 (logit cosine 0.991). Next: keep-quantized VRAM (tracked in P9). (2026-09-22) Ternary checkpoints in a folded Hadamard basis — the Prism ML ternary family (Q1_0/Q2_0upstream ids 41/42, Prism-privatePQ2_0/PTQ1_0ids 142/143: every value a per-128 f16 scale times −1/0/+1) dequantizes, andnn::hadamardimplements theprism.hadamard.*contract Ternary-Bonsai 2 declares — signs then a normalized blockwise Sylvester–Walsh–Hadamard transform on every folded projection's input, the inverse after the token gather, the tiled→grouped value-head permutation for the DeltaNet out-projection — soternary-bonsai-2-27b-pq2_0(Qwen3.8-27B at 2.13 bits/weight, 7.2 GB) is a catalog entry that serves through the same qwen35 path as the UD-Q4 27B. A checkpoint whose contract this runtime cannot honour is refused by name, never run unrotated; a table whose declared dtypes overlap (Prism's legacy group-128 file stored under the group-64 id) is refused at parse; and a folded pack is never FFN- partitioned (the down-projection transform is blockwise over the original neuron order). Verified against the Prism llama.cpp fork on the real 27B: the layer-0 DeltaNet block agrees to three digits at the pack's Q8 level. - SentencePiece
tokenizer.modelimport, both proto types —tokenizer_from_spmbuilds the HF pipeline straight from the SPM proto the Llama/Gemma/T5 families ship (a bounded hand-rolled protobuf reader, zero new dependencies). Unigram protos (T5/ALBERT/Gemma) get thePrecompiledcharsmap + whitespace-collapse normalizers, Metaspace, and a Unigram model; BPE protos (Llama-2 family) get their merge list reconstructed from vocab + scores (HF'sSentencePieceExtractoralgorithm) plus thePrepend/Replacenormalizers andByteFallback/Fuse/Stripdecode chain. The proto's specials are re-added and id-verified in both. Proven byte-identical to the same checkpoints' shippedtokenizer.json— ids and decode round-trips — on flan-t5-small (Unigram) and TinyLlama-1.1B (BPE, 61k reconstructed merges, byte-fallback cases included) across a unicode/whitespace/CJK/emoji battery (tests/real_spm.rs). - Hub downloads — streaming HuggingFace fetches into the model cache: resumable (
.part+ HTTP Range, proven byte-identical after an interrupted transfer), length-verified, shard-index aware, with a per-chunk progress callback; verified end-to-end by downloading all-MiniLM and embedding with it. - Process-lifetime model cache —
ModelSlotloads a checkpoint once per process, switches models by key, andclear()s to free VRAM; Burn'sParamisn'tSync, so access serializes behind its mutex.
- Local-first, offline, private — your own hardware is the whole story; the cloud is never a dependency.
- Use all the hardware — inventory every GPU and the CPU and run the model across them to the fullest (multi-GPU + CPU offload); quantize to fit the VRAM you actually have. Great on a laptop CPU, better on one GPU, best on several — same binary.
- Burn at the core — Burn is the one inference and training engine. Every model, every backend, and the (future) on-device fine-tune loop is built from scratch on Burn (via CubeCL) as a single backend-agnostic codebase (CPU / CUDA / Metal / Vulkan / WebGPU). No second runtime, no per-backend forks, no C/CUDA toolchain — Burn is the foundation the whole runner stands on.
- Parity or it didn't happen — a reimplementation ships only when it is numerically byte-identical to a reference.
- Performance is a gate — a change lands only when the parity + perf budgets (TTFT, decode tok/s, VRAM ceiling) hold; README perf claims link a benchmark artifact.
- README + ROADMAP are the only docs — shipped capability is described here; everything planned or next lives as a
[ ]in ROADMAP.md; git history + PRs are the record.
- Nanna — the runner is the agent: local inference for the whole agent loop, plus local embeddings for its memory. Wires Mummu in as
Provider::Local. - laurelane — on-device statement structuring + categorization; Mummu replaces its in-app Burn modules.
Reference GPU: RTX 4070 Ti SUPER 16 GB. The plan lives in ROADMAP.md.
MIT.
Unless you explicitly state otherwise, any contribution intentionally submitted for inclusion in this crate by you shall be licensed as above, without any additional terms or conditions.