A local inference engine for Apple silicon, built around the model.
Splash runs coding agents and OpenAI or Anthropic compatible applications on one Mac. It combines DFlash 2 speculative decoding, specialized Metal kernels, and automatic memory planning, with vision, tool calling, and a built-in chat page. It reuses cached prefixes and batches concurrent requests automatically.
Apple M3 or newer, macOS 26.4 or later, and Homebrew. The 4-bit examples need at least 36 GB of unified memory (48 GB recommended); 24 GB Macs can use smaller GGUF variants.
brew install incoai/tap/splash
splash serve --model unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_MThe first run downloads the model and its matching draft and starts serving
on 127.0.0.1:8000. Later starts reuse the downloads; leave room on disk for
them (storage requirements).
Once it prints Ready, leave this terminal open. Open http://127.0.0.1:8000
in your browser, or run an installed coding agent from another terminal:
splash opencode # or: splash claude / splash codex / splash hermes / splash piFor a server started with another --port, pass the same --port to the agent.
Press Ctrl+C in the server terminal to stop Splash. For LM Studio Bionic, follow its Splash setup guide.
OpenAI Chat Completions, Responses and Completions, and Anthropic Messages, with streaming, tool calls, JSON Schema output, and base64 images and PDFs:
curl http://127.0.0.1:8000/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{
"model": "unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M",
"messages": [{"role": "user", "content": "Explain speculative decoding in one sentence."}]
}'Reasoning follows the model default; "reasoning_effort": "none" turns it off
in Chat. Anthropic Messages requests reason only when they set thinking.
Reasoning settings ·
Tool calls ·
API details
Splash supports these model families, with a matching DFlash2 draft selected automatically:
| Model | GGUF example | MLX 4-bit |
|---|---|---|
| Qwen3.8-27B | unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M |
mlx-community/Qwen3.8-27B-4bit |
| Qwen3.6-35B-A3B | unsloth/Qwen3.6-35B-A3B-GGUF:UD-Q4_K_M |
mlx-community/Qwen3.6-35B-A3B-4bit |
Unsloth GGUF variants span 1–8 bits, including mixed-precision UD formats;
UD-Q8_K_XL and BF16 targets are not supported.
Prism ML Ternary Bonsai 2
is also supported in PQ2_0 (7.2 GB), including vision. Pass OWNER/REPO:VARIANT
to --model, as in the quick start. MLX targets must be affine 4-bit with
groups of 64, like the examples. Smaller variants run on
24 GB Macs.
27B variants ·
35B variants
Vision and the tokenizer come from the target model's source. Model loading and compatibility · Supported formats
Memory and context are sized automatically, up to the model's native context
window. To set your own limits or cache options, add these to splash serve:
| Option | Purpose |
|---|---|
--max-memory 28G |
Cap Metal memory use. |
--idle-release off |
Keep the model in memory while idle (default: release after 10m). |
--max-context 100K |
Set the context limit. |
--language-only |
Skip vision; serve text only. |
--kv-format bf16 |
Use BF16 KV cache. Default: 8-bit (INT8). |
--disable-ane |
Prefill on the GPU alone. Default: the 27B also uses the Neural Engine. |
--max-cache-disk 16G |
Keep cached prompts on SSD when memory runs short. Off by default. |
--persistent-cache |
Keep the SSD cache across restarts; needs --max-cache-disk. Off by default. |
Use --max-memory to leave room for other applications.
The server listens on localhost without authentication by default. For LAN
access, authentication, browser apps on other origins, and other options, see
server configuration or
splash serve --help.
KV precision ·
SSD cache
Measured with Splash 1.0 on an M5 Pro (16-core GPU, 48 GB), using its model packages and selected SPEED-Bench coding prompts over HTTP. Ratios compare with the next-fastest engine measured in that benchmark.
| Metric | Qwen3.6-35B-A3B | Qwen3.8-27B |
|---|---|---|
| Decode · short prompt | 210 tok/s (1.7×) | 74 tok/s (2.0×) |
| Prefill · 32K prompt | 2,011 tok/s (1.3×) | 363 tok/s (1.2×) |
| Cached time to first token · 32K replay | 123 ms (6.6×) | 282 ms (7.3×) |
| Aggregate decode · 4 concurrent short prompts | 357 tok/s (2.0×) | 170 tok/s (3.9×) |
Launch benchmarks · Measurement details · Run benchmarks locally
Same Unsloth UD-Q4_K_M weights on Metal. Decode speed in tok/s:
| Model | Engine | M5 Pro, 20-core GPU | M3 Max, 40-core GPU |
|---|---|---|---|
| 27B | llama.cpp | 16 | 17 |
| llama.cpp with MTP | 27 | 20 | |
| Splash | 74 | 92 | |
| 35B-A3B | llama.cpp | 69 | 66 |
| Splash | 175 | 209 |
That is 2.5–3.2× as fast on the 35B and 4.5–5.3× on the 27B (2.7–4.6× against MTP).
Closely matches llama.cpp's predictions.
| Next-token agreement ↑ | 27B | 35B-A3B |
|---|---|---|
| llama.cpp: CPU vs. GPU | 97.8% | 96.5–96.9% |
| llama.cpp: single-token vs. batched | 99.65–99.75% | 97.95% |
| Splash vs. llama.cpp | 99.30–99.45% | 97.83–98.14% |
Splash uses BF16 KV and --disable-ane in this comparison.
Benchmark details
Each supported model pairs a trained DFlash2 draft with Metal kernels for its shapes. The runtime, scheduler, cache, and API are shared. Weights are converted to the kernels' layouts as they load, with no copy on disk; kernels ship precompiled, with no Xcode or local tuning required. How Splash works
- Development: configuration, API, internals, building, and testing.
- Issues and feedback
- Apache-2.0. GGUF kernels include MIT-licensed material from llama.cpp; see third-party notices. Model weights keep their own licenses.