Run large >120B MOE LLM on a distributed network of Mac Mini/Studio's using expert parallelism with SSD streaming to reduce RAM requirements.
- Single nodes in a 4x Mac Mini cluster exceed decode tok/s over the sister project TinyTitan by 10-15% so the new engine build from scratch is performing better than expected.
- The network stack is working and performing well over raw TCP on 1GBit LAN.
- Cluster tok/s is still below a single node - the core work in this project.
MIT — see LICENSE. Copyright (c) 2026 André Borchert.
Questions, bug reports and suggestions are always welcome. You can contact André Borchert by email at 0xa0b1@gmail.com.
