freetoken
Here are 6 public repositories matching this topic...
Unofficial FreeToken fork: on one RTX 3060 12 GB, a 35B MoE at 250k of context or gpt-oss-120b; Flash-Next 125B on two. Half the RAM, image input. Runs on Turing: RTX 2060, RTX 20 series, sm_75.
-
Updated
Sep 13, 2026 - Python
Forked from FlashML-org/FreeToken, with AMD Strix Halo (gfx1151) ROCm support and kernel tuning
-
Updated
Aug 28, 2026 - Python
Reproducible Apple-Silicon (Apple MLX) deployment and one-token verification kit for the FreeToken edge-native MoE serving API — pinned, loopback-only, schema-validated evidence.
-
Updated
Aug 29, 2026 - Python
Reproducible FreeToken MoE expert-offload compatibility lab and benchmarks for NVIDIA Turing sm_75.
-
Updated
Aug 24, 2026 - Python
Top-down browser factory automation game built with a local Qwen coding model. Mine resources, automate production, power your outpost, and craft Engines in the Relay Seven world.
-
Updated
Sep 9, 2026 - TypeScript
Add this topic to your repo
To associate your repository with the freetoken topic, visit your repo's landing page and select "manage topics."