Tenant-fair LLM inference orchestration on a single GPU. No Kubernetes.
-
Updated
Sep 4, 2026 - Python
Tenant-fair LLM inference orchestration on a single GPU. No Kubernetes.
Prefill/Decode-disaggregated LLM serving library in Go — KV-cache connectors + SLO-aware elastic GPU flipping; published benchmark suite (colocated TPOT p99 inflates 133x under burst)
Real-time AI memory orchestrator for multi-model GPU serving. PERC: provably optimal KV cache eviction via fractional knapsack. 25-79% eviction cost reduction vs LRU.
To associate your repository with the gpu-serving topic, visit your repo's landing page and select "manage topics."