a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...)
-
Updated
Oct 10, 2026 - C++
a community oriented 1:1, vLLM-alike (Continuous batching, paged KV) engine in C++ with additional features (GGUF, RadixAttention, Cache-aware scheduling, ...)
A deterministic PyTorch autograd verification trap for catching silent KV-cache routing and block-alignment failures in vLLM and SGLang serving infrastructure.
Sovereign structured LLM generation engine. SGLang RadixAttention + Pydantic schema enforcement + Ollama fallback. Guaranteed JSON output. No hallucinated keys.
SGLang 在线推理服务三次挑战完整交付(学号 0102603133):HW1 部署与 Mooncake trace 压测、HW2 RadixAttention 前缀缓存测量与请求流程分析、HW3 四副本 Ray Serve 路由对比与改进(A/B/C/D)。含可复现脚本、逐请求结果与报告 PDF。
SGLang 在线推理服务三次挑战完整交付(学号 0102603133):HW1 部署与 Mooncake trace 压测、HW2 RadixAttention 前缀缓存测量与请求流程分析、HW3 四副本 Ray Serve 路由对比与改进(A/B/C/D)。含可复现脚本、逐请求结果与报告 PDF。
SGLang vs vLLM playground: FastAPI app benchmarks TTFT/throughput against your own servers, plus 2 notebooks
To associate your repository with the radixattention topic, visit your repo's landing page and select "manage topics."