📚 Documentation: docs/ — start with the user guide: launching / monitoring / resuming runs, RAG-Stack + RAG-CM optimization, system transfer, settings & cache directories. The web app has its own page.
We evaluate RAG-Stack on RAGEval and MS MARCO using two NUMA servers (SysA: four H100s; SysB: eight A100s). On average across seeds, RAG-Stack's Pareto frontiers cover 52.5% and 153.2% more of the normalized quality–performance space than the strongest end-to-end baseline on the two datasets, respectively. At budget 50, the frontiers found by RAG-Stack cover 14.8% and 13.8% more of this space than those found by the strongest optimizer baselines on the two datasets, respectively; with 20 evaluations, the transferred frontier covers 182.2% more of this space on the new system than the frontier found by re-optimizing from scratch.
The underlying Pareto results and replayable configurations are available in the end-to-end benchmark results.
# Clone, then initialize every bundled submodule. This includes the evaluator
# workspace member from https://github.com/haiqiang-zhang/RAG-Stack-Evaluator.
git clone <repo-url>
cd rag-stack
git submodule update --init --recursive
# Create environment (Python 3.12 + faiss-cpu)
conda env create -f environment.yml
conda activate rag-stack
pip install uv
# Install GenZ-LLM-Analyzer (submodule) into the active conda environment.
uv pip install -r genz_llm/requirements.txt
uv pip install -e genz_llm
# Install both local projects into that same conda environment. For API/CPU
# development and smoke tests, force the CPU Torch wheel and omit every CUDA,
# vLLM, and NIXL extra:
uv pip install --torch-backend cpu -e 'RAG-Stack-Evaluator[test]' -e .
# For measured GPU execution instead, pick the stack by your NVIDIA DRIVER
# (check: nvidia-smi). Evaluator/vLLM dependencies are owned by the
# RAG-Stack-Evaluator extra:
# driver 525–579 (this box: 560) -> cu12 (vLLM 0.18.1 + cu12 nixl backend)
# driver >=580 -> cu13 (vLLM 0.22.x + cu13 nixl backend)
# Pin the matching Torch wheel as part of the same install command.
uv pip install --torch-backend cu128 \
-e 'RAG-Stack-Evaluator[cu12]' -e '.[cu12]'
# or:
uv pip install --torch-backend cu130 \
-e 'RAG-Stack-Evaluator[cu13]' -e '.[cu13]'
Already cloned? Run
git submodule update --init --recursivebefore installing so the RAG-Stack-Evaluator workspace member and the other bundled submodules are present.
CUDA 12 vs 13 — pick the
cu12orcu13extra by your driver (above); they select the vLLM + nixl backend stack (cu12 → vLLM 0.18.1, the last release whose 1P1D disaggregation works with the cu12 nixl backend; cu13 → vLLM 0.22.x). The torch wheel is selected explicitly by the install command so its CUDA major cannot diverge from the selected vLLM stack. The install command names both editable projects explicitly becauseuv pipignores[tool.uv.sources]. (nixlpulls both backend wheels regardless; the cu12/cu13 backend is selected at runtime by your torch build.) Before every local or spawned vLLM launch, the evaluator validates an existingCUDA_HOMEagainst that Torch build and then discovers a matching Conda/system toolkit. The cu13 extra additionally installs a complete CUDA 13 wheel toolkit, which is discovered automatically on clusters without/usr/local/cuda. No activation hook, repository shell script, or manualLD_LIBRARY_PATHis required. CUDA 12's split PyPI packages do not contain annvccdriver, so the cu12 extra installs matching precompiled FlashInfer cubins. If an unsupported kernel still triggers source JIT, a matching system or Conda CUDA 12 toolkit is required and discovered automatically. A wrong-major explicit toolkit is rejected instead of being used.
After installation, run RAG-Stack with the reference configuration:
conda run -n rag-stack python -m rag_stack optimize \
--config configs/rag_stack/config_reference.yamlFor configuration, workflows, and all other features, read the documentation.
We thank the authors and contributors of RAGEval for making their RAG evaluation benchmark available.