Skip to content
haiqiang-zhangPublic

About

RAG-Stack, a framework for efficiently discovering quality--performance Pareto frontiers across diverse RAG applications and serving systems.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Repository files navigation

RAG-Stack: Co-Optimizing RAG Serving Performance and Quality

RAG-Stack architecture and optimization workflow

📚 Documentation: docs/ — start with the user guide: launching / monitoring / resuming runs, RAG-Stack + RAG-CM optimization, system transfer, settings & cache directories. The web app has its own page.

End-to-End Benchmark

End-to-end quality–performance Pareto frontiers on RAGEval and MS MARCO

We evaluate RAG-Stack on RAGEval and MS MARCO using two NUMA servers (SysA: four H100s; SysB: eight A100s). On average across seeds, RAG-Stack's Pareto frontiers cover 52.5% and 153.2% more of the normalized quality–performance space than the strongest end-to-end baseline on the two datasets, respectively. At budget 50, the frontiers found by RAG-Stack cover 14.8% and 13.8% more of this space than those found by the strongest optimizer baselines on the two datasets, respectively; with 20 evaluations, the transferred frontier covers 182.2% more of this space on the new system than the frontier found by re-optimizing from scratch.

The underlying Pareto results and replayable configurations are available in the end-to-end benchmark results.

Installation

# Clone, then initialize every bundled submodule. This includes the evaluator
# workspace member from https://github.com/haiqiang-zhang/RAG-Stack-Evaluator.
git clone <repo-url>
cd rag-stack
git submodule update --init --recursive

# Create environment (Python 3.12 + faiss-cpu)
conda env create -f environment.yml
conda activate rag-stack
pip install uv

# Install GenZ-LLM-Analyzer (submodule) into the active conda environment.
uv pip install -r genz_llm/requirements.txt
uv pip install -e genz_llm

# Install both local projects into that same conda environment. For API/CPU
# development and smoke tests, force the CPU Torch wheel and omit every CUDA,
# vLLM, and NIXL extra:
uv pip install --torch-backend cpu -e 'RAG-Stack-Evaluator[test]' -e .

# For measured GPU execution instead, pick the stack by your NVIDIA DRIVER
# (check: nvidia-smi). Evaluator/vLLM dependencies are owned by the
# RAG-Stack-Evaluator extra:
#   driver 525–579 (this box: 560) -> cu12  (vLLM 0.18.1 + cu12 nixl backend)
#   driver >=580                   -> cu13  (vLLM 0.22.x + cu13 nixl backend)
# Pin the matching Torch wheel as part of the same install command.
uv pip install --torch-backend cu128 \
  -e 'RAG-Stack-Evaluator[cu12]' -e '.[cu12]'
# or:
uv pip install --torch-backend cu130 \
  -e 'RAG-Stack-Evaluator[cu13]' -e '.[cu13]'

Already cloned? Run git submodule update --init --recursive before installing so the RAG-Stack-Evaluator workspace member and the other bundled submodules are present.

CUDA 12 vs 13 — pick the cu12 or cu13 extra by your driver (above); they select the vLLM + nixl backend stack (cu12 → vLLM 0.18.1, the last release whose 1P1D disaggregation works with the cu12 nixl backend; cu13 → vLLM 0.22.x). The torch wheel is selected explicitly by the install command so its CUDA major cannot diverge from the selected vLLM stack. The install command names both editable projects explicitly because uv pip ignores [tool.uv.sources]. (nixl pulls both backend wheels regardless; the cu12/cu13 backend is selected at runtime by your torch build.) Before every local or spawned vLLM launch, the evaluator validates an existing CUDA_HOME against that Torch build and then discovers a matching Conda/system toolkit. The cu13 extra additionally installs a complete CUDA 13 wheel toolkit, which is discovered automatically on clusters without /usr/local/cuda. No activation hook, repository shell script, or manual LD_LIBRARY_PATH is required. CUDA 12's split PyPI packages do not contain an nvcc driver, so the cu12 extra installs matching precompiled FlashInfer cubins. If an unsupported kernel still triggers source JIT, a matching system or Conda CUDA 12 toolkit is required and discovered automatically. A wrong-major explicit toolkit is rejected instead of being used.

Getting Started

After installation, run RAG-Stack with the reference configuration:

conda run -n rag-stack python -m rag_stack optimize \
    --config configs/rag_stack/config_reference.yaml

For configuration, workflows, and all other features, read the documentation.

Acknowledgements

We thank the authors and contributors of RAGEval for making their RAG evaluation benchmark available.

About

RAG-Stack, a framework for efficiently discovering quality--performance Pareto frontiers across diverse RAG applications and serving systems.

Resources

Stars

4 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages