An enterprise-grade, production-oriented multimodal RAG platform: agentic retrieval, hybrid search + reranking, OCR + Gemini Vision document intelligence, async ingestion workers, PostgreSQL/Redis/Qdrant data layer, full observability, evaluation framework, and Docker/CI/CD/K8s deployment path.
Honesty note: everything described here is implemented and validated locally. Distributed/cloud features (multi-replica Redis, Qdrant cluster, Kubernetes deployment, 1B-vector sharding) are architecture-ready but not deployed; benchmarks report only locally measured numbers.
flowchart TD
U[User] --> UI[React UI]
UI -->|NDJSON stream| API[FastAPI /api/v1]
API --> RT[Rate Limiter + Auth/Tenant]
API --> CH[Chat Orchestrator<br/>adaptive RAG]
API --> ING[Ingestion API]
ING --> Q[(Queue<br/>Redis / memory)]
Q --> W[Workers<br/>OCR · Embed · Index]
CH --> ROUTER[Query Router]
ROUTER --> RET[Hybrid Retrieval]
RET --> D[Dense<br/>Qdrant] & K[BM25 Keyword]
D & K --> F[RRF Fusion] --> RR[Reranker]
RR --> CM[Context Manager<br/>token budget]
CH --> MEM[Memory<br/>window + summary]
CM --> LLM[Gemini LLM]
LLM --> G[Groudedness Check]
G --> OUT[Answer + Citations]
W --> QS[(Qdrant)] & PG[(PostgreSQL)]
CH & API & W --> OBS[Metrics · Trace · Logs]
| Area | What's implemented |
|---|---|
| RAG | Dense + BM25 hybrid, RRF fusion, flashrank reranking, score thresholds, adaptive retry with query rewriting |
| Agentic | Lightweight query router (chitchat / general / knowledge), retrieval-quality evaluation loop (bounded) |
| Multimodal | Tesseract OCR (threaded, cached, timeout) with Gemini Vision escalation for low-confidence images; vision description at ingestion |
| Document intelligence | Structure-aware chunking (headings, tables, pages, sections), typed chunk metadata, JSON KB flattening |
| Ingestion | Async job API, state machine with audit events, backpressure (MAX_QUEUE_DEPTH), retries → dead-letter, idempotent SHA-256 dedup, versioning + incremental re-index |
| Data layer | PostgreSQL (async SQLAlchemy + Alembic) · Redis (cache/queue/rate-limit, graceful in-memory fallbacks) · Qdrant (embedded dev / server / cloud) |
| Memory | Recent-window + rolling LLM summary; token-budgeted context (default 6k tokens) |
| Security | Prompt-injection defenses, PII modes (warn/redact/block), input/image validation, rate limits, API-key auth, tenant isolation, security headers |
| Observability | Prometheus metrics, health/ready, request/job correlation, JSON logs, per-request pipeline trace |
| Quality | Golden dataset, retrieval evaluation (hit-rate/MRR), A/B strategy benchmark, 60+ offline tests |
| Delivery | Multi-stage Docker, compose stack (api/worker/pg/redis/qdrant/frontend), GitHub Actions CI/CD → GHCR, K8s manifests + HPA |
# 1. Backend
cd backend
python -m venv .venv && .venv\Scripts\activate # Windows (source .venv/bin/activate on Unix)
pip install -r requirements.txt
copy .env.example .env # add your GEMINI_API_KEY
uvicorn main:app --reload # http://localhost:8000/docs
# 2. Frontend
cd ../frontend
npm install && npm run dev # http://localhost:5173Defaults: embedded Qdrant (no server), SQLite, in-memory queue/cache — the
whole platform runs locally with only GEMINI_API_KEY set.
cp backend/.env.example backend/.env # set GEMINI_API_KEY, POSTGRES_PASSWORD
docker compose up --build # api + worker + postgres + redis + qdrant + frontend
# frontend: http://localhost:8080 api: http://localhost:8000/docs metrics: /api/v1/metricsScale workers independently: docker compose up --scale worker=4.
cd backend
python -m pytest tests -q # 60+ offline tests
python -m alembic upgrade head # migrations
# RAG evaluation (real embeddings; measured, never fabricated)
python ..\scripts\evaluate_rag.py
python ..\scripts\benchmark_rag.py # A/B: dense vs hybrid vs hybrid_rerank
python ..\scripts\benchmark_retrieval.py --queries 50 --runs 5
# Ingestion benchmark
python ..\scripts\generate_test_data.py --count 100
python ..\scripts\benchmark_ingestion.py --dataset backend\eval\generated\docs_100.jsonlMeasured locally (dev laptop, CPU ONNX, 36-entry KB): retrieval hit-rate
0.93 / MRR 0.92–1.0 (see backend/eval/reports/); hybrid retrieval p50
~6 ms (warm), rerank p50 ~33 ms; ingestion ~1.5 docs/s (CPU-bound
embedding). Reports are regenerated per run — always re-run on your hardware.
| Endpoint | Purpose |
|---|---|
POST /api/v1/chat/stream |
Streaming agentic RAG chat (NDJSON: status/sources/tokens/groundedness/usage/trace) |
POST /api/v1/ingestion/jobs (+ /upload, list, events, retry) |
Async ingestion |
GET /api/v1/documents (+ versions, delete) |
Corpus management |
POST /api/v1/images/analyze |
OCR + Vision analysis |
GET /api/v1/health · /ready · /metrics |
Liveness / readiness / Prometheus |
/api/* (legacy) |
Original chat/conversations endpoints (preserved) |
- docs/architecture.md — components & data flow
- docs/rag-pipeline.md — retrieval detail
- docs/security.md — injection defense & PII
- docs/scaling.md — 10K → 1B vectors path
- docs/evaluation.md — datasets & metrics
- docs/deployment.md — Docker/K8s/cloud notes
- docs/disaster-recovery.md — backups, RPO/RTO
MIT