Skip to content

Repository files navigation

Multimodal RAG AI Knowledge Platform

An enterprise-grade, production-oriented multimodal RAG platform: agentic retrieval, hybrid search + reranking, OCR + Gemini Vision document intelligence, async ingestion workers, PostgreSQL/Redis/Qdrant data layer, full observability, evaluation framework, and Docker/CI/CD/K8s deployment path.

Honesty note: everything described here is implemented and validated locally. Distributed/cloud features (multi-replica Redis, Qdrant cluster, Kubernetes deployment, 1B-vector sharding) are architecture-ready but not deployed; benchmarks report only locally measured numbers.

Architecture

flowchart TD
    U[User] --> UI[React UI]
    UI -->|NDJSON stream| API[FastAPI /api/v1]
    API --> RT[Rate Limiter + Auth/Tenant]
    API --> CH[Chat Orchestrator<br/>adaptive RAG]
    API --> ING[Ingestion API]
    ING --> Q[(Queue<br/>Redis / memory)]
    Q --> W[Workers<br/>OCR · Embed · Index]
    CH --> ROUTER[Query Router]
    ROUTER --> RET[Hybrid Retrieval]
    RET --> D[Dense<br/>Qdrant] & K[BM25 Keyword]
    D & K --> F[RRF Fusion] --> RR[Reranker]
    RR --> CM[Context Manager<br/>token budget]
    CH --> MEM[Memory<br/>window + summary]
    CM --> LLM[Gemini LLM]
    LLM --> G[Groudedness Check]
    G --> OUT[Answer + Citations]
    W --> QS[(Qdrant)] & PG[(PostgreSQL)]
    CH & API & W --> OBS[Metrics · Trace · Logs]
Loading

Key capabilities

Area What's implemented
RAG Dense + BM25 hybrid, RRF fusion, flashrank reranking, score thresholds, adaptive retry with query rewriting
Agentic Lightweight query router (chitchat / general / knowledge), retrieval-quality evaluation loop (bounded)
Multimodal Tesseract OCR (threaded, cached, timeout) with Gemini Vision escalation for low-confidence images; vision description at ingestion
Document intelligence Structure-aware chunking (headings, tables, pages, sections), typed chunk metadata, JSON KB flattening
Ingestion Async job API, state machine with audit events, backpressure (MAX_QUEUE_DEPTH), retries → dead-letter, idempotent SHA-256 dedup, versioning + incremental re-index
Data layer PostgreSQL (async SQLAlchemy + Alembic) · Redis (cache/queue/rate-limit, graceful in-memory fallbacks) · Qdrant (embedded dev / server / cloud)
Memory Recent-window + rolling LLM summary; token-budgeted context (default 6k tokens)
Security Prompt-injection defenses, PII modes (warn/redact/block), input/image validation, rate limits, API-key auth, tenant isolation, security headers
Observability Prometheus metrics, health/ready, request/job correlation, JSON logs, per-request pipeline trace
Quality Golden dataset, retrieval evaluation (hit-rate/MRR), A/B strategy benchmark, 60+ offline tests
Delivery Multi-stage Docker, compose stack (api/worker/pg/redis/qdrant/frontend), GitHub Actions CI/CD → GHCR, K8s manifests + HPA

Quick start (zero-config dev)

# 1. Backend
cd backend
python -m venv .venv && .venv\Scripts\activate      # Windows (source .venv/bin/activate on Unix)
pip install -r requirements.txt
copy .env.example .env                              # add your GEMINI_API_KEY
uvicorn main:app --reload                           # http://localhost:8000/docs

# 2. Frontend
cd ../frontend
npm install && npm run dev                          # http://localhost:5173

Defaults: embedded Qdrant (no server), SQLite, in-memory queue/cache — the whole platform runs locally with only GEMINI_API_KEY set.

Full stack via Docker

cp backend/.env.example backend/.env     # set GEMINI_API_KEY, POSTGRES_PASSWORD
docker compose up --build                # api + worker + postgres + redis + qdrant + frontend
# frontend: http://localhost:8080  api: http://localhost:8000/docs  metrics: /api/v1/metrics

Scale workers independently: docker compose up --scale worker=4.

Test / evaluate / benchmark

cd backend
python -m pytest tests -q                              # 60+ offline tests
python -m alembic upgrade head                         # migrations

# RAG evaluation (real embeddings; measured, never fabricated)
python ..\scripts\evaluate_rag.py
python ..\scripts\benchmark_rag.py                     # A/B: dense vs hybrid vs hybrid_rerank
python ..\scripts\benchmark_retrieval.py --queries 50 --runs 5

# Ingestion benchmark
python ..\scripts\generate_test_data.py --count 100
python ..\scripts\benchmark_ingestion.py --dataset backend\eval\generated\docs_100.jsonl

Measured locally (dev laptop, CPU ONNX, 36-entry KB): retrieval hit-rate 0.93 / MRR 0.92–1.0 (see backend/eval/reports/); hybrid retrieval p50 ~6 ms (warm), rerank p50 ~33 ms; ingestion ~1.5 docs/s (CPU-bound embedding). Reports are regenerated per run — always re-run on your hardware.

API surface (v1)

Endpoint Purpose
POST /api/v1/chat/stream Streaming agentic RAG chat (NDJSON: status/sources/tokens/groundedness/usage/trace)
POST /api/v1/ingestion/jobs (+ /upload, list, events, retry) Async ingestion
GET /api/v1/documents (+ versions, delete) Corpus management
POST /api/v1/images/analyze OCR + Vision analysis
GET /api/v1/health · /ready · /metrics Liveness / readiness / Prometheus
/api/* (legacy) Original chat/conversations endpoints (preserved)

Documentation

License

MIT

About

Production-oriented multimodal RAG platform with Agentic AI, hybrid search, OCR/Vision, document ingestion, citations, evaluation, and full-stack observability.

Topics

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages