Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
28 commits
Select commit Hold shift + click to select a range
3b3b386
chore(obs): add OpenTelemetry SDK and Arize Phoenix deps
elkaix Apr 27, 2026
2a00c47
feat(obs): add OpenTelemetry tracer init and traced_stage decorator
elkaix Apr 27, 2026
895bd1d
feat(api): add StageTelemetry DTO for per-stage timings
elkaix Apr 27, 2026
cecdba9
feat(backend): instrument RAGBackend with traced stages and StageTele…
elkaix Apr 27, 2026
ab1e65a
feat(api): emit StageTelemetry in REST and WebSocket query responses
elkaix Apr 27, 2026
4fa63e0
feat(frontend): handle telemetry event and expose on chat message
elkaix Apr 27, 2026
1254b1b
feat(frontend): add TelemetryFooter under each assistant chat message
elkaix Apr 27, 2026
af09992
feat(obs): wire init_observability in lifespan; add Phoenix service p…
elkaix Apr 27, 2026
c925492
docs(arch): document evaluation harness and observability layers
elkaix Apr 27, 2026
1f05b94
fix(frontend): use keyed Fragments in TelemetryFooter to silence Reac…
elkaix Apr 27, 2026
cf77577
fix(frontend): replace empty-config select with text hint in NewEvalR…
elkaix Apr 27, 2026
7d1ed38
build(compose): mount configs/ and eval_runs/ so the api container se…
elkaix Apr 27, 2026
63063ca
fix(frontend): pad RunsList outer container so toolbar isn't flush ag…
elkaix Apr 27, 2026
ea02081
docs(specs): add Phase 2 RAG quality matrix design
elkaix Apr 27, 2026
907f34a
docs(specs): revise Phase 2 design after spec-review findings
elkaix Apr 27, 2026
fe4f5fb
docs(plans): add Phase 2 RAG quality matrix implementation plan
elkaix Apr 27, 2026
72a0824
feat(eval): extend PipelineCfg with Phase 2 sub-configs and spend cei…
elkaix Apr 27, 2026
08532ea
feat(eval): cost ledger covers generator + judge + rewriter spend
elkaix Apr 27, 2026
c85a985
feat(eval): add BgeEmbedder as Chroma EmbeddingFunction adapter
elkaix Apr 27, 2026
be81b2b
feat(eval): add BM25HybridRetriever with RRF fusion
elkaix Apr 27, 2026
d4d9c12
feat(eval): add CrossEncoderReranker (ms-marco-MiniLM)
elkaix Apr 27, 2026
9213511
feat(eval): add QueryRewriter for LLM-based expansion with cost capture
elkaix Apr 27, 2026
35caf92
feat(eval): add RefusalHandler with similarity gate
elkaix Apr 27, 2026
1572ae3
feat(eval): wire Phase 2 levers into build_pipeline + EvalPipeline.query
elkaix Apr 27, 2026
4930996
feat(eval): add cli archive subcommand to copy small run artifacts
elkaix Apr 27, 2026
c24879a
chore(eval): add Phase 2 tier configs under configs/eval/phase2/
elkaix Apr 27, 2026
e4d243f
Merge pull request #5 from mohamed-elkholy95/feature/phase2-pipeline-…
elkaix Apr 27, 2026
b0e2609
Merge branch 'feature/eval-harness-1c' into feature/eval-harness-1d
elkaix Apr 28, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
100 changes: 100 additions & 0 deletions Architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -443,6 +443,106 @@ Run: `python -m pytest tests/ -v`

---

## Evaluation Harness

The `src/eval/` package provides a reproducible evaluation system over labeled gold sets, separate from the user-facing chat path.

### Layers

| Module | Responsibility |
|--------|----------------|
| `src/eval/schemas.py` | Pydantic contracts: `EvalQuestion`, `EvalResult`, `AggregatedMetric`, `RunMetadata`, `MetricDelta`, `CompareResult`. |
| `src/eval/pricing.py` | Hard-coded model price table + `cost_usd()` helper. |
| `src/eval/statistics.py` | `bootstrap_ci()` and `paired_permutation_test()` for run-level confidence intervals and two-run significance testing. |
| `src/eval/metrics/retrieval.py` | Recall@k, MRR@k, nDCG@k over `(gold_chunk_ids, retrieved_chunk_ids)`. |
| `src/eval/metrics/operational.py` | Per-stage latency p50/p95/p99, cost, token aggregation. |
| `src/eval/metrics/refusal.py` | Regex + LLM-judge refusal correctness for unanswerable questions. |
| `src/eval/metrics/generation.py` | Wraps existing `src/evaluation.py` LLM-as-judge functions; adds `answer_correctness` (cosine + judge mean) and `context_recall`. |
| `src/eval/datasets/squad_v2.py` | Seeded sample + frozen 200-row JSONL artifact from HuggingFace `squad_v2`. |
| `src/eval/datasets/ml_papers.py` | Hand-labeled dev set loader + manifest SHA-256 verification. |
| `src/eval/config.py` | YAML-loaded `EvalConfig`. |
| `src/eval/storage.py` | Run-directory CRUD over `eval_runs/<run_id>/`. |
| `src/eval/pipeline_factory.py` + `src/eval/_telemetry.py` | Builds an isolated RAG pipeline per (config, dataset) using ephemeral Chroma. |
| `src/eval/aggregator.py` | Per-dataset + combined `AggregatedMetric` rows from per-question results. |
| `src/eval/runner.py` | Orchestrates `git_sha`, ingest, query+score loop, aggregation, persistence. |
| `src/eval/compare.py` | Two-run diff with paired permutation tests + per-question regressions/wins. |
| `src/eval/report.py` + `templates/eval/*.html.j2` | Standalone jinja2 HTML reports. |
| `src/eval/cli.py` | `run`/`list`/`show`/`compare` argparse subcommands. |

### API + UI

`src/api/routes/eval.py` exposes:

| Method | Path | Description |
|--------|------|-------------|
| `GET` | `/api/eval/configs` | List available eval configs |
| `POST` | `/api/eval/run` | Start a new eval run (dispatched via `BackgroundTasks`) |
| `GET` | `/api/eval/runs` | List all eval runs |
| `GET` | `/api/eval/runs/{id}` | Get run metadata |
| `GET` | `/api/eval/runs/{id}/results` | Per-question results |
| `GET` | `/api/eval/runs/{id}/status` | Live status for in-progress runs |
| `GET` | `/api/eval/compare` | Two-run diff with significance tests |

Long-running runs dispatch via FastAPI `BackgroundTasks` and report progress through an in-process `RunRegistry` (`src/api/services/eval_runs.py`).

React route `/eval/*` mounts three views:
- **`RunsList`** — sortable/filterable table with multi-select compare
- **`RunDetail`** — metric chart + per-question table with lazy expand
- **`CompareView`** — side-by-side bars + Top Wins / Top Regressions cards

Charts use `recharts` with CI whiskers.

### Eval Run Directory

Each run produces `eval_runs/<run_id>/` with:
- `metadata.json` — run ID, git SHA, config name, timestamps
- `questions.jsonl` — per-question scores and retrieved chunks
- `metrics.json` — aggregated metric values with bootstrap CIs
- `cost.json` — token counts and USD costs per model
- `config.yaml` — snapshot of the config used

The `eval_runs/` directory is gitignored; the labeled dev sets in `eval_data/` are checked in.

---

## Observability

The system exports per-stage spans for every chat query via OpenTelemetry to [Arize Phoenix](https://github.com/Arize-ai/phoenix) on `localhost:6006`.

### Spans

`RAGBackend.query_with_telemetry` and `RAGBackend.stream_query` open spans:

| Span | Attributes |
|------|------------|
| `rag.retrieve` | `top_k`, `chunk_count` |
| `rag.generate` | `model`, `prompt_tokens`, `completion_tokens`, `cost_usd` |

### Telemetry Payload

The same numbers are returned to the client as a `StageTelemetry` Pydantic model (`src/api/schemas/telemetry.py`):

- REST `POST /api/query` — includes a `telemetry` field in the response JSON.
- WebSocket `/api/chat` — emits a final `{"type": "telemetry", "content": {...}}` event after the existing `done` event.

The frontend renders these as a muted footer line under each assistant chat bubble:

> *Retrieve 142ms · Generate 2.1s · 4,217 tok · $0.0083*

with a hover tooltip showing the prompt/completion token split.

### Running with Traces

Phoenix is profile-gated in `docker-compose.yml`; bare `docker compose up` does not start it.

```bash
docker compose --profile observability up
```

`init_observability()` (`src/observability.py`) is called during the FastAPI lifespan startup. It is idempotent and fail-quiet — if Phoenix is unreachable, spans become no-ops and the chat continues to work normally.

---

## Key Design Decisions

| Decision | Choice | Rationale |
Expand Down
11 changes: 11 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -176,6 +176,17 @@ docker compose up --build
| API | [localhost:8001](http://localhost:8001) |
| Swagger Docs | [localhost:8001/docs](http://localhost:8001/docs) |

### Running with traces (Phoenix)

```bash
docker compose --profile observability up
```

Phoenix UI is available at http://localhost:6006. The FastAPI backend
will export per-stage spans (`rag.retrieve`, `rag.generate`) with token
counts and cost as span attributes. If Phoenix isn't running, the app
works normally — span export silently fails.

### Manual Setup

```bash
Expand Down
18 changes: 18 additions & 0 deletions configs/eval/baseline_squad_only.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
name: "baseline_squad_only"
description: "Baseline pipeline against squad_v2_dev_200 only (ml_papers_v1 not yet labeled)."
pipeline:
chunker:
strategy: "recursive"
chunk_size: 512
chunk_overlap: 64
retriever:
top_k: 5
generator:
model: "gpt-5-mini"
reasoning_model: "gpt-4.1-nano"
eval:
datasets: ["squad_v2_dev_200"]
judge_model: "gpt-4.1-mini"
bootstrap_n: 1000
permutation_n: 10000
seed: 42
13 changes: 13 additions & 0 deletions configs/eval/phase2/phase2_baseline.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
name: "phase2_baseline"
description: "Baseline anchor for the Phase 2 matrix — identical to baseline_squad_only."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
retriever: {top_k: 5}
generator: {model: gpt-5-mini, reasoning_model: gpt-4.1-nano}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
14 changes: 14 additions & 0 deletions configs/eval/phase2/phase2b_embedder.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
name: "phase2b_embedder"
description: "Phase 2 tier 2b — swap default embedder for BAAI/bge-small-en-v1.5."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
generator: {model: gpt-5-mini, reasoning_model: gpt-4.1-nano}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
15 changes: 15 additions & 0 deletions configs/eval/phase2/phase2c_hybrid.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
name: "phase2c_hybrid"
description: "Phase 2 tier 2c — add BM25 hybrid on top of BGE."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
hybrid: {enabled: true, bm25_top_k: 20, dense_top_k: 20, rrf_k: 60}
generator: {model: gpt-5-mini, reasoning_model: gpt-4.1-nano}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
16 changes: 16 additions & 0 deletions configs/eval/phase2/phase2d_rerank.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
name: "phase2d_rerank"
description: "Phase 2 tier 2d — add cross-encoder rerank on top of hybrid."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
hybrid: {enabled: true, bm25_top_k: 20, dense_top_k: 20, rrf_k: 60}
reranker: {model: ms_marco_minilm_l6_v2, rerank_top_n: 20, final_top_k: 5}
generator: {model: gpt-5-mini, reasoning_model: gpt-4.1-nano}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
17 changes: 17 additions & 0 deletions configs/eval/phase2/phase2e_rewrite.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
name: "phase2e_rewrite"
description: "Phase 2 tier 2e — add LLM query rewriting on top of rerank."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
hybrid: {enabled: true, bm25_top_k: 20, dense_top_k: 20, rrf_k: 60}
reranker: {model: ms_marco_minilm_l6_v2, rerank_top_n: 20, final_top_k: 5}
query_rewriter: {model: gpt-4.1-nano, max_expansions: 3}
generator: {model: gpt-5-mini, reasoning_model: gpt-4.1-nano}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
18 changes: 18 additions & 0 deletions configs/eval/phase2/phase2f_models_gpt41mini.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
name: "phase2f_models_gpt41mini"
description: "Phase 2 tier 2f — answer-model comparison: gpt-4.1-mini variant."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
hybrid: {enabled: true, bm25_top_k: 20, dense_top_k: 20, rrf_k: 60}
reranker: {model: ms_marco_minilm_l6_v2, rerank_top_n: 20, final_top_k: 5}
query_rewriter: {model: gpt-4.1-nano, max_expansions: 3}
generator: {model: gpt-4.1-mini, reasoning_model: gpt-4.1-nano}
refusal_handler: {enabled: true, similarity_threshold: 0.35, no_answer_text: "I don't have enough information to answer that."}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
19 changes: 19 additions & 0 deletions configs/eval/phase2/phase2f_models_gpt5mini.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
name: "phase2f_models_gpt5mini"
description: "Phase 2 tier 2f — answer-model comparison on the full 2g stack: gpt-5-mini variant.
This config is identical to phase2g_refusal.yaml; PR-B re-uses the 2g artifact."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
hybrid: {enabled: true, bm25_top_k: 20, dense_top_k: 20, rrf_k: 60}
reranker: {model: ms_marco_minilm_l6_v2, rerank_top_n: 20, final_top_k: 5}
query_rewriter: {model: gpt-4.1-nano, max_expansions: 3}
generator: {model: gpt-5-mini, reasoning_model: gpt-4.1-nano}
refusal_handler: {enabled: true, similarity_threshold: 0.35, no_answer_text: "I don't have enough information to answer that."}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
18 changes: 18 additions & 0 deletions configs/eval/phase2/phase2f_models_haiku.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
name: "phase2f_models_haiku"
description: "Phase 2 tier 2f — answer-model comparison: claude-haiku-4-5 variant."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
hybrid: {enabled: true, bm25_top_k: 20, dense_top_k: 20, rrf_k: 60}
reranker: {model: ms_marco_minilm_l6_v2, rerank_top_n: 20, final_top_k: 5}
query_rewriter: {model: gpt-4.1-nano, max_expansions: 3}
generator: {model: claude-haiku-4-5, reasoning_model: null}
refusal_handler: {enabled: true, similarity_threshold: 0.35, no_answer_text: "I don't have enough information to answer that."}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
18 changes: 18 additions & 0 deletions configs/eval/phase2/phase2g_refusal.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
name: "phase2g_refusal"
description: "Phase 2 tier 2g — full stack with refusal handler."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
hybrid: {enabled: true, bm25_top_k: 20, dense_top_k: 20, rrf_k: 60}
reranker: {model: ms_marco_minilm_l6_v2, rerank_top_n: 20, final_top_k: 5}
query_rewriter: {model: gpt-4.1-nano, max_expansions: 3}
generator: {model: gpt-5-mini, reasoning_model: gpt-4.1-nano}
refusal_handler: {enabled: true, similarity_threshold: 0.35, no_answer_text: "I don't have enough information to answer that."}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
20 changes: 20 additions & 0 deletions docker-compose.yml
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,11 @@ services:
- ./tests:/app/tests
- ./data:/app/data
- ./books:/app/books
# WHY: The eval harness reads YAML configs from configs/eval/ and writes
# run directories to eval_runs/. Mounting both lets the dev loop
# (host CLI runs ↔ in-container API) share artifacts with no copy.
- ./configs:/app/configs
- ./eval_runs:/app/eval_runs
# WHY: Cache the ChromaDB ONNX embedding model (79MB) so it doesn't
# re-download on every container restart. First startup is slow,
# subsequent starts are instant.
Expand Down Expand Up @@ -67,5 +72,20 @@ services:
depends_on:
api:
condition: service_healthy

# PATTERN: Profile-gating — Phoenix only starts when explicitly requested
# with `docker compose --profile observability up`. A bare
# `docker compose up` starts only api + frontend, keeping the dev
# experience lightweight.
# WHY Phoenix: Arize Phoenix is an open-source, zero-cloud LLM observability
# UI that accepts OTLP spans and renders per-stage trace timelines
# with token counts and cost breakdowns.
phoenix:
image: arizephoenix/phoenix:latest
ports:
- "6006:6006"
profiles: ["observability"]
restart: unless-stopped

volumes:
chroma-model-cache:
Loading
Loading