Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
69 commits
Select commit Hold shift + click to select a range
da3de96
chore(eval): add datasets dep and gitignore eval_runs/
elkaix Apr 27, 2026
3974577
feat(eval): scaffold src/eval package with three subpackages
elkaix Apr 27, 2026
d568596
feat(eval): add Pydantic schemas for eval harness contracts
elkaix Apr 27, 2026
2ecb230
feat(eval): add model price table and cost_usd function
elkaix Apr 27, 2026
decccf3
feat(eval): add retrieval metrics (Recall@k, MRR@k, nDCG@k)
elkaix Apr 27, 2026
ca1f888
Revert "feat(eval): add Pydantic schemas for eval harness contracts"
elkaix Apr 27, 2026
0270410
feat(eval): add Pydantic schemas for eval harness contracts
elkaix Apr 27, 2026
ff33e94
feat(eval): add operational aggregators (latency/cost/tokens)
elkaix Apr 27, 2026
0a0364e
feat(eval): add bootstrap_ci for percentile confidence intervals
elkaix Apr 27, 2026
6e72673
feat(eval): add paired permutation test for two-run comparison
elkaix Apr 27, 2026
b6afa3b
feat(eval): add refusal_correctness metric (regex + LLM fallback)
elkaix Apr 27, 2026
e51b89b
feat(eval): add answer_correctness, context_recall, judge wrappers
elkaix Apr 27, 2026
c5834f7
chore(eval): pin sentence-transformers (used by answer_correctness em…
elkaix Apr 27, 2026
2af2f0b
feat(eval): add SQuAD v2 dev-set sampler and frozen 200-row artifact
elkaix Apr 27, 2026
d9a4497
feat(eval): add ML Papers v1 loader, manifest verifier, labeling guide
elkaix Apr 27, 2026
ce15b12
feat(eval): expose public package surface and add composition smoke test
elkaix Apr 27, 2026
1c9035f
chore(eval): add pyyaml and jinja2 deps for runner + reports
elkaix Apr 27, 2026
6cf9b69
feat(eval): add EvalConfig YAML loader and baseline config
elkaix Apr 27, 2026
fa0e068
feat(eval): add storage layer for eval run directories
elkaix Apr 27, 2026
3f63ef5
feat(eval): add EvalPipeline factory bridging EvalConfig to RAG compo…
elkaix Apr 27, 2026
e558fbc
refactor(eval): split telemetry helpers out of pipeline_factory
elkaix Apr 27, 2026
5a04010
feat(eval): add metric aggregator with per-dataset and combined rows
elkaix Apr 27, 2026
6122b1e
feat(eval): add EvalRunner orchestrator end-to-end
elkaix Apr 27, 2026
47b86f7
feat(eval): add two-run comparison with paired significance tests
elkaix Apr 27, 2026
baf58b7
feat(eval): add jinja2-based HTML report renderer for runs and compar…
elkaix Apr 27, 2026
d7eff7c
feat(eval): add argparse CLI for run/list/show/compare
elkaix Apr 27, 2026
8a389f4
feat(eval): wire 1B exports and add end-to-end integration smoke test
elkaix Apr 27, 2026
da2ac6d
chore(frontend): add recharts dep for eval metric charts
elkaix Apr 27, 2026
952ab50
feat(api): add eval API DTOs
elkaix Apr 27, 2026
4bf21ac
feat(api): add thread-safe in-process eval run registry
elkaix Apr 27, 2026
dca0290
feat(api): add /api/eval/* routes for runs, results, compare, configs
elkaix Apr 27, 2026
18d6edb
feat(frontend): add eval API client and TanStack Query hooks
elkaix Apr 27, 2026
6861916
feat(frontend): add MetricBars chart component for aggregated metrics
elkaix Apr 27, 2026
519967e
feat(frontend): add RunsList with sort, filter, multi-select compare
elkaix Apr 27, 2026
8424f26
feat(frontend): add RunDetail with metrics chart and per-question table
elkaix Apr 27, 2026
7929c5f
feat(frontend): add CompareView with side-by-side bars and per-questi…
elkaix Apr 27, 2026
6a51010
feat(frontend): add NewEvalRunDialog with config picker and progress …
elkaix Apr 27, 2026
2a117a1
feat(frontend): wire /eval routes and add Evaluation sidebar link
elkaix Apr 27, 2026
3b3b386
chore(obs): add OpenTelemetry SDK and Arize Phoenix deps
elkaix Apr 27, 2026
2a00c47
feat(obs): add OpenTelemetry tracer init and traced_stage decorator
elkaix Apr 27, 2026
895bd1d
feat(api): add StageTelemetry DTO for per-stage timings
elkaix Apr 27, 2026
cecdba9
feat(backend): instrument RAGBackend with traced stages and StageTele…
elkaix Apr 27, 2026
ab1e65a
feat(api): emit StageTelemetry in REST and WebSocket query responses
elkaix Apr 27, 2026
4fa63e0
feat(frontend): handle telemetry event and expose on chat message
elkaix Apr 27, 2026
1254b1b
feat(frontend): add TelemetryFooter under each assistant chat message
elkaix Apr 27, 2026
af09992
feat(obs): wire init_observability in lifespan; add Phoenix service p…
elkaix Apr 27, 2026
c925492
docs(arch): document evaluation harness and observability layers
elkaix Apr 27, 2026
1f05b94
fix(frontend): use keyed Fragments in TelemetryFooter to silence Reac…
elkaix Apr 27, 2026
cf77577
fix(frontend): replace empty-config select with text hint in NewEvalR…
elkaix Apr 27, 2026
7d1ed38
build(compose): mount configs/ and eval_runs/ so the api container se…
elkaix Apr 27, 2026
63063ca
fix(frontend): pad RunsList outer container so toolbar isn't flush ag…
elkaix Apr 27, 2026
ea02081
docs(specs): add Phase 2 RAG quality matrix design
elkaix Apr 27, 2026
907f34a
docs(specs): revise Phase 2 design after spec-review findings
elkaix Apr 27, 2026
fe4f5fb
docs(plans): add Phase 2 RAG quality matrix implementation plan
elkaix Apr 27, 2026
72a0824
feat(eval): extend PipelineCfg with Phase 2 sub-configs and spend cei…
elkaix Apr 27, 2026
08532ea
feat(eval): cost ledger covers generator + judge + rewriter spend
elkaix Apr 27, 2026
c85a985
feat(eval): add BgeEmbedder as Chroma EmbeddingFunction adapter
elkaix Apr 27, 2026
be81b2b
feat(eval): add BM25HybridRetriever with RRF fusion
elkaix Apr 27, 2026
d4d9c12
feat(eval): add CrossEncoderReranker (ms-marco-MiniLM)
elkaix Apr 27, 2026
9213511
feat(eval): add QueryRewriter for LLM-based expansion with cost capture
elkaix Apr 27, 2026
35caf92
feat(eval): add RefusalHandler with similarity gate
elkaix Apr 27, 2026
1572ae3
feat(eval): wire Phase 2 levers into build_pipeline + EvalPipeline.query
elkaix Apr 27, 2026
4930996
feat(eval): add cli archive subcommand to copy small run artifacts
elkaix Apr 27, 2026
c24879a
chore(eval): add Phase 2 tier configs under configs/eval/phase2/
elkaix Apr 27, 2026
e4d243f
Merge pull request #5 from mohamed-elkholy95/feature/phase2-pipeline-…
elkaix Apr 27, 2026
a82132f
ci(github): add Actions workflow for backend pytest + frontend build/…
elkaix Apr 27, 2026
5cdcefe
ci(github): switch to uv for faster dependency resolution
elkaix Apr 27, 2026
f5e6cd2
ci(github): set dummy API keys so LLMHandler constructors don't raise
elkaix Apr 27, 2026
6644955
ci(github): stub openai client in CI via CI_LLM_MOCK env var
elkaix Apr 27, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
121 changes: 121 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,121 @@
name: CI

# Triggers: push to main, and PRs targeting main or any feature branch.
on:
push:
branches: [main]
pull_request:
branches: [main, "feature/**"]

# Cancel superseded runs on the same ref to save runner minutes.
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: true

# Default to read-only token; no write actions performed.
permissions:
contents: read

jobs:
backend-tests:
name: Backend tests (Python ${{ matrix.python-version }})
runs-on: ubuntu-latest
timeout-minutes: 20
strategy:
fail-fast: false
matrix:
python-version: ["3.12"]

steps:
- name: Checkout
uses: actions/checkout@v4

- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python-version }}

# WHY uv instead of pip: pip's resolver backtracks for minutes on the
# transitive cohere / logfire / opentelemetry-* deps pulled in by
# arize-phoenix. uv resolves the same set in seconds and ships its own
# caching that's faster than setup-python's pip cache for this graph.
- name: Install uv
uses: astral-sh/setup-uv@v6
with:
enable-cache: true
cache-dependency-glob: requirements.txt

- name: Install dependencies
run: |
uv pip install --system -r requirements.txt
uv pip install --system pytest-cov

# WHY cache HF models: BgeEmbedder (~120MB) and the ms-marco-MiniLM cross-encoder
# (~80MB) are downloaded on first use. Without this cache, every CI run re-downloads
# them, wasting ~30s and bandwidth. Cache key is invalidated when requirements.txt
# changes (which is when transformer versions might shift).
- name: Cache Hugging Face models
uses: actions/cache@v4
with:
path: |
~/.cache/huggingface
~/.cache/torch
key: hf-${{ runner.os }}-${{ hashFiles('requirements.txt') }}
restore-keys: |
hf-${{ runner.os }}-

# WHY dummy API keys + CI_LLM_MOCK: openai.OpenAI() raises at construction
# when OPENAI_API_KEY is unset; any non-empty string lets the client construct.
# CI_LLM_MOCK=true triggers the conftest fixture that monkeypatches
# openai.OpenAI to an in-process stub, so tests that exercise the real
# backend.query path don't actually hit the OpenAI API. Local dev is
# unaffected since CI_LLM_MOCK is unset there.
# OTLP traces still try to flush to localhost:6006 and log connection-refused
# warnings — expected and harmless in CI.
- name: Run pytest
env:
PYTHONDONTWRITEBYTECODE: "1"
CI_LLM_MOCK: "true"
OPENAI_API_KEY: dummy_for_ci
ANTHROPIC_API_KEY: dummy_for_ci
GLM_API_KEY: dummy_for_ci
run: |
python -m pytest tests/ -q \
--cov=src --cov-report=term-missing --cov-report=xml \
--maxfail=5 --durations=20

- name: Upload coverage report
if: always()
uses: actions/upload-artifact@v4
with:
name: coverage-py${{ matrix.python-version }}
path: coverage.xml
if-no-files-found: ignore

frontend-build:
name: Frontend build + lint
runs-on: ubuntu-latest
timeout-minutes: 10
defaults:
run:
working-directory: frontend

steps:
- name: Checkout
uses: actions/checkout@v4

- name: Set up Node.js
uses: actions/setup-node@v4
with:
node-version: "20"
cache: npm
cache-dependency-path: frontend/package-lock.json

- name: Install dependencies
run: npm ci

- name: Lint
run: npm run lint

- name: Build
run: npm run build
3 changes: 3 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -10,3 +10,6 @@ books/
.worktrees/
data/rag.db
data/chroma/

# Eval harness run outputs
eval_runs/
100 changes: 100 additions & 0 deletions Architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -443,6 +443,106 @@ Run: `python -m pytest tests/ -v`

---

## Evaluation Harness

The `src/eval/` package provides a reproducible evaluation system over labeled gold sets, separate from the user-facing chat path.

### Layers

| Module | Responsibility |
|--------|----------------|
| `src/eval/schemas.py` | Pydantic contracts: `EvalQuestion`, `EvalResult`, `AggregatedMetric`, `RunMetadata`, `MetricDelta`, `CompareResult`. |
| `src/eval/pricing.py` | Hard-coded model price table + `cost_usd()` helper. |
| `src/eval/statistics.py` | `bootstrap_ci()` and `paired_permutation_test()` for run-level confidence intervals and two-run significance testing. |
| `src/eval/metrics/retrieval.py` | Recall@k, MRR@k, nDCG@k over `(gold_chunk_ids, retrieved_chunk_ids)`. |
| `src/eval/metrics/operational.py` | Per-stage latency p50/p95/p99, cost, token aggregation. |
| `src/eval/metrics/refusal.py` | Regex + LLM-judge refusal correctness for unanswerable questions. |
| `src/eval/metrics/generation.py` | Wraps existing `src/evaluation.py` LLM-as-judge functions; adds `answer_correctness` (cosine + judge mean) and `context_recall`. |
| `src/eval/datasets/squad_v2.py` | Seeded sample + frozen 200-row JSONL artifact from HuggingFace `squad_v2`. |
| `src/eval/datasets/ml_papers.py` | Hand-labeled dev set loader + manifest SHA-256 verification. |
| `src/eval/config.py` | YAML-loaded `EvalConfig`. |
| `src/eval/storage.py` | Run-directory CRUD over `eval_runs/<run_id>/`. |
| `src/eval/pipeline_factory.py` + `src/eval/_telemetry.py` | Builds an isolated RAG pipeline per (config, dataset) using ephemeral Chroma. |
| `src/eval/aggregator.py` | Per-dataset + combined `AggregatedMetric` rows from per-question results. |
| `src/eval/runner.py` | Orchestrates `git_sha`, ingest, query+score loop, aggregation, persistence. |
| `src/eval/compare.py` | Two-run diff with paired permutation tests + per-question regressions/wins. |
| `src/eval/report.py` + `templates/eval/*.html.j2` | Standalone jinja2 HTML reports. |
| `src/eval/cli.py` | `run`/`list`/`show`/`compare` argparse subcommands. |

### API + UI

`src/api/routes/eval.py` exposes:

| Method | Path | Description |
|--------|------|-------------|
| `GET` | `/api/eval/configs` | List available eval configs |
| `POST` | `/api/eval/run` | Start a new eval run (dispatched via `BackgroundTasks`) |
| `GET` | `/api/eval/runs` | List all eval runs |
| `GET` | `/api/eval/runs/{id}` | Get run metadata |
| `GET` | `/api/eval/runs/{id}/results` | Per-question results |
| `GET` | `/api/eval/runs/{id}/status` | Live status for in-progress runs |
| `GET` | `/api/eval/compare` | Two-run diff with significance tests |

Long-running runs dispatch via FastAPI `BackgroundTasks` and report progress through an in-process `RunRegistry` (`src/api/services/eval_runs.py`).

React route `/eval/*` mounts three views:
- **`RunsList`** — sortable/filterable table with multi-select compare
- **`RunDetail`** — metric chart + per-question table with lazy expand
- **`CompareView`** — side-by-side bars + Top Wins / Top Regressions cards

Charts use `recharts` with CI whiskers.

### Eval Run Directory

Each run produces `eval_runs/<run_id>/` with:
- `metadata.json` — run ID, git SHA, config name, timestamps
- `questions.jsonl` — per-question scores and retrieved chunks
- `metrics.json` — aggregated metric values with bootstrap CIs
- `cost.json` — token counts and USD costs per model
- `config.yaml` — snapshot of the config used

The `eval_runs/` directory is gitignored; the labeled dev sets in `eval_data/` are checked in.

---

## Observability

The system exports per-stage spans for every chat query via OpenTelemetry to [Arize Phoenix](https://github.com/Arize-ai/phoenix) on `localhost:6006`.

### Spans

`RAGBackend.query_with_telemetry` and `RAGBackend.stream_query` open spans:

| Span | Attributes |
|------|------------|
| `rag.retrieve` | `top_k`, `chunk_count` |
| `rag.generate` | `model`, `prompt_tokens`, `completion_tokens`, `cost_usd` |

### Telemetry Payload

The same numbers are returned to the client as a `StageTelemetry` Pydantic model (`src/api/schemas/telemetry.py`):

- REST `POST /api/query` — includes a `telemetry` field in the response JSON.
- WebSocket `/api/chat` — emits a final `{"type": "telemetry", "content": {...}}` event after the existing `done` event.

The frontend renders these as a muted footer line under each assistant chat bubble:

> *Retrieve 142ms · Generate 2.1s · 4,217 tok · $0.0083*

with a hover tooltip showing the prompt/completion token split.

### Running with Traces

Phoenix is profile-gated in `docker-compose.yml`; bare `docker compose up` does not start it.

```bash
docker compose --profile observability up
```

`init_observability()` (`src/observability.py`) is called during the FastAPI lifespan startup. It is idempotent and fail-quiet — if Phoenix is unreachable, spans become no-ops and the chat continues to work normally.

---

## Key Design Decisions

| Decision | Choice | Rationale |
Expand Down
11 changes: 11 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -176,6 +176,17 @@ docker compose up --build
| API | [localhost:8001](http://localhost:8001) |
| Swagger Docs | [localhost:8001/docs](http://localhost:8001/docs) |

### Running with traces (Phoenix)

```bash
docker compose --profile observability up
```

Phoenix UI is available at http://localhost:6006. The FastAPI backend
will export per-stage spans (`rag.retrieve`, `rag.generate`) with token
counts and cost as span attributes. If Phoenix isn't running, the app
works normally — span export silently fails.

### Manual Setup

```bash
Expand Down
7 changes: 7 additions & 0 deletions configs/eval/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,7 @@
# Eval configs

YAML files in this directory define complete pipeline + eval-set configurations. Each file is loaded by `src.eval.config.load_config(path)` into an `EvalConfig` Pydantic model and consumed by `EvalRunner.run()`.

To author a new config, copy `baseline.yaml`, change the fields you want to vary, and rename. Run with `python -m src.eval.cli run --config configs/eval/<your-config>.yaml`.

See `docs/superpowers/specs/2026-04-26-rag-eval-harness-phase-1-design.md` §7 for the full schema.
18 changes: 18 additions & 0 deletions configs/eval/baseline.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
name: "baseline"
description: "Current production defaults: recursive chunker, gpt-5-mini answer model, gpt-4.1-nano reasoning model."
pipeline:
chunker:
strategy: "recursive"
chunk_size: 512
chunk_overlap: 64
retriever:
top_k: 5
generator:
model: "gpt-5-mini"
reasoning_model: "gpt-4.1-nano"
eval:
datasets: ["squad_v2_dev_200", "ml_papers_v1"]
judge_model: "gpt-4.1-mini"
bootstrap_n: 1000
permutation_n: 10000
seed: 42
18 changes: 18 additions & 0 deletions configs/eval/baseline_squad_only.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
name: "baseline_squad_only"
description: "Baseline pipeline against squad_v2_dev_200 only (ml_papers_v1 not yet labeled)."
pipeline:
chunker:
strategy: "recursive"
chunk_size: 512
chunk_overlap: 64
retriever:
top_k: 5
generator:
model: "gpt-5-mini"
reasoning_model: "gpt-4.1-nano"
eval:
datasets: ["squad_v2_dev_200"]
judge_model: "gpt-4.1-mini"
bootstrap_n: 1000
permutation_n: 10000
seed: 42
13 changes: 13 additions & 0 deletions configs/eval/phase2/phase2_baseline.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
name: "phase2_baseline"
description: "Baseline anchor for the Phase 2 matrix — identical to baseline_squad_only."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
retriever: {top_k: 5}
generator: {model: gpt-5-mini, reasoning_model: gpt-4.1-nano}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
14 changes: 14 additions & 0 deletions configs/eval/phase2/phase2b_embedder.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
name: "phase2b_embedder"
description: "Phase 2 tier 2b — swap default embedder for BAAI/bge-small-en-v1.5."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
generator: {model: gpt-5-mini, reasoning_model: gpt-4.1-nano}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
15 changes: 15 additions & 0 deletions configs/eval/phase2/phase2c_hybrid.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
name: "phase2c_hybrid"
description: "Phase 2 tier 2c — add BM25 hybrid on top of BGE."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
hybrid: {enabled: true, bm25_top_k: 20, dense_top_k: 20, rrf_k: 60}
generator: {model: gpt-5-mini, reasoning_model: gpt-4.1-nano}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
16 changes: 16 additions & 0 deletions configs/eval/phase2/phase2d_rerank.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
name: "phase2d_rerank"
description: "Phase 2 tier 2d — add cross-encoder rerank on top of hybrid."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
hybrid: {enabled: true, bm25_top_k: 20, dense_top_k: 20, rrf_k: 60}
reranker: {model: ms_marco_minilm_l6_v2, rerank_top_n: 20, final_top_k: 5}
generator: {model: gpt-5-mini, reasoning_model: gpt-4.1-nano}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
17 changes: 17 additions & 0 deletions configs/eval/phase2/phase2e_rewrite.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
name: "phase2e_rewrite"
description: "Phase 2 tier 2e — add LLM query rewriting on top of rerank."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
hybrid: {enabled: true, bm25_top_k: 20, dense_top_k: 20, rrf_k: 60}
reranker: {model: ms_marco_minilm_l6_v2, rerank_top_n: 20, final_top_k: 5}
query_rewriter: {model: gpt-4.1-nano, max_expansions: 3}
generator: {model: gpt-5-mini, reasoning_model: gpt-4.1-nano}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
18 changes: 18 additions & 0 deletions configs/eval/phase2/phase2f_models_gpt41mini.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,18 @@
name: "phase2f_models_gpt41mini"
description: "Phase 2 tier 2f — answer-model comparison: gpt-4.1-mini variant."
pipeline:
chunker: {strategy: recursive, chunk_size: 512, chunk_overlap: 64}
embedder: {name: bge_small_en_v1_5}
retriever: {top_k: 5}
hybrid: {enabled: true, bm25_top_k: 20, dense_top_k: 20, rrf_k: 60}
reranker: {model: ms_marco_minilm_l6_v2, rerank_top_n: 20, final_top_k: 5}
query_rewriter: {model: gpt-4.1-nano, max_expansions: 3}
generator: {model: gpt-4.1-mini, reasoning_model: gpt-4.1-nano}
refusal_handler: {enabled: true, similarity_threshold: 0.35, no_answer_text: "I don't have enough information to answer that."}
eval:
datasets: [squad_v2_dev_200]
judge_model: gpt-4.1-mini
bootstrap_n: 1000
permutation_n: 10000
seed: 42
spend_ceiling_usd: 1.5
Loading
Loading