Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 16 additions & 0 deletions .changeset/enhance-metrics-judge-gate.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
---
"@reaatech/rag-eval-core": minor
"@reaatech/rag-eval-metrics": minor
"@reaatech/rag-eval-judge": minor
"@reaatech/rag-eval-gate": minor
"@reaatech/rag-eval-cli": patch
---

Enhance metric honesty, judge flexibility, and gate robustness ahead of first publish.

- **metrics**: extract the duplicated stop-word/stemmer/tokenizer logic into a shared `text-utils` module and drop the unused `compromise` and `natural` dependencies.
- **metrics/core**: rename the relevance scorer's lexical output to `lexical_similarity` (it was misleadingly called `semantic_similarity`). `semantic_similarity` is now only populated when an `EmbeddingProvider` is supplied, enabling true (paraphrase-aware) semantic scoring on `RelevanceScorer` and the new `AnswerCorrectnessScorer`.
- **metrics/core**: add `RetrievalScorer` (MRR, nDCG, precision/recall/hit@k from `retrieved_chunk_ids` vs. `relevant_chunk_ids`) and `AnswerCorrectnessScorer` (generated answer vs. ground truth).
- **judge/core**: support explicit `provider`, `base_url`, and `api_key` in `JudgeConfig` so OpenAI-compatible gateways, proxies, and self-hosted/local models work without relying on model-name keyword inference.
- **judge**: confidence is now the judge's self-reported certainty (parsed from the response) or, for consensus, derived from inter-judge agreement — instead of the previous circular distance-from-0.5 heuristic.
- **gate/core**: baseline-comparison gates gain a `tolerance` band so sampling noise doesn't trip CI, and all gates gain a `warn` vs. `fail` `severity`. `GateResult` now carries a `warnings` array; CI reports surface warnings without failing the build.
22 changes: 17 additions & 5 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -57,11 +57,13 @@ core ← metrics ← suite ← mcp-server, cli
| **Types & Schemas** | `@reaatech/rag-eval-core` | Domain types, Zod schemas |
| **Faithfulness Scorer** | `@reaatech/rag-eval-metrics` | Measure answer grounding in context |
| **Relevance Scorer** | `@reaatech/rag-eval-metrics` | Measure answer relevance to query |
| **Context Precision** | `@reaatech/rag-eval-metrics` | Measure retrieval ranking quality |
| **Context Precision** | `@reaatech/rag-eval-metrics` | Measure retrieval ranking quality (text-based) |
| **Context Recall** | `@reaatech/rag-eval-metrics` | Measure ground truth coverage |
| **LLM Judge** | `@reaatech/rag-eval-judge` | Calibrated quality scoring with multi-provider support |
| **Retrieval Scorer** | `@reaatech/rag-eval-metrics` | Ranking metrics (MRR, nDCG, precision/recall/hit@k) from chunk IDs |
| **Answer Correctness** | `@reaatech/rag-eval-metrics` | Measure generated answer vs. ground truth |
| **LLM Judge** | `@reaatech/rag-eval-judge` | Calibrated quality scoring across any provider or OpenAI-compatible gateway |
| **Cost Tracker** | `@reaatech/rag-eval-cost` | Per-evaluation cost calculation and budget enforcement |
| **Gate Engine** | `@reaatech/rag-eval-gate` | CI regression gates and threshold checks |
| **Gate Engine** | `@reaatech/rag-eval-gate` | CI regression gates with tolerance bands and warn/fail severity |
| **Dataset Manager** | `@reaatech/rag-eval-dataset` | Dataset loading, validation, generation |
| **Observability** | `@reaatech/rag-eval-observability` | Structured logging, OTel tracing, metrics |
| **Evaluation Suite** | `@reaatech/rag-eval-suite` | Central orchestrator tying all modules together |
Expand Down Expand Up @@ -126,7 +128,7 @@ Fast, stateless, composable operations for mid-task self-evaluation:
| Tool | Input | Output | Use Case |
|------|-------|--------|----------|
| `rag_eval.judge.faithfulness` | `{ context, generated_answer }` | `{ score, statements, supported_count }` | Check if answer is faithful to context |
| `rag_eval.judge.relevance` | `{ query, generated_answer }` | `{ score, semantic_similarity, intent_score }` | Check if answer addresses query |
| `rag_eval.judge.relevance` | `{ query, generated_answer }` | `{ score, lexical_similarity, intent_score }` | Check if answer addresses query (semantic_similarity added when an embedding provider is configured) |
| `rag_eval.judge.context_precision` | `{ query, context[], ground_truth }` | `{ score, map, ndcg }` | Check context ranking quality |
| `rag_eval.judge.context_recall` | `{ query, context[], ground_truth }` | `{ score, total_facts, covered_facts }` | Check ground truth coverage |
| `rag_eval.judge.cost_check` | `{ eval_result, budget }` | `{ within_budget, cost }` | Verify cost within budget |
Expand Down Expand Up @@ -230,6 +232,13 @@ judge:
# Primary judge model (any provider)
model: claude-opus

# Explicit provider/endpoint — required for OpenAI-compatible gateways,
# proxies, or self-hosted/local models whose names lack a provider keyword.
# Omit to infer the provider from the model name.
# provider: openai
# base_url: http://localhost:11434/v1
# api_key: ${OPENAI_API_KEY}

# Fallback models for resilience
fallback_models:
- gpt-4-turbo
Expand Down Expand Up @@ -433,11 +442,13 @@ gates:
operator: ">="
threshold: 0.85

# severity: warn records a non-blocking warning instead of failing the run
- name: min-relevance
type: threshold
metric: avg_relevance
operator: ">="
threshold: 0.80
severity: warn

- name: min-context-precision
type: threshold
Expand All @@ -461,6 +472,8 @@ gates:
type: baseline-comparison
metric: overall_score
allow_regression: false
# tolerance absorbs sampling noise so small dips don't fail CI
tolerance: 0.01
```

---
Expand Down Expand Up @@ -678,7 +691,6 @@ Before deploying a RAG evaluation pipeline to production:
## References

- **ARCHITECTURE.md** — System design deep dive and package relationships
- **DEV_PLAN.md** — Development checklist
- **README.md** — Quick start and overview
- **datasets/examples/** — Example evaluation datasets
- **MCP Specification** — https://modelcontextprotocol.io/
57 changes: 46 additions & 11 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,7 +77,7 @@ core ← metrics ← suite ← mcp-server, cli
| Package | Role | Depends On | Key Exports |
|---------|------|------------|-------------|
| `@reaatech/rag-eval-core` | Foundation types + Zod schemas | (leaf) | `EvaluationSample`, `EvalSuiteConfig`, `GateConfig`, `JudgeConfig`, `CostBreakdown`, schemas |
| `@reaatech/rag-eval-metrics` | Heuristic metric scorers | core | `FaithfulnessScorer`, `RelevanceScorer`, `ContextPrecisionScorer`, `ContextRecallScorer`, `MetricsEngine` |
| `@reaatech/rag-eval-metrics` | Metric scorers (lexical by default; semantic via an `EmbeddingProvider`) | core | `FaithfulnessScorer`, `RelevanceScorer`, `ContextPrecisionScorer`, `ContextRecallScorer`, `RetrievalScorer`, `AnswerCorrectnessScorer`, `MetricsEngine`, `text-utils` |
| `@reaatech/rag-eval-cost` | Cost tracking infrastructure | core | `CostTracker`, `Pricing`, `BudgetManager`, `CostReporter` |
| `@reaatech/rag-eval-judge` | LLM-as-judge | core, cost | `JudgeEngine`, `JudgeCalibrator`, `JudgeCostTracker`, prompts |
| `@reaatech/rag-eval-gate` | Quality gates | core | `GateEngine`, `ThresholdGates`, `BaselineGates`, `CIIntegration` |
Expand Down Expand Up @@ -207,17 +207,19 @@ core ← metrics ← suite ← mcp-server, cli
│ Input: { query, generated_answer } │
│ │
│ ┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐ │
│ │ Semantic │ │ Intent │ │ Score │ │
│ │ Similarity │ │ Coverage │ │ Aggregation │ │
│ │ │ │ │ │ │ │
│ │ - Embedding- │ │ - Decompose │ │ - Weighted │ │
│ │ based │ │ query into │ │ combination │ │
│ │ similarity │ │ intents │ │ - Semantic │ │
│ │ (cosine) │ │ - Check answer │ │ similarity + │ │
│ │ │ │ coverage │ │ intent score │ │
│ │ Similarity │ │ Intent │ │ Score │ │
│ │ │ │ Coverage │ │ Aggregation │ │
│ │ - Lexical │ │ │ │ │ │
│ │ (word/bigram) │ │ - Decompose │ │ - Weighted │ │
│ │ by default │ │ query into │ │ combination │ │
│ │ - Semantic │ │ intents │ │ - Similarity + │ │
│ │ (cosine) when │ │ - Check answer │ │ intent score │ │
│ │ embeddings │ │ coverage │ │ │ │
│ │ supplied │ │ │ │ │ │
│ └─────────────────┘ └─────────────────┘ └─────────────────┘ │
│ │
│ Output: { score, semantic_similarity, intent_score, intents } │
│ Output: { score, lexical_similarity, intent_score, │
│ semantic_similarity? } │
└─────────────────────────────────────────────────────────────────────┘
```

Expand Down Expand Up @@ -268,6 +270,40 @@ core ← metrics ← suite ← mcp-server, cli
└─────────────────────────────────────────────────────────────────────┘
```

### Retrieval Scorer

```
┌─────────────────────────────────────────────────────────────────────┐
│ Retrieval Scorer │
│ Package: @reaatech/rag-eval-metrics │
│ │
│ Input: { retrieved_chunk_ids[], relevant_chunk_ids[], k? } │
│ │
│ Ranking metrics over the retrieved order vs. the relevant set: │
│ - MRR (reciprocal rank of first relevant chunk) │
│ - nDCG (binary-relevance, log-discounted) │
│ - precision@k, recall@k, hit@k │
│ │
│ Output: { mrr, ndcg, precision_at_k, recall_at_k, hit_at_k, k } │
└─────────────────────────────────────────────────────────────────────┘
```

### Answer Correctness Scorer

```
┌─────────────────────────────────────────────────────────────────────┐
│ Answer Correctness Scorer │
│ Package: @reaatech/rag-eval-metrics │
│ │
│ Input: { generated_answer, ground_truth } │
│ │
│ - Lexical: token F1 + character bigram Dice (default) │
│ - Semantic: cosine of embeddings when an EmbeddingProvider is set │
│ │
│ Output: { score, lexical_similarity, semantic_similarity? } │
└─────────────────────────────────────────────────────────────────────┘
```

### LLM Judge with Calibration

```
Expand Down Expand Up @@ -494,7 +530,6 @@ All logs are structured JSON with standard fields:
## References

- **AGENTS.md** — Agent development guide
- **DEV_PLAN.md** — Development checklist
- **README.md** — Quick start and overview
- **datasets/examples/** — Example evaluation datasets
- **MCP Specification** — https://modelcontextprotocol.io/
39 changes: 34 additions & 5 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,10 +10,11 @@ This monorepo provides a composable suite of packages for evaluating Retrieval-A

## Features

- **Four evaluation metrics** — faithfulness, relevance, context precision, and context recall with heuristic scorers
- **LLM-as-judge** — multi-provider judging (Anthropic, OpenAI, Google) with calibration and consensus voting
- **Generation metrics** — faithfulness, relevance, and answer-correctness with fast lexical scorers; supply an `EmbeddingProvider` for paraphrase-aware semantic scoring
- **Retrieval metrics** — context precision/recall plus ranking metrics (MRR, nDCG, precision/recall/hit@k) from retrieved vs. relevant chunk IDs
- **LLM-as-judge** — multi-provider judging (Anthropic, OpenAI, Google, and any OpenAI-compatible gateway or local model) with calibration, consensus voting, and agreement-based confidence
- **Cost accounting** — per-sample and per-run token tracking with budget enforcement and alert thresholds
- **Quality gates** — threshold and baseline-comparison gates with formatted CI output and exit codes
- **Quality gates** — threshold and baseline-comparison gates with noise-tolerance bands, `warn`/`fail` severity, and formatted CI output with exit codes
- **MCP server** — three-layer tool API (`judge.*`, `suite.*`, `gate.*`) for agent-driven evaluation
- **Dataset management** — multi-format loading, Zod validation, synthetic generation, and version tracking
- **Observability** — structured Pino logging, OpenTelemetry tracing, and Prometheus-compatible metrics
Expand Down Expand Up @@ -111,12 +112,41 @@ rag-eval-pack report --results results.json --output report.md

See [`datasets/examples/`](./datasets/examples/) for sample datasets and configuration files.

## Scoring modes

Each generation metric can run at three levels of fidelity and cost — pick per use case:

| Mode | How it works | Cost | Catches paraphrase? |
| ---- | ------------ | ---- | ------------------- |
| **Lexical** (default) | Word/character overlap + intent heuristics. Deterministic, no network, no key. | Free | No — surface form only |
| **Semantic** | Cosine similarity of embeddings via an `EmbeddingProvider` you supply. | Embedding API cost | Yes |
| **LLM judge** | An LLM rates the sample with calibration and optional consensus. | Token cost | Yes, with reasoning |

Lexical scoring is reported as `lexical_similarity`; `semantic_similarity` is populated only when an embedding provider is configured. Use lexical for fast pre-commit smoke checks, semantic for paraphrase-sensitive metrics, and the judge for the final quality bar.

```typescript
import { RelevanceScorer } from "@reaatech/rag-eval-metrics";

const scorer = new RelevanceScorer({
embeddingProvider: { embed: async (texts) => myEmbedAPI(texts) },
});
```

Point the judge at any OpenAI-compatible gateway or local model via explicit config:

```typescript
const suite = new EvaluationSuite({
metrics: ["faithfulness", "relevance"],
judge: { provider: "openai", base_url: "http://localhost:11434/v1", model: "llama3.1" },
});
```

## Packages

| Package | Description |
| ------- | ----------- |
| [`@reaatech/rag-eval-core`](./packages/core) | Canonical types, Zod schemas, and domain models |
| [`@reaatech/rag-eval-metrics`](./packages/metrics) | Heuristic metric scorers (faithfulness, relevance, precision, recall) |
| [`@reaatech/rag-eval-metrics`](./packages/metrics) | Metric scorers: faithfulness, relevance, context precision/recall, retrieval ranking, and answer correctness |
| [`@reaatech/rag-eval-judge`](./packages/judge) | LLM-as-judge with calibration, consensus, and cost tracking |
| [`@reaatech/rag-eval-cost`](./packages/cost) | Pricing, budgeting, and cost reporting |
| [`@reaatech/rag-eval-gate`](./packages/gate) | Quality gates and CI regression checks |
Expand All @@ -131,7 +161,6 @@ See [`datasets/examples/`](./datasets/examples/) for sample datasets and configu
- [`ARCHITECTURE.md`](./ARCHITECTURE.md) — System design, package relationships, and data flows
- [`AGENTS.md`](./AGENTS.md) — Coding conventions, tool architecture, and development guidelines
- [`CONTRIBUTING.md`](./CONTRIBUTING.md) — Contribution workflow and release process
- [`DEV_PLAN.md`](./DEV_PLAN.md) — Development checklist and roadmap

## License

Expand Down
8 changes: 7 additions & 1 deletion packages/cli/src/commands/report.command.ts
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,13 @@ export function createReportCommand(): Command {
report = generateBasicMarkdownReport(results);
}
} else if (options.format === 'junit') {
const gateResult = results.gate_result || { passed: true, gates: [], failures: [] };
const gateResult = results.gate_result || {
passed: true,
gates: [],
failures: [],
warnings: [],
evaluated_at: new Date().toISOString(),
};
report = ci.generateJUnitXml(gateResult, results);
} else {
report = JSON.stringify(results, null, 2);
Expand Down
Loading
Loading