Skip to content

Latest commit

 

History

9 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

rag-eval-harness

CI Python 3.10+ License: MIT

Measure whether a RAG system is actually grounded — and fail the build when it stops being.

Most RAG projects ship on vibes. Someone asks the bot five questions, the answers read well, and it goes live. Then a retriever change quietly drops recall, the model starts filling the gap from its own priors, and nobody finds out until a customer is told something that was never in the knowledge base.

This harness turns that into a number you can gate a pull request on.

judge=lexical  k=5  examples=8
  hit_rate             0.875
  precision            0.175
  recall               0.875
  mrr                  0.875
  ndcg                 0.875
  groundedness         0.625
  hallucination_rate   0.375
  citation_accuracy    0.714
  citation_coverage    0.833
  context_utilisation  0.562
  answer_relevance     0.604
  answer_correctness   0.510
gate: FAIL (groundedness, hallucination_rate)

Exit code 1. The build stops.


Why another eval tool

Three things this does differently.

It runs with no API key and no dependencies. The default judge is lexical and deterministic, so the full suite runs in under a second on every commit for zero cost. An LLM judge is one flag away when you need semantic depth — but a quality gate you can only afford to run nightly is not a quality gate.

Undefined is not zero. If an example has no labelled relevant documents, its retrieval metrics are reported as undefined and excluded from the mean — not silently scored 0.0. Tools that conflate the two make a partially-labelled dataset look like a broken retriever, and the resulting number quietly poisons every decision made from it. Each metric reports how many examples it was actually observed on.

It tells you which sentence broke. An aggregate is a mood; a named sentence is a bug report. Every report ends with the weakest examples, the exact claims that had no support, the relevant documents that were never retrieved, and the contexts you paid for but never used.

### `q-003` - groundedness 0.000

> What warranty comes with electronics?

- **Unsupported:** Electronics come with a two year manufacturer warranty.
- **Relevant but not retrieved:** policy-warranty
- **Retrieved but unused:** policy-returns, policy-shipping

Install

pip install -e .           # core: standard library only
pip install -e ".[claude]" # adds the optional Claude judge

Use it

rag-eval run examples/support_bot.jsonl \
  --gate groundedness=0.85 \
  --gate hallucination_rate=0.10 \
  --markdown report.md

--gate reads as a minimum, except hallucination_rate, which is a maximum. Repeat it as often as you like.

As a library:

from rageval import Context, Example, LexicalJudge, EvalConfig, evaluate

example = Example(
    id="q1",
    question="How long is the returns window?",
    contexts=(Context("policy", "Items may be returned within 30 days."),),
    answer="You can return items within 30 days.",
    relevant_ids=frozenset({"policy"}),
    citations=("policy",),
)

report = evaluate([example], LexicalJudge(), EvalConfig(k=1))
print(report.aggregates["groundedness"])  # 1.0

Dataset format

JSONL, one example per line. Only id and question are required — every metric that depends on a missing field reports as undefined rather than zero, so a dataset with no labels still gives you usable generation metrics.

{
  "id": "q-001",
  "question": "How long do I have to return an item?",
  "contexts": [
    {"id": "policy-returns", "text": "Returns are accepted within 30 days of delivery.", "score": 0.94},
    {"id": "policy-shipping", "text": "Standard delivery takes 3 to 5 business days.", "score": 0.41}
  ],
  "answer": "Returns are accepted within 30 days of delivery.",
  "relevant_ids": ["policy-returns"],
  "citations": ["policy-returns"],
  "ground_truth": "Within 30 days of delivery."
}

contexts must be in rank order — the ordering is what the ranking metrics measure. Parse errors name the offending line.


The metrics

Retrieval

Metric Question it answers
hit_rate Did anything relevant make the top k at all?
precision How much of the top k was worth reading?
recall How much of what mattered did we find?
mrr How far down was the first good hit?
ndcg Was the good material near the top, where the model will actually read it?

precision divides by k, matching the standard IR definition, so a short result list caps the achievable score. That is the intended signal.

Generation

Metric Question it answers
groundedness What fraction of the answer's claims does the context actually support?
hallucination_rate The complement — the number that goes in front of a risk committee.
citation_accuracy Of the sources cited, how many genuinely back something up?
citation_coverage Of the sources that did the work, how many got credited?
context_utilisation How much retrieved context was actually used? Persistently low means k is too high.
answer_relevance Does the answer address the question that was asked?
answer_correctness Unigram F1 against a reference answer.

Answers are decomposed into sentence-level claims and each is judged independently. That decomposition is what makes the output actionable.

Two pairs are worth reading together. Groundedness with citation accuracy separates "made it up" from "got it right but pointed at the wrong document" — a distinction that matters enormously when a user clicks through to verify. Citation accuracy with coverage separates over-citing from under-citing.

Judges

LexicalJudge (default) ClaudeJudge
Cost Free Per call
Speed ~1s for the suite Network-bound
Deterministic Yes No
Catches paraphrase No Yes
Catches contradiction Only lexically Yes

LexicalJudge supports a claim when enough of its content words appear in one context. Requiring a single context to carry the whole claim is deliberate: support stitched together from fragments that never co-occurred is precisely how fluent fabrication passes review.

ClaudeJudge sends each claim to Claude under a constrained JSON schema, caches verdicts in-process so re-runs don't re-bill unchanged claims, and discards any supporting id the model invents — a judge citing a context it was never given would corrupt the citation metrics downstream.

export ANTHROPIC_API_KEY=...
rag-eval run dataset.jsonl --judge claude

Both satisfy the same Judge protocol, so bringing your own is a class with two methods.


In CI

- name: RAG quality gate
  run: |
    rag-eval run eval/golden.jsonl \
      --gate groundedness=0.85 \
      --gate citation_accuracy=0.90 \
      --gate hallucination_rate=0.10 \
      --markdown report.md

Exit codes: 0 every gate held, 1 a gate was breached, 2 the run could not complete.

A metric with no observations fails its gate rather than passing it. An empty gate is not a passing gate — if you asked for citation accuracy above 0.9 and nothing in the dataset cited anything, that is a problem with the run, not a pass.


Honest limitations

  • LexicalJudge cannot detect a correct paraphrase, and will accept a fluent sentence that reuses context vocabulary while inverting its meaning. It is a fast tripwire, not a semantic oracle. Use ClaudeJudge when the cost of a wrong answer exceeds the cost of a call.
  • answer_correctness is lexical F1, so a correct paraphrase scores low. Track it as a regression tripwire across runs, not as an absolute quality bar.
  • Claim decomposition is sentence-based. A sentence carrying two independent facts is judged as one unit.
  • Metrics are computed per example and averaged unweighted, so a golden set that over-represents one question type will skew the aggregate. Curate deliberately.

Development

pip install -e ".[dev]"
pytest

120 tests, no network access required.

License

MIT

About

Measure whether a RAG system is actually grounded, and fail the build when it stops being: retrieval metrics, claim-level groundedness, hallucination rate and citation accuracy.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages