Measure whether a RAG system is actually grounded — and fail the build when it stops being.
Most RAG projects ship on vibes. Someone asks the bot five questions, the answers read well, and it goes live. Then a retriever change quietly drops recall, the model starts filling the gap from its own priors, and nobody finds out until a customer is told something that was never in the knowledge base.
This harness turns that into a number you can gate a pull request on.
judge=lexical k=5 examples=8
hit_rate 0.875
precision 0.175
recall 0.875
mrr 0.875
ndcg 0.875
groundedness 0.625
hallucination_rate 0.375
citation_accuracy 0.714
citation_coverage 0.833
context_utilisation 0.562
answer_relevance 0.604
answer_correctness 0.510
gate: FAIL (groundedness, hallucination_rate)
Exit code 1. The build stops.
Three things this does differently.
It runs with no API key and no dependencies. The default judge is lexical and deterministic, so the full suite runs in under a second on every commit for zero cost. An LLM judge is one flag away when you need semantic depth — but a quality gate you can only afford to run nightly is not a quality gate.
Undefined is not zero. If an example has no labelled relevant documents, its retrieval metrics are reported as undefined and excluded from the mean — not silently scored 0.0. Tools that conflate the two make a partially-labelled dataset look like a broken retriever, and the resulting number quietly poisons every decision made from it. Each metric reports how many examples it was actually observed on.
It tells you which sentence broke. An aggregate is a mood; a named sentence is a bug report. Every report ends with the weakest examples, the exact claims that had no support, the relevant documents that were never retrieved, and the contexts you paid for but never used.
### `q-003` - groundedness 0.000
> What warranty comes with electronics?
- **Unsupported:** Electronics come with a two year manufacturer warranty.
- **Relevant but not retrieved:** policy-warranty
- **Retrieved but unused:** policy-returns, policy-shippingpip install -e . # core: standard library only
pip install -e ".[claude]" # adds the optional Claude judgerag-eval run examples/support_bot.jsonl \
--gate groundedness=0.85 \
--gate hallucination_rate=0.10 \
--markdown report.md--gate reads as a minimum, except hallucination_rate, which is a maximum. Repeat it as often as you like.
As a library:
from rageval import Context, Example, LexicalJudge, EvalConfig, evaluate
example = Example(
id="q1",
question="How long is the returns window?",
contexts=(Context("policy", "Items may be returned within 30 days."),),
answer="You can return items within 30 days.",
relevant_ids=frozenset({"policy"}),
citations=("policy",),
)
report = evaluate([example], LexicalJudge(), EvalConfig(k=1))
print(report.aggregates["groundedness"]) # 1.0JSONL, one example per line. Only id and question are required — every metric that depends on a missing field reports as undefined rather than zero, so a dataset with no labels still gives you usable generation metrics.
{
"id": "q-001",
"question": "How long do I have to return an item?",
"contexts": [
{"id": "policy-returns", "text": "Returns are accepted within 30 days of delivery.", "score": 0.94},
{"id": "policy-shipping", "text": "Standard delivery takes 3 to 5 business days.", "score": 0.41}
],
"answer": "Returns are accepted within 30 days of delivery.",
"relevant_ids": ["policy-returns"],
"citations": ["policy-returns"],
"ground_truth": "Within 30 days of delivery."
}contexts must be in rank order — the ordering is what the ranking metrics measure. Parse errors name the offending line.
| Metric | Question it answers |
|---|---|
hit_rate |
Did anything relevant make the top k at all? |
precision |
How much of the top k was worth reading? |
recall |
How much of what mattered did we find? |
mrr |
How far down was the first good hit? |
ndcg |
Was the good material near the top, where the model will actually read it? |
precision divides by k, matching the standard IR definition, so a short result list caps the achievable score. That is the intended signal.
| Metric | Question it answers |
|---|---|
groundedness |
What fraction of the answer's claims does the context actually support? |
hallucination_rate |
The complement — the number that goes in front of a risk committee. |
citation_accuracy |
Of the sources cited, how many genuinely back something up? |
citation_coverage |
Of the sources that did the work, how many got credited? |
context_utilisation |
How much retrieved context was actually used? Persistently low means k is too high. |
answer_relevance |
Does the answer address the question that was asked? |
answer_correctness |
Unigram F1 against a reference answer. |
Answers are decomposed into sentence-level claims and each is judged independently. That decomposition is what makes the output actionable.
Two pairs are worth reading together. Groundedness with citation accuracy separates "made it up" from "got it right but pointed at the wrong document" — a distinction that matters enormously when a user clicks through to verify. Citation accuracy with coverage separates over-citing from under-citing.
LexicalJudge (default) |
ClaudeJudge |
|
|---|---|---|
| Cost | Free | Per call |
| Speed | ~1s for the suite | Network-bound |
| Deterministic | Yes | No |
| Catches paraphrase | No | Yes |
| Catches contradiction | Only lexically | Yes |
LexicalJudge supports a claim when enough of its content words appear in one context. Requiring a single context to carry the whole claim is deliberate: support stitched together from fragments that never co-occurred is precisely how fluent fabrication passes review.
ClaudeJudge sends each claim to Claude under a constrained JSON schema, caches verdicts in-process so re-runs don't re-bill unchanged claims, and discards any supporting id the model invents — a judge citing a context it was never given would corrupt the citation metrics downstream.
export ANTHROPIC_API_KEY=...
rag-eval run dataset.jsonl --judge claudeBoth satisfy the same Judge protocol, so bringing your own is a class with two methods.
- name: RAG quality gate
run: |
rag-eval run eval/golden.jsonl \
--gate groundedness=0.85 \
--gate citation_accuracy=0.90 \
--gate hallucination_rate=0.10 \
--markdown report.mdExit codes: 0 every gate held, 1 a gate was breached, 2 the run could not complete.
A metric with no observations fails its gate rather than passing it. An empty gate is not a passing gate — if you asked for citation accuracy above 0.9 and nothing in the dataset cited anything, that is a problem with the run, not a pass.
LexicalJudgecannot detect a correct paraphrase, and will accept a fluent sentence that reuses context vocabulary while inverting its meaning. It is a fast tripwire, not a semantic oracle. UseClaudeJudgewhen the cost of a wrong answer exceeds the cost of a call.answer_correctnessis lexical F1, so a correct paraphrase scores low. Track it as a regression tripwire across runs, not as an absolute quality bar.- Claim decomposition is sentence-based. A sentence carrying two independent facts is judged as one unit.
- Metrics are computed per example and averaged unweighted, so a golden set that over-represents one question type will skew the aggregate. Curate deliberately.
pip install -e ".[dev]"
pytest120 tests, no network access required.
MIT