📖 Read the write-up: nogrounds
A RAG groundedness / hallucination checker. Give it an LLM answer and the context that answer was supposed to be based on; nogrounds segments the answer into atomic claims deterministically and offline, then uses an LLM judge to decide, claim by claim, whether each one is actually supported by the context. It reports a groundedness score and flags the unsupported claims — the ones the model most likely made up.
The segmentation — the part that decides what gets checked — runs with no key and no network. The judging step is an LLM call you opt into with your own key.
$ nogrounds answer.json
⏚ nogrounds — groundedness check
2/3 claims grounded · threshold 100%
✓ grounded claim 1
Redis keeps data in memory.
why: context: "Redis is an in-memory data store."
✓ grounded claim 2
It supports strings, lists, and sorted sets.
why: context lists strings, lists and sorted sets.
✗ unsupported NG001 claim 3
Redis ships with a built-in SQL query engine.
why: nothing in the context mentions SQL.
groundedness score: 67%
1 unsupported claim(s) (NG001) — likely hallucination.
A retrieval-augmented answer is only trustworthy if it stays on its sources. When a model quietly adds a fact that isn't in the retrieved context, that's the hallucination that erodes trust — and it's invisible unless you check the answer against the context. nogrounds does exactly that check: it isolates each claim and asks whether the provided context backs it up, so an unsupported claim fails loudly (non-zero exit) instead of shipping. Point it at your eval set or wire it into CI.
pip install nogroundsRequires Python 3.9+. The only runtime dependency is rich.
# a JSON file: {"answer": ..., "context": ...}
nogrounds pair.json
# explicit files
nogrounds --answer answer.txt --context context.txt
# pipe the JSON on stdin
cat pair.json | nogrounds -
# offline: show the claim segmentation + judge prompt, NO network call
nogrounds pair.json --dry-run
# relax the passing bar to 80% grounded
nogrounds pair.json --threshold 0.8| Flag | Meaning |
|---|---|
[input.json|-] |
JSON {answer, context} file, or - / omitted for stdin |
--answer FILE |
Answer text file (use with --context) |
--context FILE |
Source context text file (use with --answer) |
--threshold FLOAT |
Minimum groundedness score to pass, 0.0–1.0 (default: 1.0) |
--provider |
LLM judge: openai (default) · anthropic · mock |
--model |
Model override (else provider default or $NOGROUNDS_MODEL) |
--api-base |
Override the provider base URL (also $OPENAI_BASE_URL / $ANTHROPIC_BASE_URL) |
--dry-run |
Offline only: print segmentation + judge prompt with no network call |
--json |
Machine-readable JSON output |
--no-color |
Disable colored output |
| id | Meaning | Severity |
|---|---|---|
NG001 |
An unsupported claim — the context does not support it (likely hallucination) | blocker |
NG002 |
A partial / weakly-supported claim | warn |
| Code | Meaning |
|---|---|
0 |
Fully grounded — score ≥ threshold and no unsupported claim |
1 |
An unsupported claim (NG001), or the score is below --threshold |
2 |
Usage error — bad arguments, missing/empty input, or an unparseable judge reply |
This makes nogrounds usable as a CI / eval gate: fail the build when a RAG answer drifts off its provided sources.
nogrounds never ships or uses a key of its own. It reads your key from the environment:
export OPENAI_API_KEY=sk-... # or ANTHROPIC_API_KEY for --provider anthropic
nogrounds pair.jsonThe providers speak the standard OpenAI Chat Completions / Anthropic Messages wire
format over plain HTTP (stdlib urllib, no SDK), so you can point them at any
compatible endpoint — including a local mock — to develop and test without
spending a token:
nogrounds pair.json --api-base http://127.0.0.1:8080/v1There is also a built-in --provider mock that returns a canned all-grounded
response with no network at all — handy for demos and the test suite. See
docs/PROVIDERS.md for the expected judge JSON shape.
- The judge is an LLM, and LLMs are fallible. nogrounds uses an LLM as the judge, so its verdicts can themselves be wrong. Treat the score as a strong signal, not a proof; spot-check the flagged claims.
- Claim segmentation is heuristic. The answer is split into claims with sentence/list heuristics, not a linguistic parser. Odd punctuation or deeply nested clauses can over- or under-split.
- It only checks support against the PROVIDED context. nogrounds asks "does this context support this claim?" — not "is this claim true in the real world?" and not "is this answer correct?". A claim can be perfectly grounded in a context that is itself wrong, and a true claim can be flagged if the context doesn't happen to state it.
- You bring the key. The judging step requires your own provider key; without
one, use
--dry-runfor the offline segmentation and prompt.
These are separate tools that check different things:
- nogrounds (this tool) — checks answer ↔ context faithfulness: is the generated answer actually supported by the context it was given?
- ragmoat — checks tenant isolation in a RAG pipeline (that one tenant's retrieval can't leak another's documents).
- chunkvet — checks chunk quality of the indexed corpus (how the source is split before retrieval).
- askance — guards an LLM's output flowing into a code sink (injection into an execution/eval path).
nogrounds is the faithfulness check at the end of the pipeline — after retrieval, after generation — on the answer itself.
MIT © 2026 Jay Tank