Skip to content

Repository files navigation

nogrounds

📖 Read the write-up: nogrounds

A RAG groundedness / hallucination checker. Give it an LLM answer and the context that answer was supposed to be based on; nogrounds segments the answer into atomic claims deterministically and offline, then uses an LLM judge to decide, claim by claim, whether each one is actually supported by the context. It reports a groundedness score and flags the unsupported claims — the ones the model most likely made up.

The segmentation — the part that decides what gets checked — runs with no key and no network. The judging step is an LLM call you opt into with your own key.

$ nogrounds answer.json

⏚  nogrounds — groundedness check
2/3 claims grounded · threshold 100%

✓ grounded  claim 1
    Redis keeps data in memory.
    why: context: "Redis is an in-memory data store."

✓ grounded  claim 2
    It supports strings, lists, and sorted sets.
    why: context lists strings, lists and sorted sets.

✗ unsupported NG001  claim 3
    Redis ships with a built-in SQL query engine.
    why: nothing in the context mentions SQL.

groundedness score: 67%
1 unsupported claim(s) (NG001) — likely hallucination.

Why

A retrieval-augmented answer is only trustworthy if it stays on its sources. When a model quietly adds a fact that isn't in the retrieved context, that's the hallucination that erodes trust — and it's invisible unless you check the answer against the context. nogrounds does exactly that check: it isolates each claim and asks whether the provided context backs it up, so an unsupported claim fails loudly (non-zero exit) instead of shipping. Point it at your eval set or wire it into CI.

Install

pip install nogrounds

Requires Python 3.9+. The only runtime dependency is rich.

Usage

# a JSON file: {"answer": ..., "context": ...}
nogrounds pair.json

# explicit files
nogrounds --answer answer.txt --context context.txt

# pipe the JSON on stdin
cat pair.json | nogrounds -

# offline: show the claim segmentation + judge prompt, NO network call
nogrounds pair.json --dry-run

# relax the passing bar to 80% grounded
nogrounds pair.json --threshold 0.8

Flags

Flag Meaning
[input.json|-] JSON {answer, context} file, or - / omitted for stdin
--answer FILE Answer text file (use with --context)
--context FILE Source context text file (use with --answer)
--threshold FLOAT Minimum groundedness score to pass, 0.0–1.0 (default: 1.0)
--provider LLM judge: openai (default) · anthropic · mock
--model Model override (else provider default or $NOGROUNDS_MODEL)
--api-base Override the provider base URL (also $OPENAI_BASE_URL / $ANTHROPIC_BASE_URL)
--dry-run Offline only: print segmentation + judge prompt with no network call
--json Machine-readable JSON output
--no-color Disable colored output

Rules

id Meaning Severity
NG001 An unsupported claim — the context does not support it (likely hallucination) blocker
NG002 A partial / weakly-supported claim warn

Exit codes

Code Meaning
0 Fully grounded — score ≥ threshold and no unsupported claim
1 An unsupported claim (NG001), or the score is below --threshold
2 Usage error — bad arguments, missing/empty input, or an unparseable judge reply

This makes nogrounds usable as a CI / eval gate: fail the build when a RAG answer drifts off its provided sources.

Bring your own key (and testing against a mock)

nogrounds never ships or uses a key of its own. It reads your key from the environment:

export OPENAI_API_KEY=sk-...        # or ANTHROPIC_API_KEY for --provider anthropic
nogrounds pair.json

The providers speak the standard OpenAI Chat Completions / Anthropic Messages wire format over plain HTTP (stdlib urllib, no SDK), so you can point them at any compatible endpoint — including a local mock — to develop and test without spending a token:

nogrounds pair.json --api-base http://127.0.0.1:8080/v1

There is also a built-in --provider mock that returns a canned all-grounded response with no network at all — handy for demos and the test suite. See docs/PROVIDERS.md for the expected judge JSON shape.

Limitations (honest)

  • The judge is an LLM, and LLMs are fallible. nogrounds uses an LLM as the judge, so its verdicts can themselves be wrong. Treat the score as a strong signal, not a proof; spot-check the flagged claims.
  • Claim segmentation is heuristic. The answer is split into claims with sentence/list heuristics, not a linguistic parser. Odd punctuation or deeply nested clauses can over- or under-split.
  • It only checks support against the PROVIDED context. nogrounds asks "does this context support this claim?" — not "is this claim true in the real world?" and not "is this answer correct?". A claim can be perfectly grounded in a context that is itself wrong, and a true claim can be flagged if the context doesn't happen to state it.
  • You bring the key. The judging step requires your own provider key; without one, use --dry-run for the offline segmentation and prompt.

How nogrounds differs from my other RAG tools

These are separate tools that check different things:

  • nogrounds (this tool) — checks answer ↔ context faithfulness: is the generated answer actually supported by the context it was given?
  • ragmoat — checks tenant isolation in a RAG pipeline (that one tenant's retrieval can't leak another's documents).
  • chunkvet — checks chunk quality of the indexed corpus (how the source is split before retrieval).
  • askance — guards an LLM's output flowing into a code sink (injection into an execution/eval path).

nogrounds is the faithfulness check at the end of the pipeline — after retrieval, after generation — on the answer itself.

License

MIT © 2026 Jay Tank

About

An AI-powered RAG groundedness / hallucination checker - segments an LLM answer into atomic claims and flags those NOT supported by the retrieved context, via an LLM judge. Provider-agnostic, BYO-key, offline --dry-run. Python CLI.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages