v0.2.0. Fail-closed evaluation harness for government-facing chat systems: reproducible, provenance-stamped audit verdicts, byte-identical across Python 3.11 to 3.14, with no third-party dependencies. A silent or unreadable target scores zero rather than passing by absence. Two public projects of my own pin it by exact commit.
python accessibility evaluation provenance civic-tech reproducibility fairness govtech evaluation-framework public-sector ai-evaluation prompt-injection llm-eval evals citation-accuracy fail-closed evaluation-harness ci-gate groundedness
-
Updated
Sep 23, 2026 - Python