VerifyAI is a verification harness for testing, auditing, and evaluating responses from AI assistants and coding agents.
-
Updated
Aug 5, 2026 - Python
VerifyAI is a verification harness for testing, auditing, and evaluating responses from AI assistants and coding agents.
Policy-conformance harness for money-touching AI agents — catches over-promises against refund policy, with mechanically verified evidence (Python-derived labels, span-verified judge citations, frozen agent under test).
Quantsafe Certifier for testing whether quantized models are prone to be jailbroken and generate unsafe outputs.
Multilingual LLM safety benchmark for Urdu, Pashto, and code-switching prompts across different prompting techniques
To associate your repository with the llm-safety-evaluation topic, visit your repo's landing page and select "manage topics."