DeepEval LLM evaluation framework using Claude API — tests GEval, Hallucination, Answer Relevancy, Faithfulness, Contextual Precision/Recall, Bias, Toxicity, PII Leakage, and Tool Correctness on a mock AI customer support assistant.
python pytest eval evaluation-framework hallucination rag claude-api llm-evaluation llm-as-a-judge tool-calling deepeval safety-testing deepeval-metrics
-
Updated
Aug 28, 2026 - Python