I maintain EvalPort, an open schema (TestCase / Grader / EvalSuite / ResultSet) for making evaluation data portable between frameworks — think "one JSON shape a dashboard or CI gate can read regardless of which eval tool produced it." I read src/gauntlet/cases.py, results.py, and the five gates under src/gauntlet/gates/ (not README claims) before writing this, so this is grounded in what the code actually does, not the description.
Not a PR, not asking for a dependency or code change here — just checking whether my reading of the shapes is right before I'd build anything, and whether it's wanted at all.
What I think maps cleanly:
| gauntlet |
EvalPort |
Suite (cases.py) — name, gate, version, threshold, cases |
EvalSuite — id, version, test_cases, with gate and threshold carried in metadata (EvalPort has no suite-level pass-threshold field) |
Case — id, language, prompt |
TestCase — id, input; language as a tag |
RunResult (results.py) — target, gates, provenance |
ResultSet — suite_id, provider, with gauntlet's required provenance keys (target_version, commit, prompt_version...) folded into ResultSet.metadata |
GateResult → CaseResult |
Result → GraderResult — CaseResult.detail → reason, .observed → actual_output |
Where the grader types are a real fit, not a stretch, because the gate code is literal string matching, not fuzzy scoring:
must_contain (grounding.py, refusal.py) → EvalPort's built-in contains grader — same semantics, case-insensitive substring check.
expected (golden.py) → exact_match, though gauntlet's normalize_answer only collapses whitespace (no case-folding), which is narrower than the standard grader's ignore_case/trim_whitespace params — worth calling out rather than silently widening it.
rubric (judge gate) → llm_judge, params.prompt = case.rubric.
Where it's a genuine mismatch, not just relabeling — three things EvalPort's schema doesn't have a slot for:
must_not_contain (adversarial.py) has no standard EvalPort grader — there's no built-in "absent marker" type. Per the spec's type-openness rule I'd declare it as a non-well-known type (e.g. gauntlet_marker_absence) with params.handler set, which the spec says any runner without that handler must skip rather than misinterpret.
evaluate_refusal/evaluate_adversarial key off TargetResponse.refused / .escalated — booleans your target adapter sets, not something derivable from the response text alone. EvalPort graders score actual_output (a string); there's no first-class field for "the target self-reported refusing." I'd have to carry those flags in Result.metadata, which a generic EvalPort consumer wouldn't know to look at.
- Judge calibration is structural in gauntlet —
ADR 0001 fails a judge gate closed with no calibrated set (_calibrate_for in gates/base.py), and RunResult.verdict_withheld can mark a whole run as unscoreable (unscoreable_reason in base.py). EvalPort has no ResultSet-level "withheld" concept and no calibration semantics on llm_judge — closest is metadata.openeval.aggregation_status: "unscored", which is per-Result, not per-run. This is the piece I'd feel worst about papering over with metadata alone.
Sketch of the conversion I'd write, to be concrete about what "adapter" means here — no gauntlet-side changes, this would live entirely on the EvalPort side, reading your public JSON result pack as data:
# evalport/adapters/gauntlet-openeval-adapter — reads gauntlet's `gauntlet report --json` output
def gate_result_to_grader_results(gate: dict) -> list[dict]:
grader_type = {
"grounding": "contains", "refusal": "contains", "golden": "exact_match",
"judge": "llm_judge", "adversarial": "gauntlet_marker_absence",
"false_positive": "contains",
}[gate["gate"]]
return [
{
"grader_id": f"{gate['gate']}_gr",
"type": grader_type,
"score": 1.0 if c["passed"] else 0.0,
"passed": c["passed"],
"reason": c["detail"],
}
for c in gate["cases"]
]
This would sit alongside the existing framework adapters in evalport/adapters/ (deepeval, argilla, azure-ai-evaluation, etc. all follow the same one-directional "read the tool's native output, emit a ResultSet" shape).
Given #3 especially, I'd rather hear whether this reading is even right, and whether a lossy one-directional export is something you'd consider useful or a category error for what the evidence pack is for, before I spend time on it. No obligation either way — happy to just close this out if it's not a fit.
— Sahi, independent contributor (not affiliated with this project)
I maintain EvalPort, an open schema (
TestCase/Grader/EvalSuite/ResultSet) for making evaluation data portable between frameworks — think "one JSON shape a dashboard or CI gate can read regardless of which eval tool produced it." I readsrc/gauntlet/cases.py,results.py, and the five gates undersrc/gauntlet/gates/(not README claims) before writing this, so this is grounded in what the code actually does, not the description.Not a PR, not asking for a dependency or code change here — just checking whether my reading of the shapes is right before I'd build anything, and whether it's wanted at all.
What I think maps cleanly:
Suite(cases.py) —name,gate,version,threshold,casesEvalSuite—id,version,test_cases, withgateandthresholdcarried inmetadata(EvalPort has no suite-level pass-threshold field)Case—id,language,promptTestCase—id,input;languageas a tagRunResult(results.py) —target,gates,provenanceResultSet—suite_id,provider, with gauntlet's required provenance keys (target_version,commit,prompt_version...) folded intoResultSet.metadataGateResult→CaseResultResult→GraderResult—CaseResult.detail→reason,.observed→actual_outputWhere the grader types are a real fit, not a stretch, because the gate code is literal string matching, not fuzzy scoring:
must_contain(grounding.py,refusal.py) → EvalPort's built-incontainsgrader — same semantics, case-insensitive substring check.expected(golden.py) →exact_match, though gauntlet'snormalize_answeronly collapses whitespace (no case-folding), which is narrower than the standard grader'signore_case/trim_whitespaceparams — worth calling out rather than silently widening it.rubric(judge gate) →llm_judge,params.prompt = case.rubric.Where it's a genuine mismatch, not just relabeling — three things EvalPort's schema doesn't have a slot for:
must_not_contain(adversarial.py) has no standard EvalPort grader — there's no built-in "absent marker" type. Per the spec's type-openness rule I'd declare it as a non-well-known type (e.g.gauntlet_marker_absence) withparams.handlerset, which the spec says any runner without that handler mustskiprather than misinterpret.evaluate_refusal/evaluate_adversarialkey offTargetResponse.refused/.escalated— booleans your target adapter sets, not something derivable from the response text alone. EvalPort graders scoreactual_output(a string); there's no first-class field for "the target self-reported refusing." I'd have to carry those flags inResult.metadata, which a generic EvalPort consumer wouldn't know to look at.ADR 0001fails a judge gate closed with no calibrated set (_calibrate_foringates/base.py), andRunResult.verdict_withheldcan mark a whole run as unscoreable (unscoreable_reasoninbase.py). EvalPort has no ResultSet-level "withheld" concept and no calibration semantics onllm_judge— closest ismetadata.openeval.aggregation_status: "unscored", which is per-Result, not per-run. This is the piece I'd feel worst about papering over with metadata alone.Sketch of the conversion I'd write, to be concrete about what "adapter" means here — no gauntlet-side changes, this would live entirely on the EvalPort side, reading your public JSON result pack as data:
This would sit alongside the existing framework adapters in
evalport/adapters/(deepeval, argilla, azure-ai-evaluation, etc. all follow the same one-directional "read the tool's native output, emit a ResultSet" shape).Given #3 especially, I'd rather hear whether this reading is even right, and whether a lossy one-directional export is something you'd consider useful or a category error for what the evidence pack is for, before I spend time on it. No obligation either way — happy to just close this out if it's not a fit.
— Sahi, independent contributor (not affiliated with this project)