Skip to content

Interop check: does an EvalPort adapter for gauntlet's result packs make sense? #27

Description

@adhabnr-ux

I maintain EvalPort, an open schema (TestCase / Grader / EvalSuite / ResultSet) for making evaluation data portable between frameworks — think "one JSON shape a dashboard or CI gate can read regardless of which eval tool produced it." I read src/gauntlet/cases.py, results.py, and the five gates under src/gauntlet/gates/ (not README claims) before writing this, so this is grounded in what the code actually does, not the description.

Not a PR, not asking for a dependency or code change here — just checking whether my reading of the shapes is right before I'd build anything, and whether it's wanted at all.

What I think maps cleanly:

gauntlet EvalPort
Suite (cases.py) — name, gate, version, threshold, cases EvalSuite — id, version, test_cases, with gate and threshold carried in metadata (EvalPort has no suite-level pass-threshold field)
Case — id, language, prompt TestCase — id, input; language as a tag
RunResult (results.py) — target, gates, provenance ResultSet — suite_id, provider, with gauntlet's required provenance keys (target_version, commit, prompt_version...) folded into ResultSet.metadata
GateResult → CaseResult Result → GraderResult — CaseResult.detail → reason, .observed → actual_output

Where the grader types are a real fit, not a stretch, because the gate code is literal string matching, not fuzzy scoring:

  • must_contain (grounding.py, refusal.py) → EvalPort's built-in contains grader — same semantics, case-insensitive substring check.
  • expected (golden.py) → exact_match, though gauntlet's normalize_answer only collapses whitespace (no case-folding), which is narrower than the standard grader's ignore_case/trim_whitespace params — worth calling out rather than silently widening it.
  • rubric (judge gate) → llm_judge, params.prompt = case.rubric.

Where it's a genuine mismatch, not just relabeling — three things EvalPort's schema doesn't have a slot for:

  1. must_not_contain (adversarial.py) has no standard EvalPort grader — there's no built-in "absent marker" type. Per the spec's type-openness rule I'd declare it as a non-well-known type (e.g. gauntlet_marker_absence) with params.handler set, which the spec says any runner without that handler must skip rather than misinterpret.
  2. evaluate_refusal/evaluate_adversarial key off TargetResponse.refused / .escalated — booleans your target adapter sets, not something derivable from the response text alone. EvalPort graders score actual_output (a string); there's no first-class field for "the target self-reported refusing." I'd have to carry those flags in Result.metadata, which a generic EvalPort consumer wouldn't know to look at.
  3. Judge calibration is structural in gauntlet — ADR 0001 fails a judge gate closed with no calibrated set (_calibrate_for in gates/base.py), and RunResult.verdict_withheld can mark a whole run as unscoreable (unscoreable_reason in base.py). EvalPort has no ResultSet-level "withheld" concept and no calibration semantics on llm_judge — closest is metadata.openeval.aggregation_status: "unscored", which is per-Result, not per-run. This is the piece I'd feel worst about papering over with metadata alone.

Sketch of the conversion I'd write, to be concrete about what "adapter" means here — no gauntlet-side changes, this would live entirely on the EvalPort side, reading your public JSON result pack as data:

# evalport/adapters/gauntlet-openeval-adapter — reads gauntlet's `gauntlet report --json` output
def gate_result_to_grader_results(gate: dict) -> list[dict]:
    grader_type = {
        "grounding": "contains", "refusal": "contains", "golden": "exact_match",
        "judge": "llm_judge", "adversarial": "gauntlet_marker_absence",
        "false_positive": "contains",
    }[gate["gate"]]
    return [
        {
            "grader_id": f"{gate['gate']}_gr",
            "type": grader_type,
            "score": 1.0 if c["passed"] else 0.0,
            "passed": c["passed"],
            "reason": c["detail"],
        }
        for c in gate["cases"]
    ]

This would sit alongside the existing framework adapters in evalport/adapters/ (deepeval, argilla, azure-ai-evaluation, etc. all follow the same one-directional "read the tool's native output, emit a ResultSet" shape).

Given #3 especially, I'd rather hear whether this reading is even right, and whether a lossy one-directional export is something you'd consider useful or a category error for what the evidence pack is for, before I spend time on it. No obligation either way — happy to just close this out if it's not a fit.

— Sahi, independent contributor (not affiliated with this project)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions