report --format evalport: a second export format, and what it cannot carry - #76
Merged
Merged
Conversation
…carry EvalPort is an open schema for making evaluation data portable between frameworks. `gauntlet report results.json --format evalport --out DIR` writes one EvalPort ResultSet per gate, all sharing a run_id derived from the run rather than generated, beside a MAPPING.md generated from the documents. No gate is declared as one of EvalPort's built-in graders. Reading `must_contain` as a `contains` grader is right about the string comparison and wrong about the gate: every gate scores legibility before it scores content, so a mute target fails a Gauntlet case that a bare substring check would pass, and on the absence-phrased gates it would pass all of them. Each gate is declared under its own type with `params.handler` naming the function that produced the verdict, which is what EvalPort's type-openness rule is for. A run whose verdict was withheld is not exported. EvalPort has no ResultSet-level "this run has no verdict", so the command writes nothing and exits 4 rather than rendering a withheld verdict as a set of results. What EvalPort has no field for travels in metadata under a `gauntlet.` prefix and is listed in MAPPING.md, which is read off the export rather than typed beside it. `run_dict_from_result_sets` reads the documents back, and a test exports a real run, reads it back, and compares the bytes. Conformance is checked against EvalPort rather than against a reading of it: the published JSON Schemas, vendored and pinned by their upstream blob hashes, and evalport-sdk, EvalPort's own reference validator, as a development dependency. Neither is imported by the package.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is, and what it is not
gauntlet report results.json --format evalport --out DIRwrites a run as EvalPort documents, beside the existing--format mdand--format json.Accepted as built (owner decision, 2026-09-18) — the three open questions and their answers are at the bottom.
Why this, and not the rest of #51
#51 is mostly a Plumbline bundle exporter and importer for the cairn double-grading problem. EvalPort appears in one clause of its "Why it matters": "this is the export mechanism, and EvalPort could be a second format behind the same flag." That clause is the part with an external party waiting on it, and it is the part this delivers. The Plumbline bundle, the
gauntlet import plumbline-bundlereverse, and theDATASET.mdin #51's scope are untouched, so #51 stays open for them. Part of #51.The enquiry is #27, from EvalPort's maintainer. Re-read in full, his ask and the reply to it are specific enough to build from without guessing, and this implements what was agreed there rather than what the mapping table first proposed:
must_not_containhas no built-in equivalent; declare a non-well-known type withparams.handler, which a runner without the handler must skip rather than misinterpretgauntlet_marker_absence, handlergauntlet.gates.adversarial:evaluate_adversarialverdict_withheldis not exported as a ResultSet at allresult_setsraises, the command writes nothing and exits 4normalize_answeris whitespace-only, narrower than the standard grader'signore_case/trim_whitespace; call it out rather than silently widening itgoldenis not declared asexact_match, andMAPPING.mdsays why in those termsOne thing goes further than the thread did. The reply to #27 pointed out that
said_somethingguards the adversarial gate. Reading the gates again, every gate scores legibility before content, so the same objection defeatsmust_containtocontainsandexpectedtoexact_matchtoo: a mute target fails a Gauntlet case that a bare substring check would pass. So no gate is declared as one of EvalPort's built-in graders. Each has its own type and handler, which is what EvalPort's type-openness rule exists for, andMAPPING.mdnames the nearest built-in and why it is not that one.Shape
suite_id. A Gauntlet run puts one target through several suites, so folding them into one document would name one suite of several and let two cases from two suites collide on atest_case_idthat is only unique within a suite. They share arun_id.run_idis derived from the run, not generated. EvalPort suggests a UUID or a timestamp plus a random suffix; both would make the same results file export differently every time. It is a digest of the run, so re-exporting is byte-identical.MAPPING.mdis generated from the documents, not typed beside them, so it cannot describe a key the export stopped writing. A test asserts both directions over a fixture wide enough to reach the conditional keys.What has no EvalPort equivalent
This is the list an integrator actually needs, and the reason it is generated rather than written.
Carried in
metadataunder agauntlet.prefix, because EvalPort has no field: the gate name and its position in the run, the suite's pass-rate threshold, a golden suite's key version, a judge gate's calibration record, the run's provenance block, the results schema version, the target name, whether the whole run passed, each case's language, and the turns of a multi-turn case with the count it declared.Two of those deserve a note.
metadata.openeval.aggregationlooks like the slot forthresholdand is not one: it combines graders within a single result, not cases within a suite. And a multi-turn case has nowhere first-class to go:TestCase.inputtakes an array, but aResultcarries oneactual_outputstring.Not carried, because a Gauntlet results file never recorded them: the target's declared
refusedandescalatedbooleans, and itscitationsandcontext_ids. #27 asked about the first pair specifically and expected them inResult.metadata. They are inputs to a gate rather than outputs of one, and onlygauntlet run --recordkeeps them. Case prompts, expected answers and markers are in the suite YAML, not in the file this verb reads, which is why this writes ResultSets and no EvalSuite:TestCase.inputis required, and inventing one would be worse than omitting the document.How it was checked
Against EvalPort, rather than against a reading of it, three ways:
tests/fixtures/evalport/and pinned by their upstream Git blob hashes, so a test proves offline that each file is the published byte sequence. Run withjsonschema.evalport-sdk, EvalPort's own reference validator, as a development dependency.run_dict_from_result_setsreads a directory of these documents back into the results payload, and a test exports a run, reads it back, and compares the bytes. Both a hand-built fixture reaching every branch and a real run of the built-in suites against the toy.Three negative controls, one finding. Each sabotage was proved to land by
git hash-objectplus an occurrence count on the property itself, and the tree was restored to the baseline hash after each.100.0, never normalised to[0, 1]Resultinstead of undermetadataevalport-sdkreported the document validgauntlet.languagequietly not carriedThe second row is why both validators are here rather than one: the reference SDK does not enforce
additionalProperties: false, so it passes a document the published schema rejects. The third is why the round trip exists: schema conformance cannot see a field that was silently dropped, because the schema has no opinion about it.Branch is rebased on
main(merge c2624a1) andmake verifyis green: lockfile, ruff format and check, mypy strict, 1392 tests, 96.41% coverage,src/gauntlet/evalport.pyat 99%.make demoexercises the new format end to end. Two American-English fixes went in ahead of merge (recognise→recognizein the exporter,Licence→Licensein the vendored NOTICE), commit 68556c7.Decisions (2026-09-18), and what happened to the three draft questions
evalport-sdkandjsonschemastay in thedevgroup only; the package imports neither. The four vendored files are EvalPort's published JSON Schemas, Apache-2.0, attributed intests/fixtures/evalport/NOTICE.mdwith upstream commit and blob hashes.v0.3.0, tag 4c7f500, already an ancestor ofmain), so users today get the citation fix. This PR adds a new feature behind an existing CLI, which is a minor bump by the SemVer precedent set in 59515ef: it ships in v0.4.0. The owner steps (release-prep commit, signed tag, GitHub Release publish,pypienvironment approval) are written out ingauntlet-release-owner-steps.md; no tag was cut here.MAPPING.md, is the call to publish.Nobody outside this repository was contacted, no tag was cut or moved, and nothing was published.
Prepared with AI assistance; reviewed before submission.