Skip to content

report --format evalport: a second export format, and what it cannot carry - #76

Merged
ChelseaKR merged 3 commits into
mainfrom
feat/evalport-export-format
Sep 19, 2026
Merged

ChelseaKR merged 3 commits into
mainfrom
feat/evalport-export-format

Conversation

@ChelseaKR

@ChelseaKR ChelseaKR commented Sep 13, 2026 •

Copy link
Copy Markdown
Owner

What this is, and what it is not

gauntlet report results.json --format evalport --out DIR writes a run as EvalPort documents, beside the existing --format md and --format json.

Accepted as built (owner decision, 2026-09-18) — the three open questions and their answers are at the bottom.

Why this, and not the rest of #51

#51 is mostly a Plumbline bundle exporter and importer for the cairn double-grading problem. EvalPort appears in one clause of its "Why it matters": "this is the export mechanism, and EvalPort could be a second format behind the same flag." That clause is the part with an external party waiting on it, and it is the part this delivers. The Plumbline bundle, the gauntlet import plumbline-bundle reverse, and the DATASET.md in #51's scope are untouched, so #51 stays open for them. Part of #51.

The enquiry is #27, from EvalPort's maintainer. Re-read in full, his ask and the reply to it are specific enough to build from without guessing, and this implements what was agreed there rather than what the mapping table first proposed:

Agreed in #27 Here
must_not_contain has no built-in equivalent; declare a non-well-known type with params.handler, which a runner without the handler must skip rather than misinterpret gauntlet_marker_absence, handler gauntlet.gates.adversarial:evaluate_adversarial
The legibility predicate goes in the handler contract verbatim, not reimplemented from prose The handler is a dotted path to the function itself. A test imports every one of them, so none is a paraphrase with a colon in it
A run with a non-empty verdict_withheld is not exported as a ResultSet at all result_sets raises, the command writes nothing and exits 4
normalize_answer is whitespace-only, narrower than the standard grader's ignore_case/trim_whitespace; call it out rather than silently widening it golden is not declared as exact_match, and MAPPING.md says why in those terms

One thing goes further than the thread did. The reply to #27 pointed out that said_something guards the adversarial gate. Reading the gates again, every gate scores legibility before content, so the same objection defeats must_contain to contains and expected to exact_match too: a mute target fails a Gauntlet case that a bare substring check would pass. So no gate is declared as one of EvalPort's built-in graders. Each has its own type and handler, which is what EvalPort's type-openness rule exists for, and MAPPING.md names the nearest built-in and why it is not that one.

Shape

  • One ResultSet per gate. EvalPort's ResultSet is "the output of running an eval suite" and carries one suite_id. A Gauntlet run puts one target through several suites, so folding them into one document would name one suite of several and let two cases from two suites collide on a test_case_id that is only unique within a suite. They share a run_id.
  • run_id is derived from the run, not generated. EvalPort suggests a UUID or a timestamp plus a random suffix; both would make the same results file export differently every time. It is a digest of the run, so re-exporting is byte-identical.
  • MAPPING.md is generated from the documents, not typed beside them, so it cannot describe a key the export stopped writing. A test asserts both directions over a fixture wide enough to reach the conditional keys.

What has no EvalPort equivalent

This is the list an integrator actually needs, and the reason it is generated rather than written.

Carried in metadata under a gauntlet. prefix, because EvalPort has no field: the gate name and its position in the run, the suite's pass-rate threshold, a golden suite's key version, a judge gate's calibration record, the run's provenance block, the results schema version, the target name, whether the whole run passed, each case's language, and the turns of a multi-turn case with the count it declared.

Two of those deserve a note. metadata.openeval.aggregation looks like the slot for threshold and is not one: it combines graders within a single result, not cases within a suite. And a multi-turn case has nowhere first-class to go: TestCase.input takes an array, but a Result carries one actual_output string.

Not carried, because a Gauntlet results file never recorded them: the target's declared refused and escalated booleans, and its citations and context_ids. #27 asked about the first pair specifically and expected them in Result.metadata. They are inputs to a gate rather than outputs of one, and only gauntlet run --record keeps them. Case prompts, expected answers and markers are in the suite YAML, not in the file this verb reads, which is why this writes ResultSets and no EvalSuite: TestCase.input is required, and inventing one would be worse than omitting the document.

How it was checked

Against EvalPort, rather than against a reading of it, three ways:

  1. The published JSON Schemas, vendored verbatim under tests/fixtures/evalport/ and pinned by their upstream Git blob hashes, so a test proves offline that each file is the published byte sequence. Run with jsonschema.
  2. evalport-sdk, EvalPort's own reference validator, as a development dependency.
  3. A round trip. run_dict_from_result_sets reads a directory of these documents back into the results payload, and a test exports a run, reads it back, and compares the bytes. Both a hand-built fixture reaching every branch and a real run of the built-in suites against the toy.

Three negative controls, one finding. Each sabotage was proved to land by git hash-object plus an occurrence count on the property itself, and the tree was restored to the baseline hash after each.

Sabotage What went red What stayed green
a native score emitted as 100.0, never normalised to [0, 1] both validators, independently
producer data written at the top level of a Result instead of under metadata the JSON Schema path evalport-sdk reported the document valid
gauntlet.language quietly not carried the round trip, and the mapping table's both-directions check both conformance validators, correctly

The second row is why both validators are here rather than one: the reference SDK does not enforce additionalProperties: false, so it passes a document the published schema rejects. The third is why the round trip exists: schema conformance cannot see a field that was silently dropped, because the schema has no opinion about it.

Branch is rebased on main (merge c2624a1) and make verify is green: lockfile, ruff format and check, mypy strict, 1392 tests, 96.41% coverage, src/gauntlet/evalport.py at 99%. make demo exercises the new format end to end. Two American-English fixes went in ahead of merge (recognise → recognize in the exporter, Licence → License in the vendored NOTICE), commit 68556c7.

Decisions (2026-09-18), and what happened to the three draft questions

  1. Dev dependencies and vendored schemas: accepted. evalport-sdk and jsonschema stay in the dev group only; the package imports neither. The four vendored files are EvalPort's published JSON Schemas, Apache-2.0, attributed in tests/fixtures/evalport/NOTICE.md with upstream commit and blob hashes.
  2. Release vehicle: v0.4.0, not v0.2.0. The draft's premise is stale and was re-measured: PyPI no longer serves 0.1.0 — both 0.2.0 and 0.3.0 were published on 2026-09-13 (release v0.3.0, tag 4c7f500, already an ancestor of main), so users today get the citation fix. This PR adds a new feature behind an existing CLI, which is a minor bump by the SemVer precedent set in 59515ef: it ships in v0.4.0. The owner steps (release-prep commit, signed tag, GitHub Release publish, pypi environment approval) are written out in gauntlet-release-owner-steps.md; no tag was cut here.
  3. The mapping: accepted as built. Six Gauntlet-named grader types with dotted-path handlers, each defended in MAPPING.md, is the call to publish.

Nobody outside this repository was contacted, no tag was cut or moved, and nothing was published.

Prepared with AI assistance; reviewed before submission.

…carry

EvalPort is an open schema for making evaluation data portable between
frameworks. `gauntlet report results.json --format evalport --out DIR` writes
one EvalPort ResultSet per gate, all sharing a run_id derived from the run
rather than generated, beside a MAPPING.md generated from the documents.

No gate is declared as one of EvalPort's built-in graders. Reading
`must_contain` as a `contains` grader is right about the string comparison and
wrong about the gate: every gate scores legibility before it scores content, so
a mute target fails a Gauntlet case that a bare substring check would pass, and
on the absence-phrased gates it would pass all of them. Each gate is declared
under its own type with `params.handler` naming the function that produced the
verdict, which is what EvalPort's type-openness rule is for.

A run whose verdict was withheld is not exported. EvalPort has no
ResultSet-level "this run has no verdict", so the command writes nothing and
exits 4 rather than rendering a withheld verdict as a set of results.

What EvalPort has no field for travels in metadata under a `gauntlet.` prefix
and is listed in MAPPING.md, which is read off the export rather than typed
beside it. `run_dict_from_result_sets` reads the documents back, and a test
exports a real run, reads it back, and compares the bytes.

Conformance is checked against EvalPort rather than against a reading of it:
the published JSON Schemas, vendored and pinned by their upstream blob hashes,
and evalport-sdk, EvalPort's own reference validator, as a development
dependency. Neither is imported by the package.
@ChelseaKR
ChelseaKR marked this pull request as ready for review September 19, 2026 02:40
@ChelseaKR
ChelseaKR merged commit 5404bee into main Sep 19, 2026
8 checks passed
@ChelseaKR
ChelseaKR deleted the feat/evalport-export-format branch September 19, 2026 02:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant