Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ demo-results.json
demo-evidence.md
demo-evidence.json
demo-evidence-drift.md
demo-evalport/
gauntlet-results.json
gauntlet-evidence.md
gauntlet-evidence.json
Expand Down
29 changes: 29 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,35 @@ All notable changes will be documented here.

### Added

- **A second export format: `gauntlet report --format evalport`.** EvalPort is an
open schema for making evaluation data portable between frameworks. The new
format writes one EvalPort ResultSet per gate into `--out` as a directory, all
of them sharing a `run_id` derived from the run rather than generated, so
re-exporting the same results file writes the same bytes. Beside them it writes
a `MAPPING.md` generated from the documents themselves.

**No gate is declared as one of EvalPort's built-in graders.** Reading
`must_contain` as a `contains` grader is right about the string comparison and
wrong about the gate, because every gate scores legibility before it scores
content: a mute target fails a Gauntlet case that a bare substring check would
pass, and on the absence-phrased gates it would pass all of them. Each gate is
declared under its own type with `params.handler` naming the function that
produced the verdict, which is what EvalPort's type-openness rule is for.

**A run whose verdict was withheld is not exported.** EvalPort has no
ResultSet-level "this run has no verdict", so the command writes nothing and
exits 4 rather than rendering a withheld verdict as a set of results.

What EvalPort has no field for travels in `metadata` under a `gauntlet.` prefix
and is listed in `MAPPING.md`; what a results file never recorded is named there
too. `gauntlet.evalport.run_dict_from_result_sets` reads the documents back, and
a test exports a real run, reads it back, and compares the bytes. Conformance is
checked against EvalPort rather than against a reading of it: the published JSON
Schemas, vendored under `tests/fixtures/evalport/` and pinned by their upstream
blob hashes, and `evalport-sdk`, EvalPort's own reference validator, added as a
development dependency. Neither is imported by the package, so the export runs on
a plain install.

- **Google Analytics 4 on the documentation site, and a privacy page.** Owner
decision 2026-09-17: GA4 on every public site, with privacy copy changed to
match. `src/gauntlet/analytics.py` holds the measurement ID
Expand Down
1 change: 1 addition & 0 deletions Makefile
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,7 @@ demo:
uv run --locked gauntlet report demo-results.json --out demo-evidence.md
uv run --locked gauntlet report demo-results.json --format json --out demo-evidence.json
uv run --locked gauntlet report demo-results.json --baseline demo-results.json --out demo-evidence-drift.md
uv run --locked gauntlet report demo-results.json --format evalport --out demo-evalport

inventory:
uv run --locked gauntlet inventory --update README.md
Expand Down
50 changes: 50 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,6 +45,9 @@ uv run gauntlet run --out results.json
uv run gauntlet report results.json --out evidence.md
uv run gauntlet report results.json --format json --out evidence.json

# The same run as EvalPort ResultSets, one per gate, for a tool that reads that schema.
uv run gauntlet report results.json --format evalport --out evalport/

# Whole-run drift against an earlier run.
uv run gauntlet report results.json --baseline previous-results.json --out evidence.md

Expand Down Expand Up @@ -230,6 +233,53 @@ Each pack carries a `results_digest`: a sha256 over what the run observed, with
the clock deliberately excluded. Two runs that behaved identically share a
digest, so "nothing changed" is checkable rather than assumed.

## Exporting a run as EvalPort

[EvalPort](https://github.com/adhabnr-ux/evalport) is an open schema for making
evaluation data portable between frameworks: one JSON shape a dashboard or a CI
gate can read whichever tool produced it. `gauntlet report` writes it as a second
export format, beside the Markdown document and the JSON pack.

```sh
uv run gauntlet report results.json --format evalport --out evalport/
```

It writes one EvalPort ResultSet per gate, because a ResultSet carries a single
`suite_id` and a Gauntlet run puts one target through several suites at once. All
of them share a `run_id`, which is derived from the run rather than generated, so
re-exporting the same results file writes the same bytes. Beside them it writes a
`MAPPING.md` generated from the documents themselves, listing every field the
schema has no home for and what carries it instead.

Three things about the mapping are worth knowing before reading the output.

**No gate is exported as one of EvalPort's built-in graders.** It is tempting to
call `must_contain` a `contains` grader and `expected` an `exact_match` one, and
at the level of the string comparison that reading is right. It is wrong at the
level of the gate, because every gate scores legibility before it scores content:
a target that says nothing fails a Gauntlet case that a bare substring check would
pass, and on the absence-phrased gates it would pass every one of them. EvalPort's
type-openness rule covers exactly this, so each gate is declared under its own type
with `params.handler` naming the function that produced the verdict.

**Some things a Gauntlet results file records have no EvalPort field, and travel in
`metadata` under a `gauntlet.` prefix**: the gate name, the suite's pass-rate
threshold, a golden suite's key version, a judge gate's calibration record, each
case's language, and the turns of a multi-turn case. `MAPPING.md` lists them, and
the list is read off the export rather than typed beside it, so it cannot describe
a key the export stopped writing.

**A run whose verdict was withheld is not exported at all.** EvalPort scores each
result on its own and has no ResultSet-level "this run has no verdict", so every
available representation would assert a verdict the harness declined to reach. The
command writes nothing and exits 4, the same code the run itself exits.

The export is one-directional in the sense that matters for a consumer, and
reversible in the sense that matters for trust: `gauntlet.evalport` can read a
directory of these ResultSets back into the results payload it came from, and a
test exports a real run, reads it back, and compares the bytes. What has no
EvalPort field is carried, not dropped quietly.

## Recording a run, and grading the recording

A merge gate that reaches a live service is not deterministic, spends budget on
Expand Down
17 changes: 17 additions & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -38,10 +38,21 @@ judge = ["anthropic[bedrock]>=0.125"]

[dependency-groups]
dev = [
# evalport-sdk is EvalPort's own reference validator, and jsonschema runs the
# vendored copies of EvalPort's published JSON Schemas under
# tests/fixtures/evalport/. Both are test-only: the exporter itself has no
# dependency on either, so `gauntlet report --format evalport` runs on a plain
# install. Two validators rather than one because they check different halves.
# The schemas carry `additionalProperties: false`, which the SDK does not
# enforce; the SDK checks ResultSet rules the schemas cannot express, such as
# a duplicate (test_case_id, run_id, attempt).
"evalport-sdk>=1.3.1",
"jsonschema>=4.23",
"mypy>=1.18",
"pytest>=8",
"pytest-cov>=5",
"ruff>=0.15",
"types-jsonschema>=4.23",
"types-pyyaml>=6",
]

Expand Down Expand Up @@ -84,6 +95,12 @@ files = ["src", "tests", "examples", "real_targets"]
module = ["mrf_honest.*", "fhir_scorecard.*", "sprout", "sprout.*", "anthropic", "anthropic.*"]
ignore_missing_imports = true

# EvalPort's reference SDK ships no type information. It is imported by
# tests/test_evalport.py only.
[[tool.mypy.overrides]]
module = ["openeval", "openeval.*"]
ignore_missing_imports = true

[tool.pytest.ini_options]
minversion = "8"
testpaths = ["tests"]
Expand Down
34 changes: 32 additions & 2 deletions src/gauntlet/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -59,6 +59,8 @@
write_calibration,
)
from gauntlet.cases import Suite, builtin_suites, load_suites
from gauntlet.evalport import WithheldVerdict
from gauntlet.evalport import export as evalport_export
from gauntlet.evidence import build_evidence_pack, github_output_lines
from gauntlet.gates import judge_withheld_reason, run_suite, unscoreable_reason
from gauntlet.history import (
Expand Down Expand Up @@ -315,6 +317,8 @@ def _print_run_summary(run: RunResult, verdict: str | None = None) -> None:

def _cmd_report(args: argparse.Namespace) -> int:
run = load_run_dict(Path(args.results))
if args.format == "evalport":
return _write_evalport(run, args)
baseline = load_run_dict(Path(args.baseline)) if args.baseline else None
history = None
if args.ledger:
Expand All @@ -331,6 +335,30 @@ def _cmd_report(args: argparse.Namespace) -> int:
return 0


def _write_evalport(run: dict[str, object], args: argparse.Namespace) -> int:
"""Write the run as EvalPort documents, or say why it has none.

A withheld verdict leaves by way of exit 4, the code the run itself exits,
rather than exit 2. The harness ran; it declined to score what it saw, and
that is what the reader is told here too.
"""
if not args.out:
raise ValueError(
"--format evalport writes several documents, so it needs a directory: pass --out DIR"
)
try:
files = evalport_export(run)
except WithheldVerdict as exc:
print(f"error: {exc}", file=sys.stderr)
return EXIT_UNSCOREABLE
directory = Path(args.out)
directory.mkdir(parents=True, exist_ok=True)
for name, text in files.items():
(directory / name).write_text(text, encoding="utf-8")
print(f"wrote {len(files)} EvalPort files to {directory}")
return 0


def _append_github_output(path: Path, pack: dict[str, object]) -> None:
"""The pack's headline counts, plus the digest of the pack's own bytes.

Expand Down Expand Up @@ -681,9 +709,11 @@ def _add_report_parser(sub: argparse._SubParsersAction[argparse.ArgumentParser])
)
report_parser.add_argument(
"--format",
choices=("md", "json"),
choices=("md", "json", "evalport"),
default="md",
help="md for the human-readable document, json for the machine-readable pack",
help="md for the human-readable document, json for the machine-readable pack, "
"evalport for EvalPort ResultSets, one per gate, written into --out as a "
"directory alongside a MAPPING.md naming what the schema has no field for",
)
report_parser.add_argument("--out", help="write the evidence pack to this path")
report_parser.add_argument(
Expand Down
Loading
Loading