What
SCOPE.md lists cost guards among the disciplines this harness reimplements, and none exists. The real-target runs tracked a 100-request daily cap by hand and recorded token counts by reading the target's logs.
Add a budget gate. A suite declares limits: {max_requests, max_latency_ms_p95, max_input_tokens, max_output_tokens}. The harness measures wall latency itself; a target may report usage: {input_tokens, output_tokens, latency_ms} as an optional contract field. max_requests is enforced before the request is made: the run stops with exit 2 and the reason, and the results file is removed, honoring the existing exit-2 file guarantee. Limits that need usage the target did not report are UNVERIFIABLE: reported, never passed.
Why it matters
A feature that answers correctly at four seconds a turn, or that doubles its token use after a prompt change, has regressed in a way no content gate sees. Measuring it in the same run, with the same pack, keeps one evidence artifact per release. The absence rule holds: a missing usage is not a pass.
Scope
src/gauntlet/gates/budget.py, case-less suite (limits only), usage on TargetResponse.
- Per-run request ledger in provenance; a "Budget" table in both pack forms.
- If
budget is the only loaded gate and the target reports no usage, the run is UNSCOREABLE by the silence rule's analogue.
- Toy defect
slow_answer and paired self-test.
Out of scope
- Price tables; cost in currency only when a target reports it.
- Retry or backoff policy.
Done when
- The toy with an injected delay fails
max_latency_ms_p95 and the pack names the p95.
- A target omitting
usage reports the token limits as UNVERIFIABLE.
max_requests: 3 on a five-case suite exits 2 after three requests with no results file.
- The built-in suites' packs are byte-identical.
Pointers
src/gauntlet/targets.py, src/gauntlet/gates/base.py (unscoreable_reason), src/gauntlet/cli.py (_claim_out_path), docs/real-targets.md "Budget and model, stated first"
Proposed with AI assistance.
What
SCOPE.mdlists cost guards among the disciplines this harness reimplements, and none exists. The real-target runs tracked a 100-request daily cap by hand and recorded token counts by reading the target's logs.Add a
budgetgate. A suite declareslimits: {max_requests, max_latency_ms_p95, max_input_tokens, max_output_tokens}. The harness measures wall latency itself; a target may reportusage: {input_tokens, output_tokens, latency_ms}as an optional contract field.max_requestsis enforced before the request is made: the run stops with exit 2 and the reason, and the results file is removed, honoring the existing exit-2 file guarantee. Limits that needusagethe target did not report areUNVERIFIABLE: reported, never passed.Why it matters
A feature that answers correctly at four seconds a turn, or that doubles its token use after a prompt change, has regressed in a way no content gate sees. Measuring it in the same run, with the same pack, keeps one evidence artifact per release. The absence rule holds: a missing
usageis not a pass.Scope
src/gauntlet/gates/budget.py, case-less suite (limits only),usageonTargetResponse.budgetis the only loaded gate and the target reports no usage, the run is UNSCOREABLE by the silence rule's analogue.slow_answerand paired self-test.Out of scope
Done when
max_latency_ms_p95and the pack names the p95.usagereports the token limits as UNVERIFIABLE.max_requests: 3on a five-case suite exits 2 after three requests with no results file.Pointers
src/gauntlet/targets.py,src/gauntlet/gates/base.py(unscoreable_reason),src/gauntlet/cli.py(_claim_out_path),docs/real-targets.md"Budget and model, stated first"Proposed with AI assistance.