What
The target contract carries text, citations, context_ids, refused, and escalated. The real-target adapters under real_targets/ need to report two more things: how many claims the target withheld, and what the harness's own quote check found. Today both travel as bracketed suffixes on the observed text and as counters in the provenance block, which docs/real-targets.md calls legible but not a contract, and names as the obvious next change.
Add an optional annotations list to TargetResponse and to the HTTP response body: {"kind": str, "detail": str, "count": int | null, "ref": str | null} with a closed enum of kinds (withheld_claim, quote_verified, quote_not_found, quote_unverifiable, adapter_note). The evidence pack renders them per case and totals them per gate.
Why it matters
A reviewer reading a pack cannot currently see per-quote check outcomes without opening the recording beside it. Making annotations part of the versioned structure keeps the pack self-describing, keeps text as exactly what the target said, and gives issue #31's recordings a field to replay. Gates keep validating only the fields they check today; annotations inform, they never infer.
Scope
annotations on TargetResponse, to_dict, response_from_payload (strict: unknown kind or malformed entry is a TargetProtocolError, exit 2).
- Results and pack schema version bump; a rendered "Adapter annotations" section in both forms.
- Migrate the three adapters off text suffixes; recordings carry the field.
- A target that omits the field behaves exactly as before.
Out of scope
- Any gate changing its verdict on an annotation.
- Judge rubric inputs.
Done when
- An HTTP target returning annotations renders them in
evidence.md and evidence.json.
- A payload with an unknown kind is refused with exit 2 and no results file.
- The committed real-target packs regenerate with no bracketed suffix left in any
observed text.
- A target omitting
annotations yields a byte-identical pack to today.
Pointers
Proposed with AI assistance.
What
The target contract carries
text,citations,context_ids,refused, andescalated. The real-target adapters underreal_targets/need to report two more things: how many claims the target withheld, and what the harness's own quote check found. Today both travel as bracketed suffixes on the observed text and as counters in the provenance block, whichdocs/real-targets.mdcalls legible but not a contract, and names as the obvious next change.Add an optional
annotationslist toTargetResponseand to the HTTP response body:{"kind": str, "detail": str, "count": int | null, "ref": str | null}with a closed enum of kinds (withheld_claim,quote_verified,quote_not_found,quote_unverifiable,adapter_note). The evidence pack renders them per case and totals them per gate.Why it matters
A reviewer reading a pack cannot currently see per-quote check outcomes without opening the recording beside it. Making annotations part of the versioned structure keeps the pack self-describing, keeps
textas exactly what the target said, and gives issue #31's recordings a field to replay. Gates keep validating only the fields they check today; annotations inform, they never infer.Scope
annotationsonTargetResponse,to_dict,response_from_payload(strict: unknown kind or malformed entry is aTargetProtocolError, exit 2).Out of scope
Done when
evidence.mdandevidence.json.observedtext.annotationsyields a byte-identical pack to today.Pointers
src/gauntlet/targets.py,src/gauntlet/results.py,src/gauntlet/report.pyreal_targets/quotecheck.py,real_targets/narration.py,real_targets/rawlog.pydocs/real-targets.md("What the contract could and could not express"), New packs must record quote-check outcomes so a replay verifies rather than skips #31Proposed with AI assistance.