What
Sprout is "a public evaluation reference implementation" and it grades itself. Cairn, the sibling reference implementation, is graded by two pinned external harnesses that fail closed when unreachable. Sprout should be held to the same arrangement.
Add sprout eval export --plumbline-bundle DIR: items from the committed suites (id, lang, behavior, expected facts, forbidden certification strings, sources from the cited chunks, answering_sources from the case's cited document, load_bearing on toxicity facts), responses from the deterministic pipeline, sources.jsonl from the store, checksums. Commit plumbline/target.toml with measured floors and reasons, plumbline.pin, and a gate script; commit gauntlet_target.py (a callable target over answer.py exposing citations and context_ids) with EN/ES refusal, adversarial, grounding, and false-positive cases and gauntlet.pin. Two CI jobs, audit and gauntlet, required in the ruleset; a baseline and an audit guard as cairn has.
Why it matters
"Groundedness is 100% by construction" is a claim Sprout's own harness verifies. An auditor that did not co-evolve with the pipeline verifying the same claim is a different kind of evidence, and the gate must be red, not green, when the auditor cannot run.
Scope
- Exporter (deterministic, offline), pins, gate scripts, target adapter, CI jobs,
docs/audits/ outputs.
claims-check entries for the new numbers.
Out of scope
- Changing any suite in either harness.
- Live HTTP grading of
sprout serve (a follow-up).
Done when
- Both gates pass on the committed bundle and cases.
- A planted "safe" certification fails plumbline
adversarial and gauntlet adversarial.
- An unreachable harness makes the job fail, not skip.
- The export is byte-identical across runs.
Pointers
src/sprout/eval/, src/sprout/answer.py, ChelseaKR/cairn plumbline-gate.sh, audit_guard.py, gauntlet_target.py; ChelseaKR/plumbline DESIGN.md "Evidence bundle format (v1)"
Proposed with AI assistance.
What
Sprout is "a public evaluation reference implementation" and it grades itself. Cairn, the sibling reference implementation, is graded by two pinned external harnesses that fail closed when unreachable. Sprout should be held to the same arrangement.
Add
sprout eval export --plumbline-bundle DIR: items from the committed suites (id,lang,behavior,expectedfacts,forbiddencertification strings,sourcesfrom the cited chunks,answering_sourcesfrom the case's cited document,load_bearingon toxicity facts), responses from the deterministic pipeline,sources.jsonlfrom the store, checksums. Commitplumbline/target.tomlwith measured floors and reasons,plumbline.pin, and a gate script; commitgauntlet_target.py(a callable target overanswer.pyexposing citations andcontext_ids) with EN/ES refusal, adversarial, grounding, and false-positive cases andgauntlet.pin. Two CI jobs,auditandgauntlet, required in the ruleset; a baseline and an audit guard as cairn has.Why it matters
"Groundedness is 100% by construction" is a claim Sprout's own harness verifies. An auditor that did not co-evolve with the pipeline verifying the same claim is a different kind of evidence, and the gate must be red, not green, when the auditor cannot run.
Scope
docs/audits/outputs.claims-checkentries for the new numbers.Out of scope
sprout serve(a follow-up).Done when
adversarialand gauntletadversarial.Pointers
src/sprout/eval/,src/sprout/answer.py,ChelseaKR/cairnplumbline-gate.sh,audit_guard.py,gauntlet_target.py;ChelseaKR/plumblineDESIGN.md "Evidence bundle format (v1)"Proposed with AI assistance.