From 65d956789ebc839136a664360df8f69896e0a90d Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sat, 18 Jul 2026 20:24:25 +0300 Subject: [PATCH 01/47] docs: specify external results importer --- ...-07-18-external-results-importer-design.md | 892 ++++++++++++++++++ 1 file changed, 892 insertions(+) create mode 100644 docs/superpowers/specs/2026-07-18-external-results-importer-design.md diff --git a/docs/superpowers/specs/2026-07-18-external-results-importer-design.md b/docs/superpowers/specs/2026-07-18-external-results-importer-design.md new file mode 100644 index 0000000..2e43931 --- /dev/null +++ b/docs/superpowers/specs/2026-07-18-external-results-importer-design.md @@ -0,0 +1,892 @@ +# External Results Importer Design + +**Date:** 2026-07-18 +**Status:** Approved design; implementation not started +**Target:** ChoiceBench after v0.2.0 +**Feature branch:** `feat/external-results-importer` +**Base:** `origin/choicebench` at `a01d85b1297bd7b557e0d2a16ad51c8c1c38e866` + +## Purpose + +ChoiceBench needs a first-class way to import row-level experiment results that +were generated outside ChoiceBench. The imported data must participate in +ChoiceBench's deterministic identity, provenance, manifest, result ownership, +integrity, status accounting, reading, and evaluation systems without implying +that ChoiceBench performed the original model inference. + +The first real consumer is the immutable Stage 1 paper-data freeze at +`/home/cotenthusiast/Projects/model-generalization/paper_data_freeze`. The core +design is intentionally paper-agnostic. A thin profile translates the Stage 1 +manifests into the same generic import specifications accepted for any other +external source. + +This stage imports and validates existing row-level evidence. It does not run a +model, execute repair queues, alter historical predictions, regenerate the +freeze, or copy the 3.4 GiB raw/cache evidence forest. + +## Existing architecture and design constraints + +ChoiceBench v0.2 is manifest-first: + +- canonical JSON plus SHA-256 supplies deterministic identities; +- dataset artifacts and selected rows have verified content identities; +- models, methods, prompts, conditions, and experiments have distinct + identities; +- a condition exclusively owns `results/.csv` and its sidecar; +- immutable manifests bind the condition grid and artifact paths; +- run state accounts for operational completion, gating, and failure; +- publication-grade readers revalidate manifests, snapshots, result sidecars, + ownership fields, and question coverage; +- evaluation verifies metric implementation identity and binds consumed result + bytes and condition accounting into the evaluation identity; +- `CHOICEBENCH_HOME`, or the current directory when it is unset, is the + writable workspace; +- installed commands are strict `argparse` entry points named + `choicebench-*`. + +The importer extends these systems. It must not use the legacy filename-based +result aggregator, fabricate an inference configuration, or route rows through +a fake model backend. + +Two existing boundaries require importer-specific hardening: + +1. The generic result writer can overwrite a same-condition file after only + verifying its sidecar. Import transactions must instead perform a full + identical-content no-op or refuse divergent content. +2. The existing reader validates result question IDs but does not compare + result-side question text, options, and gold labels with the archived input + snapshot. Imported evaluable rows must be joined to a trusted snapshot and + checked field by field before they can become evaluable imported result + artifacts. + +## Chosen architecture + +The importer creates ordinary ChoiceBench run directories and manifests. It +adds explicit import records, evidence artifacts, provenance dimensions, and +lineage while reusing existing identity, locking, atomic-write, manifest, +sidecar, reader, and evaluator facilities. + +The main components are: + +1. **Strict import schema** — dataclass-based YAML loading with the same + unknown-field rejection and primitive validation style as experiment + configuration. +2. **Source-adapter protocol** — an internal boundary that returns a stable + tabular representation plus exact source-column information. CSV is the + first implementation. +3. **Expected-dataset loader** — resolves the declared question set and trusted + question/gold/option data, creates a ChoiceBench run snapshot, and binds its + digest to the import identity. +4. **Row validator and normalizer** — joins by `question_id`, validates source + evidence, and creates ChoiceBench-native-format rows only when the evidence + is eligible for evaluation. +5. **Evidence store** — atomically archives only referenced row-level source + bytes in a content-addressed store and reuses identical blobs. +6. **Lineage/overlay engine** — validates repair and offline-transformation + overlays without modifying the base import. +7. **Import transaction** — locks a run, constructs and verifies the manifest, + performs collision-safe writes, and produces a machine-readable report. +8. **Manifest-aware evaluation extensions** — account for imported evidence + and compute metrics only for eligible conditions. +9. **Stage 1 profile** — translates sealed paper manifests and queue ledgers + into generic specifications without adding paper knowledge to the core. +10. **Installed CLI** — `choicebench-import-results` with a profile switch, + dry-run mode, strict validation, output-root selection, and reports. + +Likely module boundaries are: + +```text +src/choicebench/importing/ + schema.py + adapters.py + validation.py + evidence.py + lineage.py + engine.py + profiles/stage1_paper_freeze.py +src/choicebench/cli/import_results.py +``` + +Exact filenames may change during planning if existing module cohesion calls +for a smaller layout. The adapter/profile boundary and the absence of +paper-specific branches in the reusable engine are mandatory. + +## Provenance dimensions + +Import outcome, evidence quality, paper/execution scope, and inference origin +are orthogonal. No single status field may collapse them. + +### Result origin + +Every newly written manifest records exactly one `result_origin`: + +- `native_inference` +- `external_historical_inference` +- `external_repair_inference` +- `offline_transformation` + +Missing `result_origin` may imply `native_inference` only while reading legacy +manifest schema versions. New native and imported manifests write the field +explicitly. A ChoiceBench-native-format imported result artifact retains an +external origin; the format and evaluator do not change who performed the +inference. + +### Import state + +`import_state` describes the import operation only: + +- `validated` — validation succeeded in a dry-run/validate-only report; +- `imported` — the verified artifact transaction was committed; +- `failed` — validation or the transaction failed. + +A failed transaction does not leave a partially published run. Failure details +are recorded in the import report. An already committed run may account for a +source-level failed condition through `evidence_status`; that is distinct from +an importer failure. + +### Evidence status + +`evidence_status` describes scientific usability: + +- `complete` +- `qualified` +- `partial` +- `malformed` +- `recoverable` +- `failed` + +Qualifications, limitations, damaged IDs, recoverable IDs, and failure reasons +are explicit structured evidence. The existence or row count of a CSV never +upgrades evidence status. + +### Scope disposition + +`scope_disposition` describes inclusion and preservation policy: + +- `included` +- `excluded_from_paper_matrix` +- `held` +- `superseded` + +Excluded, held, failed, and superseded entries may reference genuine preserved +source evidence and lineage. They receive no new evaluable imported result +artifact or metric, but their evidence is not erased. + +### Executability + +`executable` is an explicit boolean wherever a repair/work authorization is +represented. It is not evaluation eligibility and the importer never interprets +it as permission to run model inference. In the Stage 1 profile, historical +condition import records are non-executing, while queue records reproduce the +authoritative true/false execution classification for later stages. + +The evaluator's eligibility predicate is based on successful import, +`evidence_status in {complete, qualified}`, and an included scope disposition. +It does not use `executable`. + +## Stable identity and audit provenance + +The importer separates identity-bearing semantic provenance from audit-only +provenance. + +### Identity-bearing provenance + +The deterministic import identity includes: + +- import schema and canonicalization versions; +- source format and exact source-byte SHA-256; +- a stable logical source identifier, when declared; +- source run identifier, repository, and commit declarations when known; +- source classification: `raw`, `canonical`, `derived`, `repaired`, or + `aggregate_only`; +- dataset identity, split, expected question-set digest, snapshot revision or + fingerprint, and exact run-snapshot content digest; +- model identity, backend/provider, and explicitly known revision; +- method identity, including historical identities that must not be merged; +- prompt/template identity and generation parameters when known; +- source column mapping, parsing policy, null policy, option mapping, and extra + field policy; +- evidence status, qualifications, limitations, and stable per-row failure + evidence; +- scope disposition and applicable execution authorization; +- repair authorizations, overlay bytes, replacement IDs/reasons, and stable + lineage; +- importer or transformation implementation identity; +- the stable projection of the generic import specification. + +Changing source bytes, an expected checksum, scientific mapping, identity +metadata, expected dataset, overlay, repaired IDs, or transformation identity +creates a different full import digest and therefore a different condition or +experiment identity. + +Imported source, dataset, model, method, prompt, and condition records carry +full SHA-256 digests in addition to short filesystem-safe IDs. Readers verify +the full digest before resolving an imported condition's artifact path. + +### Audit-only provenance + +The following are recorded for traceability but excluded from import, +condition, experiment, and evaluation identities: + +- import timestamp; +- the machine-local absolute source location used during this invocation; +- an absolute `CHOICEBENCH_HOME` or CLI output root; +- temporary paths; +- hostname and other machine-local execution details; +- machine-local import-report location. + +Stage 1's historical original path is preserved as audit provenance. Its +freeze-relative canonical/raw path is stable logical provenance. The absolute +path used to locate the freeze on this machine is audit-only. + +An import-specification digest is computed from a documented stable projection +that excludes operational locations and audit fields. Moving byte-identical +sources or the specification between machines does not change identity. + +Unknown provenance stays explicit. An unknown field contains a null value and a +reason such as `not recorded by producer`; it is never replaced with an +inferred request ID, time, token count, prompt, response, provider field, +commit, seed, revision, or configuration value. + +## Generic import specification + +The generic schema follows existing strict ChoiceBench configuration patterns. +It contains these logical sections: + +- schema/protocol version; +- stable import name and optional source-run declaration; +- one or more source artifacts; +- expected dataset declarations; +- model declarations; +- method declarations; +- prompt/template declarations; +- condition declarations; +- optional repair or transformation overlays; +- metric selection; +- notes and provenance evidence; +- audit-only location bindings. + +A source artifact declaration can express: + +- user-selected source location; +- expected SHA-256; +- stable logical path/name; +- format and format version; +- source run, repository, and commit declarations; +- raw/canonical/derived/repaired/aggregate-only classification; +- exact or allowed source schema; +- source column mapping; +- null and numeric policies; +- extra-field policy; +- arbitrary safe notes and evidence. + +A condition declaration can express: + +- source and expected-dataset references; +- model/backend/provider identity; +- benchmark and split identity; +- snapshot revision/fingerprint; +- historical method identity; +- prompt/template identity; +- known generation parameters; +- the four provenance/status dimensions and `executable` where relevant; +- expected question set; +- qualifications, limitations, and exact damaged/recoverable IDs; +- optional overlay references. + +Credential-named keys are refused by the same identity canonicalization rules +used elsewhere in ChoiceBench. Arbitrary executable Python objects, pickle, and +source-controlled dynamic import targets are not supported. + +## Path and source safety + +Absolute paths are not inherently unsafe. These explicitly user-selected paths +are valid after canonicalization and containment/symlink checks: + +- an absolute `CHOICEBENCH_HOME`; +- an absolute CLI-selected output root; +- an explicitly declared read-only external source such as the Stage 1 freeze. + +The trust boundary is who selects the output, not whether the path begins at the +filesystem root. + +An untrusted import specification cannot redirect writes. Output root selection +comes only from `CHOICEBENCH_HOME` or an explicit CLI option. Specifications +contain source locations and logical artifact relationships, but never an +output destination. Every derived output path is a fixed relative path under +the selected root. + +The importer: + +- resolves and canonicalizes the selected roots before writing; +- rejects `..` traversal or any derived output escaping the output root; +- rejects an output root or run directory that resolves through unsafe + symlinks; +- refuses source symlinks when their target/containment cannot be safely + established; +- never follows a source path into a directory tree for recursive import; +- opens only explicitly referenced row-level artifacts; +- computes SHA-256 independently from the opened bytes; +- parses the same bytes that were hashed, avoiding a hash/parse time-of-check + race; +- rejects a source whose bytes differ from the specification checksum; +- never executes artifact content; +- sanitizes untrusted exception text before logging or reporting it. + +## Evidence store and deduplication + +Every referenced row-level source is archived atomically in the run's import +evidence store. Storage is addressed by the actual source-byte SHA-256, for +example: + +```text +artifacts/imports/evidence/sha256/ce/ced2c5....csv +``` + +The suffix comes from the validated adapter format, not an untrusted filename. +An integrity sidecar records the digest, size, format, and stable source +references. Identical bytes imported by multiple conditions or specifications +reuse one evidence blob. A pre-existing blob is accepted only after full byte +digest and sidecar validation; divergent overwrite is refused. + +Evidence snapshots are written to a temporary file under the destination, +flushed, verified, and atomically renamed. The evidence index is also atomic and +self-digested. No raw repository, response-cache directory, checkpoint forest, +archive collection, or other unreferenced Stage 1 file is copied. + +For malformed evidence, the byte-identical source snapshot is authoritative. +The CSV adapter retains each logical record's exact raw byte span, including +quoting, delimiters, line endings, and embedded newlines. The condition +validation artifact records each affected question ID, the byte offsets, the +SHA-256 of those exact source-row bytes, and the validation failure. A +canonical source-row digest may be recorded as additional search/index data, +but never substitutes for the raw-row-byte digest. The importer does not coerce +a malformed prediction into an apparently valid normalized prediction. + +## CSV adapter and extension fields + +The first adapter supports CSV/tabular row sources. It reads source bytes as +data, never as code, preserves header order, and produces source cells without +allowing pandas' default NaN coercion to erase the distinction between an empty +cell and a literal `NaN` string. The import specification or profile declares +the accepted null representation and numeric parsing rules. + +Variable-option questions are supported through either: + +- an explicitly mapped structured choices column; or +- an ordered set/pattern of option columns whose present values determine the + row's option count. + +The adapter never assumes four choices. A three-option ARC row remains a +three-option row; an empty trailing option is not converted into a phantom +choice. + +Every source column has one disposition: + +1. mapped to a ChoiceBench common field; +2. preserved in `external_fields_json` under a source namespace; or +3. explicitly ignored with a documented reason. + +Normal mode preserves safe unmapped fields and reports them. Strict mode +requires every column to be declared and rejects extras. Neither mode silently +drops a source column. Credential-named or secret-bearing metadata is rejected +in both modes. + +`external_fields_json` preserves the original source field names, string/null +values, and method-specific diagnostics without forcing unrelated methods into +a lossy shared schema. This covers Stage-1/Stage-2 responses, matching output, +option scores and parse flags, permutations, votes, provider finish reasons, +failure counts, historical parse data, IHS per-option fields, and other safe +diagnostics. + +Historical method identity is taken from the specification/profile, not renamed +from the source filename or simplified to a current built-in method. The Stage +1 profile preserves at least: + +- `semantic_matching_v1` +- `text_extraction` +- `two_stage_v1` +- `two_stage_v2` +- `two_stage_v3` +- `independent_hypothesis` +- `cyclic_generation_majority` +- `PriDe` + +Source-row aliases such as `cyclic`, `twostage_semantic_match`, and +`two_prompt` are validation evidence interpreted only through explicit profile +mapping. + +## Row-level validation + +Validation is by question identity, never by row count alone. All errors name +the condition, source, field, and question ID where possible. No row is silently +dropped, padded, deduplicated, reinterpreted, or repaired. + +The validator checks: + +- required mapped fields and source schema; +- unique, non-null question IDs; +- exact expected membership for complete/qualified evidence; +- declared subset membership for partial/malformed/recoverable evidence; +- exact missing and unexpected IDs; +- duplicate IDs, including all duplicate locations; +- trusted gold answer equality; +- option count and ordered option identity/text where available; +- correct-option mapping into the actual variable-size option set; +- parsed prediction validity or an exact declared damaged/failure record; +- correctness consistency when recomputable; +- method-specific numeric fields for malformed values, NaN, and infinity; +- row ownership by condition; +- source-row model, benchmark, method, and split consistency when those fields + are present; +- evidence status against observed coverage and declared damaged rows; +- all source/checksum/specification digests; +- unknown/extra column policy; +- path and symlink safety. + +The trusted expected snapshot supplies the evaluative question text, choices, +correct option, and gold answer. Source-provided versions are compared with it +but never override it. + +For `complete` and `qualified` evidence, every expected question must have one +valid evaluable prediction. A declared provider failure may occupy a source row, +but that condition is not complete until a valid authorized repair produces a +derived condition. + +For `partial`, `malformed`, or `recoverable` evidence, known defects are legal +only when the exact IDs and reasons are declared. Any additional defect fails +validation. These sources can be archived and accounted, but they do not +produce an evaluable imported result artifact. + +## ChoiceBench-native-format imported result artifacts + +Only imported conditions with: + +- `import_state=imported`; +- `evidence_status=complete` or `qualified`; and +- `scope_disposition=included` + +produce an evaluable imported result artifact under +`results/.csv`. + +Its row ownership fields use the ordinary ChoiceBench experiment, condition, +dataset, model, method, prompt, benchmark, and split identities. Its result +sidecar additionally binds the full import/source/lineage digests and explicitly +states the external `result_origin`. The artifact is ChoiceBench-native in +format and validation only; it never claims ChoiceBench executed the model. + +Non-evaluable conditions retain condition records, evidence references, exact +validation artifacts, and lineage. They do not receive placeholder result CSVs +or empty metrics. + +## Idempotence and collision safety + +An import transaction holds the normal per-run process lock for validation of +existing state through final atomic publication. + +Repeating an import with the same stable specification projection, source +bytes, expected snapshot, mapping, status dimensions, and lineage yields the +same full identities. If all existing manifest, evidence, result, and sidecar +bytes validate and match, the command reports an idempotent no-op. + +If the run ID, condition ID, source digest path, or report identity already +exists with different verified content, the importer refuses the overwrite and +names the differing identity-bearing sections. The importer does not offer a +silent reset or destructive replacement path. + +## Repair overlays and offline transformations + +An overlay declaration contains: + +- immutable base import/condition full digest; +- mandatory base evidence and condition-validation-artifact checksums; +- a base evaluable-result checksum only when the base condition has one; +- overlay source and independently computed checksum; +- exact replacement question IDs; +- a reference to a separately declared, immutable repair-authorization record; +- one replacement reason per question; +- output source classification; +- result origin; +- transformation or repair implementation identity; +- stable lineage notes. + +An overlay cannot declare or expand its own authorization set. Before overlay +validation, the importer independently validates the referenced authorization +artifact, its checksum, its scope to the base condition, its exact authorized +IDs and reasons, and its authority/executable fields. For the Stage 1 profile, +the authorization must resolve to a checksum-verified authoritative queue +record with `queue_disposition=approved`, +`execution_authority=authoritative`, and `executable=true`. + +The overlay engine requires every replacement ID to be in that independently +validated authorization record, the expected dataset, and the base expected +set. It rejects duplicate IDs, unauthorized IDs, unexpected IDs, duplicate +overlay rows, multiple overlays that replace the same ID, inconsistent +gold/options/ownership, and conflicting overlay declarations. It validates +replacement rows with the same rules as base rows. + +Applying an overlay never changes the base evidence snapshot, base condition, +or optional base result. It creates a derived condition with a new identity and +complete base -> authorization -> overlay -> resulting-artifact lineage. The +specification declares the expected derived evidence status; the importer +recomputes it from validated coverage and remaining defects and refuses a +mismatch. A declaration of complete or qualified succeeds only if all remaining +defects are resolved. + +Offline transformations use the same derived-artifact mechanism. They record +input and output checksums plus transformation code identity. A specification +cannot ask ChoiceBench to import and execute arbitrary code. A transformation +is either: + +- a registered ChoiceBench implementation whose code identity is computed by + ChoiceBench; or +- a precomputed external output whose producing code identity/digest is + declared and whose output is independently validated. + +Deterministic rematching of surviving Stage-1 responses is represented as +`result_origin=offline_transformation`, not new model inference. + +## Manifest, reader, and evaluation behavior + +Imported manifests add an identity-bearing import section and condition-level +orthogonal provenance dimensions. Audit-only provenance is stored in a separate +non-identity section whose exclusion is explicit and validated. + +Manifest validation recomputes imported child full digests and verifies their +short IDs and fixed artifact paths. Result-side identity summaries and column +digests are compared with the CSV and manifest rather than merely stored. + +The publication-grade reader: + +- validates content-addressed evidence snapshots and condition validation + artifacts; +- validates evaluable imported result artifacts through the normal sidecar and + row-ownership path; +- refuses result artifacts for ineligible conditions; +- returns evaluable rows plus manifest accounting for all imported conditions; +- preserves compatibility with legacy native v0.2 manifests. + +Evaluation: + +- computes configured metrics for included complete/qualified imported + conditions; +- includes qualifications and limitations beside qualified metrics; +- accounts for partial, malformed, recoverable, failed, excluded, held, and + superseded evidence without metrics; +- reports condition counts by each orthogonal dimension rather than one lossy + status tally; +- never estimates metrics for missing, invalid, excluded, or held rows; +- continues to validate dataset snapshots, metric implementation identity, row + ownership, and result checksums. + +Evaluation identity includes only stable semantic provenance and lineage: + +- experiment and imported condition full digests; +- evidence status and scope disposition; +- stable qualification/limitation digest; +- result/evidence/lineage digests as applicable; +- configured metric and postprocessing implementation identities; +- exact consumed evaluable result checksums. + +It excludes timestamps, absolute paths, output roots, report locations, +hostnames, and temporary paths. + +## CLI and reports + +The installed command follows current project naming conventions: + +```bash +choicebench-import-results SPEC.yaml --run-id external-study + +choicebench-import-results SPEC.yaml \ + --run-id external-study \ + --dry-run \ + --strict \ + --report /tmp/external-study-validation.json + +choicebench-import-results /path/to/paper_data_freeze \ + --profile stage1-paper-freeze \ + --run-id paper-stage1 \ + --dry-run +``` + +Required options/behavior include: + +- real import; +- `--dry-run`/`--validate-only` aliases with no result/evidence writes; +- `--output-root` as an explicit trusted alternative to `CHOICEBENCH_HOME`; +- a required safe run ID; +- `--strict` source-column/schema validation; +- repeated explicit overlay arguments or overlays in the specification; +- human-readable summaries; +- optional machine-readable JSON report; +- mandatory independent checksum verification; +- sanitized, actionable errors. + +`--output-root` is parsed before importing modules that resolve +`CHOICEBENCH_HOME`. An absolute user-selected output root is valid. The import +specification cannot set or override it. + +The machine report includes: + +- stable import/specification/experiment IDs; +- audit-only input/output locations; +- counts by import state, evidence status, scope disposition, and origin; +- evaluable versus evidence-only condition counts; +- row coverage and exact defect IDs; +- source/checksum/schema results; +- extra-column dispositions; +- overlay authorizations and lineage; +- Stage 1 queue classifications when the profile is used; +- whether the operation wrote artifacts or was an idempotent no-op. + +## Stage 1 paper profile + +The profile is selected through the same command: + +```bash +choicebench-import-results FREEZE_ROOT \ + --profile stage1-paper-freeze \ + --run-id stage1-paper-import +``` + +It reads the current sealed manifests and validator-enforced invariants. It does +not regenerate them from `build_stage1_audit.py`, which predates the final +local-only PriDe amendment. + +Authoritative inputs include, as relevant: + +- `manifests/expected_matrix.csv` +- `manifests/canonical_results_manifest.json` and CSV cross-check +- `manifests/cell_status_matrix.csv` +- `manifests/frozen_artifact_index.csv` +- `manifests/approved_rerun_queue.csv` +- `manifests/held_or_declined_reruns.csv` +- `manifests/paper_scope_excluded_reruns.csv` +- `reports/canonical_freeze_report.md` +- `checksums/checksums.sha256` + +The profile joins by explicit `cell_id` and `canonical_artifact_id`; it does not +infer scientific identity from filenames. Freeze-relative paths are resolved +beneath the explicitly selected absolute freeze root and checked against +traversal and symlinks. + +The profile produces 102 preserved condition/evidence records: + +- exactly 100 have `scope_disposition=included` and constitute the intended + paper matrix; +- exactly two hosted-Qwen PriDe records have + `scope_disposition=excluded_from_paper_matrix` and are never counted as + intended paper cells. + +This permits the hosted Qwen PriDe MMLU record to remain complete and excluded, +while the hosted Qwen PriDe ARC record remains malformed and excluded. + +The intended-matrix evidence counts must be exactly: + +- 57 complete; +- 5 qualified; +- 6 recoverable; +- 4 partial, mapped explicitly from Stage 1 + `incomplete_requires_inference`; +- 28 malformed. + +Across all 102 preserved records, the two exclusions add one complete hosted +PriDe MMLU record and one malformed hosted PriDe ARC record. Thus the preserved +evidence totals are 58 complete, 5 qualified, 6 recoverable, 4 partial, and 29 +malformed, while the scope totals remain 100 included and two excluded. + +The profile verifies all source checksums used by its generated specifications. +It preserves the eight distinct historical CSV schemas and their method-specific +fields. + +Queue authority is explicit: + +- `approved_rerun_queue.csv`: 138 question-cells, executable true; +- `held_or_declined_reruns.csv`: 687 Gemini ARC IHS question-cells, + executable false and held; +- `paper_scope_excluded_reruns.csv`: three hosted PriDe question-cells, + executable false and excluded; +- `rerun_queue.csv`: historical forensic evidence only and never execution + authority. + +The profile imports no repair output and executes none of these queues. It only +records the authorization boundary needed for future overlays. + +For expected datasets, the profile uses explicit manifest-selected frozen +canonical evidence, constructs benchmark/split question snapshots, and +cross-checks question/gold/option identity across conditions. The generated +generic specifications contain the resulting explicit expected question sets +and digests; the reusable core receives no paper-specific inference rule. + +## Security boundaries + +The importer maintains ChoiceBench's fail-closed posture: + +- no pickle or executable-object deserialization; +- no execution of artifact content; +- no arbitrary dynamic transformation import from untrusted YAML; +- no source-provided checksum trust without independent hashing; +- no untrusted specification-controlled output path; +- no traversal or unsafe symlink resolution; +- no credential-named or secret-bearing persisted metadata; +- no divergent overwrite; +- no weakening of canonicalization, manifest, dataset, or evaluation checks; +- no recursive import of repositories, caches, archives, or directory forests; +- sanitized untrusted error text; +- process locking and atomic writes for every published artifact/index. + +Checksums prove internal byte consistency, not producer authenticity. Declared +source repository/model/provider facts remain declared unless independently +resolved from supplied immutable evidence. + +## Testing strategy + +Implementation follows test-driven development with small synthetic fixtures. +No historical dataset or large artifact is copied into the repository. + +Focused tests cover: + +1. Valid generic CSV import. +2. Deterministic import identity. +3. Idempotent repeated import. +4. Divergent overwrite refusal. +5. Source checksum mismatch. +6. Source mutation after specification creation. +7. Duplicate question IDs. +8. Missing question IDs. +9. Unexpected question IDs. +10. Gold-answer mismatch. +11. Variable option counts. +12. Three-option ARC-style rows. +13. Null and NaN behavior. +14. Invalid parsed predictions. +15. Correctness mismatch. +16. Unknown and extra fields in normal/strict modes. +17. Method-specific extra-field preservation. +18. Complete imported evidence. +19. Qualified imported evidence. +20. Partial and malformed imported evidence. +21. Imported-condition evaluation. +22. Incomplete-condition accounting. +23. Base plus valid repair overlay. +24. Unauthorized repair-ID rejection. +25. Conflicting repair-overlay rejection. +26. Repair lineage and identity. +27. Offline-transformation lineage. +28. Paper-profile manifest translation. +29. Hosted PriDe exclusions are not intended paper cells. +30. Held Gemini ARC IHS remains non-executable. +31. Approved queue is not confused with the forensic queue. +32. CLI dry-run. +33. CLI real import. +34. Import and evaluation outside the repository. +35. Wheel/sdist installation retains importer functionality. +36. Security/adversarial cases matching existing release tests. +37. Orthogonal status combinations, including complete+excluded and + malformed+excluded. +38. Audit paths/timestamps do not change import or evaluation identity. +39. New manifests always write explicit `result_origin`. +40. Legacy manifests alone may infer native origin. +41. Content-addressed evidence deduplication and divergent-blob refusal. +42. Explicit absolute source/output roots are accepted while spec-controlled + output and traversal are rejected. +43. Malformed row evidence remains byte/digest exact and non-evaluable. +44. Repair overlays on evidence-only bases do not require a nonexistent base + result checksum. +45. Overlay-declared authorization cannot self-authorize replacement IDs. +46. Declared derived evidence status must equal the recomputed status. + +Existing native-run tests remain unchanged or receive compatibility coverage. +The full suite must continue to pass. + +## Stage 1 end-to-end validation + +After synthetic tests pass: + +1. Run the Stage 1 profile in dry-run mode across the entire intended matrix. +2. Confirm 100 intended and two excluded preserved cells. +3. Confirm the exact evidence-status and queue counts listed above. +4. Confirm held and excluded work stays non-executable. +5. Confirm every referenced canonical source checksum. +6. Perform a real import into a temporary absolute `CHOICEBENCH_HOME` outside + both repositories. +7. Verify the freeze has no filesystem changes. +8. Use native ChoiceBench reading/evaluation to demonstrate: + - one complete imported condition; + - one qualified imported condition; + - one incomplete or malformed evidence-only condition; + - one variable-option ARC condition. +9. Write the end-to-end report outside the repository. + +Generated imports, reports, build artifacts, temporary workspaces, and +historical evidence remain uncommitted. + +## Documentation and packaging + +User documentation will explain: + +- what external import means and why it is not native inference; +- supported CSV format and extension-field preservation; +- import schema fields and examples; +- generic, dry-run, strict, output-root, and Stage 1 profile usage; +- deterministic identity versus audit provenance; +- checksum and idempotence behavior; +- orthogonal status dimensions and evaluation eligibility; +- repair and offline-transformation lineage; +- security boundaries and limitations; +- how to implement a future adapter/profile. + +A concise changelog entry will be added because the repository records notable +features there. The package version will not change. + +Before the pull request, validation includes targeted tests, the full suite, +repository lint/type/static checks that actually exist, security/adversarial +tests, `python -m build`, `python -m twine check dist/*`, the existing isolated +wheel/sdist smoke test, the Stage 1 dry run, the isolated real import, and +representative native evaluation. + +## Review workflow + +The work follows these independent gates: + +1. architecture reconstruction and design review; +2. independent written-specification review; +3. test-driven implementation by a fresh implementer where tasks are safely + separable; +4. independent specification-compliance review; +5. independent code-quality review; +6. independent adversarial/security review; +7. fresh-context final review of the complete diff and validation evidence. + +Reviewers are read-only and do not inherit implementer conclusions. Confirmed +blockers are fixed and relevant validation is rerun before advancing. + +## Out of scope + +This stage does not: + +- modify Stage 1 source artifacts; +- run model inference; +- execute any repair, held, or excluded queue; +- start the final 2x2 experiment; +- rewrite the paper; +- alter historical predictions; +- implement unrelated evaluation modules; +- create a general data-platform abstraction; +- import response caches, checkpoints, or raw repository forests; +- commit imported outputs or historical data; +- change the package version; +- merge the eventual pull request or create a release/tag. + +## Success criteria + +The design is successful when a strict, reusable specification can import +external CSV rows into collision-safe ChoiceBench manifests and +ChoiceBench-native-format imported result artifacts; every artifact explicitly +retains external inference origin; non-evaluable evidence remains preserved and +honestly accounted; deterministic identities exclude machine-local audit data; +repair lineage is immutable and authorized; the Stage 1 profile reproduces all +sealed counts without modifying the freeze; native reading/evaluation and +installed-package workflows work outside the repository; and the feature is +submitted on its isolated branch as an unmerged pull request. From b26a46effc50a0c52d2b05538a98ee890e44bd50 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sat, 18 Jul 2026 21:19:43 +0300 Subject: [PATCH 02/47] docs: refine external importer identity design --- ...-07-18-external-results-importer-design.md | 754 ++++++++++++++---- 1 file changed, 603 insertions(+), 151 deletions(-) diff --git a/docs/superpowers/specs/2026-07-18-external-results-importer-design.md b/docs/superpowers/specs/2026-07-18-external-results-importer-design.md index 2e43931..3c08651 100644 --- a/docs/superpowers/specs/2026-07-18-external-results-importer-design.md +++ b/docs/superpowers/specs/2026-07-18-external-results-importer-design.md @@ -1,7 +1,7 @@ # External Results Importer Design **Date:** 2026-07-18 -**Status:** Approved design; implementation not started +**Status:** Independently reviewed; awaiting implementation-planning approval **Target:** ChoiceBench after v0.2.0 **Feature branch:** `feat/external-results-importer` **Base:** `origin/choicebench` at `a01d85b1297bd7b557e0d2a16ad51c8c1c38e866` @@ -32,7 +32,8 @@ ChoiceBench v0.2 is manifest-first: - dataset artifacts and selected rows have verified content identities; - models, methods, prompts, conditions, and experiments have distinct identities; -- a condition exclusively owns `results/.csv` and its sidecar; +- in manifest v2, a condition exclusively owns + `results/.csv` and its sidecar; - immutable manifests bind the condition grid and artifact paths; - run state accounts for operational completion, gating, and failure; - publication-grade readers revalidate manifests, snapshots, result sidecars, @@ -55,9 +56,9 @@ Two existing boundaries require importer-specific hardening: identical-content no-op or refuse divergent content. 2. The existing reader validates result question IDs but does not compare result-side question text, options, and gold labels with the archived input - snapshot. Imported evaluable rows must be joined to a trusted snapshot and - checked field by field before they can become evaluable imported result - artifacts. + snapshot. Imported evaluable rows must be joined to a declared expected + snapshot with an explicit trust classification and checked field by field + before they can become evaluable imported result artifacts. ## Chosen architecture @@ -74,23 +75,26 @@ The main components are: 2. **Source-adapter protocol** — an internal boundary that returns a stable tabular representation plus exact source-column information. CSV is the first implementation. -3. **Expected-dataset loader** — resolves the declared question set and trusted - question/gold/option data, creates a ChoiceBench run snapshot, and binds its - digest to the import identity. -4. **Row validator and normalizer** — joins by `question_id`, validates source +3. **Expected-dataset loader** — resolves the declared question set and + trust-qualified question/gold/option data, creates a ChoiceBench run snapshot, + and binds its digest to the realization identity. +4. **Identity/realization builder** — preserves the existing scientific + `condition_id` while assigning every imported, repaired, transformed, or + native execution record a separate immutable `realization_id`. +5. **Row validator and normalizer** — joins by `question_id`, validates source evidence, and creates ChoiceBench-native-format rows only when the evidence is eligible for evaluation. -5. **Evidence store** — atomically archives only referenced row-level source +6. **Evidence store** — atomically archives only referenced row-level source bytes in a content-addressed store and reuses identical blobs. -6. **Lineage/overlay engine** — validates repair and offline-transformation +7. **Lineage/overlay engine** — validates repair and offline-transformation overlays without modifying the base import. -7. **Import transaction** — locks a run, constructs and verifies the manifest, +8. **Import transaction** — locks a run, constructs and verifies the manifest, performs collision-safe writes, and produces a machine-readable report. -8. **Manifest-aware evaluation extensions** — account for imported evidence +9. **Manifest-aware evaluation extensions** — account for imported evidence and compute metrics only for eligible conditions. -9. **Stage 1 profile** — translates sealed paper manifests and queue ledgers +10. **Stage 1 profile** — translates sealed paper manifests and queue ledgers into generic specifications without adding paper knowledge to the core. -10. **Installed CLI** — `choicebench-import-results` with a profile switch, +11. **Installed CLI** — `choicebench-import-results` with a profile switch, dry-run mode, strict validation, output-root selection, and reports. Likely module boundaries are: @@ -116,19 +120,53 @@ paper-specific branches in the reusable engine are mandatory. Import outcome, evidence quality, paper/execution scope, and inference origin are orthogonal. No single status field may collapse them. -### Result origin +### Structured result origin and derivation + +A single scalar origin is insufficient because a repaired realization can +contain predictions from several producers. Every newly written manifest uses +a structured `result_origin` record with separate derivation and prediction +origin information: + +```yaml +result_origin: + derivation_origin: repair_overlay + prediction_origins: + - external_historical_inference + - native_inference + ordered_row_origin_digest: + lineage_component_origin_digest: +``` + +`derivation_origin` describes how the realization was created: + +- `native_execution` +- `external_import` +- `repair_overlay` +- `offline_transformation` -Every newly written manifest records exactly one `result_origin`: +`prediction_origin` describes who or what produced each prediction: - `native_inference` - `external_historical_inference` - `external_repair_inference` -- `offline_transformation` -Missing `result_origin` may imply `native_inference` only while reading legacy -manifest schema versions. New native and imported manifests write the field -explicitly. A ChoiceBench-native-format imported result artifact retains an -external origin; the format and evaluator do not change who performed the +Every evaluable row has a non-null `prediction_origin` and +`prediction_lineage_id`. The realization record and result sidecar contain the +sorted set/counts of constituent origins plus a digest of the ordered per-row +assignment. When more than one origin occurs, the per-row mapping is mandatory; +the importer never substitutes an uninformative `mixed` label. + +An offline transformation is a derivation, not new model inference. A rematched +row retains the prediction origin of its underlying model response while its +lineage points to an `offline_transformation` component and its implementation +identity. Every lineage node (base evidence, repair evidence, authorization, +and transformation) records its own origin. + +Missing `result_origin` may imply a homogeneous native realization only while +reading legacy manifest-v2 records. All newly written native and imported +manifest-v3 realizations record the structured field explicitly. A +ChoiceBench-native-format imported result artifact retains its external +prediction origins; the format and evaluator do not change who performed the inference. ### Import state @@ -139,10 +177,12 @@ inference. - `imported` — the verified artifact transaction was committed; - `failed` — validation or the transaction failed. -A failed transaction does not leave a partially published run. Failure details -are recorded in the import report. An already committed run may account for a -source-level failed condition through `evidence_status`; that is distinct from -an importer failure. +A failed transaction does not leave a partially published final run. It may +leave only a clearly marked unpublished staging directory after process or host +failure; readers never resolve staging paths. Failure details are recorded in +the import report. An already committed run may account for a source-level +failed condition through `evidence_status`; that is distinct from an importer +failure. ### Evidence status @@ -184,49 +224,163 @@ The evaluator's eligibility predicate is based on successful import, `evidence_status in {complete, qualified}`, and an included scope disposition. It does not use `executable`. -## Stable identity and audit provenance - -The importer separates identity-bearing semantic provenance from audit-only -provenance. - -### Identity-bearing provenance - -The deterministic import identity includes: - -- import schema and canonicalization versions; -- source format and exact source-byte SHA-256; -- a stable logical source identifier, when declared; -- source run identifier, repository, and commit declarations when known; -- source classification: `raw`, `canonical`, `derived`, `repaired`, or - `aggregate_only`; -- dataset identity, split, expected question-set digest, snapshot revision or - fingerprint, and exact run-snapshot content digest; -- model identity, backend/provider, and explicitly known revision; -- method identity, including historical identities that must not be merged; -- prompt/template identity and generation parameters when known; -- source column mapping, parsing policy, null policy, option mapping, and extra - field policy; -- evidence status, qualifications, limitations, and stable per-row failure - evidence; -- scope disposition and applicable execution authorization; -- repair authorizations, overlay bytes, replacement IDs/reasons, and stable - lineage; -- importer or transformation implementation identity; -- the stable projection of the generic import specification. - -Changing source bytes, an expected checksum, scientific mapping, identity -metadata, expected dataset, overlay, repaired IDs, or transformation identity -creates a different full import digest and therefore a different condition or -experiment identity. - -Imported source, dataset, model, method, prompt, and condition records carry -full SHA-256 digests in addition to short filesystem-safe IDs. Readers verify -the full digest before resolving an imported condition's artifact path. +## Identity layers and audit provenance + +The importer separates semantic condition identity, realization identity, +result artifact identity, evaluation identity, and audit-only provenance. These +layers are related but never interchangeable. + +### Semantic condition identity + +`condition_id` remains the scientific grid-cell identity already constructed by +ChoiceBench v0.2. Its canonical payload contains only scientifically meaningful +condition metadata: + +- dataset artifact and exact selection identities, benchmark, and split; +- model identity, including the backend/provider and known consumed generation + settings or revision; +- method identity, including the historical method name, effective parameters, + preflight identity, and method implementation identity where recoverable; +- prompt/template identity; +- seed, calibration/preflight identity, and applicable protocol settings. + +This matches the current model, method, prompt, dataset-selection, and condition +construction rather than defining an importer-only scientific identity. The +canonical payload has a full `condition_digest`; `condition_id` is its existing +filesystem-safe short form. + +Source CSV bytes, CSV dialect/mapping, importer code, evidence status, scope, +authorization, overlay bytes, repair IDs, and lineage are not semantic +condition fields. Changing any of them alone leaves `condition_id` unchanged. +Changing the dataset selection, model, method, prompt, generation semantics, +seed, or protocol settings changes the semantic condition. + +### Import/result realization identity + +A `realization_id` identifies one immutable evidence/provenance record that +asserts or derives results for a semantic condition. Multiple historical, +repaired, transformed, or native realizations may share one `condition_id`. +The canonical realization payload includes: + +- the full semantic `condition_digest`; +- stable import-specification digest; +- logical source records, source classifications, formats, exact source-byte + SHA-256 values, and stable source provenance; +- expected-input snapshot and question-set digests; +- adapter/importer implementation identity, complete CSV dialect, source + mapping, null/numeric/option/extra-field policies; +- validation-findings digest, evidence status, qualification/limitation/defect + digests, and scope disposition; +- structured result origin and ordered prediction-origin assignment digest; +- parent realization/evidence/result digests; +- typed authorization digest; +- overlay source bytes, replacement IDs/reasons, and transformation + input/pre-ownership-output/code digests. + +The output digest in a realization is never the final ChoiceBench CSV digest. +It is either the checksum of a separately supplied precomputed overlay/output +source, or a canonical transformation payload produced before ChoiceBench row +ownership is injected. That canonical payload excludes the child +`realization_id`, result-artifact and experiment identities, output paths, +timestamps, and machine locations. The final CSV checksum appears only in the +post-manifest result-artifact identity, sidecar, and run state. + +Lineage-component IDs included in the realization are likewise computed only +from their operation type, parent/source/evidence/authorization digests, +implementation identity, stable parameters, question ID, and pre-ownership +input/output digests. They never include the child realization, child result +artifact, or current experiment identity. The realization can therefore bind +the ordered row-to-lineage-component mapping without a self-reference. + +`import_state` is not identity-bearing: a dry-run realization that moves from +`validated` to atomically committed `imported` keeps the same planned identity. +A failed transaction publishes no realization. + +Changing source bytes, source checksum, CSV dialect, source mapping, stable +provenance, importer/adapter implementation, evidence classification, scope, +authorization, overlay, repaired IDs, prediction-origin assignment, or +transformation always changes `realization_id`. It changes `condition_id` only +when the scientific condition payload also changes. + +A repaired or transformed output that implements the same intended dataset x +model x method x prompt protocol is a derived realization of the same semantic +condition. If the repair changes the model, prompt, method algorithm, generation +semantics, dataset selection, or protocol, it belongs to a new semantic +condition. Deterministic rematching that reconstructs the declared +`semantic_matching_v1` protocol from preserved Stage-1 responses shares the +base semantic condition; a different matcher definition would change the +method and condition identities. + +### Manifest and experiment identity + +Manifest schema v3 separates `semantic_conditions` from `realizations`. +Semantic-condition records contain no realization back-reference in their +condition identity; grouping is derived from the realization table. Realizations +reference a semantic `condition_id` and own fixed result/evidence +paths plus an expected result ownership/schema contract. The immutable manifest +does not contain the result-artifact ID or digest, which cannot be known until +the exact result bytes exist. New native runs normally have one realization per +condition; imported runs may retain multiple alternative or derived +realizations. Run state is keyed by `realization_id`, not by semantic condition. + +The existing `experiment_id` remains the immutable manifest/run identity. Its +identity payload includes the semantic grid and the selected realization +records, so changing source bytes or lineage changes the experiment identity +without changing the shared scientific condition. A separate +`semantic_grid_digest` over datasets/models/methods/prompts/semantic conditions +supports cross-realization comparison. Legacy manifest-v2 records retain their +existing experiment and condition IDs and are normalized in memory as one +native realization per condition. + +### Result artifact identity + +An evaluable CSV has a separate `result_artifact_id` and full +`result_artifact_digest`. The digest binds: + +- semantic condition and realization full digests; +- manifest experiment ID; +- exact CSV SHA-256, row count, ordered columns, and columns digest; +- row-ownership summary and ordered question-identity digest; +- structured prediction-origin counts and ordered assignment digest; +- evidence and lineage digests. + +The result ID is calculated after the exact CSV bytes are produced. It is stored +in the integrity sidecar and run state, not inside the CSV or immutable manifest. +This is an explicit two-phase boundary: first the manifest/experiment fixes the +semantic conditions, realizations, output contracts, and safe paths; then the +atomic result publication computes and records the result-artifact identity. +The result-artifact digest may therefore bind the already fixed experiment ID +without an identity cycle. Manifest-v3 result paths are fixed by realization, +`results/.csv` and +`results/.artifact.json`; a second realization therefore cannot +collide with the shared semantic condition. Legacy v2 paths remain +`results/.csv`. + +### Evaluation identity + +Evaluation units are realizations grouped under semantic conditions. Alternative +realizations are never silently concatenated. The evaluation identity binds: + +- experiment digest and explicit realization-selection policy; +- every accounted semantic condition and realization full digest; +- evidence/scope/qualification/limitation and stable lineage digests; +- applicable result artifact ID/digest and exact consumed CSV SHA-256; +- metric, parser/scorer/postprocessing, and evaluator implementation identities. + +Ineligible evidence-only realizations remain identity-bound accounting units +with null result fields. Reports expose both +`conditions[condition_id].realization_ids` and realization-level metrics/status. +Changing source/mapping/provenance changes the realization and applicable +result/evaluation identities even when the semantic condition is unchanged. + +Imported sources, datasets, models, methods, prompts, semantic conditions, +realizations, and results carry full SHA-256 digests in addition to short IDs. +Readers verify the full digest before resolving any imported artifact path. ### Audit-only provenance -The following are recorded for traceability but excluded from import, -condition, experiment, and evaluation identities: +The following are recorded for traceability but excluded from semantic +condition, realization, result-artifact, experiment, and evaluation identities: - import timestamp; - the machine-local absolute source location used during this invocation; @@ -261,6 +415,8 @@ It contains these logical sections: - method declarations; - prompt/template declarations; - condition declarations; +- realization and structured-origin declarations; +- typed authorization declarations; - optional repair or transformation overlays; - metric selection; - notes and provenance evidence; @@ -272,6 +428,7 @@ A source artifact declaration can express: - expected SHA-256; - stable logical path/name; - format and format version; +- the complete identity-bearing decoding and CSV-dialect declaration; - source run, repository, and commit declarations; - raw/canonical/derived/repaired/aggregate-only classification; - exact or allowed source schema; @@ -280,6 +437,10 @@ A source artifact declaration can express: - extra-field policy; - arbitrary safe notes and evidence. +An expected-dataset declaration also states its reference kind and trust basis, +as defined below, rather than representing every reference snapshot as equally +trusted. + A condition declaration can express: - source and expected-dataset references; @@ -345,9 +506,12 @@ artifacts/imports/evidence/sha256/ce/ced2c5....csv The suffix comes from the validated adapter format, not an untrusted filename. An integrity sidecar records the digest, size, format, and stable source -references. Identical bytes imported by multiple conditions or specifications -reuse one evidence blob. A pre-existing blob is accepted only after full byte -digest and sidecar validation; divergent overwrite is refused. +references. Identical bytes referenced by multiple source declarations or +conditions within the same run reuse one evidence blob. A pre-existing blob is +accepted only after full byte digest and sidecar validation; divergent overwrite +is refused. Separate immutable runs archive their own content-addressed copy even +when bytes match; the initial design deliberately avoids a mutable global store +or cross-run hardlink trust boundary. Evidence snapshots are written to a temporary file under the destination, flushed, verified, and atomically renamed. The evidence index is also atomic and @@ -356,9 +520,9 @@ archive collection, or other unreferenced Stage 1 file is copied. For malformed evidence, the byte-identical source snapshot is authoritative. The CSV adapter retains each logical record's exact raw byte span, including -quoting, delimiters, line endings, and embedded newlines. The condition -validation artifact records each affected question ID, the byte offsets, the -SHA-256 of those exact source-row bytes, and the validation failure. A +quoting, delimiters, line endings, and embedded newlines. The +realization-validation artifact records each affected question ID, the byte +offsets, the SHA-256 of those exact source-row bytes, and the validation failure. A canonical source-row digest may be recorded as additional search/index data, but never substitutes for the raw-row-byte digest. The importer does not coerce a malformed prediction into an apparently valid normalized prediction. @@ -371,6 +535,56 @@ allowing pandas' default NaN coercion to erase the distinction between an empty cell and a literal `NaN` string. The import specification or profile declares the accepted null representation and numeric parsing rules. +### Deterministic decoding and dialect + +Parsing behavior is fully declared and identity-bearing in the realization +payload because it determines logical records and cell values. The initial CSV +adapter has these explicit defaults: + +```yaml +csv: + encoding: utf-8 + bom_policy: forbid + decoding_errors: strict + delimiter: "," + quote_character: '"' + escape_character: null + double_quote: true + line_terminators: [crlf, lf, cr] + mixed_line_terminators: allow + final_record_without_terminator: allow + blank_record_policy: reject + skip_initial_space: false + header: first_logical_record + strict_syntax: true +``` + +The initial adapter accepts only `encoding: utf-8`. `bom_policy` may instead be +explicitly set to `strip_utf8_bom`; stripping is limited to one leading UTF-8 +BOM and the byte offsets still refer to the original source. Decoding always +fails closed on invalid byte sequences; a replacement-character or ignore +policy is not supported. Delimiter, quote, and non-null escape characters are +restricted in the first adapter to distinct single-byte ASCII characters. +`double_quote` controls whether two consecutive quote characters inside a +quoted field represent one literal quote. + +The listed line terminators are the only recognized record separators and are +recognized only outside a quoted field. Their order is canonical and longest +first, so CRLF is one terminator rather than CR followed by LF. The mixed-line +policy, acceptance of an unterminated final logical record, and blank-record +policy are explicit. Raw spans include the record terminator when one is +present. A binary logical-record scanner uses the declared quote, escape, and +doubled-quote rules before text decoding, retains start/end byte offsets in the +original source, and treats newlines inside quoted fields as field bytes. It +then decodes and parses those exact spans under the same declaration. +Consequently the malformed-row raw-byte-span guarantee remains well-defined for +quoted records containing embedded CR, LF, or CRLF. The fixed UTF-8 encoding +declaration and every declared BOM, delimiter, quoting, escaping, +doubled-quote, terminator, mixed-terminator, final/blank-record, whitespace, +header, decoding-error, and strictness field are identity-bearing. A future +adapter version that supports another encoding necessarily produces a distinct +realization identity, though not a different semantic condition. + Variable-option questions are supported through either: - an explicitly mapped structured choices column; or @@ -416,11 +630,38 @@ Source-row aliases such as `cyclic`, `twostage_semantic_match`, and `two_prompt` are validation evidence interpreted only through explicit profile mapping. +## Expected-dataset reference trust + +The reusable core distinguishes two reference kinds: + +- `independent_input_snapshot` — benchmark input artifacts and selection records + that existed independently of the result rows being imported. The declaration + binds their exact checksums, selection identity, transformation chain, and any + independently known publisher revision/fingerprint. Validation may claim + question, option, and gold consistency against this snapshot, while separately + stating the authenticity limit of its provenance chain. +- `profile_derived_reference_snapshot` — a reference reconstructed only from + result-side evidence because no independent input artifact is available. It + requires at least two explicitly declared independent source groups when + possible, exact cross-source agreement for question/gold/options, a complete + derivation record, conflict refusal, and a stable derivation digest. It may + support internal consistency and membership validation, but reports must not + claim independent benchmark truth or upstream dataset authenticity. + +The reference kind, source/checksum chain, declared trust level, cross-source +policy, and derivation digest are realization- and evaluation-identity-bearing. +They do not change semantic dataset identity unless the selected questions or +their semantic content changes. Profiles may not silently upgrade a +result-derived reference to an independent snapshot merely because several +result files agree. + ## Row-level validation Validation is by question identity, never by row count alone. All errors name -the condition, source, field, and question ID where possible. No row is silently -dropped, padded, deduplicated, reinterpreted, or repaired. +the condition, realization, source, field, and question ID where possible. No +imported result row is silently dropped, padded, deduplicated, reinterpreted, or +repaired. Any declared expected-dataset derivation is a separate, identity-bound +input transformation and cannot be used to excuse duplicate result rows. The validator checks: @@ -430,7 +671,7 @@ The validator checks: - declared subset membership for partial/malformed/recoverable evidence; - exact missing and unexpected IDs; - duplicate IDs, including all duplicate locations; -- trusted gold answer equality; +- declared reference-snapshot gold answer equality, qualified by its trust kind; - option count and ordered option identity/text where available; - correct-option mapping into the actual variable-size option set; - parsed prediction validity or an exact declared damaged/failure record; @@ -444,14 +685,15 @@ The validator checks: - unknown/extra column policy; - path and symlink safety. -The trusted expected snapshot supplies the evaluative question text, choices, +The declared expected snapshot supplies the evaluative question text, choices, correct option, and gold answer. Source-provided versions are compared with it -but never override it. +but never override it. The report qualifies the resulting validation claim by +the snapshot's declared reference kind and trust chain. For `complete` and `qualified` evidence, every expected question must have one valid evaluable prediction. A declared provider failure may occupy a source row, -but that condition is not complete until a valid authorized repair produces a -derived condition. +but that realization is not complete until a valid authorized repair produces a +derived realization whose recomputed coverage supports that status. For `partial`, `malformed`, or `recoverable` evidence, known defects are legal only when the exact IDs and reasons are declared. Any additional defect fails @@ -460,63 +702,114 @@ produce an evaluable imported result artifact. ## ChoiceBench-native-format imported result artifacts -Only imported conditions with: +Only realizations with: - `import_state=imported`; - `evidence_status=complete` or `qualified`; and - `scope_disposition=included` produce an evaluable imported result artifact under -`results/.csv`. +`results/.csv`. Its row ownership fields use the ordinary ChoiceBench experiment, condition, -dataset, model, method, prompt, benchmark, and split identities. Its result -sidecar additionally binds the full import/source/lineage digests and explicitly -states the external `result_origin`. The artifact is ChoiceBench-native in -format and validation only; it never claims ChoiceBench executed the model. - -Non-evaluable conditions retain condition records, evidence references, exact -validation artifacts, and lineage. They do not receive placeholder result CSVs -or empty metrics. +dataset, model, method, prompt, benchmark, and split identities and add +`realization_id`, `prediction_origin`, and `prediction_lineage_id`. The result +sidecar additionally binds the full realization, result-artifact, source, +authorization, lineage, and ordered row-origin digests and explicitly records +the structured `result_origin`. It is a ChoiceBench-native-format imported +result artifact in format and validation only; it never claims ChoiceBench +executed the model. + +Non-evaluable realizations retain their semantic-condition reference, evidence +references, exact validation artifacts, and lineage. They do not receive +placeholder result CSVs or empty metrics. ## Idempotence and collision safety An import transaction holds the normal per-run process lock for validation of existing state through final atomic publication. -Repeating an import with the same stable specification projection, source -bytes, expected snapshot, mapping, status dimensions, and lineage yields the -same full identities. If all existing manifest, evidence, result, and sidecar -bytes validate and match, the command reports an idempotent no-op. +For a new run, it reuses ChoiceBench's parent-level +`.locks/.manifest.lock`, creates a uniquely named sibling staging +directory on the same filesystem as `runs/`, and writes the complete +manifest, content-addressed evidence, validation artifacts, evaluable results, +sidecars, and final run state there. Every file and cross-reference is validated +from the staged tree; files and directories are flushed/fsynced where the +platform supports it. Publication is one atomic, no-replace directory rename +from the staged tree to the previously absent final run path while the lock is +held. The implementation never uses a replacement rename for a run directory. +If no safe atomic no-replace publication is available on the platform, it fails +closed rather than falling back to incremental final-path writes. + +Validation failure removes only the importer-owned staging tree when safe. A +crash may leave an owner-marked staging tree, but it is not a run, is never +enumerated by readers/evaluators, and can be verified and garbage-collected by a +separate safe maintenance operation. Because the content-addressed evidence +store is inside that staging run, no externally published blob points at an +uncommitted run. Machine-readable reports written outside the run are published +separately and cannot make a failed run appear committed. + +For an already existing final run, the importer performs verify-only behavior: +it validates the entire manifest, state, evidence, result, and sidecar graph. An +exact match is an idempotent no-op; any mismatch is a refusal. It never stages +over or replaces an existing run. -If the run ID, condition ID, source digest path, or report identity already -exists with different verified content, the importer refuses the overwrite and -names the differing identity-bearing sections. The importer does not offer a -silent reset or destructive replacement path. +Repeating an import with the same stable specification projection, source +bytes, expected snapshot, mapping, status dimensions, origins, authorization, +and lineage yields the same semantic condition, realization, result-artifact, +experiment, and evaluation identities. If all existing manifest, evidence, +result, and sidecar bytes validate and match, the command reports an idempotent +no-op. + +If a run ID, realization ID, result-artifact path, source-digest path, or report +identity already exists with different verified content, the importer refuses +the overwrite and names the differing identity-bearing sections. A shared +semantic `condition_id` is not itself a collision: it may legitimately group +several separately addressed realizations. The importer does not offer a silent +reset or destructive replacement path. ## Repair overlays and offline transformations An overlay declaration contains: -- immutable base import/condition full digest; -- mandatory base evidence and condition-validation-artifact checksums; -- a base evaluable-result checksum only when the base condition has one; +- immutable base realization and semantic-condition full digests; +- mandatory base evidence and realization-validation-artifact checksums; +- a base evaluable-result checksum only when the base realization has one; - overlay source and independently computed checksum; - exact replacement question IDs; -- a reference to a separately declared, immutable repair-authorization record; +- an authorization type and reference to a separately declared, immutable, + checksum-verified authorization record; - one replacement reason per question; - output source classification; -- result origin; +- structured result origin and per-row prediction origins; - transformation or repair implementation identity; - stable lineage notes. An overlay cannot declare or expand its own authorization set. Before overlay validation, the importer independently validates the referenced authorization -artifact, its checksum, its scope to the base condition, its exact authorized -IDs and reasons, and its authority/executable fields. For the Stage 1 profile, -the authorization must resolve to a checksum-verified authoritative queue -record with `queue_disposition=approved`, -`execution_authority=authoritative`, and `executable=true`. +artifact, its checksum, its scope to the base semantic condition and realization, +its exact authorized IDs and reasons, its authority, and the declared operation +type. The transformation/overlay specification may only reference authorization; +it cannot serve as the authority for its own IDs. + +Two authorization types are distinct and non-interchangeable: + +- `inference_repair` authorizes new model inference. For the Stage 1 profile it + must resolve every condition/question pair to the checksum-verified + `approved_rerun_queue.csv` record with `queue_disposition=approved`, + `execution_authority=authoritative`, and `executable=true`. Held, excluded, + forensic, absent, or false-executable records cannot authorize inference. +- `offline_transformation` authorizes deterministic processing of preserved + evidence without model execution. Its immutable authorization source binds + the exact condition/question IDs, reasons, authority, allowed transformation + purpose, input-evidence digests, and expected-dataset digest. It explicitly + records `inference_executable=false`. It neither requires nor fabricates an + approved-rerun-queue entry. + +An authorization record can permit only its named operation type. An +`offline_transformation` record cannot authorize inference, and an +`inference_repair` record does not implicitly authorize unrelated rematching or +postprocessing. The overlay engine requires every replacement ID to be in that independently validated authorization record, the expected dataset, and the base expected @@ -525,32 +818,82 @@ overlay rows, multiple overlays that replace the same ID, inconsistent gold/options/ownership, and conflicting overlay declarations. It validates replacement rows with the same rules as base rows. -Applying an overlay never changes the base evidence snapshot, base condition, -or optional base result. It creates a derived condition with a new identity and -complete base -> authorization -> overlay -> resulting-artifact lineage. The +Applying an overlay never changes the base evidence snapshot, semantic +condition, base realization, or optional base result. It creates a derived +realization and result artifact with new identities and complete base -> +authorization -> overlay/transformation -> resulting-artifact lineage. It keeps +the shared semantic condition when the scientific protocol is unchanged. The specification declares the expected derived evidence status; the importer recomputes it from validated coverage and remaining defects and refuses a mismatch. A declaration of complete or qualified succeeds only if all remaining defects are resolved. +For a repair that combines historical base rows with newly generated rows, the +derived realization records `derivation_origin=repair_overlay`, retains +`external_historical_inference` on unchanged rows, assigns `native_inference` or +`external_repair_inference` to each replacement row as applicable, and records +the full constituent set/counts and ordered mapping. It never erases component +origins or reduces them to `mixed`. + Offline transformations use the same derived-artifact mechanism. They record -input and output checksums plus transformation code identity. A specification -cannot ask ChoiceBench to import and execute arbitrary code. A transformation -is either: +input and pre-ownership output checksums plus transformation code identity under +the non-circular rules above. The final ChoiceBench CSV checksum is added only +after the realization and experiment are fixed. A specification cannot ask +ChoiceBench to import and execute arbitrary code. A transformation is either: - a registered ChoiceBench implementation whose code identity is computed by ChoiceBench; or - a precomputed external output whose producing code identity/digest is declared and whose output is independently validated. -Deterministic rematching of surviving Stage-1 responses is represented as -`result_origin=offline_transformation`, not new model inference. +Deterministic rematching of surviving Stage-1 responses is represented with +`derivation_origin=offline_transformation`, while the rematched row retains the +origin of the inference response being rematched. It is not new model inference. + +### Stage 1 offline-transformation authority + +The six recoverable ARC `semantic_matching_v1` cells are authorized separately +from the rerun queue. The Stage 1 profile derives and content-addresses an +immutable `offline_transformation` authorization record from the +checksum-covered recoverable entries in +`manifests/canonical_results_manifest.json` (cross-checked against its CSV form). +The record identifies that manifest as the user-designated Stage 1 authority and +names these exact cells: + +- `cbp__gemini-2-5-flash__arc_challenge__semantic_matching_v1` +- `cbp__gpt-4-1-mini__arc_challenge__semantic_matching_v1` +- `cbp__llama-3-1-8b-instant__arc_challenge__semantic_matching_v1` +- `cbp__meta-llama-llama-3-1-8b-instruct__arc_challenge__semantic_matching_v1` +- `cbp__qwen-qwen2-5-7b-instruct-turbo__arc_challenge__semantic_matching_v1` +- `cbp__qwen-qwen2-5-7b-instruct__arc_challenge__semantic_matching_v1` + +The profile's manifest translation binds each historical `cell_id` to its full +ChoiceBench semantic `condition_digest`; authorization validation checks both +identities. For each it authorizes exactly these three question IDs: + +- `79e8c959bbeb74a0` +- `ad6b5d46ae54842c` +- `c30e75b011696a95` + +Thus it binds exactly six conditions and 18 condition/question pairs, their +manifest reasons, base-source checksums, expected-snapshot digest, the allowed +semantic-rematching purpose, and `inference_executable=false`. The generated +generic transformation specification references this independent authorization +record and digest; it cannot add IDs. `arc_question_audit.csv` may corroborate +row-level evidence but, because it is not itself covered by the freeze checksum +ledger, it is not the authorization trust anchor. ## Manifest, reader, and evaluation behavior -Imported manifests add an identity-bearing import section and condition-level -orthogonal provenance dimensions. Audit-only provenance is stored in a separate -non-identity section whose exclusion is explicit and validated. +Manifest v3 retains the existing models, methods, prompts, datasets, semantic +conditions, and experiment semantics, and adds an identity-bearing realization +table. Each realization references one semantic condition and carries the +source, validation, orthogonal provenance dimensions, structured origin, +authorization, evidence, lineage, and planned result contract/path. After an +evaluable artifact is atomically written, its sidecar and run-state entry—not +the immutable manifest—reference its separate result-artifact identity. +Audit-only provenance is stored in a separate non-identity section whose +exclusion is explicit and validated. Manifest validation recomputes imported child full digests and verifies their short IDs and fixed artifact paths. Result-side identity summaries and column @@ -558,33 +901,38 @@ digests are compared with the CSV and manifest rather than merely stored. The publication-grade reader: -- validates content-addressed evidence snapshots and condition validation +- validates content-addressed evidence snapshots and realization validation artifacts; - validates evaluable imported result artifacts through the normal sidecar and row-ownership path; -- refuses result artifacts for ineligible conditions; -- returns evaluable rows plus manifest accounting for all imported conditions; +- refuses result artifacts for ineligible realizations; +- selects evaluation realizations explicitly and never concatenates alternative + realizations that share a semantic condition; +- returns evaluable rows plus manifest accounting for all imported semantic + conditions and realizations; - preserves compatibility with legacy native v0.2 manifests. Evaluation: -- computes configured metrics for included complete/qualified imported - conditions; +- computes configured metrics for explicitly selected included + complete/qualified imported realizations; - includes qualifications and limitations beside qualified metrics; - accounts for partial, malformed, recoverable, failed, excluded, held, and superseded evidence without metrics; -- reports condition counts by each orthogonal dimension rather than one lossy - status tally; +- reports semantic-condition and realization counts by each orthogonal dimension + rather than one lossy status tally; - never estimates metrics for missing, invalid, excluded, or held rows; - continues to validate dataset snapshots, metric implementation identity, row ownership, and result checksums. Evaluation identity includes only stable semantic provenance and lineage: -- experiment and imported condition full digests; +- experiment and semantic-condition full digests; +- explicit realization-selection policy and every accounted realization digest; - evidence status and scope disposition; - stable qualification/limitation digest; -- result/evidence/lineage digests as applicable; +- result-artifact, evidence, authorization, structured-origin, and lineage + digests as applicable; - configured metric and postprocessing implementation identities; - exact consumed evaluable result checksums. @@ -629,10 +977,11 @@ specification cannot set or override it. The machine report includes: -- stable import/specification/experiment IDs; +- stable import-specification, semantic-grid, condition, realization, + result-artifact (when evaluable), and experiment IDs; - audit-only input/output locations; - counts by import state, evidence status, scope disposition, and origin; -- evaluable versus evidence-only condition counts; +- evaluable versus evidence-only realization and semantic-condition counts; - row coverage and exact defect IDs; - source/checksum/schema results; - extra-column dispositions; @@ -700,7 +1049,7 @@ The profile verifies all source checksums used by its generated specifications. It preserves the eight distinct historical CSV schemas and their method-specific fields. -Queue authority is explicit: +Inference-repair queue authority is explicit: - `approved_rerun_queue.csv`: 138 question-cells, executable true; - `held_or_declined_reruns.csv`: 687 Gemini ARC IHS question-cells, @@ -710,14 +1059,69 @@ Queue authority is explicit: - `rerun_queue.csv`: historical forensic evidence only and never execution authority. -The profile imports no repair output and executes none of these queues. It only -records the authorization boundary needed for future overlays. - -For expected datasets, the profile uses explicit manifest-selected frozen -canonical evidence, constructs benchmark/split question snapshots, and -cross-checks question/gold/option identity across conditions. The generated -generic specifications contain the resulting explicit expected question sets -and digests; the reusable core receives no paper-specific inference rule. +These queue classifications authorize or decline new model inference only; they +do not authorize offline transformations. The profile imports no repair output +and executes none of these queues. It records the typed authorization boundaries +needed for future overlays and the separate offline-transformation authority +defined above. + +### Stage 1 expected-dataset trust anchor + +Stage 1 does not need to reconstruct expected questions from result CSVs. Its +trust anchor is the class of pre-inference benchmark-input and split artifacts +under `raw/local_model_generalization/data/`, all bound by +`checksums/checksums.sha256`: + +- ARC raw/normalized inputs: + `data/raw/arc_challenge_raw.csv` and + `data/processed/arc_challenge_normalized.csv`; +- ARC selection and metadata: + `data/splits/arc_challenge/robustness_ids.json` and + `robustness_metadata.json`; +- MMLU raw/normalized inputs: + `data/raw/mmlu_raw.csv` and `data/processed/mmlu_normalized.csv`; +- MMLU selection and metadata: + `data/splits/benchmark/robustness_ids.json` and + `robustness_metadata.json`. + +The profile classifies these as `independent_input_snapshot` because they are +benchmark inputs independent of the result rows, with the more precise trust +label `checksum_verified_freeze_internal`. It verifies the freeze checksum +ledger against the opened bytes, parses the raw/normalized data under an +explicit adapter declaration, verifies the split-selection and metadata +digests, applies the declared selection derivation, constructs the ChoiceBench +snapshot, and then cross-checks result-side question, gold, and option evidence. + +The ARC normalized input resolves its 1,000 unique selected IDs directly and +retains 997 four-option and three three-option rows. The MMLU normalized input +contains 1,003 rows for its 1,000 unique selected IDs because each of +`79686d32dfe155ea`, `2f7aa3c7ebb98cfe`, and `74f7227e190200ac` occurs twice. +This is part of the frozen input evidence, not silently invalidated or ignored. +For MMLU the profile reproduces the archived split-construction rule from +`raw/local_model_generalization/scripts/prepare_data.py`: preserve source order +and keep the first occurrence under +`drop_duplicates(subset="question_id", keep="first")`. Before applying it, the +profile requires every duplicate occurrence to agree exactly on all parsed +fields; any conflict fails closed. The registered profile implementation +identity, ordered pre-dedup row digest, exact duplicate-ID/row digests, +first-occurrence policy, and ordered post-dedup digest are identity-bearing. +After that declared derivation, every selected MMLU ID resolves exactly once and +all 1,000 selected rows have four options. + +This chain establishes internal freeze consistency and independence from result +CSVs. It does not establish upstream publisher authenticity: the freeze does not +record a verifiable Hugging Face revision/commit/fingerprint or publisher-signed +checksum for these bytes. The profile therefore preserves unknown upstream +revision fields, reports that limitation, and never labels the snapshot as +publisher-authenticated. Byte-identical copies elsewhere in the freeze and +cross-result agreement are corroboration only, not the trust basis. If these +input artifacts were absent, the profile would have to use the weaker +`profile_derived_reference_snapshot` rules and correspondingly limited claims. + +The generated generic specifications contain the exact reference kind, trust +label, source and selection checksums, transformation/selection derivation, +expected question sets, and snapshot digests. The reusable core receives no +paper-specific inference rule. ## Security boundaries @@ -786,7 +1190,7 @@ Focused tests cover: 37. Orthogonal status combinations, including complete+excluded and malformed+excluded. 38. Audit paths/timestamps do not change import or evaluation identity. -39. New manifests always write explicit `result_origin`. +39. New manifests always write explicit structured `result_origin`. 40. Legacy manifests alone may infer native origin. 41. Content-addressed evidence deduplication and divergent-blob refusal. 42. Explicit absolute source/output roots are accepted while spec-controlled @@ -796,6 +1200,43 @@ Focused tests cover: result checksum. 45. Overlay-declared authorization cannot self-authorize replacement IDs. 46. Declared derived evidence status must equal the recomputed status. +47. Source bytes/mapping/provenance change realization and result/evaluation + identities without changing an otherwise identical semantic condition. +48. Scientifically meaningful model/dataset/method/prompt changes do change the + semantic condition. +49. Base, repaired, and transformed realizations share a semantic condition + when protocol semantics are unchanged, and alternatives are never + concatenated for evaluation. +50. Result-artifact identity changes when exact result bytes change. +51. A repaired artifact records constituent origins and the exact per-row origin + mapping; no bare `mixed` origin is accepted. +52. Offline transformation preserves underlying prediction origin while + recording a separate derivation origin and component lineage. +53. `inference_repair` requires an exact approved executable queue binding and + cannot use held, excluded, forensic, or offline authority. +54. `offline_transformation` requires its separate checksum-verified authority, + cannot self-authorize, and does not require executable inference authority. +55. The Stage 1 offline authority contains exactly six semantic-matching cells, + 18 condition/question pairs, and the three declared ARC IDs per cell. +56. Independent-input and profile-derived dataset references retain distinct + trust labels, derivations, and validation claims. +57. The Stage 1 profile anchors expected data to the checksum-verified frozen + benchmark inputs/splits and reports the unknown upstream revision. +58. Every decoding/dialect field participates in realization identity. +59. UTF-8 BOM forbid/strip behavior and strict decoding-error refusal. +60. Delimiter, quote, escape, doubled-quote, whitespace, header, and strict CSV + behavior, including invalid-declaration rejection. +61. CRLF, LF, CR, and mixed-line policies, with exact raw byte spans for quoted + records containing embedded newlines. +62. The Stage 1 MMLU snapshot verifies the three exact duplicate pairs, rejects + conflicting duplicates, applies stable first-occurrence deduplication, and + identity-binds both pre- and post-dedup ordered row digests. +63. Manifest, realization, lineage-component, transformation-output, and + result-artifact identities can be computed in order with no self-reference. +64. Failure/crash before the no-replace staging rename leaves no published final + run, and readers ignore owner-marked staging directories. +65. Existing final runs are verify-only: exact graphs no-op and any divergent + graph is refused without staging over the run. Existing native-run tests remain unchanged or receive compatibility coverage. The full suite must continue to pass. @@ -809,15 +1250,20 @@ After synthetic tests pass: 3. Confirm the exact evidence-status and queue counts listed above. 4. Confirm held and excluded work stays non-executable. 5. Confirm every referenced canonical source checksum. -6. Perform a real import into a temporary absolute `CHOICEBENCH_HOME` outside +6. Verify both frozen benchmark-input/split chains, their trust labels, all + 2,000 selected IDs, and the three variable-option ARC rows. +7. Verify the separate offline-transformation authorization has exactly six + conditions and 18 authorized condition/question pairs and grants no inference + execution. +8. Perform a real import into a temporary absolute `CHOICEBENCH_HOME` outside both repositories. -7. Verify the freeze has no filesystem changes. -8. Use native ChoiceBench reading/evaluation to demonstrate: +9. Verify the freeze has no filesystem changes. +10. Use native ChoiceBench reading/evaluation to demonstrate: - one complete imported condition; - one qualified imported condition; - one incomplete or malformed evidence-only condition; - one variable-option ARC condition. -9. Write the end-to-end report outside the repository. +11. Write the end-to-end report outside the repository. Generated imports, reports, build artifacts, temporary workspaces, and historical evidence remain uncommitted. @@ -831,9 +1277,12 @@ User documentation will explain: - import schema fields and examples; - generic, dry-run, strict, output-root, and Stage 1 profile usage; - deterministic identity versus audit provenance; +- semantic condition, realization, result-artifact, and evaluation identity; - checksum and idempotence behavior; - orthogonal status dimensions and evaluation eligibility; -- repair and offline-transformation lineage; +- mixed prediction origins, repair lineage, and separately authorized offline + transformations; +- expected-dataset trust kinds and deterministic CSV decoding/dialect behavior; - security boundaries and limitations; - how to implement a future adapter/profile. @@ -883,10 +1332,13 @@ This stage does not: The design is successful when a strict, reusable specification can import external CSV rows into collision-safe ChoiceBench manifests and -ChoiceBench-native-format imported result artifacts; every artifact explicitly -retains external inference origin; non-evaluable evidence remains preserved and -honestly accounted; deterministic identities exclude machine-local audit data; -repair lineage is immutable and authorized; the Stage 1 profile reproduces all -sealed counts without modifying the freeze; native reading/evaluation and -installed-package workflows work outside the repository; and the feature is -submitted on its isolated branch as an unmerged pull request. +ChoiceBench-native-format imported result artifacts; semantic condition, +realization, result-artifact, and evaluation identities have the specified +separation; every artifact retains structured derivation and exact constituent +prediction origins; non-evaluable evidence remains preserved and honestly +accounted; deterministic identities exclude machine-local audit data; repair +and offline-transformation lineage uses the correct immutable typed authority; +dataset trust and CSV parsing behavior are explicit; the Stage 1 profile +reproduces all sealed counts without modifying the freeze; native +reading/evaluation and installed-package workflows work outside the repository; +and the feature is submitted on its isolated branch as an unmerged pull request. From 712beb5aefc9088605d9da003e3fbc9d792f366a Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sat, 18 Jul 2026 22:34:34 +0300 Subject: [PATCH 03/47] docs: plan external results importer implementation --- .../2026-07-18-external-results-importer.md | 3260 +++++++++++++++++ 1 file changed, 3260 insertions(+) create mode 100644 docs/superpowers/plans/2026-07-18-external-results-importer.md diff --git a/docs/superpowers/plans/2026-07-18-external-results-importer.md b/docs/superpowers/plans/2026-07-18-external-results-importer.md new file mode 100644 index 0000000..1380219 --- /dev/null +++ b/docs/superpowers/plans/2026-07-18-external-results-importer.md @@ -0,0 +1,3260 @@ +# External Results Importer Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use +> `$subagent-driven-development` (recommended) or `$executing-plans` to +> implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for +> tracking. + +**Goal:** Add a reusable, provenance-preserving importer that validates +externally generated row-level results, publishes them as immutable +ChoiceBench-native-format imported result artifacts, and evaluates eligible +realizations without claiming ChoiceBench performed the original inference. + +**Architecture:** Manifest v3 separates stable scientific conditions from +source-sensitive realizations while manifest-v2 readers remain byte-compatible. +The paper-agnostic importer parses strict YAML, scans exact CSV logical records, +validates rows against an explicitly classified dataset reference, builds +identity and lineage, and atomically publishes a complete staged run. A thin +Stage 1 profile translates the immutable freeze into the same generic types. + +**Tech Stack:** Python 3.10+, dataclasses, `pathlib`, `csv`-compatible parsing, +pandas, PyYAML, SHA-256/canonical JSON through `choicebench.identity`, pytest, +setuptools, build, and Twine. The implementation adds no runtime dependency. + +## Global Constraints + +- Approved design: `docs/superpowers/specs/2026-07-18-external-results-importer-design.md`. +- Branch/worktree: `feat/external-results-importer` in + `/home/cotenthusiast/Projects/choicebench-external-results-importer`. +- Preserve the existing 808-test baseline; before every task or review-fix + commit run the targeted tests and `python -m pytest -q`, with zero failures. +- Use TDD for every behavior-changing task: add a failing behavioral test, run + it and observe the expected failure, implement the smallest slice, rerun + targeted tests, then the full suite, then commit. Test-only compatibility + characterization (Task 1) and documentation/installed-package + characterization (Task 18) intentionally begin with passing gates. +- Keep manifest-v2 validation, reading, evaluation identity, resume, and result + ownership behavior intact for legacy runs. Only a pre-existing v2 run may + continue writing v2-compatible native rows; every newly created run is v3 and + records structured result origin explicitly. +- Keep semantic `condition_id` aligned with ChoiceBench's existing benchmark + selection x model x method x prompt x seed/preflight/protocol identity. + Source bytes, parsing, mapping, status, scope, authorization, repair, and + lineage belong only to `realization_id` and downstream identities. +- Result-artifact identity is computed only after exact result bytes exist and + is stored in the v2 sidecar/run state, never in the immutable v3 manifest or + CSV. Pre-ownership transformation payloads and lineage-node IDs must not + depend on their child realization/result/experiment. +- The acyclic publication order is: semantic dataset/model/method/prompt and + condition records; source/evidence plus pre-ownership validation/lineage + digests; realization; manifest/experiment; exact evidence index and + realization-validation artifact; exact result CSV; result-artifact sidecar; + final realization-keyed run state. Only downstream records may reference an + earlier identity or exact file checksum. +- A manifest-v3 result path is + `results/.csv` with sidecar + `results/.artifact.json`. +- Published runs are immutable. A repair or offline transformation consumes a + verified base run as read-only input and publishes a different run ID, + experiment ID, realization ID, and result artifact. It never appends to, + replaces, or mutates the base run. +- Keep the generic importer free of paper names, the 100-cell matrix, ARC IDs, + Gemini/IHS/PriDe identities, and historical filename rules. All such knowledge + belongs in `importing/profiles/stage1_paper_freeze.py` and profile tests. +- The Stage 1 freeze is read-only external input. Never create, rewrite, rename, + chmod, or delete anything below + `/home/cotenthusiast/Projects/model-generalization/paper_data_freeze`. +- Tests use programmatically generated small synthetic fixtures. Do not copy + historical result rows, benchmark snapshots, cache trees, or reports into Git. +- Do not run model inference or any approved, held, excluded, or forensic queue. +- Explicit user-selected absolute source/output paths are valid after + canonicalization. Specifications cannot choose output roots. Reject traversal, + unsafe symlinks, recursive directory import, secrets, executable objects, and + divergent overwrite. +- Every high-risk gate named below gets an independent read-only review before + the next task. Resolve confirmed findings and rerun that gate's commands. +- Commit only the files named by the task. Never combine two task commits. + +## Target file map + +### New reusable importer modules + +- `src/choicebench/importing/__init__.py` — stable public importer exports. +- `src/choicebench/importing/schema.py` — strict dataclass schema and YAML loader. +- `src/choicebench/importing/csv_adapter.py` — deterministic UTF-8 logical-record + scanner and CSV adapter. +- `src/choicebench/importing/dataset_reference.py` — expected-question snapshot, + trust classification, and derivation verification. +- `src/choicebench/importing/identity.py` — semantic condition, realization, + lineage-node, origin assignment, and import-specification identities. +- `src/choicebench/importing/validation.py` — question-keyed row validation, + evidence findings, extension-field preservation, and evaluable row creation. +- `src/choicebench/importing/authorization.py` — typed authorization validation. +- `src/choicebench/importing/overlays.py` — immutable repair and offline + transformation derivation. +- `src/choicebench/importing/evidence.py` — run-local content-addressed evidence + and validation artifacts. +- `src/choicebench/importing/transaction.py` — same-filesystem staged publication + and verify-only existing-run behavior. +- `src/choicebench/importing/engine.py` — generic validate/import orchestration and + machine-readable report. +- `src/choicebench/importing/profiles/__init__.py` — profile registry. +- `src/choicebench/importing/profiles/stage1_paper_freeze.py` — Stage 1-only + translation, trust chain, statuses, queue ledgers, and schema maps. +- `src/choicebench/cli/import_results.py` — installed CLI with lazy workspace + bootstrap. + +### Existing modules changed deliberately + +- `src/choicebench/manifest.py` — version-dispatched v2/v3 validation, normalized + manifest view, v3 run state, and full child-digest/path verification. +- `src/choicebench/io/writers.py` — v2 writer compatibility plus non-circular v2 + result-artifact sidecars for v3 realizations. +- `src/choicebench/io/readers.py` — version-dispatched publication reader and + explicit realization selection. +- `src/choicebench/cli/evaluate_run.py` — legacy evaluation-v1 compatibility plus + realization-aware evaluation-v2. +- `src/choicebench/cli/run_experiment.py` — new native runs emit manifest v3 and + native realization/origin metadata; existing v2 resumes stay v2. +- `src/choicebench/infra/checkpoint.py` — realization-bound checkpoint-v2 with + legacy checkpoint-v1 resume compatibility. +- `src/choicebench/infra/artifacts.py` — no-replace directory publication and + directory fsync helper. +- `pyproject.toml` — `choicebench-import-results` entry point only; no dependency + or package-version change. +- `README.md`, `CHANGELOG.md`, and `docs/external-results-import.md` — user and + adapter documentation. + +### New and expanded tests + +- `tests/importing/conftest.py` — synthetic generic and miniature Stage 1 freeze + builders. +- `tests/importing/test_manifest_v2_compat.py` +- `tests/importing/test_manifest_v3.py` +- `tests/importing/test_schema.py` +- `tests/importing/test_csv_adapter.py` +- `tests/importing/test_dataset_reference.py` +- `tests/importing/test_identity.py` +- `tests/importing/test_result_artifact.py` +- `tests/importing/test_validation.py` +- `tests/importing/test_evidence_status.py` +- `tests/importing/test_authorization.py` +- `tests/importing/test_overlays.py` +- `tests/importing/test_evidence.py` +- `tests/importing/test_transaction.py` +- `tests/importing/test_engine.py` +- `tests/importing/test_evaluation.py` +- `tests/importing/test_native_v3.py` +- `tests/importing/test_stage1_profile.py` +- `tests/importing/test_cli.py` +- Expand `tests/test_publication_identity.py`, `tests/test_condition_grid.py`, + `tests/io/test_readers.py`, `tests/io/test_writers.py`, + `tests/scripts/test_evaluate_run.py`, `tests/scripts/test_reset_run.py`, + `tests/scripts/test_build_backend.py`, `tests/test_pride_reproduction_wiring.py`, + `tests/test_release_adversarial.py`, and `tests/test_wheel_smoke.py` only where + each task says. + +## Required review protocol + +For every task, the implementer commits only after targeted and full tests pass. +For Tasks 1, 3, 5, 6, 7, 10, 11, 13, 14, 16, 17, and 19, dispatch a fresh read-only +reviewer after the commit. Give the reviewer only the approved specification, +the task text, and the commit diff. Do not advance until all confirmed Critical +and Important findings are resolved in a focused follow-up commit and the same +validation is rerun. + +## Specification acceptance-test traceability + +The approved specification's 65 acceptance cases are assigned before any +production edit. The implementing worker must retain these numbered case IDs in +test docstrings or parametrization IDs so the final compliance reviewer can +prove that no requirement was lost: + +- Cases 1-6: Tasks 2, 7, 11, and 12 + (`test_schema.py`, `test_result_artifact.py`, `test_evidence.py`, + `test_transaction.py`, and `test_engine.py`) cover valid import, + deterministic identity, exact no-op, divergent refusal, checksum mismatch, + and source mutation. +- Cases 7-17: Tasks 3, 4, 8, and 9 (`test_csv_adapter.py`, + `test_dataset_reference.py`, `test_validation.py`, and + `test_evidence_status.py`) cover question membership, gold/options, + variable-option and three-option rows, null/NaN, predictions, correctness, + columns, and method extensions. +- Cases 18-22: Tasks 9 and 13 (`test_evidence_status.py` and + `test_evaluation.py`) cover complete/qualified/partial/malformed evidence, + imported evaluation, and non-metric accounting. +- Cases 23-27: Task 10 (`test_authorization.py` and `test_overlays.py`) covers + valid repair, unauthorized/conflicting replacements, repair identity/lineage, + and offline-transformation lineage. +- Cases 28-31: Tasks 15 and 16 (`test_stage1_profile.py`) cover profile + translation, hosted PriDe exclusion, held Gemini ARC IHS, and strict + separation of approved and forensic queues. +- Cases 32-36: Tasks 17-19 (`test_cli.py`, `test_wheel_smoke.py`, and + `test_release_adversarial.py`) cover dry-run, real import, outside-repository + operation, installed wheel/sdist behavior, and adversarial security. +- Cases 37-46: Tasks 2, 9-11, 13, and 17 cover orthogonal statuses, audit-only + identity exclusion, explicit new origin/legacy-only inference, evidence CAS, + absolute path safety, exact malformed evidence, evidence-only bases, + non-self-authorizing overlays, and derived-status recomputation. +- Cases 47-52: Tasks 5, 7, 10, and 13 cover the four identity layers, + condition-sharing alternatives, exact-byte result identity, constituent and + per-row origins, and offline derivation versus prediction origin. +- Cases 53-57: Tasks 4, 10, 15, and 16 cover typed inference/offline authority, + the exact six-condition/18-question-cell offline authority, distinct dataset + trust classes, and the frozen-input trust chain with its upstream limitation. +- Cases 58-62: Tasks 2, 3, and 16 cover every identity-bearing CSV setting, + strict UTF-8/BOM/dialect/newline behavior, exact embedded-newline spans, and + the three MMLU duplicate pairs with pre/post derivation digests. +- Cases 63-65: Tasks 5, 7, 11, and 12 cover acyclic publication identity, + crash-safe no-replace publication, staging invisibility, and existing-run + verify-only behavior. + +The mapping is a minimum, not permission to omit cross-layer integration tests +listed in the individual tasks. + +--- + +### Task 1: Freeze manifest-v2 compatibility before broad changes + +**Files:** + +- Create: `tests/importing/__init__.py` +- Create: `tests/importing/conftest.py` +- Create: `tests/importing/test_manifest_v2_compat.py` +- Inspect only: `src/choicebench/manifest.py`, `src/choicebench/io/writers.py`, + `src/choicebench/io/readers.py`, `src/choicebench/cli/evaluate_run.py` + +**Interfaces:** + +- Consumes unchanged v0.2 functions: `make_manifest`, `validate_manifest`, + `initial_run_state`, `write_run_results`, `read_manifest_results`, and + `build_evaluation_report`. +- Produces `synthetic_v2_run(tmp_path, monkeypatch) -> tuple[Path, dict, str]`, + a deterministic two-question native run with explicit source/environment + records and a hard-coded expected experiment/evaluation identity. + +- [ ] **Step 1: Confirm the untouched v0.2 baseline** + + ```bash + python -m pytest -q + ``` + + Expected: exactly 808 tests pass with zero failures. Record the duration and + commit SHA. If this does not hold, stop and diagnose the starting state before + writing any plan task files. + +- [ ] **Step 2: Add characterization tests before production changes** + + Create a literal two-question dataset/prompt/model/method/condition fixture. + Pass fixed `source` and `environment` to `build_manifest_payload`, write the + manifest, snapshots, result, sidecar, and state through the current v2 APIs, + and assert the exact current IDs. The test names and assertions must be: + + ```python + def test_v2_fixture_validates_without_rewrite(synthetic_v2_run): + run_dir, manifest, _ = synthetic_v2_run + before = {p.relative_to(run_dir): p.read_bytes() + for p in run_dir.rglob("*") if p.is_file()} + validate_manifest(manifest) + frame, loaded = read_manifest_results(run_dir) + after = {p.relative_to(run_dir): p.read_bytes() + for p in run_dir.rglob("*") if p.is_file()} + assert loaded["schema_version"] == "choicebench.manifest.v2" + assert frame["question_id"].astype(str).tolist() == ["q1", "q2"] + assert after == before + + + def test_v2_evaluation_identity_and_shape_are_frozen( + synthetic_v2_run, monkeypatch + ): + run_dir, manifest, expected_evaluation_id = synthetic_v2_run + monkeypatch.setattr(evaluate_run, "RUNS_DIR", run_dir.parent) + frame, _ = read_manifest_results(run_dir) + report = evaluate_run.build_evaluation_report( + run_dir.name, frame, manifest, reparse=False + ) + assert report["schema_version"] == "choicebench.evaluation.v1" + assert report["evaluation_id"] == expected_evaluation_id + assert set(report["conditions"]) == {"cond_legacyfixture"} + + + def test_v2_missing_result_origin_implies_native_only_in_compat_view( + synthetic_v2_run, + ): + _, manifest, _ = synthetic_v2_run + assert "result_origin" not in manifest["payload"]["conditions"][0] + ``` + + The fixture helper must assert its hard-coded experiment/evaluation IDs so an + implementer cannot update expected values casually after a regression. + +- [ ] **Step 3: Run the compatibility tests against unmodified v0.2 code** + + Run: + + ```bash + python -m pytest tests/importing/test_manifest_v2_compat.py -v + ``` + + Expected: all characterization tests PASS. If they do not, correct the + synthetic fixture; do not change production behavior in this task. + +- [ ] **Step 4: Reconfirm the baseline** + + Run: + + ```bash + python -m pytest -q + ``` + + Expected: the original 808 tests plus the new compatibility tests pass, with + zero failures. + +- [ ] **Step 5: Commit the compatibility boundary** + + ```bash + git add tests/importing/__init__.py tests/importing/conftest.py \ + tests/importing/test_manifest_v2_compat.py + git commit -m "test: freeze manifest v2 compatibility" + ``` + +**Intermediate gate:** A fresh reviewer compares the synthetic run and exact +assertions with current v2 construction, reading, and evaluation. No manifest-v3 +work begins until this test-only commit is approved. + +--- + +### Task 2: Add the strict, paper-agnostic import specification + +**Files:** + +- Create: `src/choicebench/importing/__init__.py` +- Create: `src/choicebench/importing/schema.py` +- Create: `tests/importing/test_schema.py` +- Inspect: `src/choicebench/config/schema.py`, `src/choicebench/identity.py` + +**Interfaces:** + +```python +ImportState = Literal["validated", "imported", "failed"] +EvidenceStatus = Literal[ + "complete", "qualified", "partial", "malformed", "recoverable", "failed" +] +ScopeDisposition = Literal[ + "included", "excluded_from_paper_matrix", "held", "superseded" +] +PredictionOrigin = Literal[ + "native_inference", + "external_historical_inference", + "external_repair_inference", +] +DerivationOrigin = Literal[ + "native_execution", "external_import", "repair_overlay", "offline_transformation" +] + + +@dataclass(frozen=True) +class CsvDialectSpec: + encoding: Literal["utf-8"] = "utf-8" + bom_policy: Literal["forbid", "strip_utf8_bom"] = "forbid" + decoding_errors: Literal["strict"] = "strict" + delimiter: str = "," + quote_character: str = '"' + escape_character: str | None = None + double_quote: bool = True + line_terminators: tuple[str, ...] = ("crlf", "lf", "cr") + mixed_line_terminators: Literal["allow", "forbid"] = "allow" + final_record_without_terminator: Literal["allow", "forbid"] = "allow" + blank_record_policy: Literal["reject"] = "reject" + skip_initial_space: bool = False + header: Literal["first_logical_record"] = "first_logical_record" + strict_syntax: bool = True + + +@dataclass(frozen=True) +class NumericColumnSpec: + source_column: str + value_type: Literal["integer", "float"] + null_allowed: bool + finite_only: bool = True + + +@dataclass(frozen=True) +class OptionMappingSpec: + mode: Literal["ordered_columns", "structured_json"] + ordered_columns: tuple[str, ...] + structured_column: str | None + structured_label_key: str | None + structured_text_key: str | None + + +@dataclass(frozen=True) +class SourceArtifactSpec: + source_id: str + path: Path # audit-only lookup location + logical_path: str # stable identity-bearing location + expected_sha256: str + format: Literal["csv"] + format_version: str + classification: Literal[ + "raw", "canonical", "derived", "repaired", "aggregate_only" + ] + dialect: CsvDialectSpec + columns: Mapping[str, str] + expected_columns: tuple[str, ...] + ignored_columns: Mapping[str, str] + null_values: tuple[str, ...] + numeric_columns: tuple[NumericColumnSpec, ...] + option_mapping: OptionMappingSpec + extra_field_policy: Literal["preserve_unmapped", "reject_unmapped"] + preserve_namespace: str + source_run_id: str | None + source_repository: str | None + source_commit: str | None + notes: Mapping[str, Any] + + +@dataclass(frozen=True) +class DatasetReferenceSpec: + dataset_id: str + benchmark_name: str + split: str + reference_kind: Literal[ + "independent_input_snapshot", "profile_derived_reference_snapshot" + ] + trust_label: str + source_ids: tuple[str, ...] + selection_source_id: str + expected_question_ids: tuple[str, ...] + selection_seed: int | None + selection_n_samples: int | None + subject_filter: tuple[str, ...] + selection_unknown_reasons: Mapping[str, str] + columns: Mapping[str, str] + revision: str | None + fingerprint: str | None + derivation: Mapping[str, Any] + limitations: tuple[str, ...] + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ImportModelSpec: + model_key: str + display_name: str + backend: str | None + provider: str | None + revision: str | None + effective_parameters: Mapping[str, Any] + unknown_reasons: Mapping[str, str] + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ImportMethodSpec: + method_key: str + name: str + effective_parameters: Mapping[str, Any] + implementation: Mapping[str, Any] | None + unknown_reasons: Mapping[str, str] + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ImportPromptSpec: + prompt_key: str + template_identity: str | None + template_digest: str | None + template_contents: Mapping[str, str] | None + unknown_reason: str | None + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ResultOriginSpec: + derivation_origin: DerivationOrigin + default_prediction_origin: PredictionOrigin | None + per_question_prediction_origins: Mapping[str, PredictionOrigin] + + +@dataclass(frozen=True) +class ImportConditionSpec: + condition_key: str + source_ids: tuple[str, ...] + dataset_id: str + model_key: str + method_key: str + prompt_key: str + seed: int | None + calibration_identity: Mapping[str, Any] | None + preflight_identity: Mapping[str, Any] | None + protocol_settings: Mapping[str, Any] + generation_parameters: Mapping[str, Any] + unknown_reasons: Mapping[str, str] + expected_question_ids: tuple[str, ...] + evidence_status: EvidenceStatus + scope_disposition: ScopeDisposition + executable: bool | None + qualifications: tuple[Mapping[str, Any], ...] + limitations: tuple[Mapping[str, Any], ...] + damaged_question_ids: tuple[str, ...] + recoverable_question_ids: tuple[str, ...] + result_origin: ResultOriginSpec + + +@dataclass(frozen=True) +class AuthorizationSpec: + authorization_id: str + authorization_type: Literal["inference_repair", "offline_transformation"] + source_id: str + condition_question_reasons: Mapping[str, Mapping[str, str]] + authority: str + purpose: str + executable: bool + input_evidence_digests: Mapping[str, str] + expected_snapshot_digests: Mapping[str, str] + + +@dataclass(frozen=True) +class OverlaySpec: + overlay_id: str + base_run_path: Path # audit-only lookup location + base_condition_digest: str + base_realization_id: str + base_realization_digest: str + base_evidence_digests: Mapping[str, str] + base_validation_artifact_sha256: str + base_result_sha256: str | None + source_id: str + authorization_id: str + replacement_reasons: Mapping[str, str] + result_origin: ResultOriginSpec + lineage_notes: Mapping[str, Any] + implementation: Mapping[str, Any] + input_digest: str + preownership_output_digest: str + expected_evidence_status: EvidenceStatus + + +@dataclass(frozen=True) +class ImportSpec: + schema_version: Literal["choicebench.import-spec.v1"] + import_name: str + sources: tuple[SourceArtifactSpec, ...] + datasets: tuple[DatasetReferenceSpec, ...] + models: tuple[ImportModelSpec, ...] + methods: tuple[ImportMethodSpec, ...] + prompts: tuple[ImportPromptSpec, ...] + conditions: tuple[ImportConditionSpec, ...] + authorizations: tuple[AuthorizationSpec, ...] + overlays: tuple[OverlaySpec, ...] + metrics: tuple[str, ...] # validated strictly against BUILTIN_METRICS + provenance: Mapping[str, Any] + audit: Mapping[str, Any] + + +def load_import_spec(path: Path) -> ImportSpec: ... +def stable_import_projection(spec: ImportSpec) -> dict[str, Any]: ... +def import_spec_digest(spec: ImportSpec) -> str: ... +``` + +The same file defines strict enums/dataclasses for dataset/model/method/prompt, +the orthogonal `import_state`, `evidence_status`, `scope_disposition`, structured +origins, expected question IDs, qualifications, and overlay references. The +schema has no output-root or output-path field. + +- [ ] **Step 1: Write failing strict-schema tests** + + Parameterize tests for: a valid minimal YAML; every unknown top-level/nested + key; non-mapping YAML; missing/duplicate references; strict booleans and finite + numerics; invalid SHA-256; invalid enum; arbitrary YAML object tag; recursive + credential-named keys; an attempted `output_root`; single-byte distinct ASCII + delimiter/quote/escape; fixed UTF-8/strict decoding; and explicit unknown + provenance as `{value: null, reason: "not recorded by producer"}`. + Reject every metric not present in the closed `BUILTIN_METRICS` registry, + especially `module:Class`, dotted import paths, entry points, and filesystem + paths. Import specifications can never select a dynamic import target. + Cover ordered-column and structured-JSON option declarations, typed finite + numeric policies, and preserve/reject extra-field policies; reject incomplete + or contradictory option declarations and duplicate numeric-column rules. + Validate optional native-compatibility payloads against exact current + dataset/model/method/prompt allowed keys and recomputed full/short identities; + reject arbitrary claimed IDs or partial/mismatched payloads. Stage 1 fixtures + leave these fields null rather than fabricating native equivalence. + + Add these identity assertions: + + ```python + def test_audit_locations_do_not_change_import_spec_digest( + minimal_spec, minimal_overlay_spec + ): + moved = replace( + minimal_spec, + audit={"source_path": "/different/host", "imported_at": "later"}, + ) + assert import_spec_digest(moved) == import_spec_digest(minimal_spec) + + moved_source = replace( + minimal_spec, + sources=(replace(minimal_spec.sources[0], path=Path("/other/freeze/source.csv")),), + ) + assert import_spec_digest(moved_source) == import_spec_digest(minimal_spec) + + moved_base = replace( + minimal_overlay_spec, + overlays=(replace(minimal_overlay_spec.overlays[0], + base_run_path=Path("/other/home/runs/base")),), + ) + assert import_spec_digest(moved_base) == import_spec_digest(minimal_overlay_spec) + + + @pytest.mark.parametrize( + "change", + ["mapping", "dialect", "numeric_policy", "option_policy", + "extra_field_policy", "status", "scope"], + ) + def test_stable_mapping_fields_change_import_spec_digest(minimal_spec, change): + changed = mutate_identity_field(minimal_spec, change) + assert import_spec_digest(changed) != import_spec_digest(minimal_spec) + ``` + +- [ ] **Step 2: Run tests and observe the missing module** + + Run: + + ```bash + python -m pytest tests/importing/test_schema.py -v + ``` + + Expected: FAIL during collection with + `ModuleNotFoundError: No module named 'choicebench.importing'`. + +- [ ] **Step 3: Implement strict dataclass parsing and stable projection** + + Use `yaml.safe_load`, explicit allowed-key sets, the existing credential-key + detection/canonicalization rules, and `integrity_digest`. Reject rather than + coerce booleans, integers, non-finite values, enum spellings, paths in stable + provenance, and unknown keys. Export only stable public types from + `importing/__init__.py`. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/config/test_schema.py tests/importing/test_schema.py -v + python -m pytest -q + ``` + + Expected: all schema tests pass; the original 808 tests remain green. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/importing/__init__.py \ + src/choicebench/importing/schema.py tests/importing/test_schema.py + git commit -m "feat: add strict external import specification" + ``` + +--- + +### Task 3: Implement exact UTF-8 CSV logical-record scanning + +**Files:** + +- Create: `src/choicebench/importing/csv_adapter.py` +- Create: `tests/importing/test_csv_adapter.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class OpenedSource: + source_id: str + audit_path: Path + logical_path: str + data: bytes + sha256: str + + +@dataclass(frozen=True) +class LogicalRecordSpan: + index: int + start: int + end: int + terminator: bytes + + +@dataclass(frozen=True) +class SourceRow: + values: Mapping[str, str | None] + span: LogicalRecordSpan + raw_sha256: str + + +@dataclass(frozen=True) +class AdaptedTable: + columns: tuple[str, ...] + rows: tuple[SourceRow, ...] + source_sha256: str + + +class SourceAdapter(Protocol): + def parse( + self, source: OpenedSource, declaration: SourceArtifactSpec, *, strict: bool + ) -> AdaptedTable: ... + + +def scan_csv_logical_records( + data: bytes, dialect: CsvDialectSpec +) -> tuple[LogicalRecordSpan, ...]: ... + + +def parse_csv_source( + source: OpenedSource, declaration: SourceArtifactSpec, *, strict: bool +) -> AdaptedTable: ... +``` + +- [ ] **Step 1: Write byte-literal failing tests** + + Use `Path.write_bytes` under `tmp_path`, never a committed CSV, for UTF-8 BOM, + invalid UTF-8, LF/CRLF/CR, mixed terminators, forbidden mixed terminators, + final record without terminator, blank records, doubled quotes, explicit + escape characters, quoted embedded CR/LF/CRLF, and malformed unclosed quotes. + Assert exact spans and row hashes: + + ```python + def test_embedded_newline_span_is_exact(): + data = b'id,text\r\nq1,"line one\r\nline two"\r\nq2,end\n' + spans = scan_csv_logical_records(data, CsvDialectSpec()) + assert data[spans[1].start:spans[1].end] == \ + b'q1,"line one\r\nline two"\r\n' + assert spans[1].terminator == b"\r\n" + + + def test_literal_nan_is_not_implicit_null(opened_csv, source_spec): + table = parse_csv_source(opened_csv(b"id,value\nq1,NaN\n"), source_spec, + strict=True) + assert table.rows[0].values["value"] == "NaN" + ``` + + Also cover header order, duplicate header refusal, extra columns in normal and + strict modes, explicit ignored-column reasons, empty string versus declared + null, each typed numeric policy and malformed numeric strings, ordered-column + and structured-JSON choices, three-option rows, and six-option rows. Assert + effective CLI strictness may tighten but never loosen the declared extra-field + policy and that the effective policy is realization-identity-bearing. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_csv_adapter.py -v + ``` + + Expected: FAIL because `choicebench.importing.csv_adapter` is absent. + +- [ ] **Step 3: Implement the binary scanner and adapter** + + Scan original bytes before decoding. Recognize ASCII quote/escape bytes and + CRLF longest-first; separators inside quoted fields remain field bytes. Decode + each exact logical span with strict UTF-8, then parse using the declared CSV + semantics. Never use pandas in this adapter and never normalize malformed + evidence. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_csv_adapter.py -v + python -m pytest -q + ``` + + Expected: every dialect/raw-span case passes; no baseline regression. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/importing/csv_adapter.py \ + tests/importing/test_csv_adapter.py + git commit -m "feat: add deterministic CSV source adapter" + ``` + +**Intermediate gate:** An adversarial read-only reviewer checks the state +machine against embedded newlines, escaped quotes, BOM offsets, malformed EOF, +and every identity-bearing dialect field. + +--- + +### Task 4: Build expected-dataset references and trust-qualified snapshots + +**Files:** + +- Create: `src/choicebench/importing/dataset_reference.py` +- Create: `tests/importing/test_dataset_reference.py` +- Inspect: `src/choicebench/datasets.py`, `src/choicebench/pipeline/options.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class ExpectedDataset: + dataset_id: str + benchmark_name: str + split: str + artifact_id: str + artifact_digest: str + selection_id: str + selection_digest: str + frame: pd.DataFrame + selected_question_ids: tuple[str, ...] + reference_kind: Literal[ + "independent_input_snapshot", "profile_derived_reference_snapshot" + ] + trust_label: str + question_set_digest: str + snapshot_digest: str + derivation: Mapping[str, Any] + derivation_digest: str + limitations: tuple[str, ...] + + +def build_expected_dataset( + declaration: DatasetReferenceSpec, + opened_sources: Mapping[str, OpenedSource], +) -> ExpectedDataset: ... + + +def write_expected_snapshot(staged_run: Path, dataset: ExpectedDataset) -> dict: ... +def validate_expected_snapshot(run_dir: Path, record: Mapping[str, Any]) -> None: ... +``` + +- [ ] **Step 1: Write failing trust and variable-option tests** + + Cover independently supplied input plus selection IDs; profile-derived + reference requiring two declared independent source groups; cross-source + question/gold/option disagreement refusal; full derivation digest; unknown + publisher revision limitation; duplicate selected ID refusal; missing selected + ID refusal; stable order; six options; and ARC-style three options without a + phantom empty fourth option. Recompute and verify full imported dataset- + artifact and selection digests plus their short IDs; the selection projection + follows ChoiceBench's existing artifact/content/sample-identity/seed/filter + semantics using explicit unknown/null values where historical sampling inputs + are unavailable. + + Assert two differently located/mapped/trust-classified references that produce + identical normalized semantic rows have the same artifact/selection IDs but + different reference/derivation digests. Changing gold, ordered option text, + membership, or selection order changes the applicable semantic IDs. + + ```python + def test_reference_kinds_are_not_equivalent(independent_decl, derived_decl, sources): + independent = build_expected_dataset(independent_decl, sources) + derived = build_expected_dataset(derived_decl, sources) + assert independent.reference_kind == "independent_input_snapshot" + assert derived.reference_kind == "profile_derived_reference_snapshot" + assert independent.derivation_digest != derived.derivation_digest + assert "not independently authenticated" in derived.limitations + ``` + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_dataset_reference.py -v + ``` + + Expected: FAIL because the dataset-reference module is absent. + +- [ ] **Step 3: Implement snapshot construction and self-validation** + + Use `build_option_map`, `correct_option_for_row`, + `dataset_content_digest`, and existing atomic JSON/CSV helpers. Populate the + run snapshot from the declared reference, never result-side fields. Preserve + reference kind/trust/limitations in the realization-facing reference record, + not in the semantic dataset artifact. Build full + dataset-artifact and selection records with the same semantic boundaries and + full-digest/short-ID pattern as current ChoiceBench dataset ownership; do not + put source paths, parsing/mapping policy, trust classification, or audit fields + into either semantic identity. The dataset artifact binds normalized semantic + question/gold/ordered-option content plus benchmark/split; selection binds the + exact selected semantic rows/question sample identities and known selection + semantics. Thus a source mapping/trust/provenance change that yields identical + semantic rows changes the realization but not artifact/selection/condition; + an actual semantic snapshot or selection change changes all applicable + semantic identities. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/pipeline/test_options.py \ + tests/importing/test_dataset_reference.py -v + python -m pytest -q + ``` + + Expected: trust distinctions, cross-source checks, and variable options pass; + no baseline regression. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/importing/dataset_reference.py \ + tests/importing/test_dataset_reference.py + git commit -m "feat: add trusted import dataset references" + ``` + +--- + +## Concrete compatibility resolution discovered during planning + +Current `build_execution_plan()` hashes the deterministic relative +`prompt_snapshot_path` (`artifacts/prompts/`) into `condition_id`. +The approved specification simultaneously requires v3 conditions to retain the +existing scientific ID and to exclude operational paths. Removing this field +would change every native condition ID. The implementation must therefore keep +this one prompt-ID-derived, machine-independent value as a frozen compatibility +discriminator in the v3 semantic condition payload. Tests prove it is exactly +derived from `prompt_id`; arbitrary/absolute paths remain forbidden. No other +path field is grandfathered. This is the only architecture clarification in the +plan and is directly required by inspected v0.2 code. + +--- + +### Task 5: Separate semantic condition, realization, lineage, and origin identity + +**Files:** + +- Create: `src/choicebench/importing/identity.py` +- Create: `tests/importing/test_identity.py` +- Inspect: `src/choicebench/identity.py`, + `src/choicebench/cli/run_experiment.py:980` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class ImportSemanticRecords: + dataset_artifact: Mapping[str, Any] + selection: Mapping[str, Any] + model: Mapping[str, Any] + method: Mapping[str, Any] + prompt: Mapping[str, Any] + condition: Mapping[str, Any] + + +def make_semantic_condition( + *, identity: Mapping[str, Any], fields: Mapping[str, Any] +) -> dict[str, Any]: ... + + +def build_import_semantic_identity( + *, + condition: ImportConditionSpec, + dataset: ExpectedDataset, + model: ImportModelSpec, + method: ImportMethodSpec, + prompt: ImportPromptSpec, +) -> ImportSemanticRecords: ... + + +def make_lineage_component( + *, + operation_type: str, + question_id: str, + parent_digests: Sequence[str], + source_digests: Sequence[str], + authorization_digest: str | None, + implementation: Mapping[str, Any], + parameters: Mapping[str, Any], + input_digest: str, + preownership_output_digest: str, + prediction_origin: str, +) -> dict[str, Any]: ... + + +def make_result_origin( + *, + derivation_origin: Literal[ + "native_execution", "external_import", "repair_overlay", + "offline_transformation" + ], + row_assignments: Sequence[tuple[str, str, str]], +) -> dict[str, Any]: ... + + +def importer_implementation_identity( + *, adapter: Callable[..., Any], validator: Callable[..., Any] +) -> dict[str, Any]: ... + + +def make_realization( + *, + condition_id: str, + condition_digest: str, + identity: Mapping[str, Any], + fields: Mapping[str, Any], +) -> dict[str, Any]: ... +``` + +`row_assignments` entries are `(question_id, prediction_origin, +prediction_lineage_id)`. Records contain full SHA-256 digests plus short IDs. + +- [ ] **Step 1: Write failing layered-identity tests** + + Use one base scientific identity and parameterize changes. Assert: + + ```python + @pytest.mark.parametrize( + "field", + ["source_sha256", "mapping", "dialect", "evidence_status", "scope", + "authorization_digest", "overlay_sha256", "prediction_origins"], + ) + def test_realization_changes_do_not_change_condition(base_records, field): + original_condition, original_realization = base_records + changed_condition, changed_realization = mutate_realization(base_records, field) + assert changed_condition["condition_id"] == original_condition["condition_id"] + assert changed_realization["realization_id"] != original_realization["realization_id"] + + + @pytest.mark.parametrize( + "field", ["selection_id", "model_id", "method_id", "prompt_id", "seed", + "protocol_settings"] + ) + def test_scientific_change_changes_condition(base_records, field): + changed = mutate_semantic_condition(base_records[0], field) + assert changed["condition_id"] != base_records[0]["condition_id"] + ``` + + Also test base/repair/transformation sharing one condition; rejection of bare + `mixed`; origin set/count and ordered mapping; offline rematching retaining + underlying prediction origin; lineage component exclusion of child IDs, + paths, timestamps, and final CSV hashes; audit changes having no effect; + importer package/version plus adapter/validator source identities present and + identity-bearing; the grandfathered prompt path preserving the exact current + `condition_id`; and arbitrary path injection refusal. + + Add a generic-spec integration fixture that builds dataset-artifact, + selection, external model, historical method, prompt, and condition records + from the Task 2 dataclasses. Assert the final condition payload has the exact + current keys (`benchmark` with artifact/selection IDs, `preflight`, + `model_id`, `method_id`, `prompt_id`, frozen prompt discriminator, `seed`, and + applicable `protocol_settings`). Scientific input changes alter the proper + child/full/condition digests; source bytes, dialect, mapping, status, scope, + and audit-only changes leave that payload fixed and alter only realization + and downstream identities. Unknown historical implementation/revision/prompt + values remain explicit null+reason records and are never replaced by native + ChoiceBench identities. + + Freeze these canonical child payload contracts in literal tests: + + ```python + imported_dataset_payload = { + "schema_version": "choicebench.semantic-dataset.v1", + "benchmark": benchmark_name, + "split": split, + "content_digest": normalized_semantic_content_digest, + } + imported_selection_payload = { + "artifact_id": artifact_id, + "content_digest": selected_semantic_content_digest, + "sample_identities": ordered_sample_identities, + "seed": selection_seed_or_explicit_unknown, + "n_samples": selection_n_samples_or_explicit_unknown, + "subject_filter": sorted_subject_filter, + } + imported_model_payload = { + "schema_version": "choicebench.semantic-model.v1", + "backend": backend_or_explicit_unknown, + "provider": provider_or_explicit_unknown, + "model": display_name, + "revision": revision_or_explicit_unknown, + "effective_parameters": canonical_parameters, + } + imported_method_payload = { + "schema_version": "choicebench.semantic-method.v1", + "name": exact_historical_name, + "effective_params": canonical_parameters, + "preflight": preflight_identity, + "implementation": implementation_or_explicit_unknown, + } + imported_prompt_payload = { + "schema_version": "choicebench.semantic-prompt.v1", + "template_identity": identity_or_explicit_unknown, + "template_digest": digest_or_explicit_unknown, + "template_contents": contents_or_explicit_unknown, + } + ``` + + Full digests cover these exact payloads; short prefixes remain `ds`, `sel`, + `model`, `method`, and `prompt`. These imported payloads are the documented + fallback when no exact current-native identity is recoverable. When a + synthetic external specification supplies a fully validated + `native_compatibility_identity` for every child, use the exact current native + canonical payload instead, compare its child and condition identities with + `build_execution_plan()`, and require equality. + When historical fields are unknown or the imported semantic-dataset schema is + necessarily content-derived rather than current `DatasetSpec`-derived, require + explicit documented divergence while preserving the same condition payload + structure and never pretending equivalence. Mapping, trust, and source + provenance remain realization-only in either case. + + Parameterize every `CsvDialectSpec` field individually through final + realization construction and require each parsing-affecting change to alter + `realization_id` while preserving `condition_id`. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_identity.py -v + ``` + + Expected: FAIL because the importer identity module is absent. + +- [ ] **Step 3: Implement canonical full/short identity records** + + Reuse `canonicalize`, `stable_digest`, `integrity_digest`, and `short_id`. + Reuse `choicebench.provenance.implementation_identity` for registered importer, + adapter, validator, and transformation callables; never accept their active + code identity from untrusted YAML. + Explicitly allow only the declared identity keys at every layer. Compute in + this order: semantic condition; parent-derived lineage components; + pre-ownership output; row-origin mapping; realization. Never accept a current + experiment ID, result-artifact ID, final CSV SHA, output path, or audit field + in realization identity. + + Build the imported dataset/model/method/prompt records with the same semantic + payload boundaries and full-digest/short-ID pattern as native records. + Historical method identity is declared provenance, not a claim that the + current ChoiceBench method implementation produced the rows. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/test_condition_grid.py tests/importing/test_identity.py -v + python -m pytest -q + ``` + + Expected: all layered-identity cases pass and every existing native short ID + test remains unchanged. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/importing/identity.py tests/importing/test_identity.py + git commit -m "feat: separate condition and realization identity" + ``` + +**Intermediate gate:** Independent identity/provenance review. The reviewer must +construct the identity dependency graph and confirm there is no child/self edge +and no source/provenance field in the semantic condition. + +--- + +### Task 6: Introduce manifest v3 while preserving manifest v2 exactly + +**Files:** + +- Modify: `src/choicebench/manifest.py` +- Create: `tests/importing/test_manifest_v3.py` +- Modify: `tests/importing/test_manifest_v2_compat.py` +- Modify: `tests/test_publication_identity.py` + +**Interfaces:** + +```python +PROTOCOL_VERSION = "choicebench.protocol.v2" # frozen public legacy default +PROTOCOL_V3_VERSION = "choicebench.protocol.v3" +MANIFEST_SCHEMA_VERSION = "choicebench.manifest.v2" # frozen public legacy default +MANIFEST_V3_SCHEMA_VERSION = "choicebench.manifest.v3" +RUN_STATE_SCHEMA_VERSION = "choicebench.run-state.v2" # frozen public legacy default +RUN_STATE_V3_SCHEMA_VERSION = "choicebench.run-state.v3" + + +@dataclass(frozen=True) +class ManifestView: + manifest: Mapping[str, Any] + schema_version: str + semantic_conditions: Mapping[str, Mapping[str, Any]] + realizations: Mapping[str, Mapping[str, Any]] + realization_ids_by_condition: Mapping[str, tuple[str, ...]] + legacy_v2: bool + + +def make_manifest_v3( + payload: Mapping[str, Any], *, audit: Mapping[str, Any] | None = None +) -> dict[str, Any]: ... + + +def validate_manifest_v2(manifest: Mapping[str, Any]) -> None: ... +def validate_manifest_v3(manifest: Mapping[str, Any]) -> None: ... +def validate_manifest(manifest: Mapping[str, Any]) -> None: ... +def normalize_manifest(manifest: Mapping[str, Any]) -> ManifestView: ... +``` + +Keep current `PROTOCOL_VERSION`, `MANIFEST_SCHEMA_VERSION`, +`RUN_STATE_SCHEMA_VERSION`, `build_manifest_payload()`, and `make_manifest()` as +v2-compatible public entry points even after Task 14. Importer and new native +creation call explicit v3 functions/constants. `normalize_manifest()` synthesizes exactly one in-memory +native realization for v2 without rewriting bytes or changing stored IDs. + +- [ ] **Step 1: Write failing v3 and expanded v2 compatibility tests** + + Test `semantic_conditions` and `realizations` as separate unique tables; + condition full-digest recomputation; realization full-digest recomputation; + stable `semantic_grid_digest`; realization-to-condition full-digest binding; + fixed realization result/sidecar/validation paths; semantic validation digest; + explicit structured origin; rejection + of result-artifact ID/final CSV SHA in a manifest; audit exclusion from + experiment identity; unsafe paths; duplicate IDs; and v3 run state keyed by + realization. + + Extend the Task 1 tests to ensure v2 normalization preserves stored + experiment/condition/result paths, infers native origin only in memory, and + leaves all run bytes unchanged. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_manifest_v3.py \ + tests/importing/test_manifest_v2_compat.py tests/test_publication_identity.py -v + ``` + + Expected: new v3 tests fail because the symbols/schema are absent; every Task + 1 v2 test still passes. + +- [ ] **Step 3: Implement exact schema dispatch** + + Dispatch only on the exact `schema_version`, never on missing fields. Preserve + current v2 validation code as `validate_manifest_v2`. Add child-ID/full-digest + recomputation for v3, safe fixed paths, separate audit integrity, and + realization-keyed v3 `initial_run_state`/`validate_run_state`. Add normalized + access helpers instead of rewriting callers prematurely. + +- [ ] **Step 4: Run compatibility, publication, and full tests** + + ```bash + python -m pytest tests/importing/test_manifest_v3.py \ + tests/importing/test_manifest_v2_compat.py tests/test_publication_identity.py \ + tests/test_release_adversarial.py -v + python -m pytest -q + ``` + + Expected: v2 byte/identity tests and v3 structural tests pass; no baseline + failures. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/manifest.py tests/importing/test_manifest_v3.py \ + tests/importing/test_manifest_v2_compat.py tests/test_publication_identity.py + git commit -m "feat: add manifest v3 compatibility layer" + ``` + +**Intermediate gate:** Independent manifest review covering v2 dispatch, +full-child-digest verification, path safety, audit exclusion, and the absence of +result-artifact identity from the immutable manifest. + +--- + +### Task 7: Publish non-circular realization result artifacts + +**Files:** + +- Modify: `src/choicebench/io/writers.py` +- Create: `tests/importing/test_result_artifact.py` +- Modify: `tests/io/test_writers.py` +- Modify: `tests/test_publication_identity.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class PreparedResultArtifact: + result_path: str + metadata_path: str + csv_bytes: bytes + metadata: Mapping[str, Any] + + +def prepare_manifest_result( + results: Sequence[Mapping[str, Any]], + *, + manifest: Mapping[str, Any], + realization_id: str, +) -> PreparedResultArtifact: ... + + +def publish_manifest_result( + prepared: PreparedResultArtifact, *, run_dir: Path +) -> tuple[Path, Mapping[str, Any], bool]: ... + + +def validate_result_artifact( + result_path: Path, + *, + manifest: Mapping[str, Any] | None = None, + realization: Mapping[str, Any] | None = None, +) -> dict[str, Any]: ... +``` + +The returned boolean means an identical artifact already existed. Preserve the +legacy/programmatic `write_run_results()` contract and result-artifact-v1 +validation for v2. + +Result-artifact-v2 metadata contains exact experiment/condition/realization full +digests; result file SHA-256/size/row and ordered-column/question digests; +dataset artifact/selection/expected-snapshot digests; model/method/prompt full +digests; import-spec and source-byte digests; evidence-index digest; +realization-validation semantic digest and exact file SHA-256; authorization and +lineage digests; structured derivation origin; constituent prediction-origin +set/count and ordered per-question origin/lineage digest; result-artifact full +digest/short ID; and a sidecar self-digest. Audit locations/timestamps are not +included. Imported realizations require the import/source/evidence bindings; +native-v3 realizations encode those fields as explicit not-applicable records +with reasons and still require dataset, validation, origin, lineage, and result +bindings. They are never silently omitted. + +- [ ] **Step 1: Write failing result-artifact-v2 tests** + + Test exact CSV rendering in memory; mandatory uniform condition, realization, + experiment, dataset/model/method/prompt, split, `prediction_origin`, and + `prediction_lineage_id`; result ID/digest bound to the already fixed + experiment/realization and exact file bytes; ordered columns/question/origin + digests; no result ID inside CSV/manifest; correct realization paths; full + sidecar self-validation; identical no-op; one-byte divergence refusal before + writing; missing half-pair refusal; and v1 allowed only for legacy v2. Add a + tampering/mismatch case for every v2 metadata binding above and compare it + against the manifest realization, validation artifact, evidence index, exact + CSV, and run state rather than merely trusting the sidecar's own digest. + + ```python + def test_result_identity_is_computed_after_manifest(v3_manifest, native_rows): + prepared = prepare_manifest_result( + native_rows, manifest=v3_manifest, + realization_id=native_rows[0]["realization_id"], + ) + assert "result_artifact_id" not in v3_manifest + assert prepared.metadata["experiment_id"] == v3_manifest["experiment_id"] + assert prepared.metadata["file_sha256"] == hashlib.sha256( + prepared.csv_bytes + ).hexdigest() + assert prepared.metadata["result_artifact_id"].startswith("result_") + ``` + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_result_artifact.py \ + tests/io/test_writers.py tests/test_publication_identity.py -v + ``` + + Expected: new API tests fail; legacy writer tests pass. + +- [ ] **Step 3: Implement preparation, v2 sidecar, and collision refusal** + + Render once to UTF-8 bytes, validate rows, compute the sidecar payload and + `result_artifact_digest`, then add its short ID and sidecar integrity digest. + `publish_manifest_result` may write only inside a staging tree or to an absent + pair. If either target exists, validate and compare the entire proposed graph; + return no-op only for exact equality. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_result_artifact.py \ + tests/io/test_writers.py tests/test_publication_identity.py -v + python -m pytest -q + ``` + + Expected: v1/v2 sidecar dispatch and non-circular identity cases pass; no + baseline regression. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/io/writers.py tests/importing/test_result_artifact.py \ + tests/io/test_writers.py tests/test_publication_identity.py + git commit -m "feat: publish realization result artifacts" + ``` + +**Intermediate gate:** Independent provenance/collision review. Require a +reviewer-produced dependency graph proving +condition -> realization -> experiment -> CSV/result artifact, never the +reverse. + +--- + +### Task 8: Validate source rows by question identity + +**Files:** + +- Create: `src/choicebench/importing/validation.py` +- Create: `tests/importing/test_validation.py` +- Inspect: `src/choicebench/pipeline/options.py`, + `src/choicebench/scoring/scorer.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class ValidationFinding: + code: str + source_id: str + field: str | None + question_id: str | None + record_span: LogicalRecordSpan | None + raw_row_sha256: str | None + message: str + + +@dataclass(frozen=True) +class ValidatedSourceRows: + source_id: str + rows_by_question_id: Mapping[str, SourceRow] + expected_question_ids: tuple[str, ...] + findings: tuple[ValidationFinding, ...] + findings_digest: str + + +def validate_source_rows( + table: AdaptedTable, + *, + condition: ImportConditionSpec, + expected: ExpectedDataset, +) -> ValidatedSourceRows: ... +``` + +- [ ] **Step 1: Write failing membership and scientific-content tests** + + Add one focused test per required failure: missing required field; null ID; + duplicate ID with both raw spans; missing ID; unexpected ID; gold mismatch; + option count mismatch; ordered option text mismatch; correct-option mismatch; + invalid parsed prediction for the row's actual option set; recomputable + correctness mismatch; malformed numeric/NaN/infinity; model/benchmark/split/ + method ownership mismatch; status inconsistent with observed coverage; and + source-schema/extra-column violation. + + Parameterize option coverage with three-, four-, and six-option rows. Assert no + input row is dropped, padded, deduplicated, reordered, repaired, or converted + from malformed to valid. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_validation.py -v + ``` + + Expected: FAIL because `validation.py` is absent. + +- [ ] **Step 3: Implement the fail-closed validator** + + Join only by exact string `question_id`. Compare all source-provided + question/gold/option fields with the expected snapshot, but never let source + values override it. Accumulate all deterministic findings in source-record + order and raise a single actionable validation error for undeclared defects. + Sanitize source text in messages while preserving field, question, span, and + digest evidence. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_validation.py \ + tests/pipeline/test_options.py tests/scoring/test_scorer.py -v + python -m pytest -q + ``` + + Expected: all question-keyed failures are detected and the baseline remains + green. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/importing/validation.py \ + tests/importing/test_validation.py + git commit -m "feat: validate imported rows by question identity" + ``` + +--- + +### Task 9: Preserve evidence status, malformed bytes, and extension fields + +**Files:** + +- Modify: `src/choicebench/importing/validation.py` +- Create: `tests/importing/test_evidence_status.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class RealizationValidation: + evaluable_rows: tuple[Mapping[str, Any], ...] + findings: tuple[ValidationFinding, ...] + validation_digest: str + computed_evidence_status: Literal[ + "complete", "qualified", "partial", "malformed", "recoverable", "failed" + ] + evaluable: bool + qualifications: tuple[Mapping[str, Any], ...] + limitations: tuple[Mapping[str, Any], ...] + defect_question_ids: tuple[str, ...] + + +@dataclass(frozen=True) +class PreparedValidationArtifact: + relative_path: str + json_bytes: bytes + file_sha256: str + validation_digest: str + + +def normalize_realization_rows( + validated: ValidatedSourceRows, + *, + condition: ImportConditionSpec, + expected: ExpectedDataset, + experiment_id: str | None = None, + realization_id: str | None = None, +) -> RealizationValidation: ... + + +def prepare_realization_validation_artifact( + validation: RealizationValidation, + *, + realization: Mapping[str, Any], + evidence_records: Sequence[Mapping[str, Any]], +) -> PreparedValidationArtifact: ... + + +def write_realization_validation_artifact( + staged_run: Path, prepared: PreparedValidationArtifact +) -> tuple[Path, bool]: ... + + +def validate_realization_validation_artifact( + path: Path, + *, + manifest: Mapping[str, Any], + realization: Mapping[str, Any], + expected_sha256: str, +) -> Mapping[str, Any]: ... +``` + +The first call before manifest construction produces a canonical pre-ownership +payload/digest. The second call after experiment and realization IDs are fixed +injects ownership without changing predictions or evidence decisions. + +- [ ] **Step 1: Write failing status/extension tests** + + Cover complete, qualified, partial, malformed, recoverable, and failed + evidence; complete+excluded and malformed+excluded; qualifications preserved + beside metrics; exact declared defect IDs and reasons; undeclared defect + refusal; no evaluable rows for partial/malformed/recoverable/failed/excluded/ + held/superseded or `aggregate_only` sources; unknown/extra fields preserved in namespaced + `external_fields_json`; explicit ignored fields and reasons; credential-shaped + field refusal; literal null/NaN preservation; and method-specific nested JSON + round-trip. + + For malformed evidence assert: + + ```python + assert finding.record_span == LogicalRecordSpan(2, 19, 47, b"\r\n") + assert finding.raw_row_sha256 == sha256(source_bytes[19:47]).hexdigest() + assert validation.evaluable_rows == () + ``` + + Assert deterministic JSON bytes at + `artifacts/imports/validation/.json`, self-digest validation, + exact realization/evidence/finding bindings, checksum change on any finding or + raw-row digest change, atomic temp-file publication inside staging, identical + no-op/divergent refusal, and refusal for traversal, unsafe path, + realization/evidence mismatch, tampering, or a raw-span digest that does not + reproduce from the archived evidence bytes. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_evidence_status.py -v + ``` + + Expected: FAIL because normalization/status behavior is not implemented. + +- [ ] **Step 3: Implement normalization without evidence repair** + + Fill evaluative question text/options/gold only from `ExpectedDataset`; retain + original mapped and extension fields separately. Recompute evidence status + from exact coverage/defects and refuse a declaration mismatch. Produce + evaluable rows only for imported+included complete/qualified realizations. + Never fabricate request IDs, timestamps, usage, prompts, responses, provider + fields, commits, revisions, or seeds. + + Prepare one exact validation artifact for every realization, including + evidence-only realizations. Its semantic `validation_digest` is available to + realization identity before ownership; after the realization is fixed, its + exact serialized file SHA-256 is recorded downstream in run state and any + result sidecar, never back-propagated into the realization/manifest identity. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_validation.py \ + tests/importing/test_evidence_status.py -v + python -m pytest -q + ``` + + Expected: status eligibility and exact malformed evidence pass; no baseline + regression. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/importing/validation.py \ + tests/importing/test_evidence_status.py + git commit -m "feat: preserve imported evidence and status" + ``` + +--- + +### Task 10: Validate typed authorization and derive immutable overlays + +**Files:** + +- Create: `src/choicebench/importing/authorization.py` +- Create: `src/choicebench/importing/overlays.py` +- Create: `tests/importing/test_authorization.py` +- Create: `tests/importing/test_overlays.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class ValidatedAuthorizationBundle: + authorization_id: str + authorization_digest: str + authorization_type: Literal["inference_repair", "offline_transformation"] + grants: Mapping[str, Mapping[str, str]] # condition_digest -> qid -> reason + authority: str + purpose: str + executable: bool + source_sha256: str + input_evidence_digests: Mapping[str, str] + expected_snapshot_digests: Mapping[str, str] + + +@dataclass(frozen=True) +class ValidatedAuthorization: + bundle_id: str + bundle_digest: str + authorization_type: Literal["inference_repair", "offline_transformation"] + condition_digest: str + question_reasons: Mapping[str, str] + authority: str + purpose: str + executable: bool + source_sha256: str + input_evidence_digests: Mapping[str, str] + expected_snapshot_digest: str + + +@dataclass(frozen=True) +class VerifiedBaseRealization: + condition_id: str + condition_digest: str + realization_id: str + realization_digest: str + evidence_index_digest: str + evidence_source_digests: Mapping[str, str] + validation_artifact_sha256: str + result_sha256: str | None + rows_by_question_id: Mapping[str, Mapping[str, Any]] + prediction_origins: Mapping[str, str] + + +@dataclass(frozen=True) +class DerivedRealizationPayload: + condition_id: str + condition_digest: str + preownership_rows: tuple[Mapping[str, Any], ...] + lineage_components: tuple[Mapping[str, Any], ...] + result_origin: Mapping[str, Any] + evidence_status: str + replacement_question_ids: tuple[str, ...] + + +def validate_authorization_bundle( + declaration: AuthorizationSpec, + *, + opened_source: OpenedSource, + condition_digests: Mapping[str, str], + expected: Mapping[str, ExpectedDataset], +) -> ValidatedAuthorizationBundle: ... + + +def authorization_for_condition( + bundle: ValidatedAuthorizationBundle, *, condition_digest: str +) -> ValidatedAuthorization: ... + + +def derive_overlay( + *, + base: VerifiedBaseRealization, + overlay: OverlaySpec, + authorization: ValidatedAuthorization, + overlay_table: AdaptedTable, + expected: ExpectedDataset, +) -> DerivedRealizationPayload: ... +``` + +- [ ] **Step 1: Write failing typed-authorization tests** + + `inference_repair` must require exact condition/question membership, + `queue_disposition=approved`, `execution_authority=authoritative`, and + `executable=true`. Assert refusal for held, excluded, forensic, absent, + false-executable, and offline records. `offline_transformation` must require a + separately opened checksum-verified source with exact IDs/reasons/authority/ + purpose, input-evidence digests, expected-snapshot digest, and + `inference_executable=false`; an overlay/spec cannot authorize its own IDs. + Assert the types cannot substitute for one another or bind a different base + condition/snapshot/evidence graph. + + `derive_overlay` requires exact equality between + `overlay.base_evidence_digests`, the authorization slice's applicable + `input_evidence_digests`, and the realization-scoped source/blob digest mapping + returned by `verify_import_run`; no evidence-index digest alone is accepted as + a substitute. Refuse missing, extra, or cross-realization evidence digests. + + Test both a one-condition bundle and a single content-addressed six-condition/ + 18-pair bundle. The aggregate digest binds every condition digest, question, + reason, evidence/snapshot digest, purpose, authority, and executable flag; + each condition-scoped slice retains the same immutable bundle digest and + cannot add or borrow grants from another condition. + +- [ ] **Step 2: Write failing overlay/lineage tests** + + Cover a valid evidence-only base with no result checksum; a valid evaluable + base; exact replacement subset; unauthorized, duplicate, unexpected, missing, + and conflicting replacement IDs; multiple overlays targeting one ID; + gold/options/ownership mismatch; status recomputation; base evidence/result + checksum preservation; base -> authorization -> overlay/transformation -> + resulting lineage; mixed historical/native repair origins with exact per-row + mapping; external repair origin; offline rematching preserving underlying + prediction origin; base condition/realization/validation-artifact digest + verification; and transformation pre-ownership I/O/code digests. Require a + declared origin assignment for every replacement ID: native repair rows use + `native_inference`, externally generated repair rows use + `external_repair_inference`, and offline rematching retains the base response's + prediction origin while recording `offline_transformation` as derivation. + Reject missing/extra assignments and any bare `mixed` value; identity-bind the + complete per-row mapping and stable lineage notes. + + `derive_overlay` accepts only a `VerifiedBaseRealization`, never a filesystem + path or unverified caller dictionary. Unit tests construct this frozen trusted + value and prove the pure function cannot mutate it. Task 12 integration + snapshots every base-run file before/after full-graph verification and + derivation and requires byte equality. The derived payload shares the semantic + condition only when model/dataset/method/prompt/protocol are unchanged. + +- [ ] **Step 3: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_authorization.py \ + tests/importing/test_overlays.py -v + ``` + + Expected: FAIL because authorization/overlay modules are absent. + +- [ ] **Step 4: Implement typed verification and pure derivation** + + Bundle verification consumes an independently opened authorization artifact. + The overlay function is pure with respect to the filesystem and accepts only + the trusted base value produced later by the engine's full graph verifier. It + returns a derived pre-ownership payload and never opens/writes a base run or + performs inference. Use the Task 5 lineage and origin builders, excluding + child IDs from lineage identity. + +- [ ] **Step 5: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_authorization.py \ + tests/importing/test_overlays.py tests/importing/test_identity.py -v + python -m pytest -q + ``` + + Expected: typed authority, overlay conflicts, mixed origins, and base + immutability all pass. + +- [ ] **Step 6: Commit** + + ```bash + git add src/choicebench/importing/authorization.py \ + src/choicebench/importing/overlays.py \ + tests/importing/test_authorization.py tests/importing/test_overlays.py + git commit -m "feat: add authorized immutable import overlays" + ``` + +**Intermediate gate:** Independent authorization/overlay review. The reviewer +must attempt self-authorization, cross-type authorization, duplicate/conflicting +replacement, base mutation, and a bare `mixed` origin. + +--- + +### Task 11: Add verified evidence storage and atomic staged publication + +**Files:** + +- Modify: `src/choicebench/infra/artifacts.py` +- Create: `src/choicebench/importing/evidence.py` +- Create: `src/choicebench/importing/transaction.py` +- Create: `tests/importing/test_evidence.py` +- Create: `tests/importing/test_transaction.py` +- Modify: `tests/test_release_adversarial.py` + +**Interfaces:** + +```python +def open_verified_source( + declaration: SourceArtifactSpec, + *, + containment_root: Path | None = None, +) -> OpenedSource: ... + + +def evidence_blob_path(staged_run: Path, sha256: str, format_name: str) -> Path: ... +def write_evidence_blob( + staged_run: Path, source: OpenedSource, references: Sequence[str] +) -> dict[str, Any]: ... +def validate_evidence_blob(run_dir: Path, record: Mapping[str, Any]) -> None: ... +def write_evidence_index( + staged_run: Path, records: Sequence[Mapping[str, Any]] +) -> Mapping[str, Any]: ... +def validate_evidence_index( + run_dir: Path, expected_digest: str +) -> Mapping[str, Any]: ... + + +def fsync_tree(root: Path) -> None: ... +def atomic_publish_directory_no_replace(source: Path, destination: Path) -> None: ... + + +class ImportTransaction: + def __init__(self, *, runs_dir: Path, run_id: str): ... + def __enter__(self) -> "ImportTransaction": ... + @property + def staged_run(self) -> Path: ... + def publish(self, validator: Callable[[Path], None]) -> Path: ... + def __exit__(self, exc_type, exc, traceback) -> None: ... +``` + +- [ ] **Step 1: Write failing verified-source/evidence tests** + + Cover independently computed checksum mismatch; source mutation/replacement; + same opened bytes used for hash and parse; absolute user source; profile-root + containment; traversal, directory, device, and unsafe symlink refusal; suffix + derived from validated format; within-run dedup; cross-run independent copy; + tampered blob/sidecar/index; divergent digest-path refusal; exact malformed row + evidence; no recursive import; and secret-shaped metadata refusal. + +- [ ] **Step 2: Write failing transaction tests** + + Assert the established `.locks/.manifest.lock`; same-filesystem sibling + owner-marked staging; full validator invoked before publication; destination + absent; real atomic no-replace behavior; fsync calls; failure leaves final path + absent; stale staging is never recognized as a run; safe cleanup touches only + its owner-marked directory; existing final run is verify-only and creates no + staging; identical final graph no-ops; divergent graph refuses unchanged. + + Linux implementation tests should exercise libc `renameat2(..., + RENAME_NOREPLACE)` through a small private wrapper. Mock absence of that symbol + and assert fail-closed behavior; do not fall back to `os.replace` or an + existence-check race. + +- [ ] **Step 3: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_evidence.py \ + tests/importing/test_transaction.py tests/test_release_adversarial.py -v + ``` + + Expected: FAIL because evidence/transaction APIs do not exist. + +- [ ] **Step 4: Implement same-bytes source opening and run-local CAS** + + Open only declared regular files; read once; hash `data`; pass `data` forward. + Store only referenced row-level bytes at + `artifacts/imports/evidence/sha256//.` + inside staging. Sidecars bind byte size/format/references; identical bytes in + one run reuse the blob. Write each blob, sidecar, and the self-digested + evidence index through a temporary file, flush/fsync, revalidate exact bytes, + and atomically rename within staging. Never accept a divergent existing blob, + sidecar, or index. + +- [ ] **Step 5: Implement staged no-replace publication** + + Acquire the manifest lock before examining final state. For a new run, build + under a sibling staging directory, validate/fsync the complete tree, and call + the no-replace primitive once. For an existing run, invoke the supplied full + graph verifier only; exact equality no-ops and divergence raises. A report + outside the run cannot affect commit state. + +- [ ] **Step 6: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_evidence.py \ + tests/importing/test_transaction.py tests/test_release_adversarial.py -v + python -m pytest -q + ``` + + Expected: source, CAS, crash, symlink, no-replace, and idempotence tests pass; + no baseline regression. + +- [ ] **Step 7: Commit** + + ```bash + git add src/choicebench/infra/artifacts.py \ + src/choicebench/importing/evidence.py \ + src/choicebench/importing/transaction.py \ + tests/importing/test_evidence.py tests/importing/test_transaction.py \ + tests/test_release_adversarial.py + git commit -m "feat: stage external imports atomically" + ``` + +**Intermediate gate:** Independent filesystem/security review on source opening, +staging containment, owner markers, symlink behavior, no-replace semantics, +cleanup scope, and verify-only final runs. + +--- + +### Task 12: Orchestrate generic dry-run and real imports + +**Files:** + +- Create: `src/choicebench/importing/engine.py` +- Modify: `src/choicebench/importing/__init__.py` +- Create: `tests/importing/test_engine.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class ImportRequest: + spec: ImportSpec + run_id: str + workspace_root: Path + strict: bool + overlays: tuple[OverlaySpec, ...] = () + + +@dataclass(frozen=True) +class ImportPlan: + manifest: Mapping[str, Any] + expected_datasets: Mapping[str, ExpectedDataset] + realizations: Mapping[str, RealizationValidation] + evidence_records: tuple[Mapping[str, Any], ...] + report_identity: Mapping[str, Any] + + +@dataclass(frozen=True) +class VerifiedImportRun: + manifest: Mapping[str, Any] + manifest_digest: str + realizations: Mapping[str, VerifiedBaseRealization] + + +@dataclass(frozen=True) +class ImportCounts: + conditions: Mapping[str, int] # exact keys listed below + realizations: Mapping[str, int] + rows: Mapping[str, int] + evidence_status: Mapping[str, int] + scope_disposition: Mapping[str, int] + prediction_origin: Mapping[str, int] + derivation_origin: Mapping[str, int] + source_classification: Mapping[str, int] + column_disposition: Mapping[str, int] + authorization: Mapping[str, int] + + +@dataclass(frozen=True) +class DefectReport: + total_findings: int + by_code: Mapping[str, int] + question_ids_by_code: Mapping[str, tuple[str, ...]] + findings_digest: str + + +@dataclass(frozen=True) +class ChecksumRecord: + logical_id: str + expected_sha256: str | None + actual_sha256: str + matches: bool | None + + +@dataclass(frozen=True) +class ArtifactChecksumRecord: + logical_path: str + file_sha256: str + size: int + semantic_digest: str | None + + +@dataclass(frozen=True) +class ImportChecksumReport: + all_referenced_sources_match: bool + sources: Mapping[str, ChecksumRecord] + expected_snapshots: Mapping[str, ArtifactChecksumRecord] + evidence_index: ArtifactChecksumRecord | None + validation_artifacts: Mapping[str, ArtifactChecksumRecord] + results: Mapping[str, ArtifactChecksumRecord] + + +@dataclass(frozen=True) +class ResultIdentityReport: + result_artifact_id: str + result_artifact_digest: str + file_sha256: str + + +@dataclass(frozen=True) +class ImportIdentityProjection: + import_spec_digest: str + semantic_grid_digest: str + experiment_id: str | None + experiment_digest: str | None + condition_digests: Mapping[str, str] + realization_digests: Mapping[str, str] + result_artifacts: Mapping[str, ResultIdentityReport] + authorization_bundle_digests: Mapping[str, str] + lineage_digests: Mapping[str, str] + + +@dataclass(frozen=True) +class ImportFailureReport: + code: str + source_id: str | None + field: str | None + question_id: str | None + record_span: Mapping[str, int | str] | None + raw_row_sha256: str | None + sanitized_message: str + + +@dataclass(frozen=True) +class ImportAuditReport: + source_location: str | None + output_root: str | None + imported_at: str + + +@dataclass(frozen=True) +class ImportReport: + schema_version: Literal["choicebench.import-report.v1"] + import_state: Literal["validated", "imported", "failed"] + wrote_artifacts: bool + idempotent_noop: bool + run_id: str + experiment_id: str | None + counts: ImportCounts + defects: DefectReport + checksums: ImportChecksumReport + identity_projection: ImportIdentityProjection + profile: Mapping[str, Any] + failures: tuple[ImportFailureReport, ...] + audit: ImportAuditReport + + +def build_import_plan(request: ImportRequest) -> ImportPlan: ... +def verify_import_run( + run_dir: Path, expected: ImportPlan | None = None +) -> VerifiedImportRun: ... +def execute_import(request: ImportRequest, *, dry_run: bool = False) -> ImportReport: ... +def serialize_import_report(report: ImportReport) -> dict[str, Any]: ... +``` + +The serialized JSON uses these fields exactly at top level. `profile` is `{}` +for generic imports; the Stage 1 profile uses stable `matrix`, `queues`, +`offline_authority`, `datasets`, and `source_schema_count` keys. `failures` +contains deterministic code/source/field/question/span/digest plus sanitized +message entries and is empty on success. Tracebacks, credentials, raw untrusted +text, absolute temporary paths, and timestamps appear in neither failures nor +identity projection; safe machine-local locations/timestamps may appear only +under `audit`. + +Nested report-v1 keys are also closed and exact: + +```text +counts.conditions: declared, intended, preserved +counts.realizations: planned, evaluable, evidence_only, selected +counts.rows: source, expected, validated, evaluable, defective, replaced +counts.evidence_status: complete, qualified, partial, malformed, recoverable, failed +counts.scope_disposition: included, excluded_from_paper_matrix, held, superseded +counts.prediction_origin: native_inference, external_historical_inference, + external_repair_inference +counts.derivation_origin: native_execution, external_import, repair_overlay, + offline_transformation +counts.source_classification: raw, canonical, derived, repaired, aggregate_only +counts.column_disposition: mapped, namespaced_preserved, ignored_with_reason +counts.authorization: inference_repair_bundles, offline_transformation_bundles, + granted_pairs, executable_pairs, nonexecutable_pairs +defects: total_findings, by_code, question_ids_by_code, findings_digest +checksums: all_referenced_sources_match, sources, expected_snapshots, + evidence_index, validation_artifacts, results +identity_projection: import_spec_digest, semantic_grid_digest, + experiment_id, experiment_digest, condition_digests, + realization_digests, result_artifacts, + authorization_bundle_digests, lineage_digests +identity_projection.result_artifacts[realization_id]: result_artifact_id, + result_artifact_digest, file_sha256 +profile: matrix, queues, offline_authority, datasets, source_schema_count +failures[]: code, source_id, field, question_id, record_span, + raw_row_sha256, sanitized_message +audit: source_location, output_root, imported_at +``` + +Every mapping has unknown-key rejection in report self-validation. Empty +categories remain explicit zero/empty mappings. Generic success, validation +failure, imported success, exact no-op, repair overlay, offline transformation, +and Stage 1 profile tests assert the entire nested key set and cross-total +invariants; report identity excludes `audit` and `import_state` but includes all +stable semantic/provenance/lineage projections. +`record_span` has exactly `record_index`, `start`, `end`, and +`terminator_hex`. Tests delete, add, and tamper every typed leaf in turn and +require report self-validation failure. + +- [ ] **Step 1: Write failing end-to-end synthetic engine tests** + + Cover valid generic CSV import; deterministic IDs; dry run writing nothing; + real complete/qualified artifacts; partial/malformed/recoverable evidence-only + realizations; source checksum mismatch and mutation; exact repeated import + no-op with a byte-identical tree; mapping/source/status/overlay divergence + refusal; result artifact/run-state linkage; atomic failure; import report + counts; audit paths excluded from identities; and import/evaluation outside the + repository using an absolute workspace. + + Assert run state uses `status=completed` for evaluable realizations and + `status=evidence_only` for preserved ineligible realizations. Imported + realization records explicitly use `import_state=imported`; a failed + transaction publishes no realization/run. Dry-run reports use + `import_state=validated` without changing planned realization IDs. + + Add a `load_import_spec()` -> `build_import_plan()` integration assertion for + the exact dataset/model/method/prompt/semantic-condition bridge from Task 5, + including the scientific-versus-realization-only mutation matrix. Assert a + failed validation serializes `import_state=failed`, both operation booleans + false, no experiment ID, and structured sanitized failures without publishing + a run. + + Add base-plus-overlay integration only here, after the full verifier exists: + verify the immutable base into `VerifiedImportRun`, pass the selected frozen + realization to Task 10's pure derivation, and publish the derived result under + a different run/experiment. Snapshot every base byte before/after; cover valid, + evidence-only, unauthorized, conflict, mixed-origin, offline, and divergent + cases. No Task 10 function opens the base filesystem directly. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_engine.py -v + ``` + + Expected: FAIL because engine orchestration is absent. + +- [ ] **Step 3: Implement planning and dry-run** + + Resolve sources/datasets/references, adapt rows, validate, build conditions and + realizations, resolve every importer metric only through the closed + `BUILTIN_METRICS` registry and bind its registered implementation identity, + then create manifest v3. Dry run executes every read/validation/ + identity step and returns the report without creating workspace directories, + evidence, staging, results, state, or manifest. It still renders eligible + final CSV bytes and prepares validation/result sidecars in memory after the + planned manifest is fixed, so report-v1 includes the same per-realization + result artifact IDs/digests/file SHA-256 as a real identical import. + +- [ ] **Step 4: Implement staged real publication and graph verification** + + Use `ImportTransaction`; write snapshots/evidence/validation artifacts; + normalize eligible rows with final ownership; prepare/publish result artifacts + in staging; write final realization-keyed state with exact evidence-index and + validation-artifact SHA-256 for every realization (plus result identity/digest/ + SHA-256 only when evaluable); validate the entire staged + graph; publish once. Existing final runs call `verify_import_run` and either + exact no-op or refuse. + + `verify_import_run` is the single shared full-graph verifier used for existing- + run idempotence, overlay base input, and Task 13 publication reading. It + validates manifest/state/snapshots/evidence/index/validation artifacts/results/ + sidecars, derives each realization's exact logical-source/blob SHA-256 mapping, + and returns immutable trusted values only after the whole graph passes. + +- [ ] **Step 5: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_engine.py \ + tests/importing/test_transaction.py tests/importing/test_result_artifact.py -v + python -m pytest -q + ``` + + Expected: all generic dry-run/import/idempotence/status cases pass; no baseline + regression. + +- [ ] **Step 6: Commit** + + ```bash + git add src/choicebench/importing/engine.py \ + src/choicebench/importing/__init__.py tests/importing/test_engine.py + git commit -m "feat: import external results into immutable runs" + ``` + +--- + +### Task 13: Read and evaluate explicit realizations + +**Files:** + +- Modify: `src/choicebench/io/readers.py` +- Modify: `src/choicebench/cli/evaluate_run.py` +- Create: `tests/importing/test_evaluation.py` +- Modify: `tests/io/test_readers.py` +- Modify: `tests/scripts/test_evaluate_run.py` +- Modify: `tests/test_publication_identity.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class RealizationSelection: + policy: Literal["single_evaluable_per_condition", "explicit"] + realization_ids: tuple[str, ...] + + +@dataclass(frozen=True) +class ManifestResultSet: + rows: pd.DataFrame + manifest: Mapping[str, Any] + view: ManifestView + state: Mapping[str, Any] + selection: RealizationSelection + artifacts: Mapping[str, Mapping[str, Any]] + accounting: Mapping[str, Mapping[str, Any]] + + +def read_manifest_result_set( + run_dir: Path, *, realization_ids: Sequence[str] | None = None +) -> ManifestResultSet: ... + + +def read_manifest_results(run_dir: Path) -> tuple[pd.DataFrame, dict]: ... + + +def build_evaluation_report_for_result_set( + run_id: str, + result_set: ManifestResultSet, + *, + reparse: bool, +) -> dict[str, Any]: ... +``` + +`read_manifest_results` remains the v2-compatible wrapper. For v3 it may +auto-select only when each semantic condition has at most one eligible +realization; ambiguity is a refusal. Add repeatable CLI +`choicebench-evaluate --realization-id `. + +- [ ] **Step 1: Write failing realization-reader tests** + + Cover manifest/state/evidence/snapshot/realization-validation-artifact/sidecar full-graph validation; undeclared + result refusal; result for evidence-only realization refusal; exact row + ownership; question membership and uniqueness; and fieldwise join of + `question_text`, serialized choices, `correct_option`, and correct answer text + against the expected snapshot. Include one variable-option ARC-style result. + + Add two eligible realizations for one condition and assert default refusal, + explicit selection of exactly one, unknown/ineligible ID refusal, and no + concatenation. All unselected/ineligible realizations remain in accounting. + +- [ ] **Step 2: Write failing evaluation-v2 tests** + + Assert complete/qualified selected metrics; qualifications/limitations beside + qualified metrics; partial/malformed/recoverable/failed/excluded/held/ + superseded accounting with null result/metrics; complete+excluded and + malformed+excluded orthogonality; condition grouping separate from realization + metrics; selection policy and every accounted digest in evaluation identity; + audit path/timestamp changes excluded; result byte change included; and + selection change produces a new evaluation ID. Add an importer-manifest metric + changed to `module:Class` and prove refusal occurs before `importlib` is called; + separately retain a legacy native custom-metric compatibility test. + + Report shape must include: + + ```python + assert report["conditions"][condition_id]["realization_ids"] == [real_a, real_b] + assert report["realizations"][real_a]["selected"] is True + assert report["realizations"][real_a]["benchmark_name"] == "arc_challenge" + assert report["realizations"][real_a]["evidence_status"] == "complete" + assert report["realizations"][real_a]["scope_disposition"] == "included" + assert report["realizations"][real_a]["prediction_origins"] == [ + "external_historical_inference" + ] + assert report["realizations"][real_a]["option_count_distribution"] == { + "3": 1, + "4": 1, + } + assert report["realizations"][real_b]["selected"] is False + assert report["realizations"][malformed]["metrics"] == {} + ``` + + Re-run the Task 1 v2 fixture and require the exact evaluation-v1 identity and + shape to remain unchanged. + +- [ ] **Step 3: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_evaluation.py tests/io/test_readers.py \ + tests/scripts/test_evaluate_run.py \ + tests/importing/test_manifest_v2_compat.py -v + ``` + + Expected: new v3 selection/evaluation tests fail; legacy tests pass. + +- [ ] **Step 4: Implement version-dispatched reading** + + Preserve the current v2 path in a private helper. For v3, use + `verify_import_run` rather than duplicating graph verification, then enforce + row ownership, status eligibility, and explicit selection over its trusted + realization records. + Never call `read_all_run_results`. + +- [ ] **Step 5: Implement evaluation-v2 and keep evaluation-v1** + + Branch on manifest schema. Keep the existing v1 identity/report function + byte-for-byte for legacy input. Build v2 units sorted by full realization + identity and bind selection, statuses, qualifications, evidence, + authorization, lineage, result IDs/digests/checksums, metric implementations, + and optional reparsing implementations. + + For importer-created manifests, instantiate only the manifest-bound built-in + metric name after rechecking it against `BUILTIN_METRICS` and its recorded + implementation identity; never pass an importer metric string to the current + dynamic `_load_metric` path. Preserve the existing trusted native v2/v3 custom- + metric behavior without making it reachable from an import specification. + +- [ ] **Step 6: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_evaluation.py tests/io/test_readers.py \ + tests/scripts/test_evaluate_run.py tests/test_publication_identity.py \ + tests/importing/test_manifest_v2_compat.py -v + python -m pytest -q + ``` + + Expected: v2 identity compatibility and v3 selection/accounting both pass; + the full baseline is green. + +- [ ] **Step 7: Commit** + + ```bash + git add src/choicebench/io/readers.py \ + src/choicebench/cli/evaluate_run.py tests/importing/test_evaluation.py \ + tests/io/test_readers.py tests/scripts/test_evaluate_run.py \ + tests/test_publication_identity.py + git commit -m "feat: evaluate explicit result realizations" + ``` + +**Intermediate gate:** Independent reader/evaluation review. Require tests or a +manual adversarial construction demonstrating that two alternatives cannot be +silently combined and that v2 evaluation identity remains exact. + +--- + +### Task 14: Migrate newly created native runs to manifest v3 + +**Files:** + +- Modify: `src/choicebench/cli/run_experiment.py` +- Modify: `src/choicebench/infra/checkpoint.py` +- Modify: `tests/test_condition_grid.py` +- Modify: `tests/infra/test_checkpoint.py` +- Modify: `tests/scripts/test_reset_run.py` +- Modify: `tests/scripts/test_build_backend.py` +- Modify: `tests/test_pride_reproduction_wiring.py` +- Modify: `tests/test_wheel_smoke.py` +- Create: `tests/importing/test_native_v3.py` + +**Interfaces:** + +```python +class ExecutionPlan: + selections: Sequence[BenchmarkSelection] + preflights: Mapping[tuple[int, int], Any] + semantic_conditions: Mapping[str, Mapping[str, Any]] + realizations: Mapping[str, Mapping[str, Any]] + jobs: Mapping[tuple[int, int, int], Mapping[str, Any]] + manifest: Mapping[str, Any] + + +def build_execution_plan( + config: ExperimentConfig, + *, + manifest_schema: Literal["v3", "legacy_v2_resume"] = "v3", +) -> ExecutionPlan: ... +``` + +New checkpoint-v2 records use `checkpoints/.json` and bind +experiment, condition, realization, and selection. Existing checkpoint-v1 and +programmatic behavior remain readable only in legacy-v2 resume mode. +`build_execution_plan(..., manifest_schema="v3")` calls the explicit v3 +manifest/state builders; it does not change the legacy public v2 constants or +the default behavior of `make_manifest()` for external programmatic callers. + +- [ ] **Step 1: Write failing native-v3 tests before changing the runner** + + Assert a new dry execution plan has manifest v3; unchanged current + `condition_id`; one `native_execution` realization per condition; explicit + `native_inference` row assignment/lineage; realization-keyed run state; + realization-addressed checkpoint/result/validation paths; deterministic native + validation artifact with no external-evidence claim; result-artifact-v2; and native + evaluation-v2. Assert no missing result origin in a new manifest. + + Add a pre-existing v2 run/resume test: detect the existing manifest before + candidate creation, rebuild a v2-compatible plan, verify exact experiment/ + condition/path ownership, resume/skip without rewriting completed bytes, and + never upgrade the manifest/state/sidecar in place. A new absent run may never + request `legacy_v2_resume`. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_native_v3.py \ + tests/test_condition_grid.py tests/infra/test_checkpoint.py \ + tests/scripts/test_reset_run.py -v + ``` + + Expected: new-run v3 assertions fail; all existing runner tests pass. + +- [ ] **Step 3: Split runtime jobs from semantic/realization manifest records** + + Preserve the current condition identity payload including the single frozen + prompt-path compatibility discriminator. Build one explicit native realization + per job. Runtime jobs may merge descriptive condition/realization fields for + execution, but the manifest tables stay separate and semantic records contain + no realization back-reference. + +- [ ] **Step 4: Inject explicit native row origin and use v3 result publication** + + Prefer the existing `condition_metadata` injection after each runner batch so + method implementations remain untouched. Inject condition/realization and + ownership IDs, `prediction_origin=native_inference`, and a non-null stable + `prediction_lineage_id`. Validate the completed native rows against the bound + dataset snapshot, write the same self-digested realization-validation artifact + with an empty external-evidence reference set, then publish through + `prepare_manifest_result`/`publish_manifest_result`; update realization-keyed + run state with validation SHA-256 plus result ID/digest/file checksum. Native + failure/gate accounting retains its existing evidence artifact and emits no + fake imported source/evidence record. + +- [ ] **Step 5: Add versioned checkpoint/resume behavior** + + New checkpoints bind `realization_id`; legacy-v2 resume continues using the + old condition checkpoint/result contract and implicit native origin. Do not + combine v2 checkpoint rows with v3 rows. Preserve current reset safety. + +- [ ] **Step 6: Run focused native compatibility tests** + + ```bash + python -m pytest tests/importing/test_native_v3.py \ + tests/test_condition_grid.py tests/infra/test_checkpoint.py \ + tests/scripts/test_reset_run.py tests/scripts/test_build_backend.py \ + tests/test_pride_reproduction_wiring.py tests/test_wheel_smoke.py -v + ``` + + Expected: new native runs are v3, legacy v2 resumes remain byte-compatible, + and native dummy-backend workflows pass. + +- [ ] **Step 7: Run the full suite** + + ```bash + python -m pytest -q + ``` + + Expected: all original 808 tests plus importer tests pass with zero failures. + +- [ ] **Step 8: Commit** + + ```bash + git add src/choicebench/cli/run_experiment.py \ + src/choicebench/infra/checkpoint.py \ + tests/importing/test_native_v3.py tests/test_condition_grid.py \ + tests/infra/test_checkpoint.py tests/scripts/test_reset_run.py \ + tests/scripts/test_build_backend.py tests/test_pride_reproduction_wiring.py \ + tests/test_wheel_smoke.py + git commit -m "feat: emit native manifest v3 realizations" + ``` + +**Intermediate gate:** Independent compatibility review must exercise one new +v3 native run and one pre-existing v2 resume. Stop if either changes scientific +condition IDs or rewrites legacy bytes. + +--- + +### Task 15: Translate Stage 1 manifests and typed authorities + +**Files:** + +- Create: `src/choicebench/importing/profiles/__init__.py` +- Create: `src/choicebench/importing/profiles/stage1_paper_freeze.py` +- Extend: `tests/importing/conftest.py` +- Create: `tests/importing/test_stage1_profile.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class ProfileTranslation: + spec: ImportSpec + report: Mapping[str, Any] + + +def translate_stage1_paper_freeze(freeze_root: Path) -> ProfileTranslation: ... +``` + +The profile registry exposes the name `stage1-paper-freeze`. No generic module +may import this profile or contain any Stage 1 constant. + +- [ ] **Step 1: Build a programmatic miniature freeze fixture** + + In `tests/importing/conftest.py`, write only synthetic metadata/one-row source + CSVs under `tmp_path` with the same headers and joins as: + `expected_matrix.csv`, canonical manifest JSON/CSV, status matrix, + `artifact_inventory.csv`, `frozen_artifact_index.csv`, approved/held/excluded/ + forensic queues, report, and checksum ledger. Exercise the authoritative join + chain explicitly: + + ```text + canonical_results_manifest.canonical_artifact_id + -> artifact_inventory.artifact_id + -> frozen raw path and checksum + -> frozen_artifact_index canonical path and cell_id + ``` + + A missing, duplicate, or cross-cell link at any hop fails closed. Generate + checksum values only for the miniature equivalents of files covered by the + real ledger, including `reports/canonical_freeze_report.md`; intentionally + leave only artifact inventory/index unledgered and assert those two remain + independently hashed corroboration. Assert a one-byte canonical-report change + fails its ledger check. Include one included + complete, qualified, recoverable, partial, and malformed cell; both hosted + PriDe exclusion combinations; one approved inference repair; one held record; + one excluded record; and one forensic-only record. + +- [ ] **Step 2: Write failing translation/authority tests** + + Require joins by explicit `cell_id`/`canonical_artifact_id`, never filename. + Test exact status mapping: + + ```python + STATUS_MAP = { + "canonical_complete": "complete", + "canonical_qualified": "qualified", + "recoverable_from_existing_artifacts": "recoverable", + "incomplete_requires_inference": "partial", + "malformed_requires_inference": "malformed", + } + ``` + + `excluded_from_paper_matrix` is not an evidence-status input to this map. For + the two hosted PriDe rows, assign only + `scope_disposition=excluded_from_paper_matrix`, then run ordinary row/defect + validation to compute MMLU `complete` and ARC `malformed`. Assert these two + dimensions remain independent and the records are preserved but not intended. + + Assert approved/authoritative/executable inference authority is separate from + 687-style held false, excluded false, and forensic-only records. For every + queue row, expand exact `(cell_id, question_id)` pairs; require membership in + the canonical manifest's missing/damaged IDs as applicable; require the three + queue ledgers to be pairwise disjoint; and validate `queue_disposition`, + `execution_authority`, `executable`, and `supersedes_for_execution`. Assert all + 828 classified queue pairs equal 138 approved + 687 held + 3 excluded and the + 776 forensic pairs never become authority. + + Assert method identities and all eight exact canonical source-header schemas + map explicitly, with every column assigned to mapped, namespaced-preserved, or + explicitly ignored-with-reason. Cover `semantic_matching_v1`, + `text_extraction`, `two_stage_v1/v2/v3`, `independent_hypothesis`, + `cyclic_generation_majority`, `PriDe`, and historical `baseline` without + renaming/merging. + +- [ ] **Step 3: Write the exact offline-authority test** + + The fixture must include these six historical semantic-matching cell IDs and + authorize the same three question IDs for each cell: + + ```python + RECOVERABLE_CELLS = ( + "cbp__gemini-2-5-flash__arc_challenge__semantic_matching_v1", + "cbp__gpt-4-1-mini__arc_challenge__semantic_matching_v1", + "cbp__llama-3-1-8b-instant__arc_challenge__semantic_matching_v1", + "cbp__meta-llama-llama-3-1-8b-instruct__arc_challenge__semantic_matching_v1", + "cbp__qwen-qwen2-5-7b-instruct-turbo__arc_challenge__semantic_matching_v1", + "cbp__qwen-qwen2-5-7b-instruct__arc_challenge__semantic_matching_v1", + ) + ``` + + ```python + RECOVERABLE_QUESTION_IDS = ( + "79e8c959bbeb74a0", + "ad6b5d46ae54842c", + "c30e75b011696a95", + ) + assert report["offline_authority"]["condition_count"] == 6 + assert report["offline_authority"]["question_cell_count"] == 18 + assert report["offline_authority"]["inference_executable"] is False + ``` + + Assert each cell has exactly `RECOVERABLE_QUESTION_IDS` in + `recoverable_question_ids`; preserve every question reason; and bind the + canonical source SHA-256/base-evidence digest, full semantic condition digest, + expected-snapshot digest, allowed semantic-rematching purpose, and + `inference_executable=false`. Verify the authority is a deterministic + content-addressed projection of checksum-covered canonical-manifest JSON + entries cross-checked fieldwise with the CSV, and that + `arc_question_audit.csv` cannot be its trust anchor. + + Emit one `AuthorizationSpec` bundle containing the six condition-key grants, + not six unrelated self-authorizing records. Assert its aggregate bundle digest + is reproduced by all six condition-scoped validated slices. + +- [ ] **Step 4: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_stage1_profile.py -v + ``` + + Expected: FAIL because the profile modules are absent. + +- [ ] **Step 5: Implement manifest/checksum/queue translation** + + Independently hash every opened authoritative file against + `checksums.sha256` when it has a ledger entry. Root identity/status/source + authority in the checksum-covered canonical manifest and independently + computed canonical/source-artifact bytes. Treat `artifact_inventory.csv`, + `frozen_artifact_index.csv`, and reports as corroborating cross-reference + evidence when the ledger does not cover them; record their independently + computed audit digests but never promote them to a stronger trust anchor. Use + exact method/header-signature mapping tables in the + profile; do not inspect filenames for identity. Preserve original absolute + paths only as audit metadata and freeze-relative paths as stable provenance. + +- [ ] **Step 6: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_stage1_profile.py \ + tests/importing/test_authorization.py tests/importing/test_schema.py -v + python -m pytest -q + ``` + + Expected: profile translation, exclusions, queue separation, offline authority, + and paper-agnostic core checks pass. + +- [ ] **Step 7: Commit** + + ```bash + git add src/choicebench/importing/profiles/__init__.py \ + src/choicebench/importing/profiles/stage1_paper_freeze.py \ + tests/importing/conftest.py tests/importing/test_stage1_profile.py + git commit -m "feat: translate Stage 1 import manifests" + ``` + +--- + +### Task 16: Implement the Stage 1 expected-dataset trust chain + +**Files:** + +- Modify: `src/choicebench/importing/profiles/stage1_paper_freeze.py` +- Extend: `tests/importing/test_stage1_profile.py` + +**Interfaces:** + +```python +def build_stage1_expected_datasets( + freeze_root: Path, + checksum_ledger: Mapping[str, str], +) -> tuple[ExpectedDataset, ExpectedDataset, Mapping[str, Any]]: ... +``` + +- [ ] **Step 1: Extend the synthetic freeze with input/split artifacts** + + Create small synthetic ARC raw/normalized plus selection/metadata and MMLU + raw/normalized plus selection/metadata files under these exact Stage 1 + relative paths: + + ```text + raw/local_model_generalization/data/raw/arc_challenge_raw.csv + raw/local_model_generalization/data/processed/arc_challenge_normalized.csv + raw/local_model_generalization/data/splits/arc_challenge/robustness_ids.json + raw/local_model_generalization/data/splits/arc_challenge/robustness_metadata.json + raw/local_model_generalization/data/raw/mmlu_raw.csv + raw/local_model_generalization/data/processed/mmlu_normalized.csv + raw/local_model_generalization/data/splits/benchmark/robustness_ids.json + raw/local_model_generalization/data/splits/benchmark/robustness_metadata.json + ``` + + ARC includes one three-option row. MMLU includes the exact three duplicate IDs + from the approved spec, each occurring twice with identical parsed fields. + Make selection order deliberately differ from normalized source order so the + source-order keep-first rule and final selection-order projection are tested + separately. Add every byte checksum to the synthetic ledger. + +- [ ] **Step 2: Write failing ARC/MMLU trust-chain tests** + + Assert `reference_kind=independent_input_snapshot` and + `trust_label=checksum_verified_freeze_internal`; raw/normalized/selection/ + metadata checksum verification; result independence; selection membership; + stable snapshot order; variable ARC options; unknown upstream Hugging Face + revision/authenticity limitation; and byte-identical copies/result agreement + as corroboration only. + + For MMLU assert source-order, field-identical duplicate pairs, exact duplicate + row digests, keep-first semantics equivalent to + `drop_duplicates(subset="question_id", keep="first")`, and identity-bound + pre/post ordered digests. Mutate the second occurrence of each pair in turn and + require fail-closed conflict. A different duplicate policy must change the + derivation/realization identity. + +- [ ] **Step 3: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_stage1_profile.py \ + -k "dataset or duplicate or option or trust" -v + ``` + + Expected: new trust-chain assertions fail. + +- [ ] **Step 4: Implement the frozen-input derivation** + + Verify the ledger against the exact opened bytes. Use the generic CSV adapter + and dataset-reference builder. In registered profile code, deterministically + derive the comparable semantic fields and question IDs from the raw rows and + compare them fieldwise with the checksum-bound normalized rows; report this as + a ChoiceBench revalidation, not proof that the archived normalizer was the + producer. Reproduce the archived MMLU first-occurrence rule, but first require + every duplicate occurrence to agree on all parsed fields. Record the archived + script path/checksum as provenance evidence and the ChoiceBench implementation + identity as the active verification/transformation identity. Never execute + the archived script or any other freeze content. Preserve the limitation that + this proves internal freeze consistency, not upstream publisher authenticity. + +- [ ] **Step 5: Run focused, profile, and full tests** + + ```bash + python -m pytest tests/importing/test_stage1_profile.py \ + tests/importing/test_dataset_reference.py \ + tests/importing/test_csv_adapter.py -v + python -m pytest -q + ``` + + Expected: trust labels, three-option ARC, exact duplicate derivation, and + profile translation pass; no baseline regression. + +- [ ] **Step 6: Commit** + + ```bash + git add src/choicebench/importing/profiles/stage1_paper_freeze.py \ + tests/importing/test_stage1_profile.py + git commit -m "feat: verify Stage 1 dataset trust chain" + ``` + +**Intermediate gate:** Independent Stage 1 review against the immutable freeze. +The reviewer independently checks ledger coverage, selected-ID counts, all three +MMLU duplicate pairs, variable ARC options, 100/2 scope split, status counts, +138/687/3 queue counts, and six-cell/18-row offline authority. Review is dry-run +and read-only. + +--- + +### Task 17: Add the installed CLI, reports, and security regression suite + +**Files:** + +- Create: `src/choicebench/cli/import_results.py` +- Modify: `pyproject.toml` +- Create: `tests/importing/test_cli.py` +- Modify: `tests/test_release_adversarial.py` + +**Interfaces:** + +```python +def parse_args(argv: Sequence[str] | None = None) -> argparse.Namespace: ... +def write_import_report(path: Path, report: Mapping[str, Any]) -> None: ... +def main(argv: Sequence[str] | None = None) -> int: ... +``` + +At module scope import only stdlib modules and `choicebench` package metadata; +do not import `choicebench.config.paths`, the engine, profiles, readers, or any +provider. `main()` parses and validates `--output-root`, sets +`CHOICEBENCH_HOME`, and only then imports workspace-dependent modules. + +- [ ] **Step 1: Write failing subprocess CLI tests** + + Cover generic spec import; `--dry-run` and `--validate-only` equivalence; + `--strict`; absolute `--output-root`; `CHOICEBENCH_HOME` fallback; current + directory fallback; required safe run ID; profile dispatch; repeated + `--overlay`; human summary; machine report; exact no-op; divergent refusal; + sanitized errors; and help text. Run output-root tests in fresh subprocesses + with `PYTHONPATH` controlled so module caching cannot hide import order. + + Machine report assertions use stable fields: + + ```python + assert report["counts"]["scope_disposition"]["included"] >= 1 + assert report["counts"]["evidence_status"]["complete"] >= 1 + assert report["wrote_artifacts"] is expected_write + assert report["idempotent_noop"] is expected_noop + assert report["failures"] == [] + assert report["audit"]["source_location"] == str(source.resolve()) + assert "audit" not in report["identity_projection"] + ``` + +- [ ] **Step 2: Extend adversarial/security tests before implementation** + + Add YAML executable-tag refusal; credential metadata; source traversal; + specification-controlled output; unsafe source/output symlinks; recursive + directory; source change after validation; malicious CSV error text; + malformed byte sequences; result/evidence/manifest/state tampering; divergent + report overwrite; overlay self-authorization; and response-cache forest + refusal. Explicitly accept user-selected absolute source and output roots. + +- [ ] **Step 3: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_cli.py \ + tests/test_release_adversarial.py -v + ``` + + Expected: CLI tests fail because the module/entry point is absent. + +- [ ] **Step 4: Implement bootstrap parsing and command dispatch** + + Positional input is a generic spec path unless `--profile` is present, in + which case it is the profile root. Required options/aliases are: + + ```text + choicebench-import-results INPUT --run-id ID + [--profile stage1-paper-freeze] + [--dry-run | --validate-only] + [--strict] + [--output-root PATH] + [--overlay PATH ...] + [--report PATH] + ``` + + Output-root selection comes only from CLI/environment. Canonicalize it, refuse + unsafe symlinks, set the environment, then lazily import and execute. Reports + are atomic; an identical existing report is a no-op and divergent content is + refused. Audit locations/timestamps remain outside report identity. + +- [ ] **Step 5: Register the command without changing dependencies/version** + + Add exactly: + + ```toml + choicebench-import-results = "choicebench.cli.import_results:main" + ``` + + Do not change `version = "0.2.0"` or package data. + +- [ ] **Step 6: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_cli.py \ + tests/test_release_adversarial.py -v + python -m pytest -q + ``` + + Expected: CLI/environment/security tests pass and the full baseline is green. + +- [ ] **Step 7: Commit** + + ```bash + git add src/choicebench/cli/import_results.py pyproject.toml \ + tests/importing/test_cli.py tests/test_release_adversarial.py + git commit -m "feat: add external results import command" + ``` + +**Intermediate gate:** Independent adversarial review of CLI import order, path +trust boundaries, YAML/CSV handling, report collision behavior, secret rejection, +and absence of inference execution. + +--- + +### Task 18: Document and package the importer + +**Files:** + +- Create: `docs/external-results-import.md` +- Modify: `README.md` +- Modify: `CHANGELOG.md` +- Modify: `tests/test_wheel_smoke.py` + +**Interfaces:** No new runtime interface. Documentation examples use only the +installed `choicebench-import-results` and `choicebench-evaluate` commands. + +- [ ] **Step 1: Extend the installed wheel/sdist smoke test first** + + In each installed environment, assert importer modules and console entry point + exist; run `--help`; create a two-question generic CSV/spec outside the + repository; dry-run; real import; exact repeated no-op; and evaluation. Assert + the installed wheel/sdist contains no run, report, historical, cache, or test + fixture data. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/test_wheel_smoke.py -v + ``` + + Expected: PASS because Tasks 2-17 already supplied and registered the runtime + functionality. Treat this as a packaging characterization gate before the + documentation-only changes; if it fails, diagnose and fix the responsible + earlier runtime/packaging task in a focused commit rather than hiding the + problem in documentation. + +- [ ] **Step 3: Write complete user/adapter documentation** + + `docs/external-results-import.md` must cover: external versus native inference; + supported UTF-8 CSV/dialect fields; every import-spec section; generic and + dry-run examples; `CHOICEBENCH_HOME`/absolute output; checksums and four + identity layers; orthogonal statuses; expected-dataset trust; method-specific + extension data; mixed prediction origins; immutable new-run overlays; typed + inference/offline authorization; evaluation selection/accounting; security + boundary; limitations; future adapter protocol; and a complete two-question + example. + + Link the guide from README. Add an `Unreleased` changelog entry consistent + with the existing newest-first format. State explicitly that no historical + data or model inference is included. Do not change the package version. + +- [ ] **Step 4: Run documentation-facing and full tests** + + ```bash + python -m pytest tests/test_wheel_smoke.py tests/importing/test_cli.py -v + python -m pytest -q + ``` + + Expected: wheel and sdist workflows work outside the repository; full suite + passes. + +- [ ] **Step 5: Commit** + + ```bash + git add docs/external-results-import.md README.md CHANGELOG.md \ + tests/test_wheel_smoke.py + git commit -m "docs: document and package external result imports" + ``` + +--- + +### Task 19: Full validation, independent reviews, push, and unmerged PR + +**Files:** + +- Modify only files required to resolve confirmed review findings. +- Create outside Git: Stage 1 JSON reports, temporary workspaces/builds, and PR + body. + +- [ ] **Step 1: Review the complete diff and repository cleanliness** + + ```bash + git status --short + git diff --check origin/choicebench...HEAD + git diff --stat origin/choicebench...HEAD + git diff --name-only origin/choicebench...HEAD + ``` + + Expected: only source, tests, documentation, and small synthetic test support; + no imported runs, reports, historical data, caches, credentials, build output, + or temporary files. + +- [ ] **Step 2: Run all focused importer/compatibility/security tests** + + ```bash + python -m pytest tests/importing tests/io/test_readers.py \ + tests/io/test_writers.py tests/infra/test_checkpoint.py \ + tests/test_condition_grid.py tests/test_publication_identity.py \ + tests/test_release_adversarial.py tests/scripts/test_evaluate_run.py \ + tests/scripts/test_reset_run.py -v + ``` + + Expected: all focused tests pass with zero failures. + +- [ ] **Step 3: Run the full ChoiceBench suite** + + ```bash + python -m pytest tests/ -v + ``` + + Expected: the existing 808 tests plus all new tests pass; record the exact + final count and duration for the PR. + +- [ ] **Step 4: Confirm the repository has no configured lint/type gate** + + ```bash + rg -n "ruff|mypy|pyright|black|flake8|pylint" pyproject.toml \ + .github/workflows tests + ``` + + Expected: no configured command. Do not invent a new formatter/type-check gate + in this PR; report that pytest is the repository's configured CI check. + +- [ ] **Step 5: Build and check wheel/sdist outside the repository** + + ```bash + IMPORT_BUILD_DIST="$(mktemp -d)" + python -m build --outdir "$IMPORT_BUILD_DIST" + python -m twine check "$IMPORT_BUILD_DIST"/* + python -m pytest tests/test_wheel_smoke.py -v + ``` + + Expected: build exits 0; Twine reports both distributions `PASSED`; installed + wheel and sdist generic import/evaluation smoke passes outside the repository. + +- [ ] **Step 6: Record a read-only freeze inventory before Stage 1 validation** + + ```bash + FREEZE_ROOT=/home/cotenthusiast/Projects/model-generalization/paper_data_freeze + FREEZE_METADATA_BEFORE="$(mktemp)" + FREEZE_CONTENT_BEFORE="$(mktemp)" + find "$FREEZE_ROOT" -printf '%y\t%P\t%m\t%U\t%G\t%s\t%T@\t%l\n' \ + | sort > "$FREEZE_METADATA_BEFORE" + find "$FREEZE_ROOT" -type f -print0 | sort -z | xargs -0 sha256sum \ + > "$FREEZE_CONTENT_BEFORE" + ``` + + Expected: metadata, symlink targets, and content digests for the entire freeze + are written outside both repositories. Do not execute any freeze script and + do not change file permissions or metadata. + +- [ ] **Step 7: Run the full Stage 1 strict dry run** + + ```bash + STAGE1_DRY_REPORT_DIR="$(mktemp -d)" + STAGE1_DRY_REPORT="$STAGE1_DRY_REPORT_DIR/stage1-dry-run.json" + choicebench-import-results "$FREEZE_ROOT" \ + --profile stage1-paper-freeze \ + --run-id paper-stage1-validation \ + --dry-run --strict \ + --report "$STAGE1_DRY_REPORT" + ``` + + Validate the report with: + + ```bash + jq -e ' + .profile.matrix.intended_cells == 100 and + .profile.matrix.core_method_cells == 84 and + .profile.matrix.local_pride_cells == 4 and + .profile.matrix.ihs_cells == 12 and + .profile.matrix.excluded_preserved_cells == 2 and + .counts.evidence_status.complete == 58 and + .counts.evidence_status.qualified == 5 and + .counts.evidence_status.recoverable == 6 and + .counts.evidence_status.partial == 4 and + .counts.evidence_status.malformed == 29 and + .profile.matrix.intended_evidence.complete == 57 and + .profile.matrix.intended_evidence.qualified == 5 and + .profile.matrix.intended_evidence.recoverable == 6 and + .profile.matrix.intended_evidence.partial == 4 and + .profile.matrix.intended_evidence.malformed == 28 and + .profile.queues.approved_executable_question_cells == 138 and + .profile.queues.held_nonexecutable_question_cells == 687 and + .profile.queues.excluded_nonexecutable_question_cells == 3 and + .profile.queues.forensic_question_cells == 776 and + .profile.queues.classified_question_cells == 828 and + .profile.queues.pairwise_disjoint == true and + .profile.queues.approved_authoritative_executable == true and + .profile.queues.held_executable == false and + .profile.queues.excluded_executable == false and + .profile.source_schema_count == 8 and + .profile.offline_authority.condition_count == 6 and + .profile.offline_authority.question_cell_count == 18 and + .profile.offline_authority.inference_executable == false and + .profile.datasets.arc.reference_kind == "independent_input_snapshot" and + .profile.datasets.arc.trust_label == "checksum_verified_freeze_internal" and + .profile.datasets.arc.selected_question_count == 1000 and + .profile.datasets.arc.option_count_distribution["3"] == 3 and + .profile.datasets.arc.option_count_distribution["4"] == 997 and + .profile.datasets.mmlu.reference_kind == "independent_input_snapshot" and + .profile.datasets.mmlu.trust_label == "checksum_verified_freeze_internal" and + .profile.datasets.mmlu.selected_question_count == 1000 and + .profile.datasets.mmlu.pre_dedup_row_count == 1003 and + .profile.datasets.mmlu.post_dedup_row_count == 1000 and + .profile.datasets.mmlu.duplicate_question_ids == + ["79686d32dfe155ea", "2f7aa3c7ebb98cfe", "74f7227e190200ac"] and + (.profile.datasets.mmlu.pre_dedup_digest | length) == 64 and + (.profile.datasets.mmlu.post_dedup_digest | length) == 64 and + .profile.datasets.arc.publisher_authenticated == false and + .profile.datasets.mmlu.publisher_authenticated == false and + .checksums.all_referenced_sources_match == true + ' "$STAGE1_DRY_REPORT" + ``` + + Expected: importer reports `validated`, writes no run, and every assertion is + true. Held/excluded records remain non-executable. The profile performs the + same sealed invariants through registered ChoiceBench code; it does not run + `scripts/validate_stage1_freeze.py` or any other external executable content. + +- [ ] **Step 8: Perform an isolated real Stage 1 import and evaluation** + + ```bash + STAGE1_IMPORT_ROOT="$(mktemp -d)" + STAGE1_IMPORT_REPORT_DIR="$(mktemp -d)" + STAGE1_IMPORT_REPORT="$STAGE1_IMPORT_REPORT_DIR/stage1-import.json" + choicebench-import-results "$FREEZE_ROOT" \ + --profile stage1-paper-freeze \ + --run-id paper-stage1-import \ + --strict \ + --output-root "$STAGE1_IMPORT_ROOT" \ + --report "$STAGE1_IMPORT_REPORT" + CHOICEBENCH_HOME="$STAGE1_IMPORT_ROOT" \ + choicebench-evaluate --run-id paper-stage1-import + STAGE1_EVAL_REPORT="$(find "$STAGE1_IMPORT_ROOT/reports" -maxdepth 1 \ + -type f -name 'paper-stage1-import_*_metrics.json' -print -quit)" + jq -e ' + any(.realizations[]; + .selected == true and .evidence_status == "complete" and + .scope_disposition == "included" and (.metrics | length) > 0 and + (.prediction_origins | all(. == "external_historical_inference"))) and + any(.realizations[]; + .selected == true and .evidence_status == "qualified" and + .scope_disposition == "included" and (.qualifications | length) > 0 and + (.metrics | length) > 0) and + any(.realizations[]; + (.evidence_status == "partial" or .evidence_status == "malformed") and + (.metrics | length) == 0) and + any(.realizations[]; + .benchmark_name == "arc_challenge" and .selected == true and + .option_count_distribution["3"] == 3 and (.metrics | length) > 0) and + all(.realizations[]; + if (.scope_disposition == "excluded_from_paper_matrix" or + .scope_disposition == "held" or + .scope_disposition == "superseded" or + .evidence_status == "failed") + then (.metrics | length) == 0 else true end) + ' "$STAGE1_EVAL_REPORT" + ``` + + Expected: a complete immutable run exists only below the temporary root; + included complete and qualified realizations have metrics; incomplete/ + malformed/recoverable/excluded records are accounted without metrics; at + least one evaluated ARC realization contains all three variable-option rows; + origins remain external; no inference is invoked. Record experiment, + realization, result, and evaluation IDs and the evaluation report path. + +- [ ] **Step 9: Verify the freeze stayed unchanged** + + ```bash + FREEZE_METADATA_AFTER="$(mktemp)" + FREEZE_CONTENT_AFTER="$(mktemp)" + find "$FREEZE_ROOT" -printf '%y\t%P\t%m\t%U\t%G\t%s\t%T@\t%l\n' \ + | sort > "$FREEZE_METADATA_AFTER" + find "$FREEZE_ROOT" -type f -print0 | sort -z | xargs -0 sha256sum \ + > "$FREEZE_CONTENT_AFTER" + cmp "$FREEZE_METADATA_BEFORE" "$FREEZE_METADATA_AFTER" + cmp "$FREEZE_CONTENT_BEFORE" "$FREEZE_CONTENT_AFTER" + ``` + + Expected: both `cmp` commands exit 0. The import root and reports remain + outside Git. + +- [ ] **Step 10: Run independent specification-compliance review** + + Give a fresh read-only reviewer the approved specification, implementation + plan, full diff, and validation evidence. Require a line-by-line verdict on + generic core, v2/v3 compatibility, provenance/identity, statuses, overlays, + evaluation, Stage 1 profile, CLI, tests, docs, and boundaries. Resolve all + confirmed Critical/Important findings in focused commits; rerun affected tests; + return current HEAD/evidence to that reviewer until it gives an explicit + approval on the current commit. + +- [ ] **Step 11: Run independent code-quality review** + + A different fresh reviewer inspects cohesion, API/type consistency, error + clarity, duplication, maintenance cost, performance/memory on Stage 1, and + paper leakage. Resolve confirmed blockers, rerun focused/full tests, and obtain + that reviewer's explicit approval of the resulting current commit. + +- [ ] **Step 12: Run independent adversarial/security review** + + A different reviewer attacks YAML/CSV parsing, byte spans, checksum races, + paths/symlinks, staging/no-replace, idempotence, secret handling, manifest/ + sidecar/state tampering, authorization escalation, base-run mutation, and + malformed evidence normalization. Resolve blockers and rerun security plus + full tests, then obtain that reviewer's explicit approval of current HEAD. + +- [ ] **Step 13: Run fresh-context final review** + + Give a final reviewer only the user requirements, approved spec/plan, final + diff, test/build/Twine/Stage 1 evidence, and earlier resolved-finding commits. + Require a clear PR-ready/not-ready verdict. Do not proceed on Critical or + Important findings. If it finds any, resolve them in focused commits, repeat + affected/full verification, and send the new current HEAD back to the same + fresh-context reviewer for a renewed verdict. Repeat until it explicitly says + PR-ready for the exact commit that will be pushed. + +- [ ] **Step 14: Repeat final verification after all review fixes** + + Confirm the specification, quality, security, and final reviewers all approved + the current HEAD, then repeat Steps 1–9. If this verification causes any code, + test, or documentation change, invalidate all four approvals and repeat Steps + 10–13 before proceeding. Expected: approvals and validation evidence refer to + the exact same commit. + +- [ ] **Step 15: Reconfirm remote target and branch ancestry** + + ```bash + git fetch origin + git remote show origin + git merge-base --is-ancestor origin/choicebench HEAD + git status --short --branch + ``` + + Expected: worktree clean; feature contains the synchronized integration base. + Confirm the repository's actual integration/default target from remote state + rather than assuming its name. If it differs from the branch used to start + this work or ancestry is false, stop and report instead of rebasing/retargeting + silently. + +- [ ] **Step 16: Push the feature branch** + + ```bash + git push -u origin feat/external-results-importer + ``` + + Expected: GitHub confirms the remote feature branch. Do not force-push. + +- [ ] **Step 17: Create the unmerged pull request** + + Create `/tmp/choicebench-external-results-pr.md` with `apply_patch`, containing + motivation; architecture; imported-versus-native semantics; generic importer; + Stage 1 profile; overlay/authorization design; security; tests and exact final + count; Stage 1 dry-run and real-import evidence; build/Twine; limitations; + explicit no historical data/inference statement; base/feature branch; and + exact commit hashes. Then run: + + ```bash + gh pr create \ + --base \ + --head feat/external-results-importer \ + --title "feat: add provenance-preserving external results importer" \ + --body-file /tmp/choicebench-external-results-pr.md + ``` + + Expected: GitHub returns a PR number and URL. Verify with `gh pr view`. Leave + it open; do not merge, tag, or release. If authentication prevents creation, + preserve the pushed branch and report the exact GitHub CLI error without + claiming a PR exists. + +--- + +## Plan completion condition + +Implementation is complete only when every task/checkbox is satisfied, all +independent reviews pass, the full validation evidence is current, the feature +branch is pushed, and GitHub confirms an open unmerged PR. Generated runs, +reports, historical data, caches, temporary build/workspaces, and the external +PR-body file remain intentionally excluded from Git. From fe5d90f5300f07d5f7d1be0700b5b48807b2a2bd Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sat, 18 Jul 2026 22:46:40 +0300 Subject: [PATCH 04/47] test: freeze manifest v2 compatibility --- tests/importing/__init__.py | 0 tests/importing/conftest.py | 212 +++++++++++++++++++++ tests/importing/test_manifest_v2_compat.py | 43 +++++ 3 files changed, 255 insertions(+) create mode 100644 tests/importing/__init__.py create mode 100644 tests/importing/conftest.py create mode 100644 tests/importing/test_manifest_v2_compat.py diff --git a/tests/importing/__init__.py b/tests/importing/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/tests/importing/conftest.py b/tests/importing/conftest.py new file mode 100644 index 0000000..4eb20e4 --- /dev/null +++ b/tests/importing/conftest.py @@ -0,0 +1,212 @@ +from pathlib import Path + +import pandas as pd +import pytest + +from choicebench.cli import evaluate_run +from choicebench.datasets import dataset_content_digest +from choicebench.identity import integrity_digest +from choicebench.infra.artifacts import atomic_write_text +from choicebench.io.readers import read_manifest_results +from choicebench.io.writers import validate_result_artifact, write_run_results +from choicebench.manifest import ( + build_manifest_payload, + ensure_manifest, + initial_run_state, + make_manifest, + write_run_state, +) +from choicebench.metrics import BUILTIN_METRICS +from choicebench.provenance import implementation_identity + + +EXPECTED_EXPERIMENT_ID = "exp_d1b8b8ddecd1f16f" +EXPECTED_EVALUATION_ID = "eval_a1be80926996d03a" + + +@pytest.fixture +def synthetic_v2_run(tmp_path, monkeypatch) -> tuple[Path, dict, str]: + questions = pd.DataFrame( + [ + { + "question_id": "q1", + "subject": "arithmetic", + "question_text": "What is 2 + 2?", + "choice_a": "3", + "choice_b": "4", + "choice_c": "5", + "choice_d": "6", + "correct_option": "B", + "correct_answer_text": "4", + }, + { + "question_id": "q2", + "subject": "geography", + "question_text": "What is the capital of France?", + "choice_a": "Berlin", + "choice_b": "Madrid", + "choice_c": "Paris", + "choice_d": "Rome", + "correct_option": "C", + "correct_answer_text": "Paris", + }, + ] + ) + prompt_content = "Question: {question_text}\nChoices:\n{choices}\nAnswer:" + selection_id = "sel_legacyfixture" + artifact_id = "dataset_legacyfixture" + prompt_id = "prompt_legacyfixture" + model_id = "model_legacyfixture" + method_id = "method_legacyfixture" + condition_id = "cond_legacyfixture" + + dataset_record = { + "selection_id": selection_id, + "artifact_id": artifact_id, + "benchmark": "toy", + "split": "test", + "prepared_path": "toy/test.csv", + "prepared_content_digest": "prepared-legacy-fixture", + "prepared_metadata": {"schema_version": "choicebench.dataset-artifact.v1"}, + "selected_content_digest": dataset_content_digest(questions), + "selected_question_ids": ["q1", "q2"], + "selected_sample_identities": ["sample_q1", "sample_q2"], + "run_snapshot_path": f"artifacts/datasets/{selection_id}.csv", + "row_count": 2, + } + prompt_record = { + "prompt_id": prompt_id, + "version": "legacy-v1", + "files": { + "direct_mcq": { + "content": prompt_content, + "sha256": integrity_digest(prompt_content), + } + }, + "run_snapshot_path": f"artifacts/prompts/{prompt_id}", + } + model_record = { + "model_id": model_id, + "config": {"backend": "dummy", "model_name_or_path": "dummy-model"}, + "resolved_model": None, + } + method_record = { + "method_id": method_id, + "config": {"name": "direct_mcq", "params": {}, "preflight": None}, + "implementation": {"qualified_name": "legacy.fixture:DirectMCQ"}, + } + condition_record = { + "condition_id": condition_id, + "benchmark_name": "toy", + "split": "test", + "selection_id": selection_id, + "artifact_id": artifact_id, + "model_id": model_id, + "method_id": method_id, + "prompt_id": prompt_id, + "prompt_snapshot_path": prompt_record["run_snapshot_path"], + "result_path": f"results/{condition_id}.csv", + "checkpoint_path": f"checkpoints/{condition_id}.json", + "result_metadata_path": f"results/{condition_id}.artifact.json", + "model_display_name": "dummy-model", + "resolved_model": None, + "identity": {"fixture": "native-v2"}, + } + source = { + "git_commit": "0123456789abcdef0123456789abcdef01234567", + "dirty": False, + "dirty_tracked_digest": None, + "source_tree_digest": "source-tree-legacy-fixture", + } + environment = { + "choicebench_version": "0.2.0", + "python": "3.12.0", + "implementation": "CPython", + "dependencies": {"pandas": "2.2.0", "pytest": "8.3.0"}, + } + payload = build_manifest_payload( + config={"run": {"seed": 7}, "metrics": ["accuracy"]}, + datasets=[dataset_record], + prompts=prompt_record, + models=[model_record], + methods=[method_record], + conditions=[condition_record], + source=source, + environment=environment, + ) + payload["evaluation"] = [ + { + "name": "accuracy", + "implementation": implementation_identity(BUILTIN_METRICS["accuracy"]), + } + ] + manifest = make_manifest(payload) + assert manifest["experiment_id"] == EXPECTED_EXPERIMENT_ID + + run_dir = tmp_path / "native-v2-fixture" + ensure_manifest(run_dir, manifest) + atomic_write_text( + run_dir / dataset_record["run_snapshot_path"], + questions.to_csv(index=False), + ) + atomic_write_text( + run_dir + / prompt_record["run_snapshot_path"] + / prompt_record["version"] + / "direct_mcq.txt", + prompt_content, + ) + + rows = [] + for question, parsed_choice in zip( + questions.to_dict("records"), ("B", "A"), strict=True + ): + rows.append( + { + **question, + "raw_text": f"The answer is {parsed_choice}", + "parsed_choice": parsed_choice, + "parse_status": "parse_ok", + "normalized_text": parsed_choice, + "parse_reason": "Answer successfully parsed", + "score_status": "scored", + "is_correct": parsed_choice == question["correct_option"], + "transport_status": "ok", + "experiment_id": manifest["experiment_id"], + "condition_id": condition_id, + "dataset_artifact_id": artifact_id, + "dataset_selection_id": selection_id, + "model_id": model_id, + "method_id": method_id, + "prompt_id": prompt_id, + "benchmark_split": "test", + "model_name": "dummy-model", + "method_name": "direct_mcq", + } + ) + result_path = write_run_results( + rows, + run_dir, + run_dir.name, + "direct_mcq", + "dummy-model", + "toy", + condition_id, + ) + artifact = validate_result_artifact(result_path) + state = initial_run_state(manifest) + state["conditions"][condition_id] = { + "status": "completed", + "result_path": condition_record["result_path"], + "result_sha256": artifact["file_sha256"], + } + write_run_state(run_dir, state) + + monkeypatch.setattr(evaluate_run, "RUNS_DIR", run_dir.parent) + frame, loaded = read_manifest_results(run_dir) + report = evaluate_run.build_evaluation_report( + run_dir.name, frame, loaded, reparse=False + ) + assert report["evaluation_id"] == EXPECTED_EVALUATION_ID + + return run_dir, manifest, EXPECTED_EVALUATION_ID diff --git a/tests/importing/test_manifest_v2_compat.py b/tests/importing/test_manifest_v2_compat.py new file mode 100644 index 0000000..951b46d --- /dev/null +++ b/tests/importing/test_manifest_v2_compat.py @@ -0,0 +1,43 @@ +from choicebench.cli import evaluate_run +from choicebench.io.readers import read_manifest_results +from choicebench.manifest import validate_manifest + + +def test_v2_fixture_validates_without_rewrite(synthetic_v2_run): + run_dir, manifest, _ = synthetic_v2_run + before = { + p.relative_to(run_dir): p.read_bytes() + for p in run_dir.rglob("*") + if p.is_file() + } + validate_manifest(manifest) + frame, loaded = read_manifest_results(run_dir) + after = { + p.relative_to(run_dir): p.read_bytes() + for p in run_dir.rglob("*") + if p.is_file() + } + assert loaded["schema_version"] == "choicebench.manifest.v2" + assert frame["question_id"].astype(str).tolist() == ["q1", "q2"] + assert after == before + + +def test_v2_evaluation_identity_and_shape_are_frozen( + synthetic_v2_run, monkeypatch +): + run_dir, manifest, expected_evaluation_id = synthetic_v2_run + monkeypatch.setattr(evaluate_run, "RUNS_DIR", run_dir.parent) + frame, _ = read_manifest_results(run_dir) + report = evaluate_run.build_evaluation_report( + run_dir.name, frame, manifest, reparse=False + ) + assert report["schema_version"] == "choicebench.evaluation.v1" + assert report["evaluation_id"] == expected_evaluation_id + assert set(report["conditions"]) == {"cond_legacyfixture"} + + +def test_v2_missing_result_origin_implies_native_only_in_compat_view( + synthetic_v2_run, +): + _, manifest, _ = synthetic_v2_run + assert "result_origin" not in manifest["payload"]["conditions"][0] From 9d98ce1eb67055cc5860935c3649bc1e0a3c83ec Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sat, 18 Jul 2026 22:55:43 +0300 Subject: [PATCH 05/47] test: exercise first manifest v2 reader call --- tests/importing/conftest.py | 4 ++-- tests/importing/test_manifest_v2_compat.py | 26 ++++++++++++++++++++-- 2 files changed, 26 insertions(+), 4 deletions(-) diff --git a/tests/importing/conftest.py b/tests/importing/conftest.py index 4eb20e4..83eaa8b 100644 --- a/tests/importing/conftest.py +++ b/tests/importing/conftest.py @@ -203,9 +203,9 @@ def synthetic_v2_run(tmp_path, monkeypatch) -> tuple[Path, dict, str]: write_run_state(run_dir, state) monkeypatch.setattr(evaluate_run, "RUNS_DIR", run_dir.parent) - frame, loaded = read_manifest_results(run_dir) + frame = pd.read_csv(result_path) report = evaluate_run.build_evaluation_report( - run_dir.name, frame, loaded, reparse=False + run_dir.name, frame, manifest, reparse=False ) assert report["evaluation_id"] == EXPECTED_EVALUATION_ID diff --git a/tests/importing/test_manifest_v2_compat.py b/tests/importing/test_manifest_v2_compat.py index 951b46d..9b061eb 100644 --- a/tests/importing/test_manifest_v2_compat.py +++ b/tests/importing/test_manifest_v2_compat.py @@ -1,17 +1,38 @@ +import pytest + from choicebench.cli import evaluate_run from choicebench.io.readers import read_manifest_results from choicebench.manifest import validate_manifest +from tests.importing import conftest as importing_conftest + +@pytest.fixture +def publication_reader_calls(monkeypatch): + calls = [] + original = importing_conftest.read_manifest_results -def test_v2_fixture_validates_without_rewrite(synthetic_v2_run): + def tracked_read_manifest_results(run_dir): + calls.append(run_dir) + return original(run_dir) + + monkeypatch.setattr( + importing_conftest, "read_manifest_results", tracked_read_manifest_results + ) + return calls + + +def test_v2_fixture_validates_without_rewrite( + publication_reader_calls, synthetic_v2_run +): run_dir, manifest, _ = synthetic_v2_run + assert publication_reader_calls == [] before = { p.relative_to(run_dir): p.read_bytes() for p in run_dir.rglob("*") if p.is_file() } validate_manifest(manifest) - frame, loaded = read_manifest_results(run_dir) + frame, loaded = importing_conftest.read_manifest_results(run_dir) after = { p.relative_to(run_dir): p.read_bytes() for p in run_dir.rglob("*") @@ -19,6 +40,7 @@ def test_v2_fixture_validates_without_rewrite(synthetic_v2_run): } assert loaded["schema_version"] == "choicebench.manifest.v2" assert frame["question_id"].astype(str).tolist() == ["q1", "q2"] + assert publication_reader_calls == [run_dir] assert after == before From 9d51ed620eeebac081b0836c59c883793fc1faf2 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sat, 18 Jul 2026 23:22:14 +0300 Subject: [PATCH 06/47] test: align manifest v2 IDs with package version --- tests/importing/conftest.py | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/tests/importing/conftest.py b/tests/importing/conftest.py index 83eaa8b..93025e8 100644 --- a/tests/importing/conftest.py +++ b/tests/importing/conftest.py @@ -20,8 +20,8 @@ from choicebench.provenance import implementation_identity -EXPECTED_EXPERIMENT_ID = "exp_d1b8b8ddecd1f16f" -EXPECTED_EVALUATION_ID = "eval_a1be80926996d03a" +EXPECTED_EXPERIMENT_ID = "exp_cd53979b257867df" +EXPECTED_EVALUATION_ID = "eval_93b022690174ec36" @pytest.fixture From 582dac55c7a0f4761a38ca27addfd49b955335ec Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sat, 18 Jul 2026 23:25:24 +0300 Subject: [PATCH 07/47] feat: add strict external import specification --- src/choicebench/importing/__init__.py | 51 + src/choicebench/importing/schema.py | 1501 +++++++++++++++++++++++++ tests/importing/test_schema.py | 923 +++++++++++++++ 3 files changed, 2475 insertions(+) create mode 100644 src/choicebench/importing/__init__.py create mode 100644 src/choicebench/importing/schema.py create mode 100644 tests/importing/test_schema.py diff --git a/src/choicebench/importing/__init__.py b/src/choicebench/importing/__init__.py new file mode 100644 index 0000000..378d594 --- /dev/null +++ b/src/choicebench/importing/__init__.py @@ -0,0 +1,51 @@ +"""Public types for strict external result import declarations.""" + +from choicebench.importing.schema import ( + AuthorizationSpec, + CsvDialectSpec, + DatasetReferenceSpec, + DerivationOrigin, + EvidenceStatus, + ImportConditionSpec, + ImportMethodSpec, + ImportModelSpec, + ImportPromptSpec, + ImportSpec, + ImportSpecError, + ImportState, + NumericColumnSpec, + OptionMappingSpec, + OverlaySpec, + PredictionOrigin, + ResultOriginSpec, + ScopeDisposition, + SourceArtifactSpec, + import_spec_digest, + load_import_spec, + stable_import_projection, +) + +__all__ = [ + "AuthorizationSpec", + "CsvDialectSpec", + "DatasetReferenceSpec", + "DerivationOrigin", + "EvidenceStatus", + "ImportConditionSpec", + "ImportMethodSpec", + "ImportModelSpec", + "ImportPromptSpec", + "ImportSpec", + "ImportSpecError", + "ImportState", + "NumericColumnSpec", + "OptionMappingSpec", + "OverlaySpec", + "PredictionOrigin", + "ResultOriginSpec", + "ScopeDisposition", + "SourceArtifactSpec", + "import_spec_digest", + "load_import_spec", + "stable_import_projection", +] diff --git a/src/choicebench/importing/schema.py b/src/choicebench/importing/schema.py new file mode 100644 index 0000000..eba63fb --- /dev/null +++ b/src/choicebench/importing/schema.py @@ -0,0 +1,1501 @@ +"""Strict, paper-agnostic declarations for importing external results.""" + +from __future__ import annotations + +from dataclasses import asdict, dataclass +import math +from numbers import Real +from pathlib import Path, PurePosixPath +import re +from typing import Any, Literal, Mapping, TypeAlias + +import yaml + +from choicebench.identity import canonicalize, integrity_digest, is_credential_key, short_id +from choicebench.metrics import BUILTIN_METRICS + + +ImportState: TypeAlias = Literal["validated", "imported", "failed"] +EvidenceStatus: TypeAlias = Literal[ + "complete", "qualified", "partial", "malformed", "recoverable", "failed" +] +ScopeDisposition: TypeAlias = Literal[ + "included", "excluded_from_paper_matrix", "held", "superseded" +] +PredictionOrigin: TypeAlias = Literal[ + "native_inference", + "external_historical_inference", + "external_repair_inference", +] +DerivationOrigin: TypeAlias = Literal[ + "native_execution", "external_import", "repair_overlay", "offline_transformation" +] + + +class ImportSpecError(ValueError): + """Raised when an external import declaration is unsafe or inconsistent.""" + + +@dataclass(frozen=True) +class CsvDialectSpec: + encoding: Literal["utf-8"] = "utf-8" + bom_policy: Literal["forbid", "strip_utf8_bom"] = "forbid" + decoding_errors: Literal["strict"] = "strict" + delimiter: str = "," + quote_character: str = '"' + escape_character: str | None = None + double_quote: bool = True + line_terminators: tuple[str, ...] = ("crlf", "lf", "cr") + mixed_line_terminators: Literal["allow", "forbid"] = "allow" + final_record_without_terminator: Literal["allow", "forbid"] = "allow" + blank_record_policy: Literal["reject"] = "reject" + skip_initial_space: bool = False + header: Literal["first_logical_record"] = "first_logical_record" + strict_syntax: bool = True + + +@dataclass(frozen=True) +class NumericColumnSpec: + source_column: str + value_type: Literal["integer", "float"] + null_allowed: bool + finite_only: bool = True + + +@dataclass(frozen=True) +class OptionMappingSpec: + mode: Literal["ordered_columns", "structured_json"] + ordered_columns: tuple[str, ...] + structured_column: str | None + structured_label_key: str | None + structured_text_key: str | None + + +@dataclass(frozen=True) +class SourceArtifactSpec: + source_id: str + path: Path + logical_path: str + expected_sha256: str + format: Literal["csv"] + format_version: str + classification: Literal["raw", "canonical", "derived", "repaired", "aggregate_only"] + dialect: CsvDialectSpec + columns: Mapping[str, str] + expected_columns: tuple[str, ...] + ignored_columns: Mapping[str, str] + null_values: tuple[str, ...] + numeric_columns: tuple[NumericColumnSpec, ...] + option_mapping: OptionMappingSpec + extra_field_policy: Literal["preserve_unmapped", "reject_unmapped"] + preserve_namespace: str + source_run_id: str | None + source_repository: str | None + source_commit: str | None + notes: Mapping[str, Any] + + +@dataclass(frozen=True) +class DatasetReferenceSpec: + dataset_id: str + benchmark_name: str + split: str + reference_kind: Literal[ + "independent_input_snapshot", "profile_derived_reference_snapshot" + ] + trust_label: str + source_ids: tuple[str, ...] + selection_source_id: str + expected_question_ids: tuple[str, ...] + selection_seed: int | None + selection_n_samples: int | None + subject_filter: tuple[str, ...] + selection_unknown_reasons: Mapping[str, str] + columns: Mapping[str, str] + revision: str | None + fingerprint: str | None + derivation: Mapping[str, Any] + limitations: tuple[str, ...] + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ImportModelSpec: + model_key: str + display_name: str + backend: str | None + provider: str | None + revision: str | None + effective_parameters: Mapping[str, Any] + unknown_reasons: Mapping[str, str] + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ImportMethodSpec: + method_key: str + name: str + effective_parameters: Mapping[str, Any] + implementation: Mapping[str, Any] | None + unknown_reasons: Mapping[str, str] + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ImportPromptSpec: + prompt_key: str + template_identity: str | None + template_digest: str | None + template_contents: Mapping[str, str] | None + unknown_reason: str | None + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ResultOriginSpec: + derivation_origin: DerivationOrigin + default_prediction_origin: PredictionOrigin | None + per_question_prediction_origins: Mapping[str, PredictionOrigin] + + +@dataclass(frozen=True) +class ImportConditionSpec: + condition_key: str + source_ids: tuple[str, ...] + dataset_id: str + model_key: str + method_key: str + prompt_key: str + seed: int | None + calibration_identity: Mapping[str, Any] | None + preflight_identity: Mapping[str, Any] | None + protocol_settings: Mapping[str, Any] + generation_parameters: Mapping[str, Any] + unknown_reasons: Mapping[str, str] + expected_question_ids: tuple[str, ...] + evidence_status: EvidenceStatus + scope_disposition: ScopeDisposition + executable: bool | None + qualifications: tuple[Mapping[str, Any], ...] + limitations: tuple[Mapping[str, Any], ...] + damaged_question_ids: tuple[str, ...] + recoverable_question_ids: tuple[str, ...] + result_origin: ResultOriginSpec + + +@dataclass(frozen=True) +class AuthorizationSpec: + authorization_id: str + authorization_type: Literal["inference_repair", "offline_transformation"] + source_id: str + condition_question_reasons: Mapping[str, Mapping[str, str]] + authority: str + purpose: str + executable: bool + input_evidence_digests: Mapping[str, str] + expected_snapshot_digests: Mapping[str, str] + + +@dataclass(frozen=True) +class OverlaySpec: + overlay_id: str + base_run_path: Path + base_condition_digest: str + base_realization_id: str + base_realization_digest: str + base_evidence_digests: Mapping[str, str] + base_validation_artifact_sha256: str + base_result_sha256: str | None + source_id: str + authorization_id: str + replacement_reasons: Mapping[str, str] + result_origin: ResultOriginSpec + lineage_notes: Mapping[str, Any] + implementation: Mapping[str, Any] + input_digest: str + preownership_output_digest: str + expected_evidence_status: EvidenceStatus + + +@dataclass(frozen=True) +class ImportSpec: + schema_version: Literal["choicebench.import-spec.v1"] + import_name: str + sources: tuple[SourceArtifactSpec, ...] + datasets: tuple[DatasetReferenceSpec, ...] + models: tuple[ImportModelSpec, ...] + methods: tuple[ImportMethodSpec, ...] + prompts: tuple[ImportPromptSpec, ...] + conditions: tuple[ImportConditionSpec, ...] + authorizations: tuple[AuthorizationSpec, ...] + overlays: tuple[OverlaySpec, ...] + metrics: tuple[str, ...] + provenance: Mapping[str, Any] + audit: Mapping[str, Any] + + +_TOP_LEVEL_KEYS = { + "schema_version", + "import_name", + "sources", + "datasets", + "models", + "methods", + "prompts", + "conditions", + "authorizations", + "overlays", + "metrics", + "provenance", + "audit", +} +_SOURCE_KEYS = { + "source_id", + "path", + "logical_path", + "expected_sha256", + "format", + "format_version", + "classification", + "dialect", + "columns", + "expected_columns", + "ignored_columns", + "null_values", + "numeric_columns", + "option_mapping", + "extra_field_policy", + "preserve_namespace", + "source_run_id", + "source_repository", + "source_commit", + "notes", +} +_DIALECT_KEYS = { + "encoding", + "bom_policy", + "decoding_errors", + "delimiter", + "quote_character", + "escape_character", + "double_quote", + "line_terminators", + "mixed_line_terminators", + "final_record_without_terminator", + "blank_record_policy", + "skip_initial_space", + "header", + "strict_syntax", +} +_NUMERIC_KEYS = {"source_column", "value_type", "null_allowed", "finite_only"} +_OPTION_KEYS = { + "mode", + "ordered_columns", + "structured_column", + "structured_label_key", + "structured_text_key", +} +_DATASET_KEYS = { + "dataset_id", + "benchmark_name", + "split", + "reference_kind", + "trust_label", + "source_ids", + "selection_source_id", + "expected_question_ids", + "selection_seed", + "selection_n_samples", + "subject_filter", + "selection_unknown_reasons", + "columns", + "revision", + "fingerprint", + "derivation", + "limitations", + "native_compatibility_identity", +} +_MODEL_KEYS = { + "model_key", + "display_name", + "backend", + "provider", + "revision", + "effective_parameters", + "unknown_reasons", + "native_compatibility_identity", +} +_METHOD_KEYS = { + "method_key", + "name", + "effective_parameters", + "implementation", + "unknown_reasons", + "native_compatibility_identity", +} +_PROMPT_KEYS = { + "prompt_key", + "template_identity", + "template_digest", + "template_contents", + "unknown_reason", + "native_compatibility_identity", +} +_CONDITION_KEYS = { + "condition_key", + "source_ids", + "dataset_id", + "model_key", + "method_key", + "prompt_key", + "seed", + "calibration_identity", + "preflight_identity", + "protocol_settings", + "generation_parameters", + "unknown_reasons", + "expected_question_ids", + "evidence_status", + "scope_disposition", + "executable", + "qualifications", + "limitations", + "damaged_question_ids", + "recoverable_question_ids", + "result_origin", +} +_RESULT_ORIGIN_KEYS = { + "derivation_origin", + "default_prediction_origin", + "per_question_prediction_origins", +} +_AUTHORIZATION_KEYS = { + "authorization_id", + "authorization_type", + "source_id", + "condition_question_reasons", + "authority", + "purpose", + "executable", + "input_evidence_digests", + "expected_snapshot_digests", +} +_OVERLAY_KEYS = { + "overlay_id", + "base_run_path", + "base_condition_digest", + "base_realization_id", + "base_realization_digest", + "base_evidence_digests", + "base_validation_artifact_sha256", + "base_result_sha256", + "source_id", + "authorization_id", + "replacement_reasons", + "result_origin", + "lineage_notes", + "implementation", + "input_digest", + "preownership_output_digest", + "expected_evidence_status", +} + +_EVIDENCE_STATUSES = { + "complete", "qualified", "partial", "malformed", "recoverable", "failed" +} +_SCOPE_DISPOSITIONS = { + "included", "excluded_from_paper_matrix", "held", "superseded" +} +_PREDICTION_ORIGINS = { + "native_inference", "external_historical_inference", "external_repair_inference" +} +_DERIVATION_ORIGINS = { + "native_execution", "external_import", "repair_overlay", "offline_transformation" +} +_SHA256_RE = re.compile(r"^[0-9a-f]{64}$") +_WINDOWS_ABSOLUTE_RE = re.compile(r"^[A-Za-z]:[\\/]") + + +def _mapping(value: Any, where: str) -> Mapping[str, Any]: + if not isinstance(value, Mapping): + raise ImportSpecError(f"{where} must be a YAML mapping; got {value!r}.") + bad_key = next((key for key in value if not isinstance(key, str)), None) + if bad_key is not None: + raise ImportSpecError( + f"{where} mapping keys must be strings; got {bad_key!r}." + ) + return value + + +def _exact_keys(value: Any, allowed: set[str], where: str) -> Mapping[str, Any]: + raw = _mapping(value, where) + unknown = sorted(set(raw) - allowed) + if unknown: + raise ImportSpecError(f"Unknown field(s) in {where}: {unknown}.") + missing = sorted(allowed - set(raw)) + if missing: + raise ImportSpecError(f"Missing required field(s) in {where}: {missing}.") + return raw + + +def _allowed_keys(value: Any, allowed: set[str], where: str) -> Mapping[str, Any]: + raw = _mapping(value, where) + unknown = sorted(set(raw) - allowed) + if unknown: + raise ImportSpecError(f"Unknown field(s) in {where}: {unknown}.") + return raw + + +def _nonempty(value: Any, where: str) -> str: + if not isinstance(value, str) or not value.strip(): + raise ImportSpecError(f"{where} must be a non-empty string; got {value!r}.") + return value.strip() + + +def _optional_string(value: Any, where: str) -> str | None: + if value is None: + return None + return _nonempty(value, where) + + +def _strict_bool(value: Any, where: str) -> bool: + if not isinstance(value, bool): + raise ImportSpecError(f"{where} must be true or false; got {value!r}.") + return value + + +def _optional_bool(value: Any, where: str) -> bool | None: + if value is None: + return None + return _strict_bool(value, where) + + +def _strict_int(value: Any, where: str, *, minimum: int | None = None) -> int: + if isinstance(value, bool) or not isinstance(value, int): + raise ImportSpecError(f"{where} must be an integer; got {value!r}.") + if minimum is not None and value < minimum: + raise ImportSpecError(f"{where} must be >= {minimum}; got {value!r}.") + return value + + +def _strict_number(value: Any, where: str) -> float: + if isinstance(value, bool) or not isinstance(value, Real): + raise ImportSpecError(f"{where} must be a finite number; got {value!r}.") + result = float(value) + if not math.isfinite(result): + raise ImportSpecError(f"{where} must be a finite number; got {value!r}.") + return result + + +def _optional_int(value: Any, where: str, *, minimum: int | None = None) -> int | None: + if value is None: + return None + return _strict_int(value, where, minimum=minimum) + + +def _enum(value: Any, values: set[str], where: str) -> str: + result = _nonempty(value, where) + if result not in values: + raise ImportSpecError( + f"{where} must be one of {sorted(values)}; got {result!r}." + ) + return result + + +def _sha256(value: Any, where: str) -> str: + result = _nonempty(value, where) + if not _SHA256_RE.fullmatch(result): + raise ImportSpecError(f"{where} must be a lowercase SHA-256 digest.") + return result + + +def _optional_sha256(value: Any, where: str) -> str | None: + if value is None: + return None + return _sha256(value, where) + + +def _sequence(value: Any, where: str) -> list[Any]: + if not isinstance(value, list): + raise ImportSpecError(f"{where} must be a YAML list; got {value!r}.") + return value + + +def _strings( + value: Any, + where: str, + *, + nonempty: bool = False, + unique: bool = False, + allow_empty_items: bool = False, +) -> tuple[str, ...]: + items = _sequence(value, where) + if allow_empty_items: + if not all(isinstance(item, str) for item in items): + bad = next(item for item in items if not isinstance(item, str)) + raise ImportSpecError(f"{where} entries must be strings; got {bad!r}.") + result = tuple(items) + else: + result = tuple( + _nonempty(item, f"{where}[{index}]") + for index, item in enumerate(items) + ) + if nonempty and not result: + raise ImportSpecError(f"{where} must be a non-empty list.") + if unique and len(result) != len(set(result)): + raise ImportSpecError(f"{where} contains duplicate values.") + return result + + +def _string_mapping(value: Any, where: str) -> dict[str, str]: + raw = _mapping(value, where) + return { + _nonempty(key, f"{where} key"): _nonempty(item, f"{where}.{key}") + for key, item in raw.items() + } + + +def _sha_mapping(value: Any, where: str) -> dict[str, str]: + raw = _mapping(value, where) + return { + _nonempty(key, f"{where} key"): _sha256(item, f"{where}.{key}") + for key, item in raw.items() + } + + +def _credential_keys_in(value: Any) -> set[str]: + found: set[str] = set() + if isinstance(value, Mapping): + for key, item in value.items(): + if isinstance(key, str) and is_credential_key(key): + found.add(key) + found |= _credential_keys_in(item) + elif isinstance(value, (list, tuple)): + for item in value: + found |= _credential_keys_in(item) + return found + + +def _canonical_mapping(value: Any, where: str) -> dict[str, Any]: + raw = _mapping(value, where) + try: + result = canonicalize(raw) + except (TypeError, ValueError) as exc: + raise ImportSpecError( + f"{where} must be finite, safe canonical data: {exc}" + ) from exc + return dict(result) + + +def _canonical_mapping_or_none(value: Any, where: str) -> dict[str, Any] | None: + if value is None: + return None + return _canonical_mapping(value, where) + + +def _mapping_sequence(value: Any, where: str) -> tuple[Mapping[str, Any], ...]: + return tuple( + _canonical_mapping(item, f"{where}[{index}]") + for index, item in enumerate(_sequence(value, where)) + ) + + +def _is_path_like(value: str) -> bool: + return ( + value.startswith(("/", "./", "../", "~/", "~\\")) + or _WINDOWS_ABSOLUTE_RE.match(value) is not None + ) + + +def _contains_path_like(value: Any) -> bool: + if isinstance(value, str): + return _is_path_like(value) + if isinstance(value, Mapping): + return any(_contains_path_like(item) for item in value.values()) + if isinstance(value, (list, tuple)): + return any(_contains_path_like(item) for item in value) + return False + + +def _build_provenance(value: Any) -> dict[str, Any]: + raw = _mapping(value, "provenance") + result: dict[str, Any] = {} + for key, item in raw.items(): + where = f"provenance.{key}" + record = _exact_keys(item, {"value", "reason"}, where) + reason = record["reason"] + if record["value"] is None: + reason = _nonempty(reason, f"{where}.reason") + elif reason is not None: + reason = _nonempty(reason, f"{where}.reason") + if _contains_path_like(record["value"]): + raise ImportSpecError( + f"{where}.value contains a machine-local path; put it in audit instead." + ) + canonical = _canonical_mapping( + {"value": record["value"], "reason": reason}, where + ) + result[_nonempty(key, "provenance key")] = canonical + return result + + +def _ascii_byte(value: Any, where: str, *, optional: bool = False) -> str | None: + if value is None and optional: + return None + if not isinstance(value, str) or len(value.encode("utf-8")) != 1: + raise ImportSpecError(f"{where} must be one ASCII byte.") + if ord(value) >= 128 or value in {"\x00", "\r", "\n"}: + raise ImportSpecError(f"{where} must be one usable ASCII byte.") + return value + + +def _build_dialect(value: Any, where: str) -> CsvDialectSpec: + raw = _allowed_keys(value, _DIALECT_KEYS, where) + encoding = _enum(raw.get("encoding", "utf-8"), {"utf-8"}, f"{where}.encoding") + decoding_errors = _enum( + raw.get("decoding_errors", "strict"), {"strict"}, f"{where}.decoding_errors" + ) + delimiter = _ascii_byte(raw.get("delimiter", ","), f"{where}.delimiter") + quote = _ascii_byte(raw.get("quote_character", '"'), f"{where}.quote_character") + escape = _ascii_byte( + raw.get("escape_character"), f"{where}.escape_character", optional=True + ) + characters = [item for item in (delimiter, quote, escape) if item is not None] + if len(characters) != len(set(characters)): + raise ImportSpecError( + f"{where} delimiter, quote_character, and escape_character must be distinct." + ) + line_terminators = _strings( + raw.get("line_terminators", ["crlf", "lf", "cr"]), + f"{where}.line_terminators", + nonempty=True, + unique=True, + ) + if not set(line_terminators) <= {"crlf", "lf", "cr"}: + raise ImportSpecError( + f"{where}.line_terminators may contain only crlf, lf, and cr." + ) + return CsvDialectSpec( + encoding=encoding, + bom_policy=_enum( + raw.get("bom_policy", "forbid"), + {"forbid", "strip_utf8_bom"}, + f"{where}.bom_policy", + ), + decoding_errors=decoding_errors, + delimiter=delimiter, + quote_character=quote, + escape_character=escape, + double_quote=_strict_bool(raw.get("double_quote", True), f"{where}.double_quote"), + line_terminators=line_terminators, + mixed_line_terminators=_enum( + raw.get("mixed_line_terminators", "allow"), + {"allow", "forbid"}, + f"{where}.mixed_line_terminators", + ), + final_record_without_terminator=_enum( + raw.get("final_record_without_terminator", "allow"), + {"allow", "forbid"}, + f"{where}.final_record_without_terminator", + ), + blank_record_policy=_enum( + raw.get("blank_record_policy", "reject"), + {"reject"}, + f"{where}.blank_record_policy", + ), + skip_initial_space=_strict_bool( + raw.get("skip_initial_space", False), f"{where}.skip_initial_space" + ), + header=_enum( + raw.get("header", "first_logical_record"), + {"first_logical_record"}, + f"{where}.header", + ), + strict_syntax=_strict_bool( + raw.get("strict_syntax", True), f"{where}.strict_syntax" + ), + ) + + +def _build_numeric(value: Any, where: str) -> NumericColumnSpec: + raw = _allowed_keys(value, _NUMERIC_KEYS, where) + missing = sorted({"source_column", "value_type", "null_allowed"} - set(raw)) + if missing: + raise ImportSpecError(f"Missing required field(s) in {where}: {missing}.") + return NumericColumnSpec( + source_column=_nonempty(raw["source_column"], f"{where}.source_column"), + value_type=_enum( + raw["value_type"], {"integer", "float"}, f"{where}.value_type" + ), + null_allowed=_strict_bool(raw["null_allowed"], f"{where}.null_allowed"), + finite_only=_strict_bool(raw.get("finite_only", True), f"{where}.finite_only"), + ) + + +def _build_option(value: Any, where: str) -> OptionMappingSpec: + raw = _exact_keys(value, _OPTION_KEYS, where) + mode = _enum(raw["mode"], {"ordered_columns", "structured_json"}, f"{where}.mode") + ordered = _strings( + raw["ordered_columns"], f"{where}.ordered_columns", unique=True + ) + structured_column = _optional_string( + raw["structured_column"], f"{where}.structured_column" + ) + label_key = _optional_string( + raw["structured_label_key"], f"{where}.structured_label_key" + ) + text_key = _optional_string( + raw["structured_text_key"], f"{where}.structured_text_key" + ) + if mode == "ordered_columns": + if not ordered or any( + item is not None for item in (structured_column, label_key, text_key) + ): + raise ImportSpecError( + f"{where} ordered_columns mode requires ordered columns and null structured fields." + ) + elif ordered or any(item is None for item in (structured_column, label_key, text_key)): + raise ImportSpecError( + f"{where} structured_json mode requires an empty ordered list and all structured fields." + ) + return OptionMappingSpec(mode, ordered, structured_column, label_key, text_key) + + +def _build_source(value: Any, index: int) -> SourceArtifactSpec: + where = f"sources[{index}]" + raw = _exact_keys(value, _SOURCE_KEYS, where) + logical_path = _nonempty(raw["logical_path"], f"{where}.logical_path") + pure_path = PurePosixPath(logical_path) + if pure_path.is_absolute() or ".." in pure_path.parts: + raise ImportSpecError(f"{where}.logical_path must be a safe relative logical path.") + expected_columns = _strings( + raw["expected_columns"], f"{where}.expected_columns", nonempty=True, unique=True + ) + numeric_columns = tuple( + _build_numeric(item, f"{where}.numeric_columns[{item_index}]") + for item_index, item in enumerate(_sequence(raw["numeric_columns"], f"{where}.numeric_columns")) + ) + numeric_names = [item.source_column for item in numeric_columns] + if len(numeric_names) != len(set(numeric_names)): + raise ImportSpecError(f"Duplicate numeric-column rules in {where}.numeric_columns.") + missing_numeric = sorted(set(numeric_names) - set(expected_columns)) + if missing_numeric: + raise ImportSpecError( + f"{where}.numeric_columns reference unknown expected columns {missing_numeric}." + ) + option_mapping = _build_option(raw["option_mapping"], f"{where}.option_mapping") + option_columns = ( + set(option_mapping.ordered_columns) + if option_mapping.mode == "ordered_columns" + else {option_mapping.structured_column} + ) + missing_options = sorted(item for item in option_columns if item not in expected_columns) + if missing_options: + raise ImportSpecError( + f"{where}.option_mapping references unknown expected columns {missing_options}." + ) + columns = _string_mapping(raw["columns"], f"{where}.columns") + missing_mappings = sorted(set(columns.values()) - set(expected_columns)) + if missing_mappings: + raise ImportSpecError( + f"{where}.columns references unknown expected columns {missing_mappings}." + ) + ignored_columns = _string_mapping(raw["ignored_columns"], f"{where}.ignored_columns") + missing_ignored = sorted(set(ignored_columns) - set(expected_columns)) + if missing_ignored: + raise ImportSpecError( + f"{where}.ignored_columns references unknown expected columns {missing_ignored}." + ) + path_text = _nonempty(raw["path"], f"{where}.path") + return SourceArtifactSpec( + source_id=_nonempty(raw["source_id"], f"{where}.source_id"), + path=Path(path_text), + logical_path=logical_path, + expected_sha256=_sha256(raw["expected_sha256"], f"{where}.expected_sha256"), + format=_enum(raw["format"], {"csv"}, f"{where}.format"), + format_version=_nonempty(raw["format_version"], f"{where}.format_version"), + classification=_enum( + raw["classification"], + {"raw", "canonical", "derived", "repaired", "aggregate_only"}, + f"{where}.classification", + ), + dialect=_build_dialect(raw["dialect"], f"{where}.dialect"), + columns=columns, + expected_columns=expected_columns, + ignored_columns=ignored_columns, + null_values=_strings( + raw["null_values"], + f"{where}.null_values", + unique=True, + allow_empty_items=True, + ), + numeric_columns=numeric_columns, + option_mapping=option_mapping, + extra_field_policy=_enum( + raw["extra_field_policy"], + {"preserve_unmapped", "reject_unmapped"}, + f"{where}.extra_field_policy", + ), + preserve_namespace=_nonempty( + raw["preserve_namespace"], f"{where}.preserve_namespace" + ), + source_run_id=_optional_string(raw["source_run_id"], f"{where}.source_run_id"), + source_repository=_optional_string( + raw["source_repository"], f"{where}.source_repository" + ), + source_commit=_optional_string(raw["source_commit"], f"{where}.source_commit"), + notes=_canonical_mapping(raw["notes"], f"{where}.notes"), + ) + + +def _validate_identity_claim( + raw: Mapping[str, Any], *, payload_key: str, digest_key: str, id_key: str, + prefix: str, where: str +) -> None: + payload = raw[payload_key] + if raw[digest_key] != integrity_digest(payload): + raise ImportSpecError(f"{where}.{digest_key} does not match its payload.") + if raw[id_key] != short_id(prefix, payload): + raise ImportSpecError(f"{where}.{id_key} does not match its payload.") + + +def _validate_dataset_native(value: Any, where: str) -> dict[str, Any]: + keys = { + "artifact_payload", "artifact_digest", "artifact_id", + "selection_payload", "selection_digest", "selection_id", + } + raw = _exact_keys(value, keys, where) + artifact = _exact_keys( + raw["artifact_payload"], {"spec", "content_digest", "source"}, + f"{where}.artifact_payload", + ) + spec = _exact_keys( + artifact["spec"], + { + "benchmark", "split", "hf_path", "hf_subset", "source_revision", + "normalization_version", "transforms", "output_name", + }, + f"{where}.artifact_payload.spec", + ) + normalized_artifact = { + "spec": { + "benchmark": _nonempty(spec["benchmark"], f"{where}.artifact_payload.spec.benchmark"), + "split": _nonempty(spec["split"], f"{where}.artifact_payload.spec.split"), + "hf_path": _optional_string(spec["hf_path"], f"{where}.artifact_payload.spec.hf_path"), + "hf_subset": _optional_string(spec["hf_subset"], f"{where}.artifact_payload.spec.hf_subset"), + "source_revision": _optional_string(spec["source_revision"], f"{where}.artifact_payload.spec.source_revision"), + "normalization_version": _nonempty(spec["normalization_version"], f"{where}.artifact_payload.spec.normalization_version"), + "transforms": list(_strings(spec["transforms"], f"{where}.artifact_payload.spec.transforms", unique=True)), + "output_name": _optional_string(spec["output_name"], f"{where}.artifact_payload.spec.output_name"), + }, + "content_digest": _sha256(artifact["content_digest"], f"{where}.artifact_payload.content_digest"), + "source": _canonical_mapping(artifact["source"], f"{where}.artifact_payload.source"), + } + selection = _exact_keys( + raw["selection_payload"], + {"artifact_id", "content_digest", "sample_identities", "seed", "n_samples", "subject_filter"}, + f"{where}.selection_payload", + ) + normalized_selection = { + "artifact_id": _nonempty(selection["artifact_id"], f"{where}.selection_payload.artifact_id"), + "content_digest": _sha256(selection["content_digest"], f"{where}.selection_payload.content_digest"), + "sample_identities": list(_strings(selection["sample_identities"], f"{where}.selection_payload.sample_identities", unique=False)), + "seed": _strict_int(selection["seed"], f"{where}.selection_payload.seed"), + "n_samples": _optional_int(selection["n_samples"], f"{where}.selection_payload.n_samples", minimum=1), + "subject_filter": list(_strings(selection["subject_filter"], f"{where}.selection_payload.subject_filter", unique=True)), + } + for index, digest in enumerate(normalized_selection["sample_identities"]): + _sha256(digest, f"{where}.selection_payload.sample_identities[{index}]") + if normalized_selection["subject_filter"] != sorted(normalized_selection["subject_filter"]): + raise ImportSpecError(f"{where}.selection_payload.subject_filter must be sorted.") + normalized = { + "artifact_payload": normalized_artifact, + "artifact_digest": _sha256(raw["artifact_digest"], f"{where}.artifact_digest"), + "artifact_id": _nonempty(raw["artifact_id"], f"{where}.artifact_id"), + "selection_payload": normalized_selection, + "selection_digest": _sha256(raw["selection_digest"], f"{where}.selection_digest"), + "selection_id": _nonempty(raw["selection_id"], f"{where}.selection_id"), + } + _validate_identity_claim( + normalized, payload_key="artifact_payload", digest_key="artifact_digest", + id_key="artifact_id", prefix="ds", where=where, + ) + if normalized_selection["artifact_id"] != normalized["artifact_id"]: + raise ImportSpecError(f"{where}.selection_payload.artifact_id is inconsistent.") + _validate_identity_claim( + normalized, payload_key="selection_payload", digest_key="selection_digest", + id_key="selection_id", prefix="sel", where=where, + ) + return normalized + + +def _validate_model_payload(value: Any, where: str) -> dict[str, Any]: + raw = _mapping(value, where) + backend = _enum(raw.get("backend"), {"dummy", "api", "huggingface"}, f"{where}.backend") + allowed = { + "dummy": {"backend", "model_name_or_path"}, + "api": {"backend", "provider", "model_name_or_path", "base_url", "generation_kwargs"}, + "huggingface": {"backend", "model", "device", "add_bos_token", "generation_kwargs", "loader"}, + }[backend] + _exact_keys(raw, allowed, where) + if backend == "dummy": + return { + "backend": backend, + "model_name_or_path": _nonempty( + raw["model_name_or_path"], f"{where}.model_name_or_path" + ), + } + generation_keys = ( + {"max_new_tokens", "temperature"} + if backend == "api" + else {"max_new_tokens", "temperature", "do_sample"} + ) + generation = _exact_keys( + raw["generation_kwargs"], generation_keys, f"{where}.generation_kwargs" + ) + normalized_generation: dict[str, Any] = { + "max_new_tokens": _strict_int( + generation["max_new_tokens"], + f"{where}.generation_kwargs.max_new_tokens", + minimum=1, + ), + "temperature": _strict_number( + generation["temperature"], f"{where}.generation_kwargs.temperature" + ), + } + if backend == "api": + return { + "backend": backend, + "provider": _nonempty(raw["provider"], f"{where}.provider"), + "model_name_or_path": _nonempty( + raw["model_name_or_path"], f"{where}.model_name_or_path" + ), + "base_url": _optional_string(raw["base_url"], f"{where}.base_url"), + "generation_kwargs": normalized_generation, + } + normalized_generation["do_sample"] = _strict_bool( + generation["do_sample"], f"{where}.generation_kwargs.do_sample" + ) + resolved = _mapping(raw["model"], f"{where}.model") + kind = _enum( + resolved.get("kind"), {"local", "huggingface-hub"}, f"{where}.model.kind" + ) + if kind == "local": + resolved = _exact_keys( + resolved, + {"kind", "logical_name", "content_digest", "file_count", "total_bytes"}, + f"{where}.model", + ) + normalized_model = { + "kind": kind, + "logical_name": _nonempty( + resolved["logical_name"], f"{where}.model.logical_name" + ), + "content_digest": _sha256( + resolved["content_digest"], f"{where}.model.content_digest" + ), + "file_count": _strict_int( + resolved["file_count"], f"{where}.model.file_count", minimum=1 + ), + "total_bytes": _strict_int( + resolved["total_bytes"], f"{where}.model.total_bytes", minimum=0 + ), + } + else: + resolved = _exact_keys( + resolved, + {"kind", "repo_id", "requested_revision", "resolved_commit"}, + f"{where}.model", + ) + resolved_commit = _nonempty( + resolved["resolved_commit"], f"{where}.model.resolved_commit" + ) + if not re.fullmatch(r"[0-9a-f]{40,64}", resolved_commit): + raise ImportSpecError( + f"{where}.model.resolved_commit must be a lowercase immutable commit." + ) + normalized_model = { + "kind": kind, + "repo_id": _nonempty(resolved["repo_id"], f"{where}.model.repo_id"), + "requested_revision": _optional_string( + resolved["requested_revision"], f"{where}.model.requested_revision" + ), + "resolved_commit": resolved_commit, + } + loader = _exact_keys( + raw["loader"], {"trust_remote_code", "torch_dtype"}, f"{where}.loader" + ) + return { + "backend": backend, + "model": normalized_model, + "device": _nonempty(raw["device"], f"{where}.device"), + "add_bos_token": _strict_bool(raw["add_bos_token"], f"{where}.add_bos_token"), + "generation_kwargs": normalized_generation, + "loader": { + "trust_remote_code": _strict_bool( + loader["trust_remote_code"], f"{where}.loader.trust_remote_code" + ), + "torch_dtype": _enum( + loader["torch_dtype"], + {"float16", "float32"}, + f"{where}.loader.torch_dtype", + ), + }, + } + + +_IMPLEMENTATION_KEYS = { + "qualified_name", "source_file", "source_digest", "distribution", + "distribution_version", "package_tree_digest", +} + + +def _validate_implementation(value: Any, where: str) -> dict[str, Any]: + raw = _allowed_keys(value, _IMPLEMENTATION_KEYS, where) + if "qualified_name" not in raw: + raise ImportSpecError(f"Missing required field 'qualified_name' in {where}.") + result = _canonical_mapping(raw, where) + _nonempty(result["qualified_name"], f"{where}.qualified_name") + for key in ("source_digest", "package_tree_digest"): + if key in result: + _sha256(result[key], f"{where}.{key}") + return result + + +def _validate_method_payload(value: Any, where: str) -> dict[str, Any]: + raw = _exact_keys( + value, {"name", "effective_params", "preflight", "implementation"}, where + ) + preflight = raw["preflight"] + if preflight is not None: + preflight = _exact_keys( + preflight, {"source", "split", "n"}, f"{where}.preflight" + ) + preflight = { + "source": _nonempty(preflight["source"], f"{where}.preflight.source"), + "split": _nonempty(preflight["split"], f"{where}.preflight.split"), + "n": _strict_int(preflight["n"], f"{where}.preflight.n", minimum=1), + } + return { + "name": _nonempty(raw["name"], f"{where}.name"), + "effective_params": _canonical_mapping(raw["effective_params"], f"{where}.effective_params"), + "preflight": preflight, + "implementation": _validate_implementation(raw["implementation"], f"{where}.implementation"), + } + + +def _validate_prompt_payload(value: Any, where: str) -> dict[str, Any]: + raw = _exact_keys(value, {"version", "files"}, where) + files_raw = _mapping(raw["files"], f"{where}.files") + if not files_raw: + raise ImportSpecError(f"{where}.files must not be empty.") + files: dict[str, Any] = {} + for name, item in files_raw.items(): + record = _exact_keys(item, {"sha256", "content"}, f"{where}.files.{name}") + content = _nonempty(record["content"], f"{where}.files.{name}.content") + digest = _sha256(record["sha256"], f"{where}.files.{name}.sha256") + if digest != integrity_digest(content): + raise ImportSpecError(f"{where}.files.{name}.sha256 does not match content.") + files[_nonempty(name, f"{where}.files key")] = {"sha256": digest, "content": content} + return {"version": _nonempty(raw["version"], f"{where}.version"), "files": files} + + +def _validate_simple_native( + value: Any, where: str, prefix: str, payload_validator +) -> dict[str, Any]: + id_key = f"{prefix}_id" + raw = _exact_keys(value, {"payload", "digest", id_key}, where) + normalized = { + "payload": payload_validator(raw["payload"], f"{where}.payload"), + "digest": _sha256(raw["digest"], f"{where}.digest"), + id_key: _nonempty(raw[id_key], f"{where}.{id_key}"), + } + _validate_identity_claim( + normalized, payload_key="payload", digest_key="digest", id_key=id_key, + prefix=prefix, where=where, + ) + return normalized + + +def _build_dataset(value: Any, index: int) -> DatasetReferenceSpec: + where = f"datasets[{index}]" + raw = _exact_keys(value, _DATASET_KEYS, where) + native = raw["native_compatibility_identity"] + return DatasetReferenceSpec( + dataset_id=_nonempty(raw["dataset_id"], f"{where}.dataset_id"), + benchmark_name=_nonempty(raw["benchmark_name"], f"{where}.benchmark_name"), + split=_nonempty(raw["split"], f"{where}.split"), + reference_kind=_enum( + raw["reference_kind"], + {"independent_input_snapshot", "profile_derived_reference_snapshot"}, + f"{where}.reference_kind", + ), + trust_label=_nonempty(raw["trust_label"], f"{where}.trust_label"), + source_ids=_strings(raw["source_ids"], f"{where}.source_ids", nonempty=True, unique=True), + selection_source_id=_nonempty(raw["selection_source_id"], f"{where}.selection_source_id"), + expected_question_ids=_strings(raw["expected_question_ids"], f"{where}.expected_question_ids", nonempty=True, unique=True), + selection_seed=_optional_int(raw["selection_seed"], f"{where}.selection_seed"), + selection_n_samples=_optional_int(raw["selection_n_samples"], f"{where}.selection_n_samples", minimum=1), + subject_filter=_strings(raw["subject_filter"], f"{where}.subject_filter", unique=True), + selection_unknown_reasons=_string_mapping(raw["selection_unknown_reasons"], f"{where}.selection_unknown_reasons"), + columns=_string_mapping(raw["columns"], f"{where}.columns"), + revision=_optional_string(raw["revision"], f"{where}.revision"), + fingerprint=_optional_string(raw["fingerprint"], f"{where}.fingerprint"), + derivation=_canonical_mapping(raw["derivation"], f"{where}.derivation"), + limitations=_strings(raw["limitations"], f"{where}.limitations"), + native_compatibility_identity=( + None if native is None else _validate_dataset_native(native, f"{where}.native_compatibility_identity") + ), + ) + + +def _build_model(value: Any, index: int) -> ImportModelSpec: + where = f"models[{index}]" + raw = _exact_keys(value, _MODEL_KEYS, where) + native = raw["native_compatibility_identity"] + return ImportModelSpec( + model_key=_nonempty(raw["model_key"], f"{where}.model_key"), + display_name=_nonempty(raw["display_name"], f"{where}.display_name"), + backend=_optional_string(raw["backend"], f"{where}.backend"), + provider=_optional_string(raw["provider"], f"{where}.provider"), + revision=_optional_string(raw["revision"], f"{where}.revision"), + effective_parameters=_canonical_mapping(raw["effective_parameters"], f"{where}.effective_parameters"), + unknown_reasons=_string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons"), + native_compatibility_identity=( + None if native is None else _validate_simple_native( + native, f"{where}.native_compatibility_identity", "model", _validate_model_payload + ) + ), + ) + + +def _build_method(value: Any, index: int) -> ImportMethodSpec: + where = f"methods[{index}]" + raw = _exact_keys(value, _METHOD_KEYS, where) + implementation = raw["implementation"] + native = raw["native_compatibility_identity"] + return ImportMethodSpec( + method_key=_nonempty(raw["method_key"], f"{where}.method_key"), + name=_nonempty(raw["name"], f"{where}.name"), + effective_parameters=_canonical_mapping(raw["effective_parameters"], f"{where}.effective_parameters"), + implementation=(None if implementation is None else _validate_implementation(implementation, f"{where}.implementation")), + unknown_reasons=_string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons"), + native_compatibility_identity=( + None if native is None else _validate_simple_native( + native, f"{where}.native_compatibility_identity", "method", _validate_method_payload + ) + ), + ) + + +def _build_prompt(value: Any, index: int) -> ImportPromptSpec: + where = f"prompts[{index}]" + raw = _exact_keys(value, _PROMPT_KEYS, where) + contents = raw["template_contents"] + native = raw["native_compatibility_identity"] + return ImportPromptSpec( + prompt_key=_nonempty(raw["prompt_key"], f"{where}.prompt_key"), + template_identity=_optional_string(raw["template_identity"], f"{where}.template_identity"), + template_digest=_optional_sha256(raw["template_digest"], f"{where}.template_digest"), + template_contents=(None if contents is None else _string_mapping(contents, f"{where}.template_contents")), + unknown_reason=_optional_string(raw["unknown_reason"], f"{where}.unknown_reason"), + native_compatibility_identity=( + None if native is None else _validate_simple_native( + native, f"{where}.native_compatibility_identity", "prompt", _validate_prompt_payload + ) + ), + ) + + +def _build_result_origin(value: Any, where: str) -> ResultOriginSpec: + raw = _exact_keys(value, _RESULT_ORIGIN_KEYS, where) + default = raw["default_prediction_origin"] + per_question_raw = _mapping( + raw["per_question_prediction_origins"], + f"{where}.per_question_prediction_origins", + ) + return ResultOriginSpec( + derivation_origin=_enum(raw["derivation_origin"], _DERIVATION_ORIGINS, f"{where}.derivation_origin"), + default_prediction_origin=(None if default is None else _enum(default, _PREDICTION_ORIGINS, f"{where}.default_prediction_origin")), + per_question_prediction_origins={ + _nonempty(key, f"{where}.per_question_prediction_origins key"): _enum( + item, _PREDICTION_ORIGINS, f"{where}.per_question_prediction_origins.{key}" + ) + for key, item in per_question_raw.items() + }, + ) + + +def _build_condition(value: Any, index: int) -> ImportConditionSpec: + where = f"conditions[{index}]" + raw = _exact_keys(value, _CONDITION_KEYS, where) + return ImportConditionSpec( + condition_key=_nonempty(raw["condition_key"], f"{where}.condition_key"), + source_ids=_strings(raw["source_ids"], f"{where}.source_ids", nonempty=True, unique=True), + dataset_id=_nonempty(raw["dataset_id"], f"{where}.dataset_id"), + model_key=_nonempty(raw["model_key"], f"{where}.model_key"), + method_key=_nonempty(raw["method_key"], f"{where}.method_key"), + prompt_key=_nonempty(raw["prompt_key"], f"{where}.prompt_key"), + seed=_optional_int(raw["seed"], f"{where}.seed"), + calibration_identity=_canonical_mapping_or_none(raw["calibration_identity"], f"{where}.calibration_identity"), + preflight_identity=_canonical_mapping_or_none(raw["preflight_identity"], f"{where}.preflight_identity"), + protocol_settings=_canonical_mapping(raw["protocol_settings"], f"{where}.protocol_settings"), + generation_parameters=_canonical_mapping(raw["generation_parameters"], f"{where}.generation_parameters"), + unknown_reasons=_string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons"), + expected_question_ids=_strings(raw["expected_question_ids"], f"{where}.expected_question_ids", nonempty=True, unique=True), + evidence_status=_enum(raw["evidence_status"], _EVIDENCE_STATUSES, f"{where}.evidence_status"), + scope_disposition=_enum(raw["scope_disposition"], _SCOPE_DISPOSITIONS, f"{where}.scope_disposition"), + executable=_optional_bool(raw["executable"], f"{where}.executable"), + qualifications=_mapping_sequence(raw["qualifications"], f"{where}.qualifications"), + limitations=_mapping_sequence(raw["limitations"], f"{where}.limitations"), + damaged_question_ids=_strings(raw["damaged_question_ids"], f"{where}.damaged_question_ids", unique=True), + recoverable_question_ids=_strings(raw["recoverable_question_ids"], f"{where}.recoverable_question_ids", unique=True), + result_origin=_build_result_origin(raw["result_origin"], f"{where}.result_origin"), + ) + + +def _build_authorization(value: Any, index: int) -> AuthorizationSpec: + where = f"authorizations[{index}]" + raw = _exact_keys(value, _AUTHORIZATION_KEYS, where) + reasons_raw = _mapping(raw["condition_question_reasons"], f"{where}.condition_question_reasons") + reasons = { + _nonempty(condition, f"{where}.condition_question_reasons key"): _string_mapping( + question_reasons, f"{where}.condition_question_reasons.{condition}" + ) + for condition, question_reasons in reasons_raw.items() + } + return AuthorizationSpec( + authorization_id=_nonempty(raw["authorization_id"], f"{where}.authorization_id"), + authorization_type=_enum(raw["authorization_type"], {"inference_repair", "offline_transformation"}, f"{where}.authorization_type"), + source_id=_nonempty(raw["source_id"], f"{where}.source_id"), + condition_question_reasons=reasons, + authority=_nonempty(raw["authority"], f"{where}.authority"), + purpose=_nonempty(raw["purpose"], f"{where}.purpose"), + executable=_strict_bool(raw["executable"], f"{where}.executable"), + input_evidence_digests=_sha_mapping(raw["input_evidence_digests"], f"{where}.input_evidence_digests"), + expected_snapshot_digests=_sha_mapping(raw["expected_snapshot_digests"], f"{where}.expected_snapshot_digests"), + ) + + +def _build_overlay(value: Any, index: int) -> OverlaySpec: + where = f"overlays[{index}]" + raw = _exact_keys(value, _OVERLAY_KEYS, where) + return OverlaySpec( + overlay_id=_nonempty(raw["overlay_id"], f"{where}.overlay_id"), + base_run_path=Path(_nonempty(raw["base_run_path"], f"{where}.base_run_path")), + base_condition_digest=_sha256(raw["base_condition_digest"], f"{where}.base_condition_digest"), + base_realization_id=_nonempty(raw["base_realization_id"], f"{where}.base_realization_id"), + base_realization_digest=_sha256(raw["base_realization_digest"], f"{where}.base_realization_digest"), + base_evidence_digests=_sha_mapping(raw["base_evidence_digests"], f"{where}.base_evidence_digests"), + base_validation_artifact_sha256=_sha256(raw["base_validation_artifact_sha256"], f"{where}.base_validation_artifact_sha256"), + base_result_sha256=_optional_sha256(raw["base_result_sha256"], f"{where}.base_result_sha256"), + source_id=_nonempty(raw["source_id"], f"{where}.source_id"), + authorization_id=_nonempty(raw["authorization_id"], f"{where}.authorization_id"), + replacement_reasons=_string_mapping(raw["replacement_reasons"], f"{where}.replacement_reasons"), + result_origin=_build_result_origin(raw["result_origin"], f"{where}.result_origin"), + lineage_notes=_canonical_mapping(raw["lineage_notes"], f"{where}.lineage_notes"), + implementation=_canonical_mapping(raw["implementation"], f"{where}.implementation"), + input_digest=_sha256(raw["input_digest"], f"{where}.input_digest"), + preownership_output_digest=_sha256(raw["preownership_output_digest"], f"{where}.preownership_output_digest"), + expected_evidence_status=_enum(raw["expected_evidence_status"], _EVIDENCE_STATUSES, f"{where}.expected_evidence_status"), + ) + + +def _unique(items: tuple[Any, ...], attribute: str, where: str) -> set[str]: + values = [getattr(item, attribute) for item in items] + if len(values) != len(set(values)): + raise ImportSpecError(f"Duplicate {attribute} in {where}.") + return set(values) + + +def _require_references(spec: ImportSpec) -> None: + source_ids = _unique(spec.sources, "source_id", "sources") + dataset_ids = _unique(spec.datasets, "dataset_id", "datasets") + model_keys = _unique(spec.models, "model_key", "models") + method_keys = _unique(spec.methods, "method_key", "methods") + prompt_keys = _unique(spec.prompts, "prompt_key", "prompts") + condition_keys = _unique(spec.conditions, "condition_key", "conditions") + authorization_ids = _unique(spec.authorizations, "authorization_id", "authorizations") + _unique(spec.overlays, "overlay_id", "overlays") + + for dataset in spec.datasets: + missing = sorted(set(dataset.source_ids) - source_ids) + if missing: + raise ImportSpecError( + f"Dataset {dataset.dataset_id!r} has unknown source reference(s) {missing}." + ) + if dataset.selection_source_id not in dataset.source_ids: + raise ImportSpecError( + f"Dataset {dataset.dataset_id!r} selection_source_id must reference one of its source_ids." + ) + dataset_by_id = {item.dataset_id: item for item in spec.datasets} + conditions_by_key = {item.condition_key: item for item in spec.conditions} + for condition in spec.conditions: + references = ( + (condition.dataset_id, dataset_ids, "dataset"), + (condition.model_key, model_keys, "model"), + (condition.method_key, method_keys, "method"), + (condition.prompt_key, prompt_keys, "prompt"), + ) + for reference, known, label in references: + if reference not in known: + raise ImportSpecError( + f"Condition {condition.condition_key!r} has unknown {label} reference {reference!r}." + ) + missing_sources = sorted(set(condition.source_ids) - source_ids) + if missing_sources: + raise ImportSpecError( + f"Condition {condition.condition_key!r} has unknown source reference(s) {missing_sources}." + ) + dataset_questions = set(dataset_by_id[condition.dataset_id].expected_question_ids) + expected = set(condition.expected_question_ids) + if not expected <= dataset_questions: + raise ImportSpecError( + f"Condition {condition.condition_key!r} expects question IDs absent from its dataset." + ) + damaged = set(condition.damaged_question_ids) + recoverable = set(condition.recoverable_question_ids) + if not damaged <= expected or not recoverable <= expected: + raise ImportSpecError( + f"Condition {condition.condition_key!r} damage references unknown question IDs." + ) + if not recoverable <= damaged: + raise ImportSpecError( + f"Condition {condition.condition_key!r} recoverable questions must also be damaged." + ) + origin_questions = set(condition.result_origin.per_question_prediction_origins) + if not origin_questions <= expected: + raise ImportSpecError( + f"Condition {condition.condition_key!r} has origins for unknown question IDs." + ) + + for authorization in spec.authorizations: + if authorization.source_id not in source_ids: + raise ImportSpecError( + f"Authorization {authorization.authorization_id!r} has an unknown source reference." + ) + for condition_key, reasons in authorization.condition_question_reasons.items(): + if condition_key not in condition_keys: + raise ImportSpecError( + f"Authorization {authorization.authorization_id!r} has an unknown condition reference." + ) + allowed_questions = set(conditions_by_key[condition_key].expected_question_ids) + if not set(reasons) <= allowed_questions: + raise ImportSpecError( + f"Authorization {authorization.authorization_id!r} references unknown question IDs." + ) + for overlay in spec.overlays: + if overlay.source_id not in source_ids: + raise ImportSpecError(f"Overlay {overlay.overlay_id!r} has an unknown source reference.") + if overlay.authorization_id not in authorization_ids: + raise ImportSpecError( + f"Overlay {overlay.overlay_id!r} has an unknown authorization reference." + ) + + +def _build_import_spec(raw_value: Any) -> ImportSpec: + raw = _exact_keys(raw_value, _TOP_LEVEL_KEYS, "top level") + credentials = sorted(_credential_keys_in(raw)) + if credentials: + raise ImportSpecError( + f"Import specification contains credential-named field(s) {credentials}; " + "credentials must never be stored in import declarations." + ) + sources = tuple( + _build_source(item, index) + for index, item in enumerate(_sequence(raw["sources"], "sources")) + ) + datasets = tuple( + _build_dataset(item, index) + for index, item in enumerate(_sequence(raw["datasets"], "datasets")) + ) + models = tuple( + _build_model(item, index) + for index, item in enumerate(_sequence(raw["models"], "models")) + ) + methods = tuple( + _build_method(item, index) + for index, item in enumerate(_sequence(raw["methods"], "methods")) + ) + prompts = tuple( + _build_prompt(item, index) + for index, item in enumerate(_sequence(raw["prompts"], "prompts")) + ) + conditions = tuple( + _build_condition(item, index) + for index, item in enumerate(_sequence(raw["conditions"], "conditions")) + ) + authorizations = tuple( + _build_authorization(item, index) + for index, item in enumerate(_sequence(raw["authorizations"], "authorizations")) + ) + overlays = tuple( + _build_overlay(item, index) + for index, item in enumerate(_sequence(raw["overlays"], "overlays")) + ) + for name, values in ( + ("sources", sources), ("datasets", datasets), ("models", models), + ("methods", methods), ("prompts", prompts), ("conditions", conditions), + ): + if not values: + raise ImportSpecError(f"{name} must be a non-empty list.") + metrics = _strings(raw["metrics"], "metrics", nonempty=True, unique=True) + invalid_metrics = sorted(set(metrics) - set(BUILTIN_METRICS)) + if invalid_metrics: + raise ImportSpecError( + f"Import metrics must use the closed built-in registry; unknown entries: {invalid_metrics}. " + f"Built-ins: {sorted(BUILTIN_METRICS)}." + ) + spec = ImportSpec( + schema_version=_enum( + raw["schema_version"], {"choicebench.import-spec.v1"}, "schema_version" + ), + import_name=_nonempty(raw["import_name"], "import_name"), + sources=sources, + datasets=datasets, + models=models, + methods=methods, + prompts=prompts, + conditions=conditions, + authorizations=authorizations, + overlays=overlays, + metrics=metrics, + provenance=_build_provenance(raw["provenance"]), + audit=_canonical_mapping(raw["audit"], "audit"), + ) + _require_references(spec) + return spec + + +def load_import_spec(path: Path) -> ImportSpec: + """Load a YAML import declaration using a closed, non-coercing schema.""" + source_path = Path(path) + try: + raw = yaml.safe_load(source_path.read_text(encoding="utf-8")) + except yaml.YAMLError as exc: + raise ImportSpecError(f"Could not parse {source_path} as safe YAML: {exc}") from exc + except (OSError, UnicodeError) as exc: + raise ImportSpecError(f"Could not read import specification {source_path}: {exc}") from exc + if not isinstance(raw, Mapping): + raise ImportSpecError( + f"{source_path} must contain a YAML mapping at the top level." + ) + return _build_import_spec(raw) + + +def stable_import_projection(spec: ImportSpec) -> dict[str, Any]: + """Return the identity-bearing declaration without machine-local locations.""" + if not isinstance(spec, ImportSpec): + raise TypeError(f"spec must be an ImportSpec; got {type(spec).__name__}.") + projection = asdict(spec) + projection.pop("audit", None) + for source in projection["sources"]: + source.pop("path", None) + for overlay in projection["overlays"]: + overlay.pop("base_run_path", None) + return canonicalize(projection) + + +def import_spec_digest(spec: ImportSpec) -> str: + """Return the stable full SHA-256 identity of an import declaration.""" + return integrity_digest(stable_import_projection(spec)) diff --git a/tests/importing/test_schema.py b/tests/importing/test_schema.py new file mode 100644 index 0000000..d47c85b --- /dev/null +++ b/tests/importing/test_schema.py @@ -0,0 +1,923 @@ +from __future__ import annotations + +from copy import deepcopy +from dataclasses import replace +from pathlib import Path + +import pytest +import yaml + +from choicebench.identity import integrity_digest, short_id +from choicebench.importing.schema import ( + ImportSpecError, + import_spec_digest, + load_import_spec, + stable_import_projection, +) +from choicebench.metrics import BUILTIN_METRICS + + +SHA_A = "a" * 64 +SHA_B = "b" * 64 +SHA_C = "c" * 64 + + +def _native_identity(prefix: str, payload: dict) -> dict: + return { + "payload": payload, + "digest": integrity_digest(payload), + f"{prefix}_id": short_id(prefix, payload), + } + + +@pytest.fixture +def minimal_raw(tmp_path: Path) -> dict: + return { + "schema_version": "choicebench.import-spec.v1", + "import_name": "minimal historical results", + "sources": [ + { + "source_id": "results", + "path": str(tmp_path / "results.csv"), + "logical_path": "freeze/results.csv", + "expected_sha256": SHA_A, + "format": "csv", + "format_version": "producer-v1", + "classification": "raw", + "dialect": { + "encoding": "utf-8", + "bom_policy": "forbid", + "decoding_errors": "strict", + "delimiter": ",", + "quote_character": '"', + "escape_character": None, + "double_quote": True, + "line_terminators": ["crlf", "lf", "cr"], + "mixed_line_terminators": "allow", + "final_record_without_terminator": "allow", + "blank_record_policy": "reject", + "skip_initial_space": False, + "header": "first_logical_record", + "strict_syntax": True, + }, + "columns": { + "question_id": "qid", + "prediction": "answer", + "gold": "gold", + }, + "expected_columns": [ + "qid", + "question", + "choice_a", + "choice_b", + "answer", + "gold", + "score", + ], + "ignored_columns": {"score": "producer aggregate only"}, + "null_values": ["", "NA"], + "numeric_columns": [ + { + "source_column": "score", + "value_type": "float", + "null_allowed": True, + "finite_only": True, + } + ], + "option_mapping": { + "mode": "ordered_columns", + "ordered_columns": ["choice_a", "choice_b"], + "structured_column": None, + "structured_label_key": None, + "structured_text_key": None, + }, + "extra_field_policy": "preserve_unmapped", + "preserve_namespace": "producer", + "source_run_id": None, + "source_repository": None, + "source_commit": None, + "notes": {}, + } + ], + "datasets": [ + { + "dataset_id": "dataset", + "benchmark_name": "historical-benchmark", + "split": "test", + "reference_kind": "independent_input_snapshot", + "trust_label": "producer-supplied", + "source_ids": ["results"], + "selection_source_id": "results", + "expected_question_ids": ["q1", "q2"], + "selection_seed": None, + "selection_n_samples": None, + "subject_filter": [], + "selection_unknown_reasons": { + "selection_seed": "not recorded by producer", + "selection_n_samples": "not recorded by producer", + }, + "columns": { + "question_id": "qid", + "question_text": "question", + "correct_option": "gold", + }, + "revision": None, + "fingerprint": None, + "derivation": {}, + "limitations": ["publisher revision was not recorded"], + "native_compatibility_identity": None, + } + ], + "models": [ + { + "model_key": "model", + "display_name": "historical-model", + "backend": None, + "provider": None, + "revision": None, + "effective_parameters": {}, + "unknown_reasons": { + "backend": "not recorded by producer", + "provider": "not recorded by producer", + "revision": "not recorded by producer", + }, + "native_compatibility_identity": None, + } + ], + "methods": [ + { + "method_key": "method", + "name": "historical-direct", + "effective_parameters": {}, + "implementation": None, + "unknown_reasons": { + "implementation": "not recorded by producer" + }, + "native_compatibility_identity": None, + } + ], + "prompts": [ + { + "prompt_key": "prompt", + "template_identity": None, + "template_digest": None, + "template_contents": None, + "unknown_reason": "not recorded by producer", + "native_compatibility_identity": None, + } + ], + "conditions": [ + { + "condition_key": "condition", + "source_ids": ["results"], + "dataset_id": "dataset", + "model_key": "model", + "method_key": "method", + "prompt_key": "prompt", + "seed": None, + "calibration_identity": None, + "preflight_identity": None, + "protocol_settings": {}, + "generation_parameters": {}, + "unknown_reasons": {"seed": "not recorded by producer"}, + "expected_question_ids": ["q1", "q2"], + "evidence_status": "complete", + "scope_disposition": "included", + "executable": None, + "qualifications": [], + "limitations": [], + "damaged_question_ids": [], + "recoverable_question_ids": [], + "result_origin": { + "derivation_origin": "external_import", + "default_prediction_origin": "external_historical_inference", + "per_question_prediction_origins": {}, + }, + } + ], + "authorizations": [], + "overlays": [], + "metrics": ["accuracy"], + "provenance": { + "producer_request_id": { + "value": None, + "reason": "not recorded by producer", + } + }, + "audit": { + "source_path": str(tmp_path), + "imported_at": "2026-07-18T00:00:00Z", + }, + } + + +def _load(tmp_path: Path, raw: object): + path = tmp_path / "import.yaml" + path.write_text(yaml.safe_dump(raw, sort_keys=False), encoding="utf-8") + return load_import_spec(path) + + +@pytest.fixture +def minimal_spec(tmp_path: Path, minimal_raw: dict): + return _load(tmp_path, minimal_raw) + + +@pytest.fixture +def minimal_overlay_spec(tmp_path: Path, minimal_raw: dict): + raw = deepcopy(minimal_raw) + raw["authorizations"] = [ + { + "authorization_id": "auth", + "authorization_type": "inference_repair", + "source_id": "results", + "condition_question_reasons": { + "condition": {"q2": "repair explicitly approved"} + }, + "authority": "benchmark owner", + "purpose": "repair damaged prediction", + "executable": True, + "input_evidence_digests": {"results": SHA_A}, + "expected_snapshot_digests": {"q2": SHA_B}, + } + ] + raw["overlays"] = [ + { + "overlay_id": "overlay", + "base_run_path": str(tmp_path / "base-run"), + "base_condition_digest": SHA_A, + "base_realization_id": "realization_base", + "base_realization_digest": SHA_B, + "base_evidence_digests": {"results": SHA_A}, + "base_validation_artifact_sha256": SHA_C, + "base_result_sha256": None, + "source_id": "results", + "authorization_id": "auth", + "replacement_reasons": {"q2": "malformed producer row"}, + "result_origin": { + "derivation_origin": "repair_overlay", + "default_prediction_origin": None, + "per_question_prediction_origins": { + "q2": "external_repair_inference" + }, + }, + "lineage_notes": {}, + "implementation": {"name": "approved repair"}, + "input_digest": SHA_A, + "preownership_output_digest": SHA_B, + "expected_evidence_status": "qualified", + } + ] + return _load(tmp_path, raw) + + +def test_loads_valid_minimal_yaml(minimal_spec): + assert minimal_spec.schema_version == "choicebench.import-spec.v1" + assert minimal_spec.sources[0].path.is_absolute() + assert minimal_spec.sources[0].numeric_columns[0].value_type == "float" + assert minimal_spec.sources[0].option_mapping.mode == "ordered_columns" + assert minimal_spec.metrics == ("accuracy",) + + +def test_safe_loader_rejects_non_mapping_and_python_object_tag(tmp_path: Path): + path = tmp_path / "bad.yaml" + path.write_text("- not\n- a\n- mapping\n", encoding="utf-8") + with pytest.raises(ImportSpecError, match="top level"): + load_import_spec(path) + + path.write_text("!!python/object/apply:os.system [['echo unsafe']]\n", encoding="utf-8") + with pytest.raises(ImportSpecError, match="YAML"): + load_import_spec(path) + + +@pytest.mark.parametrize( + ("section", "unknown_key"), + [ + (None, "output_root"), + ("sources", "output_path"), + ("dialect", "sniff"), + ("numeric_columns", "coerce"), + ("option_mapping", "labels"), + ("datasets", "paper_name"), + ("models", "endpoint"), + ("methods", "entry_point"), + ("prompts", "prompt_path"), + ("conditions", "import_state"), + ("result_origin", "producer"), + ("authorizations", "approved_at"), + ("overlays", "output_path"), + ], +) +def test_rejects_unknown_keys_at_every_schema_layer( + tmp_path: Path, minimal_raw: dict, section: str | None, unknown_key: str +): + raw = deepcopy(minimal_raw) + if section is None: + raw[unknown_key] = "unsafe" + elif section == "dialect": + raw["sources"][0]["dialect"][unknown_key] = True + elif section == "numeric_columns": + raw["sources"][0]["numeric_columns"][0][unknown_key] = True + elif section == "option_mapping": + raw["sources"][0]["option_mapping"][unknown_key] = [] + elif section == "result_origin": + raw["conditions"][0]["result_origin"][unknown_key] = "historical" + elif section == "authorizations": + raw["authorizations"] = deepcopy( + minimal_overlay_raw(raw)["authorizations"] + ) + raw["authorizations"][0][unknown_key] = "later" + elif section == "overlays": + overlay_raw = minimal_overlay_raw(raw) + raw["authorizations"] = overlay_raw["authorizations"] + raw["overlays"] = overlay_raw["overlays"] + raw["overlays"][0][unknown_key] = "unsafe" + else: + raw[section][0][unknown_key] = "unsafe" + with pytest.raises(ImportSpecError, match="Unknown field"): + _load(tmp_path, raw) + + +def minimal_overlay_raw(raw: dict) -> dict: + result = deepcopy(raw) + result["authorizations"] = [ + { + "authorization_id": "auth", + "authorization_type": "inference_repair", + "source_id": "results", + "condition_question_reasons": {"condition": {"q2": "approved"}}, + "authority": "owner", + "purpose": "repair", + "executable": True, + "input_evidence_digests": {"results": SHA_A}, + "expected_snapshot_digests": {"q2": SHA_B}, + } + ] + result["overlays"] = [ + { + "overlay_id": "overlay", + "base_run_path": "/audit/base", + "base_condition_digest": SHA_A, + "base_realization_id": "realization_base", + "base_realization_digest": SHA_B, + "base_evidence_digests": {"results": SHA_A}, + "base_validation_artifact_sha256": SHA_C, + "base_result_sha256": None, + "source_id": "results", + "authorization_id": "auth", + "replacement_reasons": {"q2": "damaged"}, + "result_origin": { + "derivation_origin": "repair_overlay", + "default_prediction_origin": None, + "per_question_prediction_origins": { + "q2": "external_repair_inference" + }, + }, + "lineage_notes": {}, + "implementation": {"name": "repair"}, + "input_digest": SHA_A, + "preownership_output_digest": SHA_B, + "expected_evidence_status": "qualified", + } + ] + return result + + +@pytest.mark.parametrize( + ("path", "value"), + [ + (("sources", 0, "dialect", "double_quote"), 1), + (("sources", 0, "numeric_columns", 0, "null_allowed"), "false"), + (("conditions", 0, "seed"), True), + (("conditions", 0, "executable"), 0), + (("sources", 0, "notes", "temperature"), float("nan")), + (("models", 0, "effective_parameters", "temperature"), float("inf")), + ], +) +def test_rejects_coerced_booleans_integers_and_nonfinite_numbers( + tmp_path: Path, minimal_raw: dict, path: tuple, value: object +): + raw = deepcopy(minimal_raw) + target = raw + for part in path[:-1]: + target = target[part] + target[path[-1]] = value + with pytest.raises(ImportSpecError, match="must be|finite"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + ("path", "value"), + [ + (("sources", 0, "expected_sha256"), "abc"), + (("conditions", 0, "evidence_status"), "verified"), + (("conditions", 0, "scope_disposition"), "paper"), + (("conditions", 0, "result_origin", "derivation_origin"), "native"), + (("sources", 0, "classification"), "trusted"), + ], +) +def test_rejects_invalid_sha256_and_enums( + tmp_path: Path, minimal_raw: dict, path: tuple, value: str +): + raw = deepcopy(minimal_raw) + target = raw + for part in path[:-1]: + target = target[part] + target[path[-1]] = value + with pytest.raises(ImportSpecError): + _load(tmp_path, raw) + + +@pytest.mark.parametrize("credential_key", ["api_key", "nested_access_token", "password"]) +def test_rejects_recursive_credential_named_keys( + tmp_path: Path, minimal_raw: dict, credential_key: str +): + raw = deepcopy(minimal_raw) + raw["sources"][0]["notes"] = {"outer": [{"inner": {credential_key: "secret"}}]} + with pytest.raises(ImportSpecError, match="credential-named"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + "metric", + [ + "package.metrics:UnsafeMetric", + "package.metrics.UnsafeMetric", + "entry-point://unsafe", + "/tmp/metric.py:Metric", + "./metric.py", + ], +) +def test_import_metrics_are_closed_to_builtin_registry( + tmp_path: Path, minimal_raw: dict, metric: str +): + raw = deepcopy(minimal_raw) + raw["metrics"] = [metric] + with pytest.raises(ImportSpecError, match="built-in"): + _load(tmp_path, raw) + + assert set(BUILTIN_METRICS) + + +@pytest.mark.parametrize( + ("field", "value"), + [ + ("delimiter", "::"), + ("delimiter", "é"), + ("quote_character", ","), + ("escape_character", '"'), + ("encoding", "utf-8-sig"), + ("decoding_errors", "replace"), + ], +) +def test_csv_dialect_is_fixed_utf8_strict_and_uses_distinct_ascii_bytes( + tmp_path: Path, minimal_raw: dict, field: str, value: object +): + raw = deepcopy(minimal_raw) + raw["sources"][0]["dialect"][field] = value + with pytest.raises(ImportSpecError): + _load(tmp_path, raw) + + +def test_csv_dialect_accepts_distinct_ascii_escape(tmp_path: Path, minimal_raw: dict): + raw = deepcopy(minimal_raw) + raw["sources"][0]["dialect"].update( + {"delimiter": ";", "quote_character": "'", "escape_character": "\\"} + ) + spec = _load(tmp_path, raw) + assert spec.sources[0].dialect.escape_character == "\\" + + +def test_structured_json_options_and_reject_unmapped_policy( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + source = raw["sources"][0] + source["expected_columns"].append("options_json") + source["option_mapping"] = { + "mode": "structured_json", + "ordered_columns": [], + "structured_column": "options_json", + "structured_label_key": "label", + "structured_text_key": "text", + } + source["extra_field_policy"] = "reject_unmapped" + spec = _load(tmp_path, raw) + assert spec.sources[0].option_mapping.structured_column == "options_json" + assert spec.sources[0].extra_field_policy == "reject_unmapped" + + +@pytest.mark.parametrize( + "option_mapping", + [ + { + "mode": "ordered_columns", + "ordered_columns": [], + "structured_column": None, + "structured_label_key": None, + "structured_text_key": None, + }, + { + "mode": "ordered_columns", + "ordered_columns": ["choice_a"], + "structured_column": "options_json", + "structured_label_key": None, + "structured_text_key": None, + }, + { + "mode": "structured_json", + "ordered_columns": [], + "structured_column": "options_json", + "structured_label_key": None, + "structured_text_key": "text", + }, + { + "mode": "structured_json", + "ordered_columns": ["choice_a"], + "structured_column": "options_json", + "structured_label_key": "label", + "structured_text_key": "text", + }, + ], +) +def test_rejects_incomplete_or_contradictory_option_declarations( + tmp_path: Path, minimal_raw: dict, option_mapping: dict +): + raw = deepcopy(minimal_raw) + raw["sources"][0]["option_mapping"] = option_mapping + with pytest.raises(ImportSpecError, match="option_mapping"): + _load(tmp_path, raw) + + +def test_rejects_duplicate_numeric_column_rules(tmp_path: Path, minimal_raw: dict): + raw = deepcopy(minimal_raw) + raw["sources"][0]["numeric_columns"].append( + deepcopy(raw["sources"][0]["numeric_columns"][0]) + ) + with pytest.raises(ImportSpecError, match="Duplicate numeric"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + ("section", "id_field"), + [ + ("sources", "source_id"), + ("datasets", "dataset_id"), + ("models", "model_key"), + ("methods", "method_key"), + ("prompts", "prompt_key"), + ("conditions", "condition_key"), + ], +) +def test_rejects_duplicate_declared_ids( + tmp_path: Path, minimal_raw: dict, section: str, id_field: str +): + raw = deepcopy(minimal_raw) + raw[section].append(deepcopy(raw[section][0])) + assert raw[section][0][id_field] == raw[section][1][id_field] + with pytest.raises(ImportSpecError, match="Duplicate"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + ("path", "value"), + [ + (("datasets", 0, "source_ids"), ["missing"]), + (("datasets", 0, "selection_source_id"), "missing"), + (("conditions", 0, "dataset_id"), "missing"), + (("conditions", 0, "model_key"), "missing"), + (("conditions", 0, "method_key"), "missing"), + (("conditions", 0, "prompt_key"), "missing"), + ], +) +def test_rejects_missing_references( + tmp_path: Path, minimal_raw: dict, path: tuple, value: object +): + raw = deepcopy(minimal_raw) + target = raw + for part in path[:-1]: + target = target[part] + target[path[-1]] = value + with pytest.raises(ImportSpecError, match="reference|unknown"): + _load(tmp_path, raw) + + +def test_unknown_provenance_is_explicit_null_with_reason( + tmp_path: Path, minimal_raw: dict +): + spec = _load(tmp_path, minimal_raw) + assert spec.provenance["producer_request_id"] == { + "value": None, + "reason": "not recorded by producer", + } + + for bad in ( + None, + {"value": None}, + {"value": None, "reason": ""}, + {"value": None, "reason": "unknown", "guess": "request-1"}, + ): + raw = deepcopy(minimal_raw) + raw["provenance"]["producer_request_id"] = bad + with pytest.raises(ImportSpecError, match="provenance"): + _load(tmp_path, raw) + + +def test_rejects_paths_in_stable_provenance(tmp_path: Path, minimal_raw: dict): + raw = deepcopy(minimal_raw) + raw["provenance"]["producer_path"] = { + "value": "/home/producer/results.csv", + "reason": None, + } + with pytest.raises(ImportSpecError, match="path"): + _load(tmp_path, raw) + + +def _native_compatibility_payloads() -> dict[str, dict]: + artifact_payload = { + "spec": { + "benchmark": "toy", + "split": "test", + "hf_path": None, + "hf_subset": None, + "source_revision": None, + "normalization_version": "2", + "transforms": [], + "output_name": "toy", + }, + "content_digest": SHA_A, + "source": {}, + } + selection_payload = { + "artifact_id": short_id("ds", artifact_payload), + "content_digest": SHA_B, + "sample_identities": [SHA_A, SHA_B], + "seed": 7, + "n_samples": None, + "subject_filter": [], + } + dataset_identity = { + "artifact_payload": artifact_payload, + "artifact_digest": integrity_digest(artifact_payload), + "artifact_id": short_id("ds", artifact_payload), + "selection_payload": selection_payload, + "selection_digest": integrity_digest(selection_payload), + "selection_id": short_id("sel", selection_payload), + } + model_payload = {"backend": "dummy", "model_name_or_path": "dummy-model"} + method_payload = { + "name": "direct_mcq", + "effective_params": {}, + "preflight": None, + "implementation": { + "qualified_name": "choicebench.methods.direct_mcq:DirectMCQRunner", + }, + } + prompt_payload = { + "version": "v1", + "files": { + "direct_mcq": { + "sha256": integrity_digest("prompt"), + "content": "prompt", + } + }, + } + return { + "datasets": dataset_identity, + "models": _native_identity("model", model_payload), + "methods": _native_identity("method", method_payload), + "prompts": _native_identity("prompt", prompt_payload), + } + + +def test_validates_exact_native_compatibility_identities( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + identities = _native_compatibility_payloads() + for section, identity in identities.items(): + raw[section][0]["native_compatibility_identity"] = identity + spec = _load(tmp_path, raw) + assert spec.models[0].native_compatibility_identity["model_id"].startswith( + "model_" + ) + + +@pytest.mark.parametrize( + "payload", + [ + { + "backend": "api", + "provider": "openai", + "model_name_or_path": "gpt-test", + "base_url": None, + "generation_kwargs": { + "max_new_tokens": 32, + "temperature": 0.0, + "untrusted": True, + }, + }, + { + "backend": "huggingface", + "model": { + "kind": "huggingface-hub", + "repo_id": "org/model", + "requested_revision": None, + "resolved_commit": "a" * 40, + "untrusted": True, + }, + "device": "cuda", + "add_bos_token": True, + "generation_kwargs": { + "max_new_tokens": 32, + "temperature": 0.0, + "do_sample": False, + }, + "loader": {"trust_remote_code": True, "torch_dtype": "float16"}, + }, + { + "backend": "huggingface", + "model": { + "kind": "local", + "logical_name": "model", + "content_digest": SHA_A, + "file_count": 1, + "total_bytes": 10, + }, + "device": "cuda", + "add_bos_token": True, + "generation_kwargs": { + "max_new_tokens": 32, + "temperature": 0.0, + "do_sample": False, + }, + "loader": {"trust_remote_code": 1, "torch_dtype": "float16"}, + }, + ], +) +def test_rejects_unknown_or_coerced_nested_native_model_fields( + tmp_path: Path, minimal_raw: dict, payload: dict +): + raw = deepcopy(minimal_raw) + raw["models"][0]["native_compatibility_identity"] = _native_identity( + "model", payload + ) + with pytest.raises(ImportSpecError, match="native_compatibility_identity"): + _load(tmp_path, raw) + + +def test_rejects_unknown_native_method_preflight_field( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + payload = deepcopy(_native_compatibility_payloads()["methods"]["payload"]) + payload["preflight"] = { + "source": "benchmark", + "split": "validation", + "n": 100, + "untrusted": True, + } + raw["methods"][0]["native_compatibility_identity"] = _native_identity( + "method", payload + ) + with pytest.raises(ImportSpecError, match="native_compatibility_identity"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize("section", ["datasets", "models", "methods", "prompts"]) +@pytest.mark.parametrize("mutation", ["unknown", "partial", "mismatched_digest", "mismatched_id"]) +def test_rejects_unverified_native_compatibility_claims( + tmp_path: Path, minimal_raw: dict, section: str, mutation: str +): + raw = deepcopy(minimal_raw) + identity = deepcopy(_native_compatibility_payloads()[section]) + if mutation == "unknown": + identity["claimed_native_id"] = "native" + elif mutation == "partial": + identity.pop(next(iter(identity))) + elif mutation == "mismatched_digest": + digest_key = "artifact_digest" if section == "datasets" else "digest" + identity[digest_key] = SHA_C + else: + id_key = { + "datasets": "selection_id", + "models": "model_id", + "methods": "method_id", + "prompts": "prompt_id", + }[section] + identity[id_key] = "claimed_arbitrarily" + raw[section][0]["native_compatibility_identity"] = identity + with pytest.raises(ImportSpecError, match="native_compatibility_identity"): + _load(tmp_path, raw) + + +def test_audit_locations_do_not_change_import_spec_digest( + minimal_spec, minimal_overlay_spec +): + moved = replace( + minimal_spec, + audit={"source_path": "/different/host", "imported_at": "later"}, + ) + assert import_spec_digest(moved) == import_spec_digest(minimal_spec) + + moved_source = replace( + minimal_spec, + sources=( + replace( + minimal_spec.sources[0], path=Path("/other/freeze/source.csv") + ), + ), + ) + assert import_spec_digest(moved_source) == import_spec_digest(minimal_spec) + + moved_base = replace( + minimal_overlay_spec, + overlays=( + replace( + minimal_overlay_spec.overlays[0], + base_run_path=Path("/other/home/runs/base"), + ), + ), + ) + assert import_spec_digest(moved_base) == import_spec_digest( + minimal_overlay_spec + ) + + +def mutate_identity_field(spec, change: str): + source = spec.sources[0] + condition = spec.conditions[0] + if change == "mapping": + return replace(spec, sources=(replace(source, columns={**source.columns, "x": "y"}),)) + if change == "dialect": + return replace( + spec, + sources=(replace(source, dialect=replace(source.dialect, delimiter=";")),), + ) + if change == "numeric_policy": + numeric = source.numeric_columns[0] + return replace( + spec, + sources=( + replace( + source, + numeric_columns=(replace(numeric, null_allowed=False),), + ), + ), + ) + if change == "option_policy": + option = source.option_mapping + return replace( + spec, + sources=( + replace( + source, + option_mapping=replace( + option, ordered_columns=tuple(reversed(option.ordered_columns)) + ), + ), + ), + ) + if change == "extra_field_policy": + return replace( + spec, + sources=(replace(source, extra_field_policy="reject_unmapped"),), + ) + if change == "status": + return replace( + spec, + conditions=(replace(condition, evidence_status="qualified"),), + ) + if change == "scope": + return replace( + spec, + conditions=( + replace(condition, scope_disposition="excluded_from_paper_matrix"), + ), + ) + raise AssertionError(change) + + +@pytest.mark.parametrize( + "change", + [ + "mapping", + "dialect", + "numeric_policy", + "option_policy", + "extra_field_policy", + "status", + "scope", + ], +) +def test_stable_mapping_fields_change_import_spec_digest(minimal_spec, change): + changed = mutate_identity_field(minimal_spec, change) + assert import_spec_digest(changed) != import_spec_digest(minimal_spec) + + +def test_stable_projection_is_json_native_and_has_no_audit_locations( + minimal_overlay_spec, +): + projection = stable_import_projection(minimal_overlay_spec) + assert "audit" not in projection + assert "path" not in projection["sources"][0] + assert "base_run_path" not in projection["overlays"][0] From 439220e404973fd1494f452424875ffdfd36b68d Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sat, 18 Jul 2026 23:43:51 +0300 Subject: [PATCH 08/47] fix: tighten external import specification --- src/choicebench/importing/schema.py | 232 +++++++++++++++++++--- tests/importing/test_schema.py | 290 ++++++++++++++++++++++++++-- 2 files changed, 477 insertions(+), 45 deletions(-) diff --git a/src/choicebench/importing/schema.py b/src/choicebench/importing/schema.py index eba63fb..083ca73 100644 --- a/src/choicebench/importing/schema.py +++ b/src/choicebench/importing/schema.py @@ -555,6 +555,26 @@ def _string_mapping(value: Any, where: str) -> dict[str, str]: } +def _validate_unknown_reasons( + values: Mapping[str, Any], reasons: Mapping[str, str], where: str +) -> None: + unknown = sorted(set(reasons) - set(values)) + if unknown: + raise ImportSpecError(f"{where} contains unknown reason key(s) {unknown}.") + missing = sorted(key for key, value in values.items() if value is None and key not in reasons) + if missing: + raise ImportSpecError( + f"{where} must document every unknown nullable field; missing {missing}." + ) + contradictory = sorted( + key for key, value in values.items() if value is not None and key in reasons + ) + if contradictory: + raise ImportSpecError( + f"{where} gives unknown reasons for known field(s) {contradictory}." + ) + + def _sha_mapping(value: Any, where: str) -> dict[str, str]: raw = _mapping(value, where) return { @@ -806,6 +826,36 @@ def _build_source(value: Any, index: int) -> SourceArtifactSpec: raise ImportSpecError( f"{where}.ignored_columns references unknown expected columns {missing_ignored}." ) + mapped_values = list(columns.values()) + if len(mapped_values) != len(set(mapped_values)): + raise ImportSpecError( + f"{where}.columns assigns one source column to multiple mapped dispositions." + ) + mapped_columns = set(mapped_values) + ignored_set = set(ignored_columns) + conflicts = sorted( + (mapped_columns & option_columns) + | (mapped_columns & ignored_set) + | (option_columns & ignored_set) + ) + if conflicts: + raise ImportSpecError( + f"{where} source column disposition conflicts for {conflicts}; mapped, " + "option, and ignored columns must be pairwise disjoint." + ) + extra_field_policy = _enum( + raw["extra_field_policy"], + {"preserve_unmapped", "reject_unmapped"}, + f"{where}.extra_field_policy", + ) + disposed_columns = mapped_columns | option_columns | ignored_set + if extra_field_policy == "reject_unmapped": + undisposed = sorted(set(expected_columns) - disposed_columns) + if undisposed: + raise ImportSpecError( + f"{where} reject_unmapped requires one explicit disposition per source " + f"column; missing {undisposed}." + ) path_text = _nonempty(raw["path"], f"{where}.path") return SourceArtifactSpec( source_id=_nonempty(raw["source_id"], f"{where}.source_id"), @@ -831,11 +881,7 @@ def _build_source(value: Any, index: int) -> SourceArtifactSpec: ), numeric_columns=numeric_columns, option_mapping=option_mapping, - extra_field_policy=_enum( - raw["extra_field_policy"], - {"preserve_unmapped", "reject_unmapped"}, - f"{where}.extra_field_policy", - ), + extra_field_policy=extra_field_policy, preserve_namespace=_nonempty( raw["preserve_namespace"], f"{where}.preserve_namespace" ), @@ -1087,19 +1133,40 @@ def _validate_method_payload(value: Any, where: str) -> dict[str, Any]: def _validate_prompt_payload(value: Any, where: str) -> dict[str, Any]: raw = _exact_keys(value, {"version", "files"}, where) files_raw = _mapping(raw["files"], f"{where}.files") - if not files_raw: - raise ImportSpecError(f"{where}.files must not be empty.") + template_names = {"direct_mcq", "free_text", "option_matching"} + if set(files_raw) != template_names: + raise ImportSpecError( + f"{where}.files must contain exactly the current template names " + f"{sorted(template_names)}." + ) files: dict[str, Any] = {} - for name, item in files_raw.items(): + for name in ("direct_mcq", "free_text", "option_matching"): + item = files_raw[name] record = _exact_keys(item, {"sha256", "content"}, f"{where}.files.{name}") - content = _nonempty(record["content"], f"{where}.files.{name}.content") + content = record["content"] + if not isinstance(content, str): + raise ImportSpecError(f"{where}.files.{name}.content must be a string.") digest = _sha256(record["sha256"], f"{where}.files.{name}.sha256") if digest != integrity_digest(content): raise ImportSpecError(f"{where}.files.{name}.sha256 does not match content.") - files[_nonempty(name, f"{where}.files key")] = {"sha256": digest, "content": content} + files[name] = {"sha256": digest, "content": content} return {"version": _nonempty(raw["version"], f"{where}.version"), "files": files} +def _validate_prompt_native(value: Any, where: str) -> dict[str, Any]: + raw = _exact_keys(value, {"prompt_id", "version", "files"}, where) + payload = _validate_prompt_payload( + {"version": raw["version"], "files": raw["files"]}, where + ) + prompt_id = _nonempty(raw["prompt_id"], f"{where}.prompt_id") + expected_id = f"prompt_{integrity_digest(payload)[:16]}" + if prompt_id != expected_id: + raise ImportSpecError( + f"{where}.prompt_id does not match prompt_bundle_identity semantics." + ) + return {"prompt_id": prompt_id, **payload} + + def _validate_simple_native( value: Any, where: str, prefix: str, payload_validator ) -> dict[str, Any]: @@ -1121,6 +1188,21 @@ def _build_dataset(value: Any, index: int) -> DatasetReferenceSpec: where = f"datasets[{index}]" raw = _exact_keys(value, _DATASET_KEYS, where) native = raw["native_compatibility_identity"] + selection_seed = _optional_int(raw["selection_seed"], f"{where}.selection_seed") + selection_n_samples = _optional_int( + raw["selection_n_samples"], f"{where}.selection_n_samples", minimum=1 + ) + selection_unknown_reasons = _string_mapping( + raw["selection_unknown_reasons"], f"{where}.selection_unknown_reasons" + ) + _validate_unknown_reasons( + { + "selection_seed": selection_seed, + "selection_n_samples": selection_n_samples, + }, + selection_unknown_reasons, + f"{where}.selection_unknown_reasons", + ) return DatasetReferenceSpec( dataset_id=_nonempty(raw["dataset_id"], f"{where}.dataset_id"), benchmark_name=_nonempty(raw["benchmark_name"], f"{where}.benchmark_name"), @@ -1134,10 +1216,10 @@ def _build_dataset(value: Any, index: int) -> DatasetReferenceSpec: source_ids=_strings(raw["source_ids"], f"{where}.source_ids", nonempty=True, unique=True), selection_source_id=_nonempty(raw["selection_source_id"], f"{where}.selection_source_id"), expected_question_ids=_strings(raw["expected_question_ids"], f"{where}.expected_question_ids", nonempty=True, unique=True), - selection_seed=_optional_int(raw["selection_seed"], f"{where}.selection_seed"), - selection_n_samples=_optional_int(raw["selection_n_samples"], f"{where}.selection_n_samples", minimum=1), + selection_seed=selection_seed, + selection_n_samples=selection_n_samples, subject_filter=_strings(raw["subject_filter"], f"{where}.subject_filter", unique=True), - selection_unknown_reasons=_string_mapping(raw["selection_unknown_reasons"], f"{where}.selection_unknown_reasons"), + selection_unknown_reasons=selection_unknown_reasons, columns=_string_mapping(raw["columns"], f"{where}.columns"), revision=_optional_string(raw["revision"], f"{where}.revision"), fingerprint=_optional_string(raw["fingerprint"], f"{where}.fingerprint"), @@ -1153,14 +1235,23 @@ def _build_model(value: Any, index: int) -> ImportModelSpec: where = f"models[{index}]" raw = _exact_keys(value, _MODEL_KEYS, where) native = raw["native_compatibility_identity"] + backend = _optional_string(raw["backend"], f"{where}.backend") + provider = _optional_string(raw["provider"], f"{where}.provider") + revision = _optional_string(raw["revision"], f"{where}.revision") + unknown_reasons = _string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons") + _validate_unknown_reasons( + {"backend": backend, "provider": provider, "revision": revision}, + unknown_reasons, + f"{where}.unknown_reasons", + ) return ImportModelSpec( model_key=_nonempty(raw["model_key"], f"{where}.model_key"), display_name=_nonempty(raw["display_name"], f"{where}.display_name"), - backend=_optional_string(raw["backend"], f"{where}.backend"), - provider=_optional_string(raw["provider"], f"{where}.provider"), - revision=_optional_string(raw["revision"], f"{where}.revision"), + backend=backend, + provider=provider, + revision=revision, effective_parameters=_canonical_mapping(raw["effective_parameters"], f"{where}.effective_parameters"), - unknown_reasons=_string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons"), + unknown_reasons=unknown_reasons, native_compatibility_identity=( None if native is None else _validate_simple_native( native, f"{where}.native_compatibility_identity", "model", _validate_model_payload @@ -1174,12 +1265,23 @@ def _build_method(value: Any, index: int) -> ImportMethodSpec: raw = _exact_keys(value, _METHOD_KEYS, where) implementation = raw["implementation"] native = raw["native_compatibility_identity"] + implementation = ( + None + if implementation is None + else _validate_implementation(implementation, f"{where}.implementation") + ) + unknown_reasons = _string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons") + _validate_unknown_reasons( + {"implementation": implementation}, + unknown_reasons, + f"{where}.unknown_reasons", + ) return ImportMethodSpec( method_key=_nonempty(raw["method_key"], f"{where}.method_key"), name=_nonempty(raw["name"], f"{where}.name"), effective_parameters=_canonical_mapping(raw["effective_parameters"], f"{where}.effective_parameters"), - implementation=(None if implementation is None else _validate_implementation(implementation, f"{where}.implementation")), - unknown_reasons=_string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons"), + implementation=implementation, + unknown_reasons=unknown_reasons, native_compatibility_identity=( None if native is None else _validate_simple_native( native, f"{where}.native_compatibility_identity", "method", _validate_method_payload @@ -1193,16 +1295,39 @@ def _build_prompt(value: Any, index: int) -> ImportPromptSpec: raw = _exact_keys(value, _PROMPT_KEYS, where) contents = raw["template_contents"] native = raw["native_compatibility_identity"] + template_identity = _optional_string( + raw["template_identity"], f"{where}.template_identity" + ) + template_digest = _optional_sha256( + raw["template_digest"], f"{where}.template_digest" + ) + template_contents = ( + None + if contents is None + else _string_mapping(contents, f"{where}.template_contents") + ) + unknown_reason = _optional_string(raw["unknown_reason"], f"{where}.unknown_reason") + has_unknown = any( + item is None for item in (template_identity, template_digest, template_contents) + ) + if has_unknown and unknown_reason is None: + raise ImportSpecError( + f"{where}.unknown_reason must explain unrecoverable prompt fields." + ) + if not has_unknown and unknown_reason is not None: + raise ImportSpecError( + f"{where}.unknown_reason contradicts fully known prompt fields." + ) return ImportPromptSpec( prompt_key=_nonempty(raw["prompt_key"], f"{where}.prompt_key"), - template_identity=_optional_string(raw["template_identity"], f"{where}.template_identity"), - template_digest=_optional_sha256(raw["template_digest"], f"{where}.template_digest"), - template_contents=(None if contents is None else _string_mapping(contents, f"{where}.template_contents")), - unknown_reason=_optional_string(raw["unknown_reason"], f"{where}.unknown_reason"), + template_identity=template_identity, + template_digest=template_digest, + template_contents=template_contents, + unknown_reason=unknown_reason, native_compatibility_identity=( - None if native is None else _validate_simple_native( - native, f"{where}.native_compatibility_identity", "prompt", _validate_prompt_payload - ) + None + if native is None + else _validate_prompt_native(native, f"{where}.native_compatibility_identity") ), ) @@ -1229,6 +1354,23 @@ def _build_result_origin(value: Any, where: str) -> ResultOriginSpec: def _build_condition(value: Any, index: int) -> ImportConditionSpec: where = f"conditions[{index}]" raw = _exact_keys(value, _CONDITION_KEYS, where) + seed = _optional_int(raw["seed"], f"{where}.seed") + calibration_identity = _canonical_mapping_or_none( + raw["calibration_identity"], f"{where}.calibration_identity" + ) + preflight_identity = _canonical_mapping_or_none( + raw["preflight_identity"], f"{where}.preflight_identity" + ) + unknown_reasons = _string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons") + _validate_unknown_reasons( + { + "seed": seed, + "calibration_identity": calibration_identity, + "preflight_identity": preflight_identity, + }, + unknown_reasons, + f"{where}.unknown_reasons", + ) return ImportConditionSpec( condition_key=_nonempty(raw["condition_key"], f"{where}.condition_key"), source_ids=_strings(raw["source_ids"], f"{where}.source_ids", nonempty=True, unique=True), @@ -1236,12 +1378,12 @@ def _build_condition(value: Any, index: int) -> ImportConditionSpec: model_key=_nonempty(raw["model_key"], f"{where}.model_key"), method_key=_nonempty(raw["method_key"], f"{where}.method_key"), prompt_key=_nonempty(raw["prompt_key"], f"{where}.prompt_key"), - seed=_optional_int(raw["seed"], f"{where}.seed"), - calibration_identity=_canonical_mapping_or_none(raw["calibration_identity"], f"{where}.calibration_identity"), - preflight_identity=_canonical_mapping_or_none(raw["preflight_identity"], f"{where}.preflight_identity"), + seed=seed, + calibration_identity=calibration_identity, + preflight_identity=preflight_identity, protocol_settings=_canonical_mapping(raw["protocol_settings"], f"{where}.protocol_settings"), generation_parameters=_canonical_mapping(raw["generation_parameters"], f"{where}.generation_parameters"), - unknown_reasons=_string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons"), + unknown_reasons=unknown_reasons, expected_question_ids=_strings(raw["expected_question_ids"], f"{where}.expected_question_ids", nonempty=True, unique=True), evidence_status=_enum(raw["evidence_status"], _EVIDENCE_STATUSES, f"{where}.evidence_status"), scope_disposition=_enum(raw["scope_disposition"], _SCOPE_DISPOSITIONS, f"{where}.scope_disposition"), @@ -1264,14 +1406,26 @@ def _build_authorization(value: Any, index: int) -> AuthorizationSpec: ) for condition, question_reasons in reasons_raw.items() } + authorization_type = _enum( + raw["authorization_type"], + {"inference_repair", "offline_transformation"}, + f"{where}.authorization_type", + ) + executable = _strict_bool(raw["executable"], f"{where}.executable") + expected_executable = authorization_type == "inference_repair" + if executable is not expected_executable: + raise ImportSpecError( + f"{where}.authorization_type={authorization_type!r} requires " + f"executable={expected_executable!r}." + ) return AuthorizationSpec( authorization_id=_nonempty(raw["authorization_id"], f"{where}.authorization_id"), - authorization_type=_enum(raw["authorization_type"], {"inference_repair", "offline_transformation"}, f"{where}.authorization_type"), + authorization_type=authorization_type, source_id=_nonempty(raw["source_id"], f"{where}.source_id"), condition_question_reasons=reasons, authority=_nonempty(raw["authority"], f"{where}.authority"), purpose=_nonempty(raw["purpose"], f"{where}.purpose"), - executable=_strict_bool(raw["executable"], f"{where}.executable"), + executable=executable, input_evidence_digests=_sha_mapping(raw["input_evidence_digests"], f"{where}.input_evidence_digests"), expected_snapshot_digests=_sha_mapping(raw["expected_snapshot_digests"], f"{where}.expected_snapshot_digests"), ) @@ -1368,6 +1522,20 @@ def _require_references(spec: ImportSpec) -> None: raise ImportSpecError( f"Condition {condition.condition_key!r} has origins for unknown question IDs." ) + evaluable = ( + condition.evidence_status in {"complete", "qualified"} + and condition.scope_disposition == "included" + ) + if ( + evaluable + and condition.result_origin.default_prediction_origin is None + and origin_questions != expected + ): + missing_origins = sorted(expected - origin_questions) + raise ImportSpecError( + f"Condition {condition.condition_key!r} has evaluable rows without a " + f"prediction origin: {missing_origins}." + ) for authorization in spec.authorizations: if authorization.source_id not in source_ids: diff --git a/tests/importing/test_schema.py b/tests/importing/test_schema.py index d47c85b..ef8c13e 100644 --- a/tests/importing/test_schema.py +++ b/tests/importing/test_schema.py @@ -15,6 +15,7 @@ stable_import_projection, ) from choicebench.metrics import BUILTIN_METRICS +from choicebench.pipeline.prompt_builder import prompt_bundle_identity SHA_A = "a" * 64 @@ -179,7 +180,11 @@ def minimal_raw(tmp_path: Path) -> dict: "preflight_identity": None, "protocol_settings": {}, "generation_parameters": {}, - "unknown_reasons": {"seed": "not recorded by producer"}, + "unknown_reasons": { + "seed": "not recorded by producer", + "calibration_identity": "not recorded by producer", + "preflight_identity": "not recorded by producer", + }, "expected_question_ids": ["q1", "q2"], "evidence_status": "complete", "scope_disposition": "included", @@ -501,11 +506,58 @@ def test_structured_json_options_and_reject_unmapped_policy( "structured_text_key": "text", } source["extra_field_policy"] = "reject_unmapped" + source["ignored_columns"].update( + { + "question": "reference snapshot owns the question text", + "choice_a": "superseded by structured choices", + "choice_b": "superseded by structured choices", + } + ) spec = _load(tmp_path, raw) assert spec.sources[0].option_mapping.structured_column == "options_json" assert spec.sources[0].extra_field_policy == "reject_unmapped" +@pytest.mark.parametrize( + "conflict", + [ + "duplicate_mapping", + "mapped_option", + "mapped_ignored", + "option_ignored", + "undeclared_strict_column", + ], +) +def test_each_source_column_has_exactly_one_disposition( + tmp_path: Path, minimal_raw: dict, conflict: str +): + raw = deepcopy(minimal_raw) + source = raw["sources"][0] + if conflict == "duplicate_mapping": + source["columns"]["raw_prediction"] = "answer" + elif conflict == "mapped_option": + source["columns"]["prediction"] = "choice_a" + elif conflict == "mapped_ignored": + source["columns"]["prediction"] = "score" + elif conflict == "option_ignored": + source["ignored_columns"]["choice_a"] = "cannot also be an option" + else: + source["extra_field_policy"] = "reject_unmapped" + with pytest.raises(ImportSpecError, match="disposition|source column"): + _load(tmp_path, raw) + + +def test_reject_unmapped_accepts_exactly_partitioned_columns( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + source = raw["sources"][0] + source["extra_field_policy"] = "reject_unmapped" + source["ignored_columns"]["question"] = "question comes from reference snapshot" + spec = _load(tmp_path, raw) + assert spec.sources[0].extra_field_policy == "reject_unmapped" + + @pytest.mark.parametrize( "option_mapping", [ @@ -622,6 +674,169 @@ def test_unknown_provenance_is_explicit_null_with_reason( _load(tmp_path, raw) +@pytest.mark.parametrize( + "case", + [ + "dataset_missing", + "dataset_contradictory", + "dataset_unknown_key", + "model_missing", + "model_contradictory", + "model_unknown_key", + "method_missing", + "method_contradictory", + "prompt_missing", + "prompt_contradictory", + "condition_missing", + "condition_contradictory", + "condition_unknown_key", + ], +) +def test_nullable_scientific_fields_require_exact_unknown_reasons( + tmp_path: Path, minimal_raw: dict, case: str +): + raw = deepcopy(minimal_raw) + if case == "dataset_missing": + raw["datasets"][0]["selection_unknown_reasons"].pop("selection_seed") + elif case == "dataset_contradictory": + raw["datasets"][0]["selection_seed"] = 7 + elif case == "dataset_unknown_key": + raw["datasets"][0]["selection_unknown_reasons"]["revision"] = "unknown" + elif case == "model_missing": + raw["models"][0]["unknown_reasons"].pop("backend") + elif case == "model_contradictory": + raw["models"][0]["backend"] = "dummy" + elif case == "model_unknown_key": + raw["models"][0]["unknown_reasons"]["temperature"] = "unknown" + elif case == "method_missing": + raw["methods"][0]["unknown_reasons"].pop("implementation") + elif case == "method_contradictory": + raw["methods"][0]["implementation"] = { + "qualified_name": "historical:Method" + } + elif case == "prompt_missing": + raw["prompts"][0]["unknown_reason"] = None + elif case == "prompt_contradictory": + raw["prompts"][0].update( + { + "template_identity": "historical-template", + "template_digest": SHA_A, + "template_contents": {"prompt": "contents"}, + } + ) + elif case == "condition_missing": + raw["conditions"][0]["unknown_reasons"].pop("calibration_identity") + elif case == "condition_contradictory": + raw["conditions"][0]["seed"] = 7 + else: + raw["conditions"][0]["unknown_reasons"]["model"] = "unknown" + with pytest.raises(ImportSpecError, match="unknown.reason|unknown_reasons|unknown_reason"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + ("authorization_type", "executable"), + [("inference_repair", True), ("offline_transformation", False)], +) +def test_authorization_type_binds_executability( + tmp_path: Path, + minimal_raw: dict, + authorization_type: str, + executable: bool, +): + raw = deepcopy(minimal_raw) + raw["authorizations"] = [ + { + "authorization_id": "auth", + "authorization_type": authorization_type, + "source_id": "results", + "condition_question_reasons": {"condition": {"q1": "approved"}}, + "authority": "benchmark owner", + "purpose": "bounded operation", + "executable": executable, + "input_evidence_digests": {"results": SHA_A}, + "expected_snapshot_digests": {"q1": SHA_B}, + } + ] + spec = _load(tmp_path, raw) + assert spec.authorizations[0].executable is executable + + +@pytest.mark.parametrize( + ("authorization_type", "executable"), + [("inference_repair", False), ("offline_transformation", True)], +) +def test_rejects_authorization_executability_mismatch( + tmp_path: Path, + minimal_raw: dict, + authorization_type: str, + executable: bool, +): + raw = deepcopy(minimal_raw) + raw["authorizations"] = [ + { + "authorization_id": "auth", + "authorization_type": authorization_type, + "source_id": "results", + "condition_question_reasons": {"condition": {"q1": "approved"}}, + "authority": "benchmark owner", + "purpose": "bounded operation", + "executable": executable, + "input_evidence_digests": {"results": SHA_A}, + "expected_snapshot_digests": {"q1": SHA_B}, + } + ] + with pytest.raises(ImportSpecError, match="authorization_type|executable"): + _load(tmp_path, raw) + + +def test_included_evaluable_condition_requires_prediction_origin_coverage( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + origin = raw["conditions"][0]["result_origin"] + origin["default_prediction_origin"] = None + origin["per_question_prediction_origins"] = { + "q1": "external_historical_inference" + } + with pytest.raises(ImportSpecError, match="prediction origin"): + _load(tmp_path, raw) + + +def test_exact_per_question_origins_cover_evaluable_condition( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + origin = raw["conditions"][0]["result_origin"] + origin["default_prediction_origin"] = None + origin["per_question_prediction_origins"] = { + "q1": "external_historical_inference", + "q2": "external_repair_inference", + } + spec = _load(tmp_path, raw) + assert spec.conditions[0].result_origin.default_prediction_origin is None + + +@pytest.mark.parametrize( + ("status", "scope"), + [ + ("partial", "included"), + ("complete", "excluded_from_paper_matrix"), + ], +) +def test_non_evaluable_evidence_may_omit_prediction_origins( + tmp_path: Path, minimal_raw: dict, status: str, scope: str +): + raw = deepcopy(minimal_raw) + condition = raw["conditions"][0] + condition["evidence_status"] = status + condition["scope_disposition"] = scope + condition["result_origin"]["default_prediction_origin"] = None + condition["result_origin"]["per_question_prediction_origins"] = {} + spec = _load(tmp_path, raw) + assert spec.conditions[0].evidence_status == status + + def test_rejects_paths_in_stable_provenance(tmp_path: Path, minimal_raw: dict): raw = deepcopy(minimal_raw) raw["provenance"]["producer_path"] = { @@ -672,20 +887,11 @@ def _native_compatibility_payloads() -> dict[str, dict]: "qualified_name": "choicebench.methods.direct_mcq:DirectMCQRunner", }, } - prompt_payload = { - "version": "v1", - "files": { - "direct_mcq": { - "sha256": integrity_digest("prompt"), - "content": "prompt", - } - }, - } return { "datasets": dataset_identity, "models": _native_identity("model", model_payload), "methods": _native_identity("method", method_payload), - "prompts": _native_identity("prompt", prompt_payload), + "prompts": prompt_bundle_identity("v1"), } @@ -700,6 +906,61 @@ def test_validates_exact_native_compatibility_identities( assert spec.models[0].native_compatibility_identity["model_id"].startswith( "model_" ) + assert ( + spec.prompts[0].native_compatibility_identity + == prompt_bundle_identity("v1") + ) + + +def test_native_prompt_identity_uses_exact_contents_without_redaction( + tmp_path: Path, minimal_raw: dict +): + prompt_root = tmp_path / "prompts" + version_dir = prompt_root / "credential-shaped-science" + version_dir.mkdir(parents=True) + contents = { + "direct_mcq": "Question: {question}\napi_key=scientific-label\nAnswer:", + "free_text": "Question: {question}\npassword=ordinary-text\nAnswer:", + "option_matching": "Question: {question}\n{options}\nAnswer:", + } + for name, content in contents.items(): + (version_dir / f"{name}.txt").write_text(content, encoding="utf-8") + identity = prompt_bundle_identity("credential-shaped-science", prompt_root) + raw = deepcopy(minimal_raw) + raw["prompts"][0]["native_compatibility_identity"] = identity + spec = _load(tmp_path, raw) + assert spec.prompts[0].native_compatibility_identity == identity + + +@pytest.mark.parametrize("mutation", ["missing_template", "extra_template", "short_id"]) +def test_rejects_non_native_prompt_bundle_shapes( + tmp_path: Path, minimal_raw: dict, mutation: str +): + identity = deepcopy(prompt_bundle_identity("v1")) + if mutation == "missing_template": + identity["files"].pop("free_text") + payload = {"version": identity["version"], "files": identity["files"]} + identity["prompt_id"] = f"prompt_{integrity_digest(payload)[:16]}" + elif mutation == "extra_template": + identity["files"]["unexpected"] = { + "content": "extra", + "sha256": integrity_digest("extra"), + } + payload = {"version": identity["version"], "files": identity["files"]} + identity["prompt_id"] = f"prompt_{integrity_digest(payload)[:16]}" + else: + content = identity["files"]["direct_mcq"]["content"] + content += "\napi_key=scientific-label" + identity["files"]["direct_mcq"] = { + "content": content, + "sha256": integrity_digest(content), + } + payload = {"version": identity["version"], "files": identity["files"]} + identity["prompt_id"] = short_id("prompt", payload) + raw = deepcopy(minimal_raw) + raw["prompts"][0]["native_compatibility_identity"] = identity + with pytest.raises(ImportSpecError, match="native_compatibility_identity"): + _load(tmp_path, raw) @pytest.mark.parametrize( @@ -795,8 +1056,11 @@ def test_rejects_unverified_native_compatibility_claims( elif mutation == "partial": identity.pop(next(iter(identity))) elif mutation == "mismatched_digest": - digest_key = "artifact_digest" if section == "datasets" else "digest" - identity[digest_key] = SHA_C + if section == "prompts": + identity["files"]["direct_mcq"]["sha256"] = SHA_C + else: + digest_key = "artifact_digest" if section == "datasets" else "digest" + identity[digest_key] = SHA_C else: id_key = { "datasets": "selection_id", From 659849603d9c5a991f437beec0c27dd672e81a25 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sat, 18 Jul 2026 23:51:38 +0300 Subject: [PATCH 09/47] fix: accept external implementation identities --- src/choicebench/importing/schema.py | 13 +++++- tests/importing/test_schema.py | 72 +++++++++++++++++++++++++++++ 2 files changed, 84 insertions(+), 1 deletion(-) diff --git a/src/choicebench/importing/schema.py b/src/choicebench/importing/schema.py index 083ca73..cf67300 100644 --- a/src/choicebench/importing/schema.py +++ b/src/choicebench/importing/schema.py @@ -1092,7 +1092,7 @@ def _validate_model_payload(value: Any, where: str) -> dict[str, Any]: _IMPLEMENTATION_KEYS = { "qualified_name", "source_file", "source_digest", "distribution", - "distribution_version", "package_tree_digest", + "distribution_version", "package_tree_digest", "package_file_count", } @@ -1102,9 +1102,20 @@ def _validate_implementation(value: Any, where: str) -> dict[str, Any]: raise ImportSpecError(f"Missing required field 'qualified_name' in {where}.") result = _canonical_mapping(raw, where) _nonempty(result["qualified_name"], f"{where}.qualified_name") + for key in ("source_file", "distribution", "distribution_version"): + if key in result: + _nonempty(result[key], f"{where}.{key}") for key in ("source_digest", "package_tree_digest"): if key in result: _sha256(result[key], f"{where}.{key}") + has_package_digest = "package_tree_digest" in result + has_package_count = "package_file_count" in result + if has_package_digest != has_package_count: + raise ImportSpecError( + f"{where}.package_tree_digest and {where}.package_file_count must appear together." + ) + if has_package_count: + _strict_int(result["package_file_count"], f"{where}.package_file_count", minimum=1) return result diff --git a/tests/importing/test_schema.py b/tests/importing/test_schema.py index ef8c13e..b62e7a3 100644 --- a/tests/importing/test_schema.py +++ b/tests/importing/test_schema.py @@ -16,6 +16,7 @@ ) from choicebench.metrics import BUILTIN_METRICS from choicebench.pipeline.prompt_builder import prompt_bundle_identity +from choicebench.provenance import implementation_identity SHA_A = "a" * 64 @@ -912,6 +913,77 @@ def test_validates_exact_native_compatibility_identities( ) +def test_accepts_external_package_implementation_identity_in_native_method( + tmp_path: Path, minimal_raw: dict +): + implementation = implementation_identity(yaml.YAMLObject) + assert implementation["package_file_count"] > 0 + payload = deepcopy(_native_compatibility_payloads()["methods"]["payload"]) + payload["implementation"] = implementation + identity = _native_identity("method", payload) + raw = deepcopy(minimal_raw) + raw["methods"][0]["native_compatibility_identity"] = identity + + spec = _load(tmp_path, raw) + + assert spec.methods[0].native_compatibility_identity == identity + + +@pytest.mark.parametrize("package_file_count", [0, True]) +def test_rejects_non_positive_or_non_strict_package_file_count( + tmp_path: Path, minimal_raw: dict, package_file_count: object +): + implementation = implementation_identity(yaml.YAMLObject) + implementation["package_file_count"] = package_file_count + payload = deepcopy(_native_compatibility_payloads()["methods"]["payload"]) + payload["implementation"] = implementation + raw = deepcopy(minimal_raw) + raw["methods"][0]["native_compatibility_identity"] = _native_identity( + "method", payload + ) + + with pytest.raises( + ImportSpecError, + match=r"package_file_count.*(?:must be an integer|must be >= 1)", + ): + _load(tmp_path, raw) + + +@pytest.mark.parametrize("missing_key", ["package_tree_digest", "package_file_count"]) +def test_requires_package_tree_digest_and_file_count_together( + tmp_path: Path, minimal_raw: dict, missing_key: str +): + implementation = implementation_identity(yaml.YAMLObject) + implementation.pop(missing_key) + payload = deepcopy(_native_compatibility_payloads()["methods"]["payload"]) + payload["implementation"] = implementation + raw = deepcopy(minimal_raw) + raw["methods"][0]["native_compatibility_identity"] = _native_identity( + "method", payload + ) + + with pytest.raises(ImportSpecError, match="package_tree_digest.*package_file_count.*together"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + "field", ["source_file", "distribution", "distribution_version"] +) +def test_rejects_empty_optional_implementation_string( + tmp_path: Path, minimal_raw: dict, field: str +): + implementation = {"qualified_name": "external:Target", field: ""} + payload = deepcopy(_native_compatibility_payloads()["methods"]["payload"]) + payload["implementation"] = implementation + raw = deepcopy(minimal_raw) + raw["methods"][0]["native_compatibility_identity"] = _native_identity( + "method", payload + ) + + with pytest.raises(ImportSpecError, match=field): + _load(tmp_path, raw) + + def test_native_prompt_identity_uses_exact_contents_without_redaction( tmp_path: Path, minimal_raw: dict ): From 732f8c36dacc49ce611cddee25ec758184828bc0 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 00:12:39 +0300 Subject: [PATCH 10/47] feat: add deterministic CSV source adapter --- src/choicebench/importing/csv_adapter.py | 403 +++++++++++++++++++++ tests/importing/test_csv_adapter.py | 435 +++++++++++++++++++++++ 2 files changed, 838 insertions(+) create mode 100644 src/choicebench/importing/csv_adapter.py create mode 100644 tests/importing/test_csv_adapter.py diff --git a/src/choicebench/importing/csv_adapter.py b/src/choicebench/importing/csv_adapter.py new file mode 100644 index 0000000..c8bb8a5 --- /dev/null +++ b/src/choicebench/importing/csv_adapter.py @@ -0,0 +1,403 @@ +"""Deterministic byte-preserving adapter for external CSV result sources.""" + +from __future__ import annotations + +import csv +from dataclasses import asdict, dataclass +from hashlib import sha256 +import io +import json +import math +from pathlib import Path +import re +from typing import Mapping, Protocol + +from choicebench.identity import canonicalize, is_credential_key +from choicebench.importing.schema import CsvDialectSpec, SourceArtifactSpec + + +_UTF8_BOM = b"\xef\xbb\xbf" +_TERMINATORS = {"crlf": b"\r\n", "lf": b"\n", "cr": b"\r"} +_INTEGER_RE = re.compile(r"[+-]?\d+\Z") +_FLOAT_RE = re.compile( + r"[+-]?(?:(?:\d+(?:\.\d*)?)|(?:\.\d+))(?:[eE][+-]?\d+)?\Z" +) +_NONFINITE_FLOAT_RE = re.compile( + r"[+-]?(?:nan|inf(?:inity)?)\Z", flags=re.IGNORECASE +) + + +class CsvAdapterError(ValueError): + """Raised when source CSV bytes do not satisfy their declaration.""" + + +@dataclass(frozen=True) +class OpenedSource: + source_id: str + audit_path: Path + logical_path: str + data: bytes + sha256: str + + +@dataclass(frozen=True) +class LogicalRecordSpan: + index: int + start: int + end: int + terminator: bytes + + +@dataclass(frozen=True) +class SourceRow: + values: Mapping[str, str | None] + span: LogicalRecordSpan + raw_sha256: str + + +@dataclass(frozen=True) +class AdaptedTable: + columns: tuple[str, ...] + rows: tuple[SourceRow, ...] + source_sha256: str + + +class SourceAdapter(Protocol): + def parse( + self, source: OpenedSource, declaration: SourceArtifactSpec, *, strict: bool + ) -> AdaptedTable: ... + + +def _dialect_bytes(dialect: CsvDialectSpec) -> tuple[int, int, int | None]: + values = (dialect.delimiter, dialect.quote_character, dialect.escape_character) + encoded: list[int | None] = [] + for name, value in zip(("delimiter", "quote", "escape"), values, strict=True): + if value is None: + encoded.append(None) + continue + raw = value.encode("utf-8") + if len(raw) != 1 or raw[0] >= 128 or raw in {b"\x00", b"\r", b"\n"}: + raise CsvAdapterError(f"CSV {name} must be one usable ASCII byte.") + encoded.append(raw[0]) + non_null = [value for value in encoded if value is not None] + if len(non_null) != len(set(non_null)): + raise CsvAdapterError("CSV delimiter, quote, and escape bytes must be distinct.") + return encoded[0], encoded[1], encoded[2] # type: ignore[return-value] + + +def _declared_terminators(dialect: CsvDialectSpec) -> tuple[bytes, ...]: + try: + declared = tuple(_TERMINATORS[name] for name in dialect.line_terminators) + except KeyError as exc: + raise CsvAdapterError(f"Unsupported CSV line terminator {exc.args[0]!r}.") from exc + if not declared or len(declared) != len(set(declared)): + raise CsvAdapterError("CSV line terminators must be a non-empty unique list.") + return tuple(sorted(declared, key=len, reverse=True)) + + +def scan_csv_logical_records( + data: bytes, dialect: CsvDialectSpec +) -> tuple[LogicalRecordSpan, ...]: + """Return exact source-byte spans for logical CSV records.""" + if not isinstance(data, bytes): + raise TypeError(f"data must be bytes; got {type(data).__name__}.") + delimiter, quote, escape = _dialect_bytes(dialect) + terminators = _declared_terminators(dialect) + spans: list[LogicalRecordSpan] = [] + observed_terminators: set[bytes] = set() + start = 0 + if data.startswith(_UTF8_BOM): + if dialect.bom_policy == "forbid": + raise CsvAdapterError("Source contains a forbidden UTF-8 BOM.") + index = len(_UTF8_BOM) + else: + index = 0 + in_quotes = False + at_field_start = True + + while index < len(data): + byte = data[index] + if in_quotes: + if escape is not None and byte == escape: + if index + 1 >= len(data): + raise CsvAdapterError( + f"CSV escape byte at byte {index} has no following byte." + ) + index += 2 + continue + if byte == quote: + if ( + dialect.double_quote + and index + 1 < len(data) + and data[index + 1] == quote + ): + index += 2 + continue + in_quotes = False + index += 1 + continue + + if escape is not None and byte == escape: + if index + 1 >= len(data): + raise CsvAdapterError( + f"CSV escape byte at byte {index} has no following byte." + ) + at_field_start = False + index += 2 + continue + terminator = next( + (candidate for candidate in terminators if data.startswith(candidate, index)), + None, + ) + if terminator is not None: + end = index + len(terminator) + if index == start: + raise CsvAdapterError(f"CSV blank logical record at byte {start} is forbidden.") + observed_terminators.add(terminator) + if ( + dialect.mixed_line_terminators == "forbid" + and len(observed_terminators) > 1 + ): + raise CsvAdapterError("CSV contains forbidden mixed line terminators.") + spans.append(LogicalRecordSpan(len(spans), start, end, terminator)) + start = end + index = end + at_field_start = True + continue + if byte == delimiter: + at_field_start = True + elif byte == quote and at_field_start: + in_quotes = True + at_field_start = False + elif not (dialect.skip_initial_space and at_field_start and byte == 0x20): + at_field_start = False + index += 1 + + if in_quotes: + raise CsvAdapterError( + f"CSV has an unclosed quote in logical record starting at byte {start}." + ) + if start < len(data): + if dialect.final_record_without_terminator == "forbid": + raise CsvAdapterError("CSV final logical record has no declared terminator.") + spans.append(LogicalRecordSpan(len(spans), start, len(data), b"")) + return tuple(spans) + + +def effective_csv_adapter_projection( + declaration: SourceArtifactSpec, *, strict: bool +) -> dict[str, object]: + """Return parsing controls that must participate in realization identity.""" + effective_extra_policy = ( + "reject_unmapped" + if strict or declaration.extra_field_policy == "reject_unmapped" + else "preserve_unmapped" + ) + return canonicalize( + { + "adapter": "choicebench.csv.v1", + "dialect": asdict(declaration.dialect), + "declared_extra_field_policy": declaration.extra_field_policy, + "cli_strict": strict, + "effective_extra_field_policy": effective_extra_policy, + } + ) + + +def _decode_record( + data: bytes, + span: LogicalRecordSpan, + dialect: CsvDialectSpec, + *, + strip_bom: bool, +) -> list[str]: + record = data[span.start : span.end - len(span.terminator) if span.terminator else span.end] + if strip_bom: + record = record[len(_UTF8_BOM) :] + try: + text = record.decode("utf-8", errors="strict") + except UnicodeDecodeError as exc: + raise CsvAdapterError( + f"Invalid UTF-8 in CSV logical record {span.index} at source byte " + f"{span.start + exc.start}." + ) from exc + try: + parsed = list( + csv.reader( + io.StringIO(text, newline=""), + delimiter=dialect.delimiter, + quotechar=dialect.quote_character, + escapechar=dialect.escape_character, + doublequote=dialect.double_quote, + skipinitialspace=dialect.skip_initial_space, + strict=dialect.strict_syntax, + ) + ) + except csv.Error as exc: + raise CsvAdapterError( + f"Malformed CSV syntax in logical record {span.index}: {type(exc).__name__}." + ) from exc + if len(parsed) != 1: + raise CsvAdapterError( + f"CSV logical record {span.index} parsed into {len(parsed)} physical rows." + ) + return parsed[0] + + +def _validate_numeric( + value: str | None, *, source_column: str, value_type: str, + null_allowed: bool, finite_only: bool, record_index: int +) -> None: + if value is None: + if not null_allowed: + raise CsvAdapterError( + f"CSV null is forbidden for numeric column {source_column!r} " + f"in logical record {record_index}." + ) + return + if value_type == "integer": + if _INTEGER_RE.fullmatch(value) is None: + raise CsvAdapterError( + f"CSV integer column {source_column!r} has a malformed value in " + f"logical record {record_index}." + ) + return + if _FLOAT_RE.fullmatch(value) is not None: + number = float(value) + elif _NONFINITE_FLOAT_RE.fullmatch(value) is not None: + number = float(value) + else: + raise CsvAdapterError( + f"CSV finite float column {source_column!r} has a malformed value in " + f"logical record {record_index}." + ) + if finite_only and not math.isfinite(number): + raise CsvAdapterError( + f"CSV finite float column {source_column!r} has a non-finite value in " + f"logical record {record_index}." + ) + + +def _validate_structured_choices( + value: str | None, declaration: SourceArtifactSpec, *, record_index: int +) -> None: + mapping = declaration.option_mapping + if mapping.mode != "structured_json": + return + if value is None: + raise CsvAdapterError( + f"CSV structured choices are null in logical record {record_index}." + ) + try: + choices = json.loads(value) + except (json.JSONDecodeError, UnicodeError) as exc: + raise CsvAdapterError( + f"CSV structured choices are malformed in logical record {record_index}." + ) from exc + if not isinstance(choices, list) or not choices: + raise CsvAdapterError( + f"CSV structured choices must be a non-empty list in logical record {record_index}." + ) + assert mapping.structured_label_key is not None + assert mapping.structured_text_key is not None + labels: set[str] = set() + for choice in choices: + if not isinstance(choice, Mapping): + raise CsvAdapterError( + f"CSV structured choices contain a non-object in logical record {record_index}." + ) + label = choice.get(mapping.structured_label_key) + text = choice.get(mapping.structured_text_key) + if not isinstance(label, str) or not label or not isinstance(text, str): + raise CsvAdapterError( + "CSV structured choices lack string label/text fields in logical " + f"record {record_index}." + ) + if label in labels: + raise CsvAdapterError( + f"CSV structured choices contain duplicate labels in logical record {record_index}." + ) + labels.add(label) + + +def parse_csv_source( + source: OpenedSource, declaration: SourceArtifactSpec, *, strict: bool +) -> AdaptedTable: + """Validate and parse one already-opened, checksum-bound CSV source.""" + if source.source_id != declaration.source_id: + raise CsvAdapterError("Opened source_id does not match its source declaration.") + if source.logical_path != declaration.logical_path: + raise CsvAdapterError("Opened source logical_path does not match its declaration.") + actual_sha256 = sha256(source.data).hexdigest() + if source.sha256 != actual_sha256: + raise CsvAdapterError("Opened source checksum does not match its bytes.") + if declaration.expected_sha256 != actual_sha256: + raise CsvAdapterError("Source checksum does not match the import declaration.") + if declaration.format != "csv": + raise CsvAdapterError(f"CSV adapter cannot parse format {declaration.format!r}.") + if source.data.startswith(_UTF8_BOM) and declaration.dialect.bom_policy == "forbid": + raise CsvAdapterError("Source contains a forbidden UTF-8 BOM.") + + spans = scan_csv_logical_records(source.data, declaration.dialect) + if not spans: + raise CsvAdapterError("CSV source has no header logical record.") + header = tuple( + _decode_record( + source.data, + spans[0], + declaration.dialect, + strip_bom=( + declaration.dialect.bom_policy == "strip_utf8_bom" + and source.data.startswith(_UTF8_BOM) + ), + ) + ) + duplicates = sorted({column for column in header if header.count(column) > 1}) + if duplicates: + raise CsvAdapterError(f"CSV duplicate header columns are forbidden: {duplicates}.") + credential_columns = sorted(column for column in header if is_credential_key(column)) + if credential_columns: + raise CsvAdapterError( + f"CSV credential-named columns are forbidden: {credential_columns}." + ) + missing = sorted(set(declaration.expected_columns) - set(header)) + if missing: + raise CsvAdapterError(f"CSV is missing declared source columns: {missing}.") + extras = sorted(set(header) - set(declaration.expected_columns)) + effective = effective_csv_adapter_projection(declaration, strict=strict) + if extras and effective["effective_extra_field_policy"] == "reject_unmapped": + raise CsvAdapterError(f"CSV has unexpected source columns: {extras}.") + + rows: list[SourceRow] = [] + for span in spans[1:]: + cells = _decode_record( + source.data, span, declaration.dialect, strip_bom=False + ) + if len(cells) != len(header): + raise CsvAdapterError( + f"CSV logical record {span.index} field count {len(cells)} does not " + f"match header field count {len(header)}." + ) + values: dict[str, str | None] = { + column: None if cell in declaration.null_values else cell + for column, cell in zip(header, cells, strict=True) + } + for numeric in declaration.numeric_columns: + _validate_numeric( + values[numeric.source_column], + source_column=numeric.source_column, + value_type=numeric.value_type, + null_allowed=numeric.null_allowed, + finite_only=numeric.finite_only, + record_index=span.index, + ) + if declaration.option_mapping.mode == "structured_json": + assert declaration.option_mapping.structured_column is not None + _validate_structured_choices( + values[declaration.option_mapping.structured_column], + declaration, + record_index=span.index, + ) + raw = source.data[span.start : span.end] + rows.append(SourceRow(values, span, sha256(raw).hexdigest())) + return AdaptedTable(header, tuple(rows), actual_sha256) diff --git a/tests/importing/test_csv_adapter.py b/tests/importing/test_csv_adapter.py new file mode 100644 index 0000000..9eed0e5 --- /dev/null +++ b/tests/importing/test_csv_adapter.py @@ -0,0 +1,435 @@ +from __future__ import annotations + +from dataclasses import replace +from hashlib import sha256 +from pathlib import Path + +import pytest + +from choicebench.importing.csv_adapter import ( + CsvAdapterError, + OpenedSource, + effective_csv_adapter_projection, + parse_csv_source, + scan_csv_logical_records, +) +from choicebench.importing.schema import ( + CsvDialectSpec, + NumericColumnSpec, + OptionMappingSpec, + SourceArtifactSpec, +) + + +def _opened(tmp_path: Path, data: bytes) -> OpenedSource: + path = tmp_path / "source.csv" + path.write_bytes(data) + return OpenedSource( + source_id="source", + audit_path=path, + logical_path="freeze/source.csv", + data=data, + sha256=sha256(data).hexdigest(), + ) + + +def _source_spec( + data: bytes, + *, + expected_columns: tuple[str, ...] = ("id", "value"), + columns: dict[str, str] | None = None, + ignored_columns: dict[str, str] | None = None, + null_values: tuple[str, ...] = (), + numeric_columns: tuple[NumericColumnSpec, ...] = (), + option_mapping: OptionMappingSpec | None = None, + extra_field_policy: str = "preserve_unmapped", + dialect: CsvDialectSpec | None = None, +) -> SourceArtifactSpec: + return SourceArtifactSpec( + source_id="source", + path=Path("/explicit/read-only/source.csv"), + logical_path="freeze/source.csv", + expected_sha256=sha256(data).hexdigest(), + format="csv", + format_version="producer-v1", + classification="raw", + dialect=dialect or CsvDialectSpec(), + columns=columns or {"question_id": "id", "prediction": "value"}, + expected_columns=expected_columns, + ignored_columns=ignored_columns or {}, + null_values=null_values, + numeric_columns=numeric_columns, + option_mapping=option_mapping + or OptionMappingSpec("ordered_columns", ("value",), None, None, None), + extra_field_policy=extra_field_policy, # type: ignore[arg-type] + preserve_namespace="producer", + source_run_id=None, + source_repository=None, + source_commit=None, + notes={}, + ) + + +def _parse( + tmp_path: Path, + data: bytes, + *, + strict: bool = False, + **spec_kwargs: object, +): + return parse_csv_source( + _opened(tmp_path, data), _source_spec(data, **spec_kwargs), strict=strict + ) + + +@pytest.mark.parametrize( + ("data", "terminators"), + [ + (b"id,value\nq1,x\n", (b"\n", b"\n")), + (b"id,value\r\nq1,x\r\n", (b"\r\n", b"\r\n")), + (b"id,value\rq1,x\r", (b"\r", b"\r")), + (b"id,value\r\nq1,x\nq2,y\r", (b"\r\n", b"\n", b"\r")), + ], +) +def test_scanner_recognizes_declared_terminators(data: bytes, terminators: tuple[bytes, ...]): + spans = scan_csv_logical_records(data, CsvDialectSpec()) + assert tuple(span.terminator for span in spans) == terminators + assert b"".join(data[span.start : span.end] for span in spans) == data + + +def test_embedded_newline_span_is_exact(): + data = b'id,text\r\nq1,"line one\r\nline two"\r\nq2,end\n' + spans = scan_csv_logical_records(data, CsvDialectSpec()) + assert data[spans[1].start : spans[1].end] == b'q1,"line one\r\nline two"\r\n' + assert spans[1].terminator == b"\r\n" + + +@pytest.mark.parametrize("embedded", [b"\r", b"\n", b"\r\n"]) +def test_each_embedded_terminator_remains_inside_quoted_record(embedded: bytes): + data = b'id,text\nq1,"a' + embedded + b'b"\n' + spans = scan_csv_logical_records(data, CsvDialectSpec()) + assert len(spans) == 2 + assert data[spans[1].start : spans[1].end] == b'q1,"a' + embedded + b'b"\n' + + +def test_doubled_quote_does_not_end_quoted_field(): + data = b'id,text\nq1,"a""b\nc"\n' + spans = scan_csv_logical_records(data, CsvDialectSpec(double_quote=True)) + assert len(spans) == 2 + assert data[spans[1].start : spans[1].end] == b'q1,"a""b\nc"\n' + + +def test_explicit_escape_protects_quote_and_newline(): + data = b'id,text\nq1,"a\\"b\nc"\n' + dialect = CsvDialectSpec(escape_character="\\", double_quote=False) + spans = scan_csv_logical_records(data, dialect) + assert len(spans) == 2 + assert data[spans[1].start : spans[1].end] == b'q1,"a\\"b\nc"\n' + + +def test_explicit_escape_outside_quotes_protects_one_lf_byte(): + data = b"id,text\nq1,a\\\nb\n" + dialect = CsvDialectSpec(escape_character="\\") + spans = scan_csv_logical_records(data, dialect) + assert len(spans) == 2 + assert data[spans[1].start : spans[1].end] == b"q1,a\\\nb\n" + + +def test_unclosed_quote_is_rejected_with_exact_record_start(): + data = b'id,text\nq1,"unterminated\n' + with pytest.raises(CsvAdapterError, match=r"unclosed quote.*byte 8"): + scan_csv_logical_records(data, CsvDialectSpec()) + + +def test_final_record_without_terminator_is_preserved(): + data = b"id,value\nq1,x" + spans = scan_csv_logical_records(data, CsvDialectSpec()) + assert spans[-1].terminator == b"" + assert data[spans[-1].start : spans[-1].end] == b"q1,x" + + +def test_forbidden_final_record_without_terminator_is_rejected(): + data = b"id,value\nq1,x" + dialect = CsvDialectSpec(final_record_without_terminator="forbid") + with pytest.raises(CsvAdapterError, match="final logical record"): + scan_csv_logical_records(data, dialect) + + +def test_forbidden_mixed_terminators_are_rejected(): + data = b"id,value\r\nq1,x\n" + dialect = CsvDialectSpec(mixed_line_terminators="forbid") + with pytest.raises(CsvAdapterError, match="mixed line terminators"): + scan_csv_logical_records(data, dialect) + + +@pytest.mark.parametrize("blank", [b"\n", b"\r", b"\r\n"]) +def test_blank_logical_records_are_rejected(blank: bytes): + data = b"id,value\n" + blank + b"q1,x\n" + with pytest.raises(CsvAdapterError, match="blank logical record"): + scan_csv_logical_records(data, CsvDialectSpec()) + + +def test_bom_is_forbidden_by_default(tmp_path: Path): + data = b"\xef\xbb\xbfid,value\nq1,x\n" + with pytest.raises(CsvAdapterError, match="UTF-8 BOM"): + _parse(tmp_path, data) + + +def test_bom_strip_preserves_original_byte_offsets(tmp_path: Path): + data = b"\xef\xbb\xbfid,value\nq1,x\n" + dialect = CsvDialectSpec(bom_policy="strip_utf8_bom") + table = _parse(tmp_path, data, dialect=dialect) + assert table.columns == ("id", "value") + assert table.rows[0].span.start == len(b"\xef\xbb\xbfid,value\n") + assert table.rows[0].values == {"id": "q1", "value": "x"} + + +def test_bom_strip_keeps_a_quoted_first_header_field_in_quote_state(): + data = b'\xef\xbb\xbf"id\ncontinued",value\nq1,x\n' + dialect = CsvDialectSpec(bom_policy="strip_utf8_bom") + spans = scan_csv_logical_records(data, dialect) + assert len(spans) == 2 + assert data[spans[0].start : spans[0].end] == b'\xef\xbb\xbf"id\ncontinued",value\n' + + +def test_invalid_utf8_is_rejected_without_replacement(tmp_path: Path): + data = b"id,value\nq1,\xff\n" + with pytest.raises(CsvAdapterError, match="UTF-8.*logical record 1"): + _parse(tmp_path, data) + + +def test_header_order_and_raw_row_hash_are_preserved(tmp_path: Path): + data = b"value,id\r\nx,q1\r\n" + table = _parse( + tmp_path, + data, + expected_columns=("value", "id"), + columns={"question_id": "id", "prediction": "value"}, + ) + assert table.columns == ("value", "id") + assert table.rows[0].raw_sha256 == sha256(b"x,q1\r\n").hexdigest() + + +def test_duplicate_header_is_rejected(tmp_path: Path): + data = b"id,id\nq1,x\n" + with pytest.raises(CsvAdapterError, match="duplicate header"): + _parse(tmp_path, data, expected_columns=("id",)) + + +@pytest.mark.parametrize("credential_column", ["api_key", "nested_access_token", "password"]) +def test_credential_shaped_header_is_rejected_in_every_mode( + tmp_path: Path, credential_column: str +): + data = f"id,value,{credential_column}\nq1,x,secret\n".encode() + for strict in (False, True): + with pytest.raises(CsvAdapterError, match="credential-named"): + _parse(tmp_path, data, strict=strict) + + +def test_normal_mode_preserves_unmapped_extra_columns(tmp_path: Path): + data = b"id,value,diagnostic\nq1,x,kept\n" + table = _parse(tmp_path, data, strict=False) + assert table.rows[0].values["diagnostic"] == "kept" + + +def test_cli_strict_mode_tightens_preserve_unmapped_policy(tmp_path: Path): + data = b"id,value,diagnostic\nq1,x,kept\n" + with pytest.raises(CsvAdapterError, match="unexpected source columns.*diagnostic"): + _parse(tmp_path, data, strict=True) + + +def test_cli_normal_mode_cannot_loosen_declared_reject_policy(tmp_path: Path): + data = b"id,value,diagnostic\nq1,x,kept\n" + with pytest.raises(CsvAdapterError, match="unexpected source columns.*diagnostic"): + _parse(tmp_path, data, strict=False, extra_field_policy="reject_unmapped") + + +def test_effective_strictness_is_identity_bearing(): + data = b"id,value\nq1,x\n" + declaration = _source_spec(data) + assert effective_csv_adapter_projection( + declaration, strict=False + ) != effective_csv_adapter_projection(declaration, strict=True) + assert effective_csv_adapter_projection( + replace(declaration, dialect=replace(declaration.dialect, delimiter=";")), + strict=False, + ) != effective_csv_adapter_projection(declaration, strict=False) + + +def test_declared_ignored_column_requires_reason_and_is_retained(tmp_path: Path): + data = b"id,value,score\nq1,x,0.5\n" + table = _parse( + tmp_path, + data, + expected_columns=("id", "value", "score"), + ignored_columns={"score": "producer aggregate only"}, + ) + assert table.rows[0].values["score"] == "0.5" + + +def test_missing_declared_column_is_rejected(tmp_path: Path): + data = b"id,value\nq1,x\n" + with pytest.raises(CsvAdapterError, match="missing declared source columns.*score"): + _parse(tmp_path, data, expected_columns=("id", "value", "score")) + + +def test_empty_string_is_distinct_from_null_unless_declared(tmp_path: Path): + data = b"id,value\nq1,\n" + assert _parse(tmp_path, data).rows[0].values["value"] == "" + assert _parse(tmp_path, data, null_values=("",)).rows[0].values["value"] is None + + +def test_literal_nan_is_not_implicit_null(tmp_path: Path): + data = b"id,value\nq1,NaN\n" + table = _parse(tmp_path, data, strict=True) + assert table.rows[0].values["value"] == "NaN" + + +@pytest.mark.parametrize("value", [b"1", b"-2", b"+3"]) +def test_integer_numeric_policy_accepts_exact_integer_strings(tmp_path: Path, value: bytes): + data = b"id,value\nq1," + value + b"\n" + numeric = (NumericColumnSpec("value", "integer", False),) + assert _parse(tmp_path, data, numeric_columns=numeric).rows[0].values["value"] == value.decode() + + +@pytest.mark.parametrize("value", [b"1.0", b"1e2", b"x", b" 1"]) +def test_integer_numeric_policy_rejects_malformed_strings(tmp_path: Path, value: bytes): + data = b"id,value\nq1," + value + b"\n" + numeric = (NumericColumnSpec("value", "integer", False),) + with pytest.raises(CsvAdapterError, match="integer.*value"): + _parse(tmp_path, data, numeric_columns=numeric) + + +@pytest.mark.parametrize("value", [b"1", b"-1.25", b"1e2"]) +def test_float_numeric_policy_accepts_finite_numbers(tmp_path: Path, value: bytes): + data = b"id,value\nq1," + value + b"\n" + numeric = (NumericColumnSpec("value", "float", False),) + _parse(tmp_path, data, numeric_columns=numeric) + + +@pytest.mark.parametrize("value", [b"NaN", b"inf", b"-Infinity", b"x", b" 1"]) +def test_finite_float_policy_rejects_nonfinite_or_malformed_values(tmp_path: Path, value: bytes): + data = b"id,value\nq1," + value + b"\n" + numeric = (NumericColumnSpec("value", "float", False, finite_only=True),) + with pytest.raises(CsvAdapterError, match="finite float.*value"): + _parse(tmp_path, data, numeric_columns=numeric) + + +def test_nonfinite_float_is_allowed_only_when_explicit(tmp_path: Path): + data = b"id,value\nq1,NaN\n" + numeric = (NumericColumnSpec("value", "float", False, finite_only=False),) + assert _parse(tmp_path, data, numeric_columns=numeric).rows[0].values["value"] == "NaN" + + +def test_numeric_null_policy_is_explicit(tmp_path: Path): + data = b"id,value\nq1,NA\n" + with pytest.raises(CsvAdapterError, match="null.*value"): + _parse( + tmp_path, + data, + null_values=("NA",), + numeric_columns=(NumericColumnSpec("value", "float", False),), + ) + table = _parse( + tmp_path, + data, + null_values=("NA",), + numeric_columns=(NumericColumnSpec("value", "float", True),), + ) + assert table.rows[0].values["value"] is None + + +@pytest.mark.parametrize("option_count", [3, 6]) +def test_ordered_option_mapping_supports_variable_option_counts( + tmp_path: Path, option_count: int +): + option_columns = tuple(f"choice_{index}" for index in range(option_count)) + header = ("id", *option_columns) + data = ( + ",".join(header) + + "\nq1," + + ",".join(f"v{index}" for index in range(option_count)) + + "\n" + ).encode() + table = _parse( + tmp_path, + data, + expected_columns=header, + columns={"question_id": "id"}, + option_mapping=OptionMappingSpec("ordered_columns", option_columns, None, None, None), + ) + assert tuple(table.rows[0].values[name] for name in option_columns) == tuple( + f"v{index}" for index in range(option_count) + ) + + +def test_empty_trailing_ordered_option_stays_null_not_phantom_choice(tmp_path: Path): + data = b"id,a,b,c,d\nq1,A,B,C,\n" + table = _parse( + tmp_path, + data, + expected_columns=("id", "a", "b", "c", "d"), + columns={"question_id": "id"}, + null_values=("",), + option_mapping=OptionMappingSpec("ordered_columns", ("a", "b", "c", "d"), None, None, None), + ) + assert [table.rows[0].values[name] for name in ("a", "b", "c", "d")] == ["A", "B", "C", None] + + +def test_structured_json_choices_accept_variable_option_list(tmp_path: Path): + data = ( + b'id,choices\nq1,"[{""label"":""A"",""text"":""one""},' + b'{""label"":""B"",""text"":""two""},' + b'{""label"":""C"",""text"":""three""}]"\n' + ) + mapping = OptionMappingSpec("structured_json", (), "choices", "label", "text") + table = _parse( + tmp_path, + data, + expected_columns=("id", "choices"), + columns={"question_id": "id"}, + option_mapping=mapping, + ) + assert table.rows[0].values["choices"].startswith('[{"label":"A"') + + +@pytest.mark.parametrize( + "payload", + [b"not-json", b"{}", b'[{"label":"A"}]', b'[{"label":"A","text":1}]'], +) +def test_structured_json_choices_reject_malformed_payloads(tmp_path: Path, payload: bytes): + escaped = payload.replace(b'"', b'""') + data = b'id,choices\nq1,"' + escaped + b'"\n' + mapping = OptionMappingSpec("structured_json", (), "choices", "label", "text") + with pytest.raises(CsvAdapterError, match="structured choices"): + _parse( + tmp_path, + data, + expected_columns=("id", "choices"), + columns={"question_id": "id"}, + option_mapping=mapping, + ) + + +def test_row_with_more_fields_than_header_is_rejected(tmp_path: Path): + data = b"id,value\nq1,x,unexpected\n" + with pytest.raises(CsvAdapterError, match="field count"): + _parse(tmp_path, data) + + +def test_source_checksum_and_declaration_source_id_are_verified(tmp_path: Path): + data = b"id,value\nq1,x\n" + source = _opened(tmp_path, data) + bad_digest = replace(_source_spec(data), expected_sha256="0" * 64) + with pytest.raises(CsvAdapterError, match="checksum"): + parse_csv_source(source, bad_digest, strict=False) + bad_source_id = replace(_source_spec(data), source_id="different") + with pytest.raises(CsvAdapterError, match="source_id"): + parse_csv_source(source, bad_source_id, strict=False) + + +def test_source_logical_path_must_match_declaration(tmp_path: Path): + data = b"id,value\nq1,x\n" + source = replace(_opened(tmp_path, data), logical_path="other/source.csv") + with pytest.raises(CsvAdapterError, match="logical_path"): + parse_csv_source(source, _source_spec(data), strict=False) From c23c239777c5a267b0903e46451091649254a874 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 00:22:24 +0300 Subject: [PATCH 11/47] fix: enforce declared CSV line terminators --- src/choicebench/importing/csv_adapter.py | 22 +++++++++--- tests/importing/test_csv_adapter.py | 44 ++++++++++++++++++++++++ 2 files changed, 62 insertions(+), 4 deletions(-) diff --git a/src/choicebench/importing/csv_adapter.py b/src/choicebench/importing/csv_adapter.py index c8bb8a5..810b5c0 100644 --- a/src/choicebench/importing/csv_adapter.py +++ b/src/choicebench/importing/csv_adapter.py @@ -18,6 +18,7 @@ _UTF8_BOM = b"\xef\xbb\xbf" _TERMINATORS = {"crlf": b"\r\n", "lf": b"\n", "cr": b"\r"} +_ACTUAL_TERMINATORS = (b"\r\n", b"\n", b"\r") _INTEGER_RE = re.compile(r"[+-]?\d+\Z") _FLOAT_RE = re.compile( r"[+-]?(?:(?:\d+(?:\.\d*)?)|(?:\.\d+))(?:[eE][+-]?\d+)?\Z" @@ -102,7 +103,7 @@ def scan_csv_logical_records( if not isinstance(data, bytes): raise TypeError(f"data must be bytes; got {type(data).__name__}.") delimiter, quote, escape = _dialect_bytes(dialect) - terminators = _declared_terminators(dialect) + declared_terminators = set(_declared_terminators(dialect)) spans: list[LogicalRecordSpan] = [] observed_terminators: set[bytes] = set() start = 0 @@ -112,6 +113,7 @@ def scan_csv_logical_records( index = len(_UTF8_BOM) else: index = 0 + content_start = index in_quotes = False at_field_start = True @@ -146,12 +148,21 @@ def scan_csv_logical_records( index += 2 continue terminator = next( - (candidate for candidate in terminators if data.startswith(candidate, index)), + ( + candidate + for candidate in _ACTUAL_TERMINATORS + if data.startswith(candidate, index) + ), None, ) if terminator is not None: + if terminator not in declared_terminators: + raise CsvAdapterError( + f"CSV contains undeclared line terminator {terminator!r} " + f"at byte {index}." + ) end = index + len(terminator) - if index == start: + if index == content_start: raise CsvAdapterError(f"CSV blank logical record at byte {start} is forbidden.") observed_terminators.add(terminator) if ( @@ -162,6 +173,7 @@ def scan_csv_logical_records( spans.append(LogicalRecordSpan(len(spans), start, end, terminator)) start = end index = end + content_start = end at_field_start = True continue if byte == delimiter: @@ -212,14 +224,16 @@ def _decode_record( strip_bom: bool, ) -> list[str]: record = data[span.start : span.end - len(span.terminator) if span.terminator else span.end] + stripped_prefix_size = 0 if strip_bom: + stripped_prefix_size = len(_UTF8_BOM) record = record[len(_UTF8_BOM) :] try: text = record.decode("utf-8", errors="strict") except UnicodeDecodeError as exc: raise CsvAdapterError( f"Invalid UTF-8 in CSV logical record {span.index} at source byte " - f"{span.start + exc.start}." + f"{span.start + stripped_prefix_size + exc.start}." ) from exc try: parsed = list( diff --git a/tests/importing/test_csv_adapter.py b/tests/importing/test_csv_adapter.py index 9eed0e5..196cef7 100644 --- a/tests/importing/test_csv_adapter.py +++ b/tests/importing/test_csv_adapter.py @@ -162,6 +162,35 @@ def test_forbidden_mixed_terminators_are_rejected(): scan_csv_logical_records(data, dialect) +@pytest.mark.parametrize( + ("declared", "actual"), + [ + (("lf",), b"\r\n"), + (("lf",), b"\r"), + (("crlf",), b"\n"), + (("crlf",), b"\r"), + (("cr",), b"\r\n"), + (("cr",), b"\n"), + ], +) +def test_actual_undeclared_terminator_is_rejected( + declared: tuple[str, ...], actual: bytes +): + data = b"id,value" + actual + b"q1,x" + actual + dialect = CsvDialectSpec(line_terminators=declared) + with pytest.raises(CsvAdapterError, match="undeclared line terminator"): + scan_csv_logical_records(data, dialect) + + +def test_mixed_policy_compares_actual_crlf_and_lf_terminators(): + data = b"id,value\r\nq1,x\n" + dialect = CsvDialectSpec( + line_terminators=("crlf", "lf"), mixed_line_terminators="forbid" + ) + with pytest.raises(CsvAdapterError, match="mixed line terminators"): + scan_csv_logical_records(data, dialect) + + @pytest.mark.parametrize("blank", [b"\n", b"\r", b"\r\n"]) def test_blank_logical_records_are_rejected(blank: bytes): data = b"id,value\n" + blank + b"q1,x\n" @@ -169,6 +198,14 @@ def test_blank_logical_records_are_rejected(blank: bytes): scan_csv_logical_records(data, CsvDialectSpec()) +@pytest.mark.parametrize("terminator", [b"\n", b"\r", b"\r\n"]) +def test_bom_strip_rejects_blank_first_logical_record_directly(terminator: bytes): + data = b"\xef\xbb\xbf" + terminator + b"q1,x" + terminator + dialect = CsvDialectSpec(bom_policy="strip_utf8_bom") + with pytest.raises(CsvAdapterError, match="blank logical record"): + scan_csv_logical_records(data, dialect) + + def test_bom_is_forbidden_by_default(tmp_path: Path): data = b"\xef\xbb\xbfid,value\nq1,x\n" with pytest.raises(CsvAdapterError, match="UTF-8 BOM"): @@ -198,6 +235,13 @@ def test_invalid_utf8_is_rejected_without_replacement(tmp_path: Path): _parse(tmp_path, data) +def test_invalid_utf8_offset_in_bom_stripped_header_uses_original_bytes(tmp_path: Path): + data = b"\xef\xbb\xbfid,\xff\nq1,x\n" + dialect = CsvDialectSpec(bom_policy="strip_utf8_bom") + with pytest.raises(CsvAdapterError, match=r"logical record 0 at source byte 6"): + _parse(tmp_path, data, dialect=dialect) + + def test_header_order_and_raw_row_hash_are_preserved(tmp_path: Path): data = b"value,id\r\nx,q1\r\n" table = _parse( From 4edb9770db3941b55176cb39399ffa69886ca0e7 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 00:50:31 +0300 Subject: [PATCH 12/47] feat: add trusted import dataset references --- .../importing/dataset_reference.py | 551 ++++++++++++++++++ tests/importing/test_dataset_reference.py | 436 ++++++++++++++ 2 files changed, 987 insertions(+) create mode 100644 src/choicebench/importing/dataset_reference.py create mode 100644 tests/importing/test_dataset_reference.py diff --git a/src/choicebench/importing/dataset_reference.py b/src/choicebench/importing/dataset_reference.py new file mode 100644 index 0000000..b4e9e4c --- /dev/null +++ b/src/choicebench/importing/dataset_reference.py @@ -0,0 +1,551 @@ +"""Trust-qualified expected-dataset snapshots for external result imports.""" + +from __future__ import annotations + +from dataclasses import dataclass +from hashlib import sha256 +from io import BytesIO +import json +from pathlib import Path, PurePosixPath +from typing import Any, Literal, Mapping + +import pandas as pd + +from choicebench.datasets import ( + NORMALIZATION_VERSION, + dataset_content_digest, + dataset_sample_identities, + validate_normalized_dataset, +) +from choicebench.identity import canonicalize, integrity_digest, short_id +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.schema import DatasetReferenceSpec +from choicebench.infra.artifacts import atomic_write_json, atomic_write_text +from choicebench.pipeline.options import build_option_map, correct_option_for_row + + +class DatasetReferenceError(ValueError): + """Raised when an expected-dataset reference is incomplete or inconsistent.""" + + +@dataclass(frozen=True) +class ExpectedDataset: + dataset_id: str + benchmark_name: str + split: str + artifact_id: str + artifact_digest: str + selection_id: str + selection_digest: str + selection_semantics: Mapping[str, Any] + frame: pd.DataFrame + selected_question_ids: tuple[str, ...] + reference_kind: Literal[ + "independent_input_snapshot", "profile_derived_reference_snapshot" + ] + trust_label: str + question_set_digest: str + snapshot_digest: str + derivation: Mapping[str, Any] + derivation_digest: str + limitations: tuple[str, ...] + + +_SNAPSHOT_SCHEMA = "choicebench.expected-dataset.v1" +_REFERENCE_SCHEMA = "choicebench.dataset-reference.v1" +_REQUIRED_FIELDS = ("question_id", "question_text", "correct_option") + + +def _source_frame(source_id: str, source: OpenedSource) -> pd.DataFrame: + if source.source_id != source_id: + raise DatasetReferenceError( + f"Expected dataset source key {source_id!r} does not match opened source_id " + f"{source.source_id!r}." + ) + actual = sha256(source.data).hexdigest() + if source.sha256 != actual: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} checksum does not match its bytes." + ) + try: + return pd.read_csv( + BytesIO(source.data), dtype=str, keep_default_na=False, na_filter=False + ) + except (UnicodeError, pd.errors.ParserError, pd.errors.EmptyDataError) as exc: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} is not a valid UTF-8 CSV." + ) from exc + + +def _choice_keys(columns: Mapping[str, str]) -> tuple[str, ...]: + keys = tuple( + sorted( + (key for key in columns if key.startswith("choice_") and len(key) == 8), + key=lambda key: key.removeprefix("choice_"), + ) + ) + return keys + + +def _normalize_source( + declaration: DatasetReferenceSpec, source_id: str, source: OpenedSource +) -> pd.DataFrame: + raw = _source_frame(source_id, source) + mapping = dict(declaration.columns) + missing_common = sorted(set(_REQUIRED_FIELDS) - set(mapping)) + if missing_common: + raise DatasetReferenceError( + f"Expected dataset mapping is missing required fields {missing_common}." + ) + missing_source = sorted(set(mapping.values()) - set(raw.columns)) + if missing_source: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} is missing mapped columns {missing_source}." + ) + if len(mapping.values()) != len(set(mapping.values())): + raise DatasetReferenceError("Expected dataset mapping reuses a source column.") + + records: list[dict[str, Any]] = [] + choice_keys = _choice_keys(mapping) + has_structured_choices = "choices_json" in mapping + if not choice_keys and not has_structured_choices: + raise DatasetReferenceError("Expected dataset mapping declares no ordered options.") + + ordinary_keys = sorted( + key for key in mapping if key not in choice_keys and key != "choices_json" + ) + for row_number, raw_row in raw.iterrows(): + record = {key: raw_row[mapping[key]] for key in ordinary_keys} + question_id = str(record["question_id"]).strip() + if not question_id: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has an empty question_id at row " + f"{row_number + 2}." + ) + record["question_id"] = question_id + record["correct_option"] = str(record["correct_option"]).strip().upper() + + if has_structured_choices: + raw_choices = raw_row[mapping["choices_json"]] + try: + choices = json.loads(raw_choices) + except (json.JSONDecodeError, TypeError) as exc: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has malformed ordered options " + f"for question_id {question_id!r}." + ) from exc + if not isinstance(choices, list): + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has malformed ordered options " + f"for question_id {question_id!r}." + ) + normalized_choices = [] + for index, item in enumerate(choices): + if not isinstance(item, Mapping) or not isinstance(item.get("text"), str): + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has malformed ordered options " + f"for question_id {question_id!r}." + ) + normalized_choices.append( + { + "text": item["text"], + "source_index": int(item.get("source_index", index)), + } + ) + else: + normalized_choices = [ + {"text": raw_row[mapping[key]], "source_index": index} + for index, key in enumerate(choice_keys) + if str(raw_row[mapping[key]]).strip() + ] + record["choices_json"] = json.dumps( + normalized_choices, ensure_ascii=True, separators=(",", ":") + ) + try: + options = build_option_map(record) + record["correct_option"] = correct_option_for_row(record, options) + except (TypeError, ValueError, json.JSONDecodeError) as exc: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has invalid ordered options or " + f"correct_option for question_id {question_id!r}: {exc}" + ) from exc + records.append(record) + + frame = pd.DataFrame(records) + if frame.empty: + raise DatasetReferenceError(f"Expected dataset source {source_id!r} has no rows.") + duplicated = frame["question_id"].astype(str).duplicated(keep=False) + if duplicated.any(): + duplicate_ids = sorted(frame.loc[duplicated, "question_id"].astype(str).unique()) + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has duplicate question_id values " + f"{duplicate_ids}." + ) + return frame + + +def _selected_frame( + frame: pd.DataFrame, expected_question_ids: tuple[str, ...] +) -> pd.DataFrame: + if len(expected_question_ids) != len(set(expected_question_ids)): + duplicates = sorted( + question_id + for question_id in set(expected_question_ids) + if expected_question_ids.count(question_id) > 1 + ) + raise DatasetReferenceError( + f"Expected dataset declaration has duplicate selected question ID(s) {duplicates}." + ) + indexed = frame.set_index(frame["question_id"].astype(str), drop=False) + missing = [question_id for question_id in expected_question_ids if question_id not in indexed.index] + if missing: + raise DatasetReferenceError( + f"Expected dataset is missing selected question ID(s) {missing}." + ) + selected = indexed.loc[list(expected_question_ids)].reset_index(drop=True) + return selected + + +def _semantic_disagreement(left: pd.DataFrame, right: pd.DataFrame) -> str | None: + for left_row, right_row in zip( + left.to_dict("records"), right.to_dict("records"), strict=True + ): + question_id = str(left_row["question_id"]) + if left_row["question_text"] != right_row["question_text"]: + return f"question_text disagreement for question_id {question_id!r}" + if left_row["correct_option"] != right_row["correct_option"]: + return f"correct_option disagreement for question_id {question_id!r}" + if build_option_map(left_row) != build_option_map(right_row): + return f"ordered options disagreement for question_id {question_id!r}" + return None + + +def _profile_groups(declaration: DatasetReferenceSpec) -> tuple[tuple[str, ...], ...]: + raw = declaration.derivation.get("independent_source_groups") + if not isinstance(raw, (list, tuple)): + raise DatasetReferenceError( + "Profile-derived reference requires at least two independent source groups." + ) + groups: list[tuple[str, ...]] = [] + for item in raw: + if not isinstance(item, (list, tuple)) or not item: + raise DatasetReferenceError( + "Profile-derived reference requires at least two independent source groups." + ) + groups.append(tuple(str(source_id) for source_id in item)) + if len(groups) < 2: + raise DatasetReferenceError( + "Profile-derived reference requires at least two independent source groups." + ) + flattened = [source_id for group in groups for source_id in group] + if len(flattened) != len(set(flattened)) or set(flattened) != set(declaration.source_ids): + raise DatasetReferenceError( + "Profile-derived reference independent source groups must be disjoint and " + "cover every declared source." + ) + return tuple(groups) + + +def _artifact_payload(dataset: ExpectedDataset | pd.DataFrame, declaration: DatasetReferenceSpec) -> dict[str, Any]: + frame = dataset.frame if isinstance(dataset, ExpectedDataset) else dataset + return { + "spec": { + "benchmark": declaration.benchmark_name, + "split": declaration.split, + "hf_path": None, + "hf_subset": None, + "source_revision": None, + "normalization_version": NORMALIZATION_VERSION, + "transforms": [], + "output_name": declaration.benchmark_name, + }, + "content_digest": dataset_content_digest(frame), + "source": {}, + } + + +def _selection_payload( + frame: pd.DataFrame, declaration: DatasetReferenceSpec, artifact_id: str +) -> dict[str, Any]: + return { + "artifact_id": artifact_id, + "content_digest": dataset_content_digest(frame), + "sample_identities": dataset_sample_identities(frame), + "seed": declaration.selection_seed, + "n_samples": declaration.selection_n_samples, + "subject_filter": sorted(declaration.subject_filter), + } + + +def build_expected_dataset( + declaration: DatasetReferenceSpec, + opened_sources: Mapping[str, OpenedSource], +) -> ExpectedDataset: + """Build one semantic snapshot plus its trust-qualified reference record.""" + missing_sources = sorted(set(declaration.source_ids) - set(opened_sources)) + if missing_sources: + raise DatasetReferenceError( + f"Expected dataset declaration is missing opened source(s) {missing_sources}." + ) + if declaration.selection_source_id not in declaration.source_ids: + raise DatasetReferenceError("selection_source_id is not a declared dataset source.") + + normalized = { + source_id: _normalize_source(declaration, source_id, opened_sources[source_id]) + for source_id in declaration.source_ids + } + selected = { + source_id: _selected_frame(frame, declaration.expected_question_ids) + for source_id, frame in normalized.items() + } + if declaration.reference_kind == "profile_derived_reference_snapshot": + _profile_groups(declaration) + baseline = selected[declaration.selection_source_id] + for source_id in declaration.source_ids: + disagreement = _semantic_disagreement(baseline, selected[source_id]) + if disagreement is not None: + raise DatasetReferenceError( + f"Profile-derived reference source {source_id!r} has {disagreement}." + ) + + frame = selected[declaration.selection_source_id].copy() + try: + validate_normalized_dataset(frame, source="expected dataset reference") + except Exception as exc: + raise DatasetReferenceError(f"Expected dataset reference is invalid: {exc}") from exc + + artifact_payload = _artifact_payload(frame, declaration) + artifact_digest = integrity_digest(artifact_payload) + artifact_id = short_id("ds", artifact_payload) + selection_payload = _selection_payload(frame, declaration, artifact_id) + selection_digest = integrity_digest(selection_payload) + selection_id = short_id("sel", selection_payload) + selection_semantics = canonicalize( + { + "seed": declaration.selection_seed, + "n_samples": declaration.selection_n_samples, + "subject_filter": sorted(declaration.subject_filter), + } + ) + + limitations = list(declaration.limitations) + if declaration.revision is None and "publisher revision was not recorded" not in limitations: + limitations.append("publisher revision was not recorded") + if ( + declaration.reference_kind == "profile_derived_reference_snapshot" + and "not independently authenticated" not in limitations + ): + limitations.append("not independently authenticated") + + source_chain = [ + { + "source_id": source_id, + "logical_path": opened_sources[source_id].logical_path, + "sha256": opened_sources[source_id].sha256, + } + for source_id in declaration.source_ids + ] + derivation = canonicalize( + { + "schema_version": _REFERENCE_SCHEMA, + "dataset_id": declaration.dataset_id, + "reference_kind": declaration.reference_kind, + "trust_label": declaration.trust_label, + "source_chain": source_chain, + "selection_source_id": declaration.selection_source_id, + "columns": dict(declaration.columns), + "revision": declaration.revision, + "fingerprint": declaration.fingerprint, + "declared_derivation": dict(declaration.derivation), + "selected_question_ids": list(declaration.expected_question_ids), + "selection_semantics": selection_semantics, + "limitations": limitations, + } + ) + derivation_digest = integrity_digest(derivation) + question_set_digest = integrity_digest(list(declaration.expected_question_ids)) + snapshot_digest = integrity_digest( + { + "artifact_digest": artifact_digest, + "selection_digest": selection_digest, + "question_set_digest": question_set_digest, + "derivation_digest": derivation_digest, + "reference_kind": declaration.reference_kind, + "trust_label": declaration.trust_label, + } + ) + return ExpectedDataset( + dataset_id=declaration.dataset_id, + benchmark_name=declaration.benchmark_name, + split=declaration.split, + artifact_id=artifact_id, + artifact_digest=artifact_digest, + selection_id=selection_id, + selection_digest=selection_digest, + selection_semantics=selection_semantics, + frame=frame, + selected_question_ids=tuple(declaration.expected_question_ids), + reference_kind=declaration.reference_kind, + trust_label=declaration.trust_label, + question_set_digest=question_set_digest, + snapshot_digest=snapshot_digest, + derivation=derivation, + derivation_digest=derivation_digest, + limitations=tuple(limitations), + ) + + +def _safe_relative_path(value: Any, field: str) -> Path: + if not isinstance(value, str): + raise DatasetReferenceError(f"Expected dataset {field} must be a relative path.") + pure = PurePosixPath(value) + if pure.is_absolute() or ".." in pure.parts: + raise DatasetReferenceError(f"Expected dataset {field} is unsafe.") + return Path(*pure.parts) + + +def _snapshot_record(dataset: ExpectedDataset) -> dict[str, Any]: + snapshot_path = f"artifacts/datasets/{dataset.selection_id}.csv" + metadata_path = f"artifacts/datasets/{dataset.selection_id}.reference.json" + csv_bytes = dataset.frame.to_csv(index=False, lineterminator="\n").encode("utf-8") + record: dict[str, Any] = { + "schema_version": _SNAPSHOT_SCHEMA, + "dataset_id": dataset.dataset_id, + "benchmark_name": dataset.benchmark_name, + "split": dataset.split, + "artifact_id": dataset.artifact_id, + "artifact_digest": dataset.artifact_digest, + "selection_id": dataset.selection_id, + "selection_digest": dataset.selection_digest, + "selection_semantics": dataset.selection_semantics, + "selected_question_ids": list(dataset.selected_question_ids), + "question_set_digest": dataset.question_set_digest, + "snapshot_digest": dataset.snapshot_digest, + "reference_kind": dataset.reference_kind, + "trust_label": dataset.trust_label, + "derivation": dataset.derivation, + "derivation_digest": dataset.derivation_digest, + "limitations": list(dataset.limitations), + "run_snapshot_path": snapshot_path, + "reference_metadata_path": metadata_path, + "snapshot_sha256": sha256(csv_bytes).hexdigest(), + "row_count": len(dataset.frame), + } + record["record_digest"] = integrity_digest(record) + return record + + +def write_expected_snapshot(staged_run: Path, dataset: ExpectedDataset) -> dict[str, Any]: + """Atomically write the selected semantic CSV and its self-digesting record.""" + root = Path(staged_run) + record = _snapshot_record(dataset) + snapshot_path = root / _safe_relative_path(record["run_snapshot_path"], "run_snapshot_path") + metadata_path = root / _safe_relative_path( + record["reference_metadata_path"], "reference_metadata_path" + ) + atomic_write_text(snapshot_path, dataset.frame.to_csv(index=False, lineterminator="\n")) + atomic_write_json(metadata_path, record) + return record + + +def validate_expected_snapshot(run_dir: Path, record: Mapping[str, Any]) -> None: + """Recompute every snapshot, semantic identity, and reference digest.""" + root = Path(run_dir) + raw_record = dict(record) + claimed_digest = raw_record.pop("record_digest", None) + if claimed_digest != integrity_digest(raw_record): + raise DatasetReferenceError("Expected dataset reference record integrity failed.") + snapshot_path = root / _safe_relative_path(record.get("run_snapshot_path"), "run_snapshot_path") + metadata_path = root / _safe_relative_path( + record.get("reference_metadata_path"), "reference_metadata_path" + ) + try: + persisted = json.loads(metadata_path.read_text(encoding="utf-8")) + except (OSError, json.JSONDecodeError, UnicodeError) as exc: + raise DatasetReferenceError("Expected dataset reference metadata is unreadable.") from exc + if canonicalize(persisted) != canonicalize(dict(record)): + raise DatasetReferenceError("Expected dataset reference metadata integrity failed.") + try: + snapshot_bytes = snapshot_path.read_bytes() + except OSError as exc: + raise DatasetReferenceError("Expected dataset snapshot is unreadable.") from exc + if sha256(snapshot_bytes).hexdigest() != record.get("snapshot_sha256"): + raise DatasetReferenceError("Expected dataset snapshot integrity failed.") + try: + frame = pd.read_csv(BytesIO(snapshot_bytes), dtype=str, keep_default_na=False, na_filter=False) + validate_normalized_dataset(frame, source="expected dataset snapshot") + except Exception as exc: + raise DatasetReferenceError("Expected dataset snapshot integrity failed.") from exc + selected_ids = tuple(frame["question_id"].astype(str)) + if selected_ids != tuple(record.get("selected_question_ids", ())): + raise DatasetReferenceError("Expected dataset snapshot question ownership is invalid.") + if len(frame) != record.get("row_count"): + raise DatasetReferenceError("Expected dataset snapshot row count is invalid.") + if integrity_digest(list(selected_ids)) != record.get("question_set_digest"): + raise DatasetReferenceError("Expected dataset snapshot question-set identity is invalid.") + derivation = record.get("derivation") + if integrity_digest(derivation) != record.get("derivation_digest"): + raise DatasetReferenceError("Expected dataset derivation integrity failed.") + + artifact_payload = { + "spec": { + "benchmark": record.get("benchmark_name"), + "split": record.get("split"), + "hf_path": None, + "hf_subset": None, + "source_revision": None, + "normalization_version": NORMALIZATION_VERSION, + "transforms": [], + "output_name": record.get("benchmark_name"), + }, + "content_digest": dataset_content_digest(frame), + "source": {}, + } + artifact_digest = integrity_digest(artifact_payload) + artifact_id = short_id("ds", artifact_payload) + if artifact_digest != record.get("artifact_digest") or artifact_id != record.get("artifact_id"): + raise DatasetReferenceError("Expected dataset artifact identity is invalid.") + selection_semantics = record.get("selection_semantics") + if ( + not isinstance(selection_semantics, Mapping) + or set(selection_semantics) != {"seed", "n_samples", "subject_filter"} + or ( + selection_semantics["seed"] is not None + and not isinstance(selection_semantics["seed"], int) + ) + or ( + selection_semantics["n_samples"] is not None + and not isinstance(selection_semantics["n_samples"], int) + ) + or not isinstance(selection_semantics["subject_filter"], list) + or not all( + isinstance(subject, str) + for subject in selection_semantics["subject_filter"] + ) + ): + raise DatasetReferenceError("Expected dataset selection semantics are invalid.") + selection_payload = { + "artifact_id": artifact_id, + "content_digest": dataset_content_digest(frame), + "sample_identities": dataset_sample_identities(frame), + "seed": selection_semantics["seed"], + "n_samples": selection_semantics["n_samples"], + "subject_filter": selection_semantics["subject_filter"], + } + if ( + integrity_digest(selection_payload) != record.get("selection_digest") + or short_id("sel", selection_payload) != record.get("selection_id") + ): + raise DatasetReferenceError("Expected dataset selection identity is invalid.") + expected_snapshot_digest = integrity_digest( + { + "artifact_digest": record.get("artifact_digest"), + "selection_digest": record.get("selection_digest"), + "question_set_digest": record.get("question_set_digest"), + "derivation_digest": record.get("derivation_digest"), + "reference_kind": record.get("reference_kind"), + "trust_label": record.get("trust_label"), + } + ) + if expected_snapshot_digest != record.get("snapshot_digest"): + raise DatasetReferenceError("Expected dataset reference snapshot identity is invalid.") diff --git a/tests/importing/test_dataset_reference.py b/tests/importing/test_dataset_reference.py new file mode 100644 index 0000000..c3722a9 --- /dev/null +++ b/tests/importing/test_dataset_reference.py @@ -0,0 +1,436 @@ +from __future__ import annotations + +from dataclasses import replace +from hashlib import sha256 +from pathlib import Path + +import pandas as pd +import pytest + +from choicebench.datasets import ( + NORMALIZATION_VERSION, + dataset_content_digest, + dataset_sample_identities, +) +from choicebench.identity import integrity_digest, short_id +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.dataset_reference import ( + DatasetReferenceError, + build_expected_dataset, + validate_expected_snapshot, + write_expected_snapshot, +) +from choicebench.importing.schema import DatasetReferenceSpec +from choicebench.pipeline.options import build_option_map + + +def _csv_bytes(rows: list[dict[str, str]]) -> bytes: + frame = pd.DataFrame(rows) + return frame.to_csv(index=False, lineterminator="\n").encode("utf-8") + + +def _source( + source_id: str, + rows: list[dict[str, str]], + *, + logical_path: str | None = None, + audit_path: Path | None = None, +) -> OpenedSource: + data = _csv_bytes(rows) + return OpenedSource( + source_id=source_id, + audit_path=audit_path or Path(f"/machine-a/{source_id}.csv"), + logical_path=logical_path or f"publisher/{source_id}.csv", + data=data, + sha256=sha256(data).hexdigest(), + ) + + +def _rows(*, six_options: bool = False) -> list[dict[str, str]]: + option_count = 6 if six_options else 3 + questions = ( + ("q1", "First?", "B", ("one", "two", "three", "four", "five", "six")), + ("q2", "Second?", "A", ("red", "green", "blue", "cyan", "magenta", "yellow")), + ("q3", "Third?", "C", ("cat", "dog", "owl", "fox", "yak", "eel")), + ) + result: list[dict[str, str]] = [] + for question_id, question, gold, options in questions: + row = {"qid": question_id, "stem": question, "gold": gold} + row.update( + {f"option_{index + 1}": option for index, option in enumerate(options[:option_count])} + ) + result.append(row) + return result + + +def _columns(option_count: int = 3) -> dict[str, str]: + columns = { + "question_id": "qid", + "question_text": "stem", + "correct_option": "gold", + } + columns.update( + {f"choice_{chr(97 + index)}": f"option_{index + 1}" for index in range(option_count)} + ) + return columns + + +def _declaration( + *, + source_ids: tuple[str, ...] = ("reference",), + selection_source_id: str = "reference", + expected_question_ids: tuple[str, ...] = ("q2", "q1"), + reference_kind: str = "independent_input_snapshot", + trust_label: str = "publisher-input", + columns: dict[str, str] | None = None, + revision: str | None = "publisher-r1", + derivation: dict | None = None, + limitations: tuple[str, ...] = (), +) -> DatasetReferenceSpec: + return DatasetReferenceSpec( + dataset_id="historical-test", + benchmark_name="synthetic", + split="test", + reference_kind=reference_kind, # type: ignore[arg-type] + trust_label=trust_label, + source_ids=source_ids, + selection_source_id=selection_source_id, + expected_question_ids=expected_question_ids, + selection_seed=None, + selection_n_samples=None, + subject_filter=(), + selection_unknown_reasons={ + "selection_seed": "not recorded by producer", + "selection_n_samples": "not recorded by producer", + }, + columns=columns or _columns(), + revision=revision, + fingerprint="publisher-fingerprint", + derivation=derivation or {}, + limitations=limitations, + native_compatibility_identity=None, + ) + + +def _derived_declaration(**changes: object) -> DatasetReferenceSpec: + base = _declaration( + source_ids=("results-a", "results-b"), + selection_source_id="results-a", + reference_kind="profile_derived_reference_snapshot", + trust_label="cross-source-internal-consistency", + derivation={ + "method": "exact selected-row agreement", + "independent_source_groups": [["results-a"], ["results-b"]], + }, + ) + return replace(base, **changes) + + +def _semantic_payloads(dataset) -> tuple[dict, dict]: + artifact_payload = { + "spec": { + "benchmark": dataset.benchmark_name, + "split": dataset.split, + "hf_path": None, + "hf_subset": None, + "source_revision": None, + "normalization_version": NORMALIZATION_VERSION, + "transforms": [], + "output_name": dataset.benchmark_name, + }, + "content_digest": dataset_content_digest(dataset.frame), + "source": {}, + } + selection_payload = { + "artifact_id": short_id("ds", artifact_payload), + "content_digest": dataset_content_digest(dataset.frame), + "sample_identities": dataset_sample_identities(dataset.frame), + "seed": None, + "n_samples": None, + "subject_filter": [], + } + return artifact_payload, selection_payload + + +def test_independent_reference_builds_full_semantic_identities_and_stable_order(): + dataset = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + + artifact_payload, selection_payload = _semantic_payloads(dataset) + assert dataset.selected_question_ids == ("q2", "q1") + assert dataset.frame["question_id"].tolist() == ["q2", "q1"] + assert dataset.artifact_digest == integrity_digest(artifact_payload) + assert dataset.artifact_id == short_id("ds", artifact_payload) + assert dataset.selection_digest == integrity_digest(selection_payload) + assert dataset.selection_id == short_id("sel", selection_payload) + assert dataset.question_set_digest == integrity_digest(["q2", "q1"]) + assert dataset.reference_kind == "independent_input_snapshot" + + +def test_reference_kinds_are_not_equivalent(): + rows = _rows() + independent = build_expected_dataset( + _declaration(), {"reference": _source("reference", rows)} + ) + derived = build_expected_dataset( + _derived_declaration(), + { + "results-a": _source("results-a", rows), + "results-b": _source("results-b", rows), + }, + ) + + assert independent.artifact_id == derived.artifact_id + assert independent.selection_id == derived.selection_id + assert independent.reference_kind == "independent_input_snapshot" + assert derived.reference_kind == "profile_derived_reference_snapshot" + assert independent.derivation_digest != derived.derivation_digest + assert independent.snapshot_digest != derived.snapshot_digest + assert "not independently authenticated" in derived.limitations + + +def test_profile_derived_reference_requires_two_independent_source_groups(): + declaration = _derived_declaration( + derivation={ + "method": "exact agreement", + "independent_source_groups": [["results-a", "results-b"]], + } + ) + with pytest.raises(DatasetReferenceError, match="two independent source groups"): + build_expected_dataset( + declaration, + { + "results-a": _source("results-a", _rows()), + "results-b": _source("results-b", _rows()), + }, + ) + + +@pytest.mark.parametrize( + ("field", "replacement", "message"), + [ + ("stem", "Disagrees?", "question_text"), + ("gold", "C", "correct_option"), + ("option_2", "different option", "ordered options"), + ], +) +def test_profile_derived_reference_refuses_cross_source_semantic_disagreement( + field: str, replacement: str, message: str +): + changed = _rows() + changed[0][field] = replacement + with pytest.raises(DatasetReferenceError, match=message): + build_expected_dataset( + _derived_declaration(expected_question_ids=("q1", "q2")), + { + "results-a": _source("results-a", _rows()), + "results-b": _source("results-b", changed), + }, + ) + + +def test_derivation_digest_binds_complete_stable_reference_chain_not_audit_path(): + declaration = _declaration() + first = build_expected_dataset( + declaration, + { + "reference": _source( + "reference", _rows(), audit_path=Path("/machine-a/input.csv") + ) + }, + ) + relocated = build_expected_dataset( + declaration, + { + "reference": _source( + "reference", _rows(), audit_path=Path("/machine-b/input.csv") + ) + }, + ) + changed_chain = build_expected_dataset( + declaration, + { + "reference": _source( + "reference", _rows(), logical_path="publisher/mirror/input.csv" + ) + }, + ) + + assert first.derivation_digest == integrity_digest(first.derivation) + assert relocated.derivation_digest == first.derivation_digest + assert changed_chain.derivation_digest != first.derivation_digest + + +def test_unknown_publisher_revision_is_an_explicit_limitation(): + dataset = build_expected_dataset( + _declaration(revision=None), {"reference": _source("reference", _rows())} + ) + assert "publisher revision was not recorded" in dataset.limitations + + +def test_duplicate_selected_question_id_is_refused(): + with pytest.raises(DatasetReferenceError, match="duplicate selected question ID.*q1"): + build_expected_dataset( + _declaration(expected_question_ids=("q1", "q1")), + {"reference": _source("reference", _rows())}, + ) + + +def test_missing_selected_question_id_is_refused(): + with pytest.raises(DatasetReferenceError, match="missing selected question ID.*absent"): + build_expected_dataset( + _declaration(expected_question_ids=("q1", "absent")), + {"reference": _source("reference", _rows())}, + ) + + +def test_duplicate_source_question_id_is_refused(): + rows = _rows() + rows.append(dict(rows[0])) + with pytest.raises(DatasetReferenceError, match="duplicate question_id.*q1"): + build_expected_dataset( + _declaration(expected_question_ids=("q1",)), + {"reference": _source("reference", rows)}, + ) + + +def test_six_options_are_preserved_in_order(): + dataset = build_expected_dataset( + _declaration(columns=_columns(6), expected_question_ids=("q1",)), + {"reference": _source("reference", _rows(six_options=True))}, + ) + assert build_option_map(dataset.frame.iloc[0].to_dict()) == { + "A": "one", + "B": "two", + "C": "three", + "D": "four", + "E": "five", + "F": "six", + } + + +def test_arc_style_three_options_have_no_phantom_fourth_option(): + dataset = build_expected_dataset( + _declaration(expected_question_ids=("q1",)), + {"reference": _source("reference", _rows())}, + ) + assert build_option_map(dataset.frame.iloc[0].to_dict()) == { + "A": "one", + "B": "two", + "C": "three", + } + assert "choice_d" not in dataset.frame.columns + + +def test_source_mapping_trust_and_logical_provenance_are_not_semantic_identity(): + ordinary = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + renamed_rows = [ + { + "id": row["qid"], + "prompt": row["stem"], + "answer": row["gold"], + "a": row["option_1"], + "b": row["option_2"], + "c": row["option_3"], + } + for row in _rows() + ] + remapped = build_expected_dataset( + _declaration( + trust_label="archive-copy", + columns={ + "question_id": "id", + "question_text": "prompt", + "correct_option": "answer", + "choice_a": "a", + "choice_b": "b", + "choice_c": "c", + }, + ), + { + "reference": _source( + "reference", renamed_rows, logical_path="archive/remapped.csv" + ) + }, + ) + + assert remapped.artifact_id == ordinary.artifact_id + assert remapped.artifact_digest == ordinary.artifact_digest + assert remapped.selection_id == ordinary.selection_id + assert remapped.selection_digest == ordinary.selection_digest + assert remapped.derivation_digest != ordinary.derivation_digest + assert remapped.snapshot_digest != ordinary.snapshot_digest + + +def test_semantic_content_membership_and_order_change_applicable_identities(): + base_rows = _rows() + base = build_expected_dataset( + _declaration(), {"reference": _source("reference", base_rows)} + ) + changed_gold_rows = [dict(row) for row in base_rows] + changed_gold_rows[1]["gold"] = "B" + changed_gold = build_expected_dataset( + _declaration(), {"reference": _source("reference", changed_gold_rows)} + ) + changed_option_rows = [dict(row) for row in base_rows] + changed_option_rows[1]["option_1"] = "scarlet" + changed_option = build_expected_dataset( + _declaration(), {"reference": _source("reference", changed_option_rows)} + ) + changed_membership = build_expected_dataset( + _declaration(expected_question_ids=("q2", "q3")), + {"reference": _source("reference", base_rows)}, + ) + changed_order = build_expected_dataset( + _declaration(expected_question_ids=("q1", "q2")), + {"reference": _source("reference", base_rows)}, + ) + + for changed in (changed_gold, changed_option, changed_membership, changed_order): + assert changed.artifact_id != base.artifact_id + assert changed.selection_id != base.selection_id + + +def test_snapshot_is_written_atomically_and_self_validates(tmp_path: Path): + dataset = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + record = write_expected_snapshot(tmp_path, dataset) + + assert record["artifact_digest"] == dataset.artifact_digest + assert record["selection_digest"] == dataset.selection_digest + assert record["derivation_digest"] == dataset.derivation_digest + assert record["reference_kind"] == dataset.reference_kind + assert record["selected_question_ids"] == ["q2", "q1"] + assert (tmp_path / record["run_snapshot_path"]).is_file() + assert (tmp_path / record["reference_metadata_path"]).is_file() + validate_expected_snapshot(tmp_path, record) + + snapshot = tmp_path / record["run_snapshot_path"] + snapshot.write_text(snapshot.read_text(encoding="utf-8") + "tampered", encoding="utf-8") + with pytest.raises(DatasetReferenceError, match="snapshot.*integrity"): + validate_expected_snapshot(tmp_path, record) + + +def test_snapshot_self_validation_preserves_known_selection_semantics(tmp_path: Path): + declaration = replace( + _declaration(), + selection_seed=17, + selection_n_samples=2, + subject_filter=("science",), + selection_unknown_reasons={}, + ) + dataset = build_expected_dataset( + declaration, {"reference": _source("reference", _rows())} + ) + record = write_expected_snapshot(tmp_path, dataset) + + assert record["selection_semantics"] == { + "seed": 17, + "n_samples": 2, + "subject_filter": ["science"], + } + validate_expected_snapshot(tmp_path, record) From b9bb0b7482109df4705580c18f776173e7ecd7b7 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 01:09:10 +0300 Subject: [PATCH 13/47] fix: harden imported dataset identity handoff --- .../importing/dataset_reference.py | 391 ++++++++++++++++-- tests/importing/test_dataset_reference.py | 202 ++++++++- 2 files changed, 538 insertions(+), 55 deletions(-) diff --git a/src/choicebench/importing/dataset_reference.py b/src/choicebench/importing/dataset_reference.py index b4e9e4c..6282d35 100644 --- a/src/choicebench/importing/dataset_reference.py +++ b/src/choicebench/importing/dataset_reference.py @@ -12,7 +12,6 @@ import pandas as pd from choicebench.datasets import ( - NORMALIZATION_VERSION, dataset_content_digest, dataset_sample_identities, validate_normalized_dataset, @@ -35,9 +34,13 @@ class ExpectedDataset: split: str artifact_id: str artifact_digest: str + artifact_payload: Mapping[str, Any] selection_id: str selection_digest: str + selection_payload: Mapping[str, Any] selection_semantics: Mapping[str, Any] + selection_unknown_reasons: Mapping[str, str] + identity_mode: Literal["imported_semantic_fallback", "native_compatibility"] frame: pd.DataFrame selected_question_ids: tuple[str, ...] reference_kind: Literal[ @@ -54,6 +57,44 @@ class ExpectedDataset: _SNAPSHOT_SCHEMA = "choicebench.expected-dataset.v1" _REFERENCE_SCHEMA = "choicebench.dataset-reference.v1" _REQUIRED_FIELDS = ("question_id", "question_text", "correct_option") +_NATIVE_DATASET_SPEC_KEYS = { + "benchmark", + "split", + "hf_path", + "hf_subset", + "source_revision", + "normalization_version", + "transforms", + "output_name", +} +_SNAPSHOT_RECORD_KEYS = { + "schema_version", + "dataset_id", + "benchmark_name", + "split", + "artifact_id", + "artifact_digest", + "artifact_payload", + "selection_id", + "selection_digest", + "selection_payload", + "selection_semantics", + "selection_unknown_reasons", + "identity_mode", + "selected_question_ids", + "question_set_digest", + "snapshot_digest", + "reference_kind", + "trust_label", + "derivation", + "derivation_digest", + "limitations", + "run_snapshot_path", + "reference_metadata_path", + "snapshot_sha256", + "row_count", + "record_digest", +} def _source_frame(source_id: str, source: OpenedSource) -> pd.DataFrame: @@ -140,23 +181,57 @@ def _normalize_source( f"for question_id {question_id!r}." ) normalized_choices = [] + source_indices: set[int] = set() for index, item in enumerate(choices): if not isinstance(item, Mapping) or not isinstance(item.get("text"), str): raise DatasetReferenceError( f"Expected dataset source {source_id!r} has malformed ordered options " f"for question_id {question_id!r}." ) + raw_source_index = item.get("source_index", index) + try: + if isinstance(raw_source_index, bool): + raise ValueError + source_index = int(raw_source_index) + except (TypeError, ValueError) as exc: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has an invalid source_index " + f"for question_id {question_id!r}." + ) from exc + if source_index < 0 or source_index in source_indices: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has an invalid source_index " + f"for question_id {question_id!r}." + ) + if not item["text"].strip(): + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has an empty structured " + f"option for question_id {question_id!r}." + ) + source_indices.add(source_index) normalized_choices.append( { "text": item["text"], - "source_index": int(item.get("source_index", index)), + "source_index": source_index, } ) else: + choice_values = [str(raw_row[mapping[key]]) for key in choice_keys] + nonempty_indices = [ + index for index, value in enumerate(choice_values) if value.strip() + ] + if nonempty_indices: + last_nonempty = nonempty_indices[-1] + if any(not value.strip() for value in choice_values[:last_nonempty]): + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has an empty option followed " + f"by a populated option for question_id {question_id!r}." + ) + choice_values = choice_values[: last_nonempty + 1] normalized_choices = [ - {"text": raw_row[mapping[key]], "source_index": index} - for index, key in enumerate(choice_keys) - if str(raw_row[mapping[key]]).strip() + {"text": value, "source_index": index} + for index, value in enumerate(choice_values) + if value.strip() ] record["choices_json"] = json.dumps( normalized_choices, ensure_ascii=True, separators=(",", ":") @@ -246,21 +321,14 @@ def _profile_groups(declaration: DatasetReferenceSpec) -> tuple[tuple[str, ...], return tuple(groups) -def _artifact_payload(dataset: ExpectedDataset | pd.DataFrame, declaration: DatasetReferenceSpec) -> dict[str, Any]: - frame = dataset.frame if isinstance(dataset, ExpectedDataset) else dataset +def _fallback_artifact_payload( + frame: pd.DataFrame, declaration: DatasetReferenceSpec +) -> dict[str, Any]: return { - "spec": { - "benchmark": declaration.benchmark_name, - "split": declaration.split, - "hf_path": None, - "hf_subset": None, - "source_revision": None, - "normalization_version": NORMALIZATION_VERSION, - "transforms": [], - "output_name": declaration.benchmark_name, - }, + "schema_version": "choicebench.semantic-dataset.v1", + "benchmark": declaration.benchmark_name, + "split": declaration.split, "content_digest": dataset_content_digest(frame), - "source": {}, } @@ -277,6 +345,98 @@ def _selection_payload( } +def _validated_identity_records( + frame: pd.DataFrame, declaration: DatasetReferenceSpec +) -> tuple[ + Mapping[str, Any], + str, + str, + Mapping[str, Any], + str, + str, + Literal["imported_semantic_fallback", "native_compatibility"], +]: + native = declaration.native_compatibility_identity + if native is None: + artifact_payload = canonicalize(_fallback_artifact_payload(frame, declaration)) + artifact_digest = integrity_digest(artifact_payload) + artifact_id = short_id("ds", artifact_payload) + selection_payload = canonicalize( + _selection_payload(frame, declaration, artifact_id) + ) + return ( + artifact_payload, + artifact_digest, + artifact_id, + selection_payload, + integrity_digest(selection_payload), + short_id("sel", selection_payload), + "imported_semantic_fallback", + ) + + required = { + "artifact_payload", + "artifact_digest", + "artifact_id", + "selection_payload", + "selection_digest", + "selection_id", + } + if not isinstance(native, Mapping) or set(native) != required: + raise DatasetReferenceError( + "Expected dataset native compatibility identity has invalid fields." + ) + artifact_payload = canonicalize(native["artifact_payload"]) + selection_payload = canonicalize(native["selection_payload"]) + artifact_digest = integrity_digest(artifact_payload) + artifact_id = short_id("ds", artifact_payload) + selection_digest = integrity_digest(selection_payload) + selection_id = short_id("sel", selection_payload) + if ( + native["artifact_digest"] != artifact_digest + or native["artifact_id"] != artifact_id + or native["selection_digest"] != selection_digest + or native["selection_id"] != selection_id + ): + raise DatasetReferenceError( + "Expected dataset native compatibility identity claims are invalid." + ) + if not isinstance(artifact_payload, Mapping) or set(artifact_payload) != { + "spec", + "content_digest", + "source", + }: + raise DatasetReferenceError( + "Expected dataset native compatibility artifact payload is invalid." + ) + spec = artifact_payload.get("spec") + if ( + not isinstance(spec, Mapping) + or set(spec) != _NATIVE_DATASET_SPEC_KEYS + or not isinstance(artifact_payload.get("source"), Mapping) + or spec.get("benchmark") != declaration.benchmark_name + or spec.get("split") != declaration.split + or artifact_payload.get("content_digest") != dataset_content_digest(frame) + ): + raise DatasetReferenceError( + "Expected dataset native compatibility content does not own the selected rows." + ) + expected_selection = _selection_payload(frame, declaration, artifact_id) + if selection_payload != canonicalize(expected_selection): + raise DatasetReferenceError( + "Expected dataset native compatibility selection does not own the selected rows." + ) + return ( + artifact_payload, + artifact_digest, + artifact_id, + selection_payload, + selection_digest, + selection_id, + "native_compatibility", + ) + + def build_expected_dataset( declaration: DatasetReferenceSpec, opened_sources: Mapping[str, OpenedSource], @@ -314,12 +474,31 @@ def build_expected_dataset( except Exception as exc: raise DatasetReferenceError(f"Expected dataset reference is invalid: {exc}") from exc - artifact_payload = _artifact_payload(frame, declaration) - artifact_digest = integrity_digest(artifact_payload) - artifact_id = short_id("ds", artifact_payload) - selection_payload = _selection_payload(frame, declaration, artifact_id) - selection_digest = integrity_digest(selection_payload) - selection_id = short_id("sel", selection_payload) + unknown_reasons = canonicalize(dict(declaration.selection_unknown_reasons)) + expected_unknown_keys = { + field + for field, value in ( + ("selection_seed", declaration.selection_seed), + ("selection_n_samples", declaration.selection_n_samples), + ) + if value is None + } + if set(unknown_reasons) != expected_unknown_keys or not all( + isinstance(reason, str) and reason.strip() + for reason in unknown_reasons.values() + ): + raise DatasetReferenceError( + "Expected dataset selection unknown reasons do not match null semantics." + ) + ( + artifact_payload, + artifact_digest, + artifact_id, + selection_payload, + selection_digest, + selection_id, + identity_mode, + ) = _validated_identity_records(frame, declaration) selection_semantics = canonicalize( { "seed": declaration.selection_seed, @@ -359,6 +538,8 @@ def build_expected_dataset( "declared_derivation": dict(declaration.derivation), "selected_question_ids": list(declaration.expected_question_ids), "selection_semantics": selection_semantics, + "selection_unknown_reasons": unknown_reasons, + "identity_mode": identity_mode, "limitations": limitations, } ) @@ -380,9 +561,13 @@ def build_expected_dataset( split=declaration.split, artifact_id=artifact_id, artifact_digest=artifact_digest, + artifact_payload=artifact_payload, selection_id=selection_id, selection_digest=selection_digest, + selection_payload=selection_payload, selection_semantics=selection_semantics, + selection_unknown_reasons=unknown_reasons, + identity_mode=identity_mode, frame=frame, selected_question_ids=tuple(declaration.expected_question_ids), reference_kind=declaration.reference_kind, @@ -404,6 +589,27 @@ def _safe_relative_path(value: Any, field: str) -> Path: return Path(*pure.parts) +def _contained_path(root: Path, value: Any, field: str) -> Path: + relative = _safe_relative_path(value, field) + try: + canonical_root = root.resolve(strict=True) + except OSError as exc: + raise DatasetReferenceError("Expected dataset output root is unreadable.") from exc + candidate = canonical_root / relative + current = canonical_root + for part in relative.parts: + current = current / part + if current.is_symlink(): + raise DatasetReferenceError( + f"Expected dataset {field} contains an unsafe symlink." + ) + try: + candidate.resolve(strict=False).relative_to(canonical_root) + except (OSError, ValueError) as exc: + raise DatasetReferenceError(f"Expected dataset {field} is unsafe.") from exc + return candidate + + def _snapshot_record(dataset: ExpectedDataset) -> dict[str, Any]: snapshot_path = f"artifacts/datasets/{dataset.selection_id}.csv" metadata_path = f"artifacts/datasets/{dataset.selection_id}.reference.json" @@ -415,9 +621,13 @@ def _snapshot_record(dataset: ExpectedDataset) -> dict[str, Any]: "split": dataset.split, "artifact_id": dataset.artifact_id, "artifact_digest": dataset.artifact_digest, + "artifact_payload": dataset.artifact_payload, "selection_id": dataset.selection_id, "selection_digest": dataset.selection_digest, + "selection_payload": dataset.selection_payload, "selection_semantics": dataset.selection_semantics, + "selection_unknown_reasons": dataset.selection_unknown_reasons, + "identity_mode": dataset.identity_mode, "selected_question_ids": list(dataset.selected_question_ids), "question_set_digest": dataset.question_set_digest, "snapshot_digest": dataset.snapshot_digest, @@ -439,8 +649,9 @@ def write_expected_snapshot(staged_run: Path, dataset: ExpectedDataset) -> dict[ """Atomically write the selected semantic CSV and its self-digesting record.""" root = Path(staged_run) record = _snapshot_record(dataset) - snapshot_path = root / _safe_relative_path(record["run_snapshot_path"], "run_snapshot_path") - metadata_path = root / _safe_relative_path( + snapshot_path = _contained_path(root, record["run_snapshot_path"], "run_snapshot_path") + metadata_path = _contained_path( + root, record["reference_metadata_path"], "reference_metadata_path" ) atomic_write_text(snapshot_path, dataset.frame.to_csv(index=False, lineterminator="\n")) @@ -448,16 +659,79 @@ def write_expected_snapshot(staged_run: Path, dataset: ExpectedDataset) -> dict[ return record +def _validate_snapshot_record_shape(record: Mapping[str, Any]) -> None: + if set(record) != _SNAPSHOT_RECORD_KEYS: + raise DatasetReferenceError("Expected dataset reference record fields are invalid.") + if record["schema_version"] != _SNAPSHOT_SCHEMA: + raise DatasetReferenceError("Expected dataset reference schema is unsupported.") + nonempty_strings = { + "dataset_id", + "benchmark_name", + "split", + "artifact_id", + "artifact_digest", + "selection_id", + "selection_digest", + "question_set_digest", + "snapshot_digest", + "trust_label", + "derivation_digest", + "run_snapshot_path", + "reference_metadata_path", + "snapshot_sha256", + "record_digest", + } + if any( + not isinstance(record[field], str) or not record[field] + for field in nonempty_strings + ): + raise DatasetReferenceError("Expected dataset reference record fields are invalid.") + if record["reference_kind"] not in { + "independent_input_snapshot", + "profile_derived_reference_snapshot", + } or record["identity_mode"] not in { + "imported_semantic_fallback", + "native_compatibility", + }: + raise DatasetReferenceError("Expected dataset reference record fields are invalid.") + if not all( + isinstance(record[field], Mapping) + for field in ( + "artifact_payload", + "selection_payload", + "selection_semantics", + "selection_unknown_reasons", + "derivation", + ) + ): + raise DatasetReferenceError("Expected dataset reference record fields are invalid.") + selected = record["selected_question_ids"] + limitations = record["limitations"] + if ( + not isinstance(selected, list) + or not selected + or not all(isinstance(value, str) and value for value in selected) + or len(selected) != len(set(selected)) + or not isinstance(limitations, list) + or not all(isinstance(value, str) for value in limitations) + or isinstance(record["row_count"], bool) + or not isinstance(record["row_count"], int) + or record["row_count"] < 1 + ): + raise DatasetReferenceError("Expected dataset reference record fields are invalid.") + + def validate_expected_snapshot(run_dir: Path, record: Mapping[str, Any]) -> None: """Recompute every snapshot, semantic identity, and reference digest.""" + _validate_snapshot_record_shape(record) root = Path(run_dir) raw_record = dict(record) claimed_digest = raw_record.pop("record_digest", None) if claimed_digest != integrity_digest(raw_record): raise DatasetReferenceError("Expected dataset reference record integrity failed.") - snapshot_path = root / _safe_relative_path(record.get("run_snapshot_path"), "run_snapshot_path") - metadata_path = root / _safe_relative_path( - record.get("reference_metadata_path"), "reference_metadata_path" + snapshot_path = _contained_path(root, record["run_snapshot_path"], "run_snapshot_path") + metadata_path = _contained_path( + root, record["reference_metadata_path"], "reference_metadata_path" ) try: persisted = json.loads(metadata_path.read_text(encoding="utf-8")) @@ -487,25 +761,38 @@ def validate_expected_snapshot(run_dir: Path, record: Mapping[str, Any]) -> None if integrity_digest(derivation) != record.get("derivation_digest"): raise DatasetReferenceError("Expected dataset derivation integrity failed.") - artifact_payload = { - "spec": { - "benchmark": record.get("benchmark_name"), - "split": record.get("split"), - "hf_path": None, - "hf_subset": None, - "source_revision": None, - "normalization_version": NORMALIZATION_VERSION, - "transforms": [], - "output_name": record.get("benchmark_name"), - }, - "content_digest": dataset_content_digest(frame), - "source": {}, - } + artifact_payload = canonicalize(record["artifact_payload"]) + if record["identity_mode"] == "imported_semantic_fallback": + expected_artifact_payload = { + "schema_version": "choicebench.semantic-dataset.v1", + "benchmark": record["benchmark_name"], + "split": record["split"], + "content_digest": dataset_content_digest(frame), + } + if artifact_payload != expected_artifact_payload: + raise DatasetReferenceError("Expected dataset artifact identity is invalid.") + else: + if not isinstance(artifact_payload, Mapping) or set(artifact_payload) != { + "spec", + "content_digest", + "source", + }: + raise DatasetReferenceError("Expected dataset artifact identity is invalid.") + native_spec = artifact_payload["spec"] + if ( + not isinstance(native_spec, Mapping) + or set(native_spec) != _NATIVE_DATASET_SPEC_KEYS + or not isinstance(artifact_payload["source"], Mapping) + or native_spec.get("benchmark") != record["benchmark_name"] + or native_spec.get("split") != record["split"] + or artifact_payload["content_digest"] != dataset_content_digest(frame) + ): + raise DatasetReferenceError("Expected dataset artifact identity is invalid.") artifact_digest = integrity_digest(artifact_payload) artifact_id = short_id("ds", artifact_payload) if artifact_digest != record.get("artifact_digest") or artifact_id != record.get("artifact_id"): raise DatasetReferenceError("Expected dataset artifact identity is invalid.") - selection_semantics = record.get("selection_semantics") + selection_semantics = record["selection_semantics"] if ( not isinstance(selection_semantics, Mapping) or set(selection_semantics) != {"seed", "n_samples", "subject_filter"} @@ -532,6 +819,22 @@ def validate_expected_snapshot(run_dir: Path, record: Mapping[str, Any]) -> None "n_samples": selection_semantics["n_samples"], "subject_filter": selection_semantics["subject_filter"], } + if canonicalize(record["selection_payload"]) != selection_payload: + raise DatasetReferenceError("Expected dataset selection identity is invalid.") + unknown_reasons = record["selection_unknown_reasons"] + expected_unknown_keys = { + field + for field, value in ( + ("selection_seed", selection_semantics["seed"]), + ("selection_n_samples", selection_semantics["n_samples"]), + ) + if value is None + } + if set(unknown_reasons) != expected_unknown_keys or not all( + isinstance(reason, str) and reason.strip() + for reason in unknown_reasons.values() + ): + raise DatasetReferenceError("Expected dataset selection semantics are invalid.") if ( integrity_digest(selection_payload) != record.get("selection_digest") or short_id("sel", selection_payload) != record.get("selection_id") diff --git a/tests/importing/test_dataset_reference.py b/tests/importing/test_dataset_reference.py index c3722a9..c9757dd 100644 --- a/tests/importing/test_dataset_reference.py +++ b/tests/importing/test_dataset_reference.py @@ -2,6 +2,7 @@ from dataclasses import replace from hashlib import sha256 +import json from pathlib import Path import pandas as pd @@ -128,18 +129,10 @@ def _derived_declaration(**changes: object) -> DatasetReferenceSpec: def _semantic_payloads(dataset) -> tuple[dict, dict]: artifact_payload = { - "spec": { - "benchmark": dataset.benchmark_name, - "split": dataset.split, - "hf_path": None, - "hf_subset": None, - "source_revision": None, - "normalization_version": NORMALIZATION_VERSION, - "transforms": [], - "output_name": dataset.benchmark_name, - }, + "schema_version": "choicebench.semantic-dataset.v1", + "benchmark": dataset.benchmark_name, + "split": dataset.split, "content_digest": dataset_content_digest(dataset.frame), - "source": {}, } selection_payload = { "artifact_id": short_id("ds", artifact_payload), @@ -162,8 +155,15 @@ def test_independent_reference_builds_full_semantic_identities_and_stable_order( assert dataset.frame["question_id"].tolist() == ["q2", "q1"] assert dataset.artifact_digest == integrity_digest(artifact_payload) assert dataset.artifact_id == short_id("ds", artifact_payload) + assert dataset.artifact_payload == artifact_payload assert dataset.selection_digest == integrity_digest(selection_payload) assert dataset.selection_id == short_id("sel", selection_payload) + assert dataset.selection_payload == selection_payload + assert dataset.identity_mode == "imported_semantic_fallback" + assert dataset.selection_unknown_reasons == { + "selection_seed": "not recorded by producer", + "selection_n_samples": "not recorded by producer", + } assert dataset.question_set_digest == integrity_digest(["q2", "q1"]) assert dataset.reference_kind == "independent_input_snapshot" @@ -434,3 +434,183 @@ def test_snapshot_self_validation_preserves_known_selection_semantics(tmp_path: "subject_filter": ["science"], } validate_expected_snapshot(tmp_path, record) + + +def test_validated_native_compatibility_identity_is_preserved_exactly(): + fallback = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + native_artifact_payload = { + "spec": { + "benchmark": "synthetic", + "split": "test", + "hf_path": "publisher/synthetic", + "hf_subset": None, + "source_revision": "immutable-r1", + "normalization_version": NORMALIZATION_VERSION, + "transforms": [], + "output_name": "synthetic", + }, + "content_digest": dataset_content_digest(fallback.frame), + "source": {"revision": "immutable-r1"}, + } + artifact_digest = integrity_digest(native_artifact_payload) + artifact_id = short_id("ds", native_artifact_payload) + native_selection_payload = { + "artifact_id": artifact_id, + "content_digest": dataset_content_digest(fallback.frame), + "sample_identities": dataset_sample_identities(fallback.frame), + "seed": 23, + "n_samples": 2, + "subject_filter": ["science"], + } + native = { + "artifact_payload": native_artifact_payload, + "artifact_digest": artifact_digest, + "artifact_id": artifact_id, + "selection_payload": native_selection_payload, + "selection_digest": integrity_digest(native_selection_payload), + "selection_id": short_id("sel", native_selection_payload), + } + declaration = replace( + _declaration(), + selection_seed=23, + selection_n_samples=2, + subject_filter=("science",), + selection_unknown_reasons={}, + native_compatibility_identity=native, + ) + + dataset = build_expected_dataset( + declaration, {"reference": _source("reference", _rows())} + ) + + assert dataset.identity_mode == "native_compatibility" + assert dataset.artifact_payload == native_artifact_payload + assert dataset.selection_payload == native_selection_payload + assert dataset.artifact_id == artifact_id + assert dataset.selection_id == native["selection_id"] + + +def test_native_compatibility_identity_must_own_the_selected_semantic_rows(): + fallback = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + artifact_payload = { + "spec": { + "benchmark": "synthetic", + "split": "test", + "hf_path": None, + "hf_subset": None, + "source_revision": None, + "normalization_version": NORMALIZATION_VERSION, + "transforms": [], + "output_name": "synthetic", + }, + "content_digest": "0" * 64, + "source": {}, + } + artifact_id = short_id("ds", artifact_payload) + selection_payload = { + "artifact_id": artifact_id, + "content_digest": dataset_content_digest(fallback.frame), + "sample_identities": dataset_sample_identities(fallback.frame), + "seed": 1, + "n_samples": 2, + "subject_filter": [], + } + native = { + "artifact_payload": artifact_payload, + "artifact_digest": integrity_digest(artifact_payload), + "artifact_id": artifact_id, + "selection_payload": selection_payload, + "selection_digest": integrity_digest(selection_payload), + "selection_id": short_id("sel", selection_payload), + } + declaration = replace( + _declaration(), + selection_seed=1, + selection_n_samples=2, + selection_unknown_reasons={}, + native_compatibility_identity=native, + ) + + with pytest.raises(DatasetReferenceError, match="native compatibility.*content"): + build_expected_dataset( + declaration, {"reference": _source("reference", _rows())} + ) + + +@pytest.mark.parametrize("source_index", ["not-an-integer", -1, 1]) +def test_structured_choice_source_indices_fail_closed(source_index): + rows = [ + { + "qid": "q1", + "stem": "First?", + "gold": "A", + "choices": json.dumps( + [ + {"text": "one", "source_index": source_index}, + {"text": "two", "source_index": 1}, + ] + ), + } + ] + declaration = _declaration( + expected_question_ids=("q1",), + columns={ + "question_id": "qid", + "question_text": "stem", + "correct_option": "gold", + "choices_json": "choices", + }, + ) + + with pytest.raises(DatasetReferenceError, match="source_index"): + build_expected_dataset( + declaration, {"reference": _source("reference", rows)} + ) + + +def test_ordered_options_reject_an_empty_middle_value(): + rows = _rows() + rows[0]["option_2"] = "" + + with pytest.raises(DatasetReferenceError, match="empty option.*followed"): + build_expected_dataset( + _declaration(expected_question_ids=("q1",)), + {"reference": _source("reference", rows)}, + ) + + +@pytest.mark.parametrize("mutation", ["schema", "extra_field"]) +def test_snapshot_record_schema_is_exact_and_fail_closed(tmp_path: Path, mutation: str): + dataset = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + record = write_expected_snapshot(tmp_path, dataset) + if mutation == "schema": + record["schema_version"] = "choicebench.expected-dataset.v999" + else: + record["unexpected"] = "not allowed" + unsigned = dict(record) + unsigned.pop("record_digest") + record["record_digest"] = integrity_digest(unsigned) + + with pytest.raises(DatasetReferenceError, match="schema|fields"): + validate_expected_snapshot(tmp_path, record) + + +def test_snapshot_validation_rejects_symlinked_artifact_paths(tmp_path: Path): + dataset = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + record = write_expected_snapshot(tmp_path, dataset) + snapshot = tmp_path / record["run_snapshot_path"] + outside = tmp_path.parent / f"{tmp_path.name}-outside.csv" + outside.write_bytes(snapshot.read_bytes()) + snapshot.unlink() + snapshot.symlink_to(outside) + + with pytest.raises(DatasetReferenceError, match="symlink|unsafe"): + validate_expected_snapshot(tmp_path, record) From 68f1892cbfa54fb3224673d8e3e9d78b0b8b2413 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 01:15:03 +0300 Subject: [PATCH 14/47] fix: separate dataset artifact and selection identity --- .../importing/dataset_reference.py | 61 ++++++++++++++----- tests/importing/test_dataset_reference.py | 54 +++++++++++++--- 2 files changed, 94 insertions(+), 21 deletions(-) diff --git a/src/choicebench/importing/dataset_reference.py b/src/choicebench/importing/dataset_reference.py index 6282d35..dd7ffff 100644 --- a/src/choicebench/importing/dataset_reference.py +++ b/src/choicebench/importing/dataset_reference.py @@ -41,6 +41,7 @@ class ExpectedDataset: selection_semantics: Mapping[str, Any] selection_unknown_reasons: Mapping[str, str] identity_mode: Literal["imported_semantic_fallback", "native_compatibility"] + artifact_frame: pd.DataFrame frame: pd.DataFrame selected_question_ids: tuple[str, ...] reference_kind: Literal[ @@ -151,6 +152,14 @@ def _normalize_source( has_structured_choices = "choices_json" in mapping if not choice_keys and not has_structured_choices: raise DatasetReferenceError("Expected dataset mapping declares no ordered options.") + if choice_keys: + expected_keys = tuple( + f"choice_{chr(ord('a') + index)}" for index in range(len(choice_keys)) + ) + if choice_keys != expected_keys: + raise DatasetReferenceError( + "Expected dataset ordered choice keys must be contiguous from choice_a." + ) ordinary_keys = sorted( key for key in mapping if key not in choice_keys and key != "choices_json" @@ -346,7 +355,9 @@ def _selection_payload( def _validated_identity_records( - frame: pd.DataFrame, declaration: DatasetReferenceSpec + artifact_frame: pd.DataFrame, + selected_frame: pd.DataFrame, + declaration: DatasetReferenceSpec, ) -> tuple[ Mapping[str, Any], str, @@ -358,11 +369,13 @@ def _validated_identity_records( ]: native = declaration.native_compatibility_identity if native is None: - artifact_payload = canonicalize(_fallback_artifact_payload(frame, declaration)) + artifact_payload = canonicalize( + _fallback_artifact_payload(artifact_frame, declaration) + ) artifact_digest = integrity_digest(artifact_payload) artifact_id = short_id("ds", artifact_payload) selection_payload = canonicalize( - _selection_payload(frame, declaration, artifact_id) + _selection_payload(selected_frame, declaration, artifact_id) ) return ( artifact_payload, @@ -416,12 +429,13 @@ def _validated_identity_records( or not isinstance(artifact_payload.get("source"), Mapping) or spec.get("benchmark") != declaration.benchmark_name or spec.get("split") != declaration.split - or artifact_payload.get("content_digest") != dataset_content_digest(frame) + or artifact_payload.get("content_digest") + != dataset_content_digest(artifact_frame) ): raise DatasetReferenceError( "Expected dataset native compatibility content does not own the selected rows." ) - expected_selection = _selection_payload(frame, declaration, artifact_id) + expected_selection = _selection_payload(selected_frame, declaration, artifact_id) if selection_payload != canonicalize(expected_selection): raise DatasetReferenceError( "Expected dataset native compatibility selection does not own the selected rows." @@ -454,6 +468,15 @@ def build_expected_dataset( source_id: _normalize_source(declaration, source_id, opened_sources[source_id]) for source_id in declaration.source_ids } + for source_id, source_frame in normalized.items(): + try: + validate_normalized_dataset( + source_frame, source=f"expected dataset source {source_id!r}" + ) + except Exception as exc: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} is invalid: {exc}" + ) from exc selected = { source_id: _selected_frame(frame, declaration.expected_question_ids) for source_id, frame in normalized.items() @@ -468,6 +491,7 @@ def build_expected_dataset( f"Profile-derived reference source {source_id!r} has {disagreement}." ) + artifact_frame = normalized[declaration.selection_source_id].copy() frame = selected[declaration.selection_source_id].copy() try: validate_normalized_dataset(frame, source="expected dataset reference") @@ -498,7 +522,7 @@ def build_expected_dataset( selection_digest, selection_id, identity_mode, - ) = _validated_identity_records(frame, declaration) + ) = _validated_identity_records(artifact_frame, frame, declaration) selection_semantics = canonicalize( { "seed": declaration.selection_seed, @@ -568,6 +592,7 @@ def build_expected_dataset( selection_semantics=selection_semantics, selection_unknown_reasons=unknown_reasons, identity_mode=identity_mode, + artifact_frame=artifact_frame, frame=frame, selected_question_ids=tuple(declaration.expected_question_ids), reference_kind=declaration.reference_kind, @@ -763,13 +788,15 @@ def validate_expected_snapshot(run_dir: Path, record: Mapping[str, Any]) -> None artifact_payload = canonicalize(record["artifact_payload"]) if record["identity_mode"] == "imported_semantic_fallback": - expected_artifact_payload = { - "schema_version": "choicebench.semantic-dataset.v1", - "benchmark": record["benchmark_name"], - "split": record["split"], - "content_digest": dataset_content_digest(frame), - } - if artifact_payload != expected_artifact_payload: + if ( + not isinstance(artifact_payload, Mapping) + or set(artifact_payload) + != {"schema_version", "benchmark", "split", "content_digest"} + or artifact_payload["schema_version"] + != "choicebench.semantic-dataset.v1" + or artifact_payload["benchmark"] != record["benchmark_name"] + or artifact_payload["split"] != record["split"] + ): raise DatasetReferenceError("Expected dataset artifact identity is invalid.") else: if not isinstance(artifact_payload, Mapping) or set(artifact_payload) != { @@ -785,9 +812,15 @@ def validate_expected_snapshot(run_dir: Path, record: Mapping[str, Any]) -> None or not isinstance(artifact_payload["source"], Mapping) or native_spec.get("benchmark") != record["benchmark_name"] or native_spec.get("split") != record["split"] - or artifact_payload["content_digest"] != dataset_content_digest(frame) ): raise DatasetReferenceError("Expected dataset artifact identity is invalid.") + content_digest = artifact_payload.get("content_digest") + if ( + not isinstance(content_digest, str) + or len(content_digest) != 64 + or any(character not in "0123456789abcdef" for character in content_digest) + ): + raise DatasetReferenceError("Expected dataset artifact identity is invalid.") artifact_digest = integrity_digest(artifact_payload) artifact_id = short_id("ds", artifact_payload) if artifact_digest != record.get("artifact_digest") or artifact_id != record.get("artifact_id"): diff --git a/tests/importing/test_dataset_reference.py b/tests/importing/test_dataset_reference.py index c9757dd..4b528a5 100644 --- a/tests/importing/test_dataset_reference.py +++ b/tests/importing/test_dataset_reference.py @@ -132,7 +132,7 @@ def _semantic_payloads(dataset) -> tuple[dict, dict]: "schema_version": "choicebench.semantic-dataset.v1", "benchmark": dataset.benchmark_name, "split": dataset.split, - "content_digest": dataset_content_digest(dataset.frame), + "content_digest": dataset_content_digest(dataset.artifact_frame), } selection_payload = { "artifact_id": short_id("ds", artifact_payload), @@ -389,9 +389,12 @@ def test_semantic_content_membership_and_order_change_applicable_identities(): {"reference": _source("reference", base_rows)}, ) - for changed in (changed_gold, changed_option, changed_membership, changed_order): + for changed in (changed_gold, changed_option): assert changed.artifact_id != base.artifact_id assert changed.selection_id != base.selection_id + for changed in (changed_membership, changed_order): + assert changed.artifact_id == base.artifact_id + assert changed.selection_id != base.selection_id def test_snapshot_is_written_atomically_and_self_validates(tmp_path: Path): @@ -437,8 +440,9 @@ def test_snapshot_self_validation_preserves_known_selection_semantics(tmp_path: def test_validated_native_compatibility_identity_is_preserved_exactly(): - fallback = build_expected_dataset( - _declaration(), {"reference": _source("reference", _rows())} + full_artifact = build_expected_dataset( + _declaration(expected_question_ids=("q1", "q2", "q3")), + {"reference": _source("reference", _rows())}, ) native_artifact_payload = { "spec": { @@ -451,15 +455,15 @@ def test_validated_native_compatibility_identity_is_preserved_exactly(): "transforms": [], "output_name": "synthetic", }, - "content_digest": dataset_content_digest(fallback.frame), + "content_digest": dataset_content_digest(full_artifact.frame), "source": {"revision": "immutable-r1"}, } artifact_digest = integrity_digest(native_artifact_payload) artifact_id = short_id("ds", native_artifact_payload) native_selection_payload = { "artifact_id": artifact_id, - "content_digest": dataset_content_digest(fallback.frame), - "sample_identities": dataset_sample_identities(fallback.frame), + "content_digest": dataset_content_digest(full_artifact.frame.iloc[[1, 0]]), + "sample_identities": dataset_sample_identities(full_artifact.frame.iloc[[1, 0]]), "seed": 23, "n_samples": 2, "subject_filter": ["science"], @@ -490,6 +494,8 @@ def test_validated_native_compatibility_identity_is_preserved_exactly(): assert dataset.selection_payload == native_selection_payload assert dataset.artifact_id == artifact_id assert dataset.selection_id == native["selection_id"] + assert len(dataset.artifact_frame) == 3 + assert len(dataset.frame) == 2 def test_native_compatibility_identity_must_own_the_selected_semantic_rows(): @@ -541,6 +547,22 @@ def test_native_compatibility_identity_must_own_the_selected_semantic_rows(): ) +def test_unselected_semantic_content_changes_artifact_and_selection_identity(): + rows = _rows() + original = build_expected_dataset( + _declaration(), {"reference": _source("reference", rows)} + ) + changed_rows = [dict(row) for row in rows] + changed_rows[2]["stem"] = "Changed unselected question?" + changed = build_expected_dataset( + _declaration(), {"reference": _source("reference", changed_rows)} + ) + + assert dataset_content_digest(original.frame) == dataset_content_digest(changed.frame) + assert original.artifact_id != changed.artifact_id + assert original.selection_id != changed.selection_id + + @pytest.mark.parametrize("source_index", ["not-an-integer", -1, 1]) def test_structured_choice_source_indices_fail_closed(source_index): rows = [ @@ -583,6 +605,24 @@ def test_ordered_options_reject_an_empty_middle_value(): ) +def test_ordered_option_mapping_requires_contiguous_semantic_keys(): + declaration = _declaration( + expected_question_ids=("q1",), + columns={ + "question_id": "qid", + "question_text": "stem", + "correct_option": "gold", + "choice_a": "option_1", + "choice_c": "option_3", + }, + ) + + with pytest.raises(DatasetReferenceError, match="contiguous"): + build_expected_dataset( + declaration, {"reference": _source("reference", _rows())} + ) + + @pytest.mark.parametrize("mutation", ["schema", "extra_field"]) def test_snapshot_record_schema_is_exact_and_fail_closed(tmp_path: Path, mutation: str): dataset = build_expected_dataset( From 3b3664a5614760a76ebfdd348da10c42bdae9248 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 01:40:18 +0300 Subject: [PATCH 15/47] fix: preserve normalized dataset integer semantics --- .../importing/dataset_reference.py | 17 +++++++++++++++ tests/importing/test_dataset_reference.py | 21 +++++++++++++++++++ 2 files changed, 38 insertions(+) diff --git a/src/choicebench/importing/dataset_reference.py b/src/choicebench/importing/dataset_reference.py index dd7ffff..21ee66e 100644 --- a/src/choicebench/importing/dataset_reference.py +++ b/src/choicebench/importing/dataset_reference.py @@ -174,6 +174,23 @@ def _normalize_source( ) record["question_id"] = question_id record["correct_option"] = str(record["correct_option"]).strip().upper() + for integer_field in ("correct_index", "n_choices"): + if integer_field not in record: + continue + raw_integer = str(record[integer_field]).strip() + try: + parsed_integer = int(raw_integer) + except ValueError as exc: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has malformed {integer_field} " + f"for question_id {question_id!r}." + ) from exc + if str(parsed_integer) != raw_integer or parsed_integer < 0: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has malformed {integer_field} " + f"for question_id {question_id!r}." + ) + record[integer_field] = parsed_integer if has_structured_choices: raw_choices = raw_row[mapping["choices_json"]] diff --git a/tests/importing/test_dataset_reference.py b/tests/importing/test_dataset_reference.py index 4b528a5..f4e62c1 100644 --- a/tests/importing/test_dataset_reference.py +++ b/tests/importing/test_dataset_reference.py @@ -623,6 +623,27 @@ def test_ordered_option_mapping_requires_contiguous_semantic_keys(): ) +def test_current_normalized_integer_fields_preserve_native_semantic_types(): + rows = _rows() + for row in rows: + row["correct_index"] = str(ord(row["gold"]) - ord("A")) + row["n_choices"] = "3" + declaration = _declaration( + columns={ + **_columns(), + "correct_index": "correct_index", + "n_choices": "n_choices", + } + ) + + dataset = build_expected_dataset( + declaration, {"reference": _source("reference", rows)} + ) + + assert dataset.artifact_frame["correct_index"].tolist() == [1, 0, 2] + assert dataset.artifact_frame["n_choices"].tolist() == [3, 3, 3] + + @pytest.mark.parametrize("mutation", ["schema", "extra_field"]) def test_snapshot_record_schema_is_exact_and_fail_closed(tmp_path: Path, mutation: str): dataset = build_expected_dataset( From b05ba090c36ac889309814204c71fe00995ec11e Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 01:47:08 +0300 Subject: [PATCH 16/47] feat: separate condition and realization identity --- src/choicebench/importing/identity.py | 627 +++++++++++++++++++ tests/importing/test_identity.py | 862 ++++++++++++++++++++++++++ 2 files changed, 1489 insertions(+) create mode 100644 src/choicebench/importing/identity.py create mode 100644 tests/importing/test_identity.py diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py new file mode 100644 index 0000000..8684a08 --- /dev/null +++ b/src/choicebench/importing/identity.py @@ -0,0 +1,627 @@ +"""Layered semantic, realization, lineage, and prediction-origin identities.""" + +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any, Callable, Literal, Mapping, Sequence + +from choicebench import __version__ +from choicebench.identity import canonicalize, integrity_digest, short_id +from choicebench.importing.dataset_reference import ExpectedDataset +from choicebench.importing.schema import ( + ImportConditionSpec, + ImportMethodSpec, + ImportModelSpec, + ImportPromptSpec, +) +from choicebench.provenance import implementation_identity + + +class ImportIdentityError(ValueError): + """Raised when an importer identity is cyclic, unsafe, or inconsistent.""" + + +@dataclass(frozen=True) +class ImportSemanticRecords: + dataset_artifact: Mapping[str, Any] + selection: Mapping[str, Any] + model: Mapping[str, Any] + method: Mapping[str, Any] + prompt: Mapping[str, Any] + condition: Mapping[str, Any] + + +_CONDITION_KEYS = { + "benchmark", + "preflight", + "model_id", + "method_id", + "prompt_id", + "prompt_snapshot_path", + "seed", +} +_DERIVATION_ORIGINS = { + "native_execution", + "external_import", + "repair_overlay", + "offline_transformation", +} +_PREDICTION_ORIGINS = { + "native_inference", + "external_historical_inference", + "external_repair_inference", +} +_FORBIDDEN_IDENTITY_KEYS = { + "audit_path", + "absolute_path", + "choicebench_home", + "experiment_id", + "final_csv_sha256", + "hostname", + "import_timestamp", + "output_path", + "output_root", + "realization_id", + "realization_digest", + "report_path", + "result_artifact_id", + "result_artifact_digest", + "result_sha256", + "source_path", + "temporary_path", + "timestamp", +} + + +def _unknown(reason: str | None, field: str) -> dict[str, Any]: + if not isinstance(reason, str) or not reason.strip(): + raise ImportIdentityError(f"Missing {field} requires an explicit unknown reason.") + return {"value": None, "reason": reason} + + +def _nullable_semantic(value: Any, reason: str | None, field: str) -> Any: + if value is not None: + return value + if isinstance(reason, str) and reason.strip().casefold() == "not applicable": + return None + return _unknown(reason, field) + + +def _identity_record(prefix: str, payload: Mapping[str, Any]) -> tuple[str, str, dict]: + identity = canonicalize(payload) + return short_id(prefix, identity), integrity_digest(identity), identity + + +def _verify_claim( + *, prefix: str, payload: Mapping[str, Any], claimed_id: Any, claimed_digest: Any +) -> tuple[str, str, dict]: + if not isinstance(payload, Mapping): + raise ImportIdentityError(f"Invalid validated native {prefix} identity payload.") + identity_id, digest, identity = _identity_record(prefix, payload) + if claimed_id != identity_id or claimed_digest != digest: + raise ImportIdentityError(f"Invalid validated native {prefix} identity claim.") + return identity_id, digest, identity + + +def _reject_forbidden(value: Any, where: str) -> None: + if isinstance(value, Mapping): + for key, item in value.items(): + if key in _FORBIDDEN_IDENTITY_KEYS: + raise ImportIdentityError(f"{where} contains forbidden field {key!r}.") + _reject_forbidden(item, where) + elif isinstance(value, (list, tuple)): + for item in value: + _reject_forbidden(item, where) + + +def _semantic_child_records( + *, + condition: ImportConditionSpec, + dataset: ExpectedDataset, + model: ImportModelSpec, + method: ImportMethodSpec, + prompt: ImportPromptSpec, +) -> tuple[dict, dict, dict, dict, dict]: + dataset_record = { + "artifact_id": dataset.artifact_id, + "artifact_digest": dataset.artifact_digest, + "identity": canonicalize(dataset.artifact_payload), + "identity_mode": dataset.identity_mode, + } + if dataset.artifact_digest != integrity_digest(dataset_record["identity"]): + raise ImportIdentityError("Expected dataset artifact digest is inconsistent.") + if dataset.artifact_id != short_id("ds", dataset_record["identity"]): + raise ImportIdentityError("Expected dataset artifact ID is inconsistent.") + selection_record = { + "selection_id": dataset.selection_id, + "selection_digest": dataset.selection_digest, + "identity": canonicalize(dataset.selection_payload), + } + if dataset.selection_digest != integrity_digest(selection_record["identity"]): + raise ImportIdentityError("Expected dataset selection digest is inconsistent.") + if dataset.selection_id != short_id("sel", selection_record["identity"]): + raise ImportIdentityError("Expected dataset selection ID is inconsistent.") + + if model.native_compatibility_identity is not None: + native = model.native_compatibility_identity + if not isinstance(native, Mapping) or set(native) != { + "payload", + "digest", + "model_id", + }: + raise ImportIdentityError("Invalid validated native model identity fields.") + model_id, model_digest, model_payload = _verify_claim( + prefix="model", + payload=native["payload"], + claimed_id=native["model_id"], + claimed_digest=native["digest"], + ) + if model.backend is not None and model_payload.get("backend") != model.backend: + raise ImportIdentityError( + "Validated native model backend conflicts with its declaration." + ) + if model.provider is not None and model_payload.get("provider") != model.provider: + raise ImportIdentityError( + "Validated native model provider conflicts with its declaration." + ) + if ( + "model_name_or_path" in model_payload + and model_payload["model_name_or_path"] != model.display_name + ): + raise ImportIdentityError( + "Validated native model name conflicts with its declaration." + ) + if condition.generation_parameters: + native_generation = model_payload.get("generation_kwargs") + if not isinstance(native_generation, Mapping) or any( + native_generation.get(name) != value + for name, value in condition.generation_parameters.items() + ): + raise ImportIdentityError( + "Condition generation parameters are not bound by the validated " + "native model identity." + ) + model_mode = "native_compatibility" + else: + effective_parameters = dict(model.effective_parameters) + for name, value in condition.generation_parameters.items(): + if name in effective_parameters and effective_parameters[name] != value: + raise ImportIdentityError( + f"Generation parameter {name!r} conflicts with model effective parameters." + ) + effective_parameters[name] = value + model_payload = { + "schema_version": "choicebench.semantic-model.v1", + "backend": ( + model.backend + if model.backend is not None + else _unknown(model.unknown_reasons.get("backend"), "model backend") + ), + "provider": ( + model.provider + if model.provider is not None + else _unknown(model.unknown_reasons.get("provider"), "model provider") + ), + "model": model.display_name, + "revision": ( + model.revision + if model.revision is not None + else _unknown(model.unknown_reasons.get("revision"), "model revision") + ), + "effective_parameters": effective_parameters, + } + model_id, model_digest, model_payload = _identity_record("model", model_payload) + model_mode = "imported_semantic_fallback" + model_record = { + "model_id": model_id, + "model_digest": model_digest, + "identity": model_payload, + "identity_mode": model_mode, + } + + if method.native_compatibility_identity is not None: + native = method.native_compatibility_identity + if not isinstance(native, Mapping) or set(native) != { + "payload", + "digest", + "method_id", + }: + raise ImportIdentityError("Invalid validated native method identity fields.") + method_id, method_digest, method_payload = _verify_claim( + prefix="method", + payload=native["payload"], + claimed_id=native["method_id"], + claimed_digest=native["digest"], + ) + declared_preflight = _nullable_semantic( + condition.preflight_identity, + condition.unknown_reasons.get("preflight_identity"), + "method preflight identity", + ) + if ( + method_payload.get("name") != method.name + or canonicalize(method_payload.get("effective_params")) + != canonicalize(method.effective_parameters) + or canonicalize(method_payload.get("preflight")) + != canonicalize(declared_preflight) + or ( + method.implementation is not None + and canonicalize(method_payload.get("implementation")) + != canonicalize(method.implementation) + ) + ): + raise ImportIdentityError( + "Validated native method identity conflicts with its semantic declaration." + ) + method_mode = "native_compatibility" + else: + method_payload = { + "schema_version": "choicebench.semantic-method.v1", + "name": method.name, + "effective_params": dict(method.effective_parameters), + "preflight": _nullable_semantic( + condition.preflight_identity, + condition.unknown_reasons.get("preflight_identity"), + "method preflight identity", + ), + "implementation": ( + method.implementation + if method.implementation is not None + else _unknown( + method.unknown_reasons.get("implementation"), + "method implementation", + ) + ), + } + method_id, method_digest, method_payload = _identity_record( + "method", method_payload + ) + method_mode = "imported_semantic_fallback" + method_record = { + "method_id": method_id, + "method_digest": method_digest, + "identity": method_payload, + "identity_mode": method_mode, + } + + if prompt.native_compatibility_identity is not None: + native = prompt.native_compatibility_identity + if not isinstance(native, Mapping) or set(native) != { + "prompt_id", + "version", + "files", + }: + raise ImportIdentityError("Invalid validated native prompt identity fields.") + prompt_payload = canonicalize( + {"version": native["version"], "files": native["files"]} + ) + prompt_id = f"prompt_{integrity_digest(prompt_payload)[:16]}" + if native["prompt_id"] != prompt_id: + raise ImportIdentityError("Invalid validated native prompt identity claim.") + declared_contents = prompt.template_contents + native_files = prompt_payload.get("files") + native_contents = ( + { + name: item.get("content") + for name, item in native_files.items() + if isinstance(item, Mapping) + } + if isinstance(native_files, Mapping) + else None + ) + if ( + declared_contents is not None + and canonicalize(declared_contents) != canonicalize(native_contents) + ): + raise ImportIdentityError( + "Validated native prompt contents conflict with their declaration." + ) + if ( + prompt.template_identity is not None + and prompt.template_identity != prompt_payload.get("version") + ): + raise ImportIdentityError( + "Validated native prompt identity conflicts with its declaration." + ) + if ( + prompt.template_digest is not None + and prompt.template_digest != integrity_digest(native_files) + ): + raise ImportIdentityError( + "Validated native prompt digest conflicts with its declaration." + ) + prompt_digest = integrity_digest(prompt_payload) + prompt_mode = "native_compatibility" + else: + unknown = lambda field: _unknown(prompt.unknown_reason, f"prompt {field}") + prompt_payload = { + "schema_version": "choicebench.semantic-prompt.v1", + "template_identity": ( + prompt.template_identity + if prompt.template_identity is not None + else unknown("template identity") + ), + "template_digest": ( + prompt.template_digest + if prompt.template_digest is not None + else unknown("template digest") + ), + "template_contents": ( + prompt.template_contents + if prompt.template_contents is not None + else unknown("template contents") + ), + } + prompt_id, prompt_digest, prompt_payload = _identity_record( + "prompt", prompt_payload + ) + prompt_mode = "imported_semantic_fallback" + prompt_record = { + "prompt_id": prompt_id, + "prompt_digest": prompt_digest, + "identity": prompt_payload, + "identity_mode": prompt_mode, + } + return dataset_record, selection_record, model_record, method_record, prompt_record + + +def make_semantic_condition( + *, identity: Mapping[str, Any], fields: Mapping[str, Any] +) -> dict[str, Any]: + """Create one scientific condition without realization provenance.""" + allowed = set(_CONDITION_KEYS) + if "protocol_settings" in identity: + allowed.add("protocol_settings") + if set(identity) != allowed: + raise ImportIdentityError("Semantic condition identity fields are invalid.") + prompt_id = identity.get("prompt_id") + expected_prompt_path = f"artifacts/prompts/{prompt_id}" + if identity.get("prompt_snapshot_path") != expected_prompt_path: + raise ImportIdentityError( + "Semantic condition prompt snapshot path must be derived from prompt_id." + ) + _reject_forbidden(identity, "Semantic condition identity") + payload = canonicalize(identity) + condition_id = short_id("cond", payload) + condition_digest = integrity_digest(payload) + extra = canonicalize(fields) + reserved = {"condition_id", "condition_digest", "identity"} + if not isinstance(extra, Mapping) or reserved & set(extra): + raise ImportIdentityError("Semantic condition fields overwrite identity ownership.") + return { + "condition_id": condition_id, + "condition_digest": condition_digest, + "identity": payload, + **extra, + } + + +def build_import_semantic_identity( + *, + condition: ImportConditionSpec, + dataset: ExpectedDataset, + model: ImportModelSpec, + method: ImportMethodSpec, + prompt: ImportPromptSpec, +) -> ImportSemanticRecords: + """Build semantic children and a current-shape scientific condition.""" + if tuple(condition.expected_question_ids) != dataset.selected_question_ids: + raise ImportIdentityError("Condition question set does not match dataset selection.") + dataset_record, selection, model_record, method_record, prompt_record = ( + _semantic_child_records( + condition=condition, + dataset=dataset, + model=model, + method=method, + prompt=prompt, + ) + ) + seed = ( + condition.seed + if condition.seed is not None + else _unknown(condition.unknown_reasons.get("seed"), "condition seed") + ) + preflight = _nullable_semantic( + condition.calibration_identity, + condition.unknown_reasons.get("calibration_identity"), + "condition calibration identity", + ) + condition_payload: dict[str, Any] = { + "benchmark": { + "name": dataset.benchmark_name, + "split": dataset.split, + "selection_id": selection["selection_id"], + "artifact_id": dataset_record["artifact_id"], + }, + "preflight": preflight, + "model_id": model_record["model_id"], + "method_id": method_record["method_id"], + "prompt_id": prompt_record["prompt_id"], + "prompt_snapshot_path": f"artifacts/prompts/{prompt_record['prompt_id']}", + "seed": seed, + } + if condition.protocol_settings: + condition_payload["protocol_settings"] = condition.protocol_settings + condition_record = make_semantic_condition( + identity=condition_payload, + fields={ + "condition_key": condition.condition_key, + "expected_question_ids": list(condition.expected_question_ids), + }, + ) + return ImportSemanticRecords( + dataset_artifact=dataset_record, + selection=selection, + model=model_record, + method=method_record, + prompt=prompt_record, + condition=condition_record, + ) + + +def _validate_digest(value: Any, field: str, *, optional: bool = False) -> str | None: + if optional and value is None: + return None + if ( + not isinstance(value, str) + or len(value) != 64 + or any(character not in "0123456789abcdef" for character in value) + ): + raise ImportIdentityError(f"{field} must be a lowercase SHA-256 digest.") + return value + + +def make_lineage_component( + *, + operation_type: str, + question_id: str, + parent_digests: Sequence[str], + source_digests: Sequence[str], + authorization_digest: str | None, + implementation: Mapping[str, Any], + parameters: Mapping[str, Any], + input_digest: str, + preownership_output_digest: str, + prediction_origin: str, +) -> dict[str, Any]: + """Build a parent-derived lineage node without any child/self edge.""" + if not isinstance(operation_type, str) or not operation_type: + raise ImportIdentityError("Lineage operation_type must be non-empty.") + if not isinstance(question_id, str) or not question_id: + raise ImportIdentityError("Lineage question_id must be non-empty.") + if prediction_origin not in _PREDICTION_ORIGINS: + raise ImportIdentityError(f"Invalid prediction origin {prediction_origin!r}.") + for index, digest in enumerate(parent_digests): + _validate_digest(digest, f"parent_digests[{index}]") + for index, digest in enumerate(source_digests): + _validate_digest(digest, f"source_digests[{index}]") + _validate_digest(authorization_digest, "authorization_digest", optional=True) + _validate_digest(input_digest, "input_digest") + _validate_digest(preownership_output_digest, "preownership_output_digest") + _reject_forbidden(implementation, "Lineage implementation") + _reject_forbidden(parameters, "Lineage parameters") + payload = canonicalize( + { + "schema_version": "choicebench.lineage-component.v1", + "operation_type": operation_type, + "question_id": question_id, + "parent_digests": sorted(parent_digests), + "source_digests": sorted(source_digests), + "authorization_digest": authorization_digest, + "implementation": implementation, + "parameters": parameters, + "input_digest": input_digest, + "preownership_output_digest": preownership_output_digest, + "prediction_origin": prediction_origin, + } + ) + return { + "lineage_id": short_id("lin", payload), + "lineage_digest": integrity_digest(payload), + "identity": payload, + } + + +def make_result_origin( + *, + derivation_origin: Literal[ + "native_execution", "external_import", "repair_overlay", "offline_transformation" + ], + row_assignments: Sequence[tuple[str, str, str]], +) -> dict[str, Any]: + """Record constituent and ordered per-row prediction origins.""" + if derivation_origin not in _DERIVATION_ORIGINS: + raise ImportIdentityError(f"Invalid derivation origin {derivation_origin!r}.") + if not row_assignments: + raise ImportIdentityError("Result origin requires row assignments.") + rows: list[dict[str, str]] = [] + seen: set[str] = set() + counts: dict[str, int] = {} + for question_id, prediction_origin, lineage_id in row_assignments: + if not isinstance(question_id, str) or not question_id or question_id in seen: + raise ImportIdentityError("Result origin question IDs must be non-empty and unique.") + if prediction_origin not in _PREDICTION_ORIGINS: + raise ImportIdentityError(f"Invalid prediction origin {prediction_origin!r}.") + if not isinstance(lineage_id, str) or not lineage_id: + raise ImportIdentityError("Result origin lineage IDs must be non-empty.") + seen.add(question_id) + counts[prediction_origin] = counts.get(prediction_origin, 0) + 1 + rows.append( + { + "question_id": question_id, + "prediction_origin": prediction_origin, + "prediction_lineage_id": lineage_id, + } + ) + rows = canonicalize(rows) + lineage_projection = [ + { + "prediction_origin": row["prediction_origin"], + "prediction_lineage_id": row["prediction_lineage_id"], + } + for row in rows + ] + return { + "derivation_origin": derivation_origin, + "prediction_origins": sorted(counts), + "prediction_origin_counts": {key: counts[key] for key in sorted(counts)}, + "row_assignments": rows, + "ordered_row_origin_digest": integrity_digest(rows), + "lineage_component_origin_digest": integrity_digest(lineage_projection), + } + + +def importer_implementation_identity( + *, adapter: Callable[..., Any], validator: Callable[..., Any] +) -> dict[str, Any]: + """Bind installed ChoiceBench plus registered adapter and validator code.""" + return canonicalize( + { + "schema_version": "choicebench.importer-implementation.v1", + "package": {"name": "choicebench", "version": __version__}, + "adapter": implementation_identity(adapter), + "validator": implementation_identity(validator), + } + ) + + +def make_realization( + *, + condition_id: str, + condition_digest: str, + identity: Mapping[str, Any], + fields: Mapping[str, Any], +) -> dict[str, Any]: + """Create an immutable realization identity distinct from its condition.""" + if not isinstance(condition_id, str) or not condition_id.startswith("cond_"): + raise ImportIdentityError("Realization condition_id is invalid.") + _validate_digest(condition_digest, "condition_digest") + if not isinstance(identity, Mapping) or not identity: + raise ImportIdentityError("Realization identity must be a non-empty mapping.") + _reject_forbidden(identity, "Realization identity") + payload = canonicalize( + { + "schema_version": "choicebench.realization.v1", + "condition_id": condition_id, + "condition_digest": condition_digest, + "realization": identity, + } + ) + extra = canonicalize(fields) + reserved = { + "realization_id", + "realization_digest", + "condition_id", + "condition_digest", + "identity", + } + if not isinstance(extra, Mapping) or reserved & set(extra): + raise ImportIdentityError("Realization fields overwrite identity ownership.") + return { + "realization_id": short_id("real", payload), + "realization_digest": integrity_digest(payload), + "condition_id": condition_id, + "condition_digest": condition_digest, + "identity": payload, + **extra, + } diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py new file mode 100644 index 0000000..5ef4ecd --- /dev/null +++ b/tests/importing/test_identity.py @@ -0,0 +1,862 @@ +from __future__ import annotations + +from dataclasses import asdict, replace +from hashlib import sha256 +import importlib +from pathlib import Path + +import pandas as pd +import pytest + +from choicebench.config.schema import ( + BenchmarkConfig, + ExperimentConfig, + MethodConfig, + ModelConfig, + RunConfig, +) +from choicebench.datasets import dataset_content_digest +from choicebench.identity import integrity_digest, short_id +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.dataset_reference import build_expected_dataset +from choicebench.importing.identity import ( + ImportIdentityError, + build_import_semantic_identity, + importer_implementation_identity, + make_lineage_component, + make_realization, + make_result_origin, + make_semantic_condition, +) +from choicebench.importing.schema import ( + DatasetReferenceSpec, + CsvDialectSpec, + ImportConditionSpec, + ImportMethodSpec, + ImportModelSpec, + ImportPromptSpec, + ResultOriginSpec, +) + + +def _adapter(value: str) -> str: + return value + + +def _validator(value: str) -> bool: + return bool(value) + + +def _dataset(): + frame = pd.DataFrame( + [ + {"qid": "q1", "stem": "One?", "gold": "A", "a": "x", "b": "y"}, + {"qid": "q2", "stem": "Two?", "gold": "B", "a": "m", "b": "n"}, + {"qid": "q3", "stem": "Three?", "gold": "A", "a": "i", "b": "j"}, + ] + ) + data = frame.to_csv(index=False, lineterminator="\n").encode() + source = OpenedSource( + source_id="benchmark", + audit_path=Path("/machine-a/input.csv"), + logical_path="inputs/benchmark.csv", + data=data, + sha256=sha256(data).hexdigest(), + ) + declaration = DatasetReferenceSpec( + dataset_id="dataset", + benchmark_name="benchmark", + split="test", + reference_kind="independent_input_snapshot", + trust_label="declared-independent-input", + source_ids=("benchmark",), + selection_source_id="benchmark", + expected_question_ids=("q2", "q1"), + selection_seed=None, + selection_n_samples=2, + subject_filter=(), + selection_unknown_reasons={"selection_seed": "not recorded"}, + columns={ + "question_id": "qid", + "question_text": "stem", + "correct_option": "gold", + "choice_a": "a", + "choice_b": "b", + }, + revision=None, + fingerprint=None, + derivation={"source_role": "benchmark_input"}, + limitations=("publisher revision unknown",), + native_compatibility_identity=None, + ) + return build_expected_dataset(declaration, {"benchmark": source}) + + +def _model() -> ImportModelSpec: + return ImportModelSpec( + model_key="model", + display_name="historical-model", + backend=None, + provider=None, + revision=None, + effective_parameters={"temperature": 0}, + unknown_reasons={ + "backend": "not recorded", + "provider": "not recorded", + "revision": "not recorded", + }, + native_compatibility_identity=None, + ) + + +def _method() -> ImportMethodSpec: + return ImportMethodSpec( + method_key="method", + name="semantic_matching_v1", + effective_parameters={"matching": "semantic"}, + implementation=None, + unknown_reasons={"implementation": "historical code unavailable"}, + native_compatibility_identity=None, + ) + + +def _prompt() -> ImportPromptSpec: + return ImportPromptSpec( + prompt_key="prompt", + template_identity=None, + template_digest=None, + template_contents=None, + unknown_reason="raw prompt was not preserved", + native_compatibility_identity=None, + ) + + +def _condition() -> ImportConditionSpec: + return ImportConditionSpec( + condition_key="condition", + source_ids=("results",), + dataset_id="dataset", + model_key="model", + method_key="method", + prompt_key="prompt", + seed=None, + calibration_identity=None, + preflight_identity=None, + protocol_settings={}, + generation_parameters={"max_tokens": 64}, + unknown_reasons={ + "seed": "not recorded", + "calibration_identity": "not applicable", + "preflight_identity": "not recorded", + }, + expected_question_ids=("q2", "q1"), + evidence_status="complete", + scope_disposition="included", + executable=None, + qualifications=(), + limitations=(), + damaged_question_ids=(), + recoverable_question_ids=(), + result_origin=ResultOriginSpec( + derivation_origin="external_import", + default_prediction_origin="external_historical_inference", + per_question_prediction_origins={}, + ), + ) + + +def _records(**condition_changes): + condition = replace(_condition(), **condition_changes) + return build_import_semantic_identity( + condition=condition, + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +def test_fallback_child_payloads_are_exact_and_unknowns_are_explicit(): + records = _records() + + assert records.dataset_artifact["identity"] == { + "benchmark": "benchmark", + "content_digest": records.dataset_artifact["identity"]["content_digest"], + "schema_version": "choicebench.semantic-dataset.v1", + "split": "test", + } + assert records.model["identity"] == { + "backend": {"reason": "not recorded", "value": None}, + "effective_parameters": {"max_tokens": 64, "temperature": 0}, + "model": "historical-model", + "provider": {"reason": "not recorded", "value": None}, + "revision": {"reason": "not recorded", "value": None}, + "schema_version": "choicebench.semantic-model.v1", + } + assert records.method["identity"] == { + "effective_params": {"matching": "semantic"}, + "implementation": { + "reason": "historical code unavailable", + "value": None, + }, + "name": "semantic_matching_v1", + "preflight": {"reason": "not recorded", "value": None}, + "schema_version": "choicebench.semantic-method.v1", + } + unknown_prompt = {"reason": "raw prompt was not preserved", "value": None} + assert records.prompt["identity"] == { + "schema_version": "choicebench.semantic-prompt.v1", + "template_contents": unknown_prompt, + "template_digest": unknown_prompt, + "template_identity": unknown_prompt, + } + + +def test_child_records_have_full_digests_and_existing_short_prefixes(): + records = _records() + for record, prefix, id_key, digest_key in ( + (records.dataset_artifact, "ds", "artifact_id", "artifact_digest"), + (records.selection, "sel", "selection_id", "selection_digest"), + (records.model, "model", "model_id", "model_digest"), + (records.method, "method", "method_id", "method_digest"), + (records.prompt, "prompt", "prompt_id", "prompt_digest"), + ): + assert record[digest_key] == integrity_digest(record["identity"]) + assert record[id_key] == short_id(prefix, record["identity"]) + + +def test_condition_payload_matches_current_choicebench_keys_and_frozen_prompt_path(): + records = _records() + identity = records.condition["identity"] + + assert set(identity) == { + "benchmark", + "preflight", + "model_id", + "method_id", + "prompt_id", + "prompt_snapshot_path", + "seed", + } + assert identity["benchmark"] == { + "name": "benchmark", + "split": "test", + "artifact_id": records.dataset_artifact["artifact_id"], + "selection_id": records.selection["selection_id"], + } + assert identity["prompt_snapshot_path"] == ( + f"artifacts/prompts/{records.prompt['prompt_id']}" + ) + assert records.condition["condition_digest"] == integrity_digest(identity) + assert records.condition["condition_id"] == short_id("cond", identity) + + +@pytest.mark.parametrize( + ("field", "value"), + [ + ("seed", 7), + ("protocol_settings", {"threshold": 3}), + ("generation_parameters", {"max_tokens": 65}), + ], +) +def test_scientific_condition_changes_change_condition_identity(field, value): + assert _records(**{field: value}).condition["condition_id"] != _records().condition[ + "condition_id" + ] + + +@pytest.mark.parametrize( + "field", + ["selection_id", "artifact_id", "model_id", "method_id", "prompt_id"], +) +def test_semantic_child_change_changes_condition_identity(field): + original = _records().condition + identity = dict(original["identity"]) + if field in {"selection_id", "artifact_id"}: + identity["benchmark"] = dict(identity["benchmark"]) + identity["benchmark"][field] = f"{field}_changed" + else: + identity[field] = f"{field}_changed" + if field == "prompt_id": + identity["prompt_snapshot_path"] = ( + f"artifacts/prompts/{identity['prompt_id']}" + ) + changed = make_semantic_condition(identity=identity, fields={}) + assert changed["condition_id"] != original["condition_id"] + + +def test_make_semantic_condition_rejects_arbitrary_prompt_path_injection(): + identity = dict(_records().condition["identity"]) + identity["prompt_snapshot_path"] = "/tmp/attacker/prompt" + + with pytest.raises(ImportIdentityError, match="prompt.*path"): + make_semantic_condition(identity=identity, fields={"condition_key": "x"}) + + +def test_native_compatibility_children_are_preserved_without_fabrication(): + fallback = _records() + model_payload = {"backend": "dummy", "model_name_or_path": "native"} + method_payload = { + "name": "direct", + "effective_params": {}, + "preflight": None, + "implementation": {"qualified_name": "choicebench.methods:Direct"}, + } + native_model = { + "payload": model_payload, + "digest": integrity_digest(model_payload), + "model_id": short_id("model", model_payload), + } + native_method = { + "payload": method_payload, + "digest": integrity_digest(method_payload), + "method_id": short_id("method", method_payload), + } + prompt_payload = { + "version": "v1", + "files": { + name: {"sha256": integrity_digest(name), "content": name} + for name in ("direct_mcq", "free_text", "option_matching") + }, + } + native_prompt = { + "prompt_id": short_id("prompt", prompt_payload), + **prompt_payload, + } + records = build_import_semantic_identity( + condition=replace( + _condition(), + generation_parameters={}, + unknown_reasons={ + **_condition().unknown_reasons, + "preflight_identity": "not applicable", + }, + ), + dataset=_dataset(), + model=replace( + _model(), + display_name="native", + backend="dummy", + effective_parameters={}, + unknown_reasons={"provider": "not applicable", "revision": "not applicable"}, + native_compatibility_identity=native_model, + ), + method=replace( + _method(), + name="direct", + effective_parameters={}, + implementation=method_payload["implementation"], + unknown_reasons={}, + native_compatibility_identity=native_method, + ), + prompt=replace( + _prompt(), + template_identity="v1", + template_digest=integrity_digest(prompt_payload["files"]), + template_contents={ + name: item["content"] for name, item in prompt_payload["files"].items() + }, + unknown_reason=None, + native_compatibility_identity=native_prompt, + ), + ) + + assert records.model["identity"] == model_payload + assert records.model["model_id"] == native_model["model_id"] + assert records.method["identity"] == method_payload + assert records.prompt["identity"] == prompt_payload + assert records.condition["condition_id"] != fallback.condition["condition_id"] + + +def test_native_model_identity_refuses_unbound_generation_parameters(): + payload = {"backend": "dummy", "model_name_or_path": "historical-model"} + native = { + "payload": payload, + "digest": integrity_digest(payload), + "model_id": short_id("model", payload), + } + with pytest.raises(ImportIdentityError, match="generation"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=replace(_model(), native_compatibility_identity=native), + method=_method(), + prompt=_prompt(), + ) + + +def test_native_child_claims_must_match_their_semantic_declarations(): + model_payload = {"backend": "dummy", "model_name_or_path": "historical-model"} + native_model = { + "payload": model_payload, + "digest": integrity_digest(model_payload), + "model_id": short_id("model", model_payload), + } + with pytest.raises(ImportIdentityError, match="backend"): + build_import_semantic_identity( + condition=replace(_condition(), generation_parameters={}), + dataset=_dataset(), + model=replace( + _model(), backend="api", native_compatibility_identity=native_model + ), + method=_method(), + prompt=_prompt(), + ) + + method_payload = { + "name": "direct_mcq", + "effective_params": {}, + "preflight": None, + "implementation": {"qualified_name": "choicebench.methods:Direct"}, + } + native_method = { + "payload": method_payload, + "digest": integrity_digest(method_payload), + "method_id": short_id("method", method_payload), + } + with pytest.raises(ImportIdentityError, match="method.*declaration"): + build_import_semantic_identity( + condition=replace( + _condition(), + preflight_identity=None, + unknown_reasons={ + **_condition().unknown_reasons, + "preflight_identity": "not applicable", + }, + ), + dataset=_dataset(), + model=_model(), + method=replace(_method(), native_compatibility_identity=native_method), + prompt=_prompt(), + ) + + +def test_native_prompt_contents_must_match_the_semantic_declaration(): + payload = { + "version": "v1", + "files": { + name: {"sha256": integrity_digest(name), "content": name} + for name in ("direct_mcq", "free_text", "option_matching") + }, + } + native = {"prompt_id": short_id("prompt", payload), **payload} + with pytest.raises(ImportIdentityError, match="prompt.*contents"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=replace( + _prompt(), + template_identity="v1", + template_digest=integrity_digest(payload["files"]), + template_contents={"direct_mcq": "different"}, + unknown_reason=None, + native_compatibility_identity=native, + ), + ) + + +def test_fully_validated_native_children_reproduce_build_execution_plan_condition(): + import choicebench.cli.run_experiment as run_experiment + + run_experiment = importlib.reload(run_experiment) + config = ExperimentConfig( + "native-compat", + [ModelConfig("dummy", "same")], + [BenchmarkConfig("toy", n_samples=2)], + [MethodConfig("direct_mcq")], + ["accuracy"], + RunConfig(seed=42), + ) + plan = run_experiment.build_execution_plan(config) + native_condition = next(iter(plan.conditions.values())) + native_selection = plan.selections[0] + artifact = native_selection.artifact + artifact_payload = { + "spec": artifact.metadata["spec"], + "content_digest": artifact.content_digest, + "source": artifact.metadata["source"], + } + selection_payload = { + "artifact_id": artifact.artifact_id, + "content_digest": dataset_content_digest(native_selection.questions), + "sample_identities": list(native_selection.sample_identities), + "seed": 42, + "n_samples": 2, + "subject_filter": [], + } + native_dataset = { + "artifact_payload": artifact_payload, + "artifact_digest": integrity_digest(artifact_payload), + "artifact_id": artifact.artifact_id, + "selection_payload": selection_payload, + "selection_digest": integrity_digest(selection_payload), + "selection_id": native_selection.selection_id, + } + source_bytes = artifact.dataframe.to_csv(index=False, lineterminator="\n").encode() + source = OpenedSource( + source_id="native-input", + audit_path=artifact.path, + logical_path="prepared/toy.csv", + data=source_bytes, + sha256=sha256(source_bytes).hexdigest(), + ) + dataset = build_expected_dataset( + DatasetReferenceSpec( + dataset_id="toy", + benchmark_name="toy", + split="test", + reference_kind="independent_input_snapshot", + trust_label="choicebench-verified", + source_ids=("native-input",), + selection_source_id="native-input", + expected_question_ids=tuple( + native_selection.questions["question_id"].astype(str) + ), + selection_seed=42, + selection_n_samples=2, + subject_filter=(), + selection_unknown_reasons={}, + columns={column: column for column in artifact.dataframe.columns}, + revision=None, + fingerprint=None, + derivation={"kind": "native compatibility fixture"}, + limitations=(), + native_compatibility_identity=native_dataset, + ), + {"native-input": source}, + ) + native_model_record = plan.manifest["payload"]["models"][0] + native_model_payload = { + "backend": "dummy", + "model_name_or_path": "same", + } + native_method_record = plan.manifest["payload"]["methods"][0] + native_method_payload = { + "name": native_method_record["config"]["name"], + "effective_params": native_method_record["config"]["params"], + "preflight": native_method_record["config"]["preflight"], + "implementation": native_method_record["implementation"], + } + native_prompt_record = plan.manifest["payload"]["prompts"] + condition = replace( + _condition(), + dataset_id="toy", + seed=42, + generation_parameters={}, + unknown_reasons={ + "calibration_identity": "not applicable", + "preflight_identity": "not applicable", + }, + expected_question_ids=tuple( + native_selection.questions["question_id"].astype(str) + ), + ) + records = build_import_semantic_identity( + condition=condition, + dataset=dataset, + model=replace( + _model(), + display_name="same", + backend="dummy", + provider=None, + revision=None, + unknown_reasons={"provider": "not applicable", "revision": "not applicable"}, + effective_parameters={}, + native_compatibility_identity={ + "payload": native_model_payload, + "digest": integrity_digest(native_model_payload), + "model_id": native_model_record["model_id"], + }, + ), + method=replace( + _method(), + name="direct_mcq", + effective_parameters=native_method_record["config"]["params"], + implementation=native_method_record["implementation"], + unknown_reasons={}, + native_compatibility_identity={ + "payload": native_method_payload, + "digest": integrity_digest(native_method_payload), + "method_id": native_method_record["method_id"], + }, + ), + prompt=replace( + _prompt(), + template_identity=native_prompt_record["version"], + template_digest=integrity_digest(native_prompt_record["files"]), + template_contents={ + name: item["content"] + for name, item in native_prompt_record["files"].items() + }, + unknown_reason=None, + native_compatibility_identity={ + "prompt_id": native_prompt_record["prompt_id"], + "version": native_prompt_record["version"], + "files": native_prompt_record["files"], + }, + ), + ) + + assert records.condition["identity"] == native_condition["identity"] + assert records.condition["condition_id"] == native_condition["condition_id"] + + +def _realization_identity() -> dict: + origin = make_result_origin( + derivation_origin="external_import", + row_assignments=( + ("q2", "external_historical_inference", "lin_a"), + ("q1", "external_historical_inference", "lin_b"), + ), + ) + return { + "import_spec_digest": "1" * 64, + "source_sha256": "2" * 64, + "mapping": {"question_id": "qid"}, + "dialect": {"delimiter": ","}, + "evidence_status": "complete", + "scope": "included", + "authorization_digest": None, + "overlay_sha256": None, + "prediction_origins": origin, + } + + +@pytest.mark.parametrize( + ("field", "value"), + [ + ("source_sha256", "3" * 64), + ("mapping", {"question_id": "id"}), + ("dialect", {"delimiter": ";"}), + ("evidence_status", "qualified"), + ("scope", "superseded"), + ("authorization_digest", "4" * 64), + ("overlay_sha256", "5" * 64), + ("prediction_origins", {"ordered_row_origin_digest": "6" * 64}), + ], +) +def test_realization_changes_leave_condition_fixed(field, value): + condition = _records().condition + original = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=_realization_identity(), + fields={"audit": {"source_path": "/machine-a/results.csv"}}, + ) + changed_identity = _realization_identity() + changed_identity[field] = value + changed = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=changed_identity, + fields={"audit": {"source_path": "/machine-b/results.csv"}}, + ) + + assert changed["condition_id"] == original["condition_id"] + assert changed["realization_id"] != original["realization_id"] + + +def test_audit_only_realization_fields_do_not_change_identity(): + condition = _records().condition + first = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=_realization_identity(), + fields={"audit": {"source_path": "/machine-a/a.csv", "timestamp": "now"}}, + ) + moved = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=_realization_identity(), + fields={"audit": {"source_path": "/machine-b/a.csv", "timestamp": "later"}}, + ) + assert first["realization_id"] == moved["realization_id"] + assert first["realization_digest"] == moved["realization_digest"] + + +@pytest.mark.parametrize( + ("field", "value"), + [ + ("encoding", "utf-8-sig"), + ("bom_policy", "strip_utf8_bom"), + ("decoding_errors", "replace"), + ("delimiter", ";"), + ("quote_character", "'"), + ("escape_character", "\\"), + ("double_quote", False), + ("line_terminators", ["lf"]), + ("mixed_line_terminators", "forbid"), + ("final_record_without_terminator", "forbid"), + ("blank_record_policy", "allow"), + ("skip_initial_space", True), + ("header", "explicit"), + ("strict_syntax", False), + ], +) +def test_every_csv_decoding_and_dialect_field_changes_realization_not_condition( + field, value +): + condition = _records().condition + base_dialect = asdict(CsvDialectSpec()) + first_identity = _realization_identity() + first_identity["dialect"] = base_dialect + changed_identity = _realization_identity() + changed_identity["dialect"] = {**base_dialect, field: value} + first = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=first_identity, + fields={}, + ) + changed = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=changed_identity, + fields={}, + ) + assert first["condition_id"] == changed["condition_id"] + assert first["realization_id"] != changed["realization_id"] + + +def test_base_repair_and_transformation_are_distinct_realizations_of_one_condition(): + condition = _records().condition + identities = [] + for derivation in ("external_import", "repair_overlay", "offline_transformation"): + identity = _realization_identity() + identity["derivation_origin"] = derivation + identities.append( + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + ) + assert {item["condition_id"] for item in identities} == { + condition["condition_id"] + } + assert len({item["realization_id"] for item in identities}) == 3 + + +def test_realization_refuses_post_publication_or_machine_local_identity_fields(): + condition = _records().condition + for forbidden in ( + "result_artifact_id", + "experiment_id", + "output_path", + "source_path", + "timestamp", + ): + identity = _realization_identity() + identity[forbidden] = "forbidden" + with pytest.raises(ImportIdentityError, match=forbidden): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_lineage_component_is_parent_derived_and_has_no_child_edge(): + component = make_lineage_component( + operation_type="offline_transformation", + question_id="q1", + parent_digests=("2" * 64, "1" * 64), + source_digests=("4" * 64, "3" * 64), + authorization_digest="5" * 64, + implementation={"qualified_name": "choicebench.transforms:rematch"}, + parameters={"matcher": "v1"}, + input_digest="6" * 64, + preownership_output_digest="7" * 64, + prediction_origin="external_historical_inference", + ) + + assert component["lineage_digest"] == integrity_digest(component["identity"]) + assert component["lineage_id"] == short_id("lin", component["identity"]) + assert component["identity"]["parent_digests"] == ["1" * 64, "2" * 64] + assert not ({"realization_id", "result_artifact_id", "experiment_id"} & set(component["identity"])) + + +def test_lineage_refuses_child_paths_timestamps_and_final_csv_hashes(): + for forbidden in ("realization_id", "output_path", "timestamp", "final_csv_sha256"): + with pytest.raises(ImportIdentityError, match=forbidden): + make_lineage_component( + operation_type="offline_transformation", + question_id="q1", + parent_digests=(), + source_digests=("1" * 64,), + authorization_digest=None, + implementation={"qualified_name": "x:y"}, + parameters={forbidden: "bad"}, + input_digest="2" * 64, + preownership_output_digest="3" * 64, + prediction_origin="external_historical_inference", + ) + + +def test_result_origin_records_constituents_counts_and_ordered_per_row_mapping(): + origin = make_result_origin( + derivation_origin="repair_overlay", + row_assignments=( + ("q1", "external_historical_inference", "lin_base"), + ("q2", "native_inference", "lin_repair"), + ("q3", "native_inference", "lin_repair_2"), + ), + ) + + assert origin["prediction_origins"] == [ + "external_historical_inference", + "native_inference", + ] + assert origin["prediction_origin_counts"] == { + "external_historical_inference": 1, + "native_inference": 2, + } + assert [item["question_id"] for item in origin["row_assignments"]] == [ + "q1", + "q2", + "q3", + ] + assert origin["ordered_row_origin_digest"] == integrity_digest( + origin["row_assignments"] + ) + + +def test_result_origin_rejects_bare_mixed_duplicate_rows_and_invalid_origins(): + for assignments in ( + (("q1", "mixed", "lin_a"),), + ( + ("q1", "external_historical_inference", "lin_a"), + ("q1", "native_inference", "lin_b"), + ), + ): + with pytest.raises(ImportIdentityError): + make_result_origin( + derivation_origin="repair_overlay", row_assignments=assignments + ) + + +def test_offline_transformation_keeps_underlying_prediction_origin(): + origin = make_result_origin( + derivation_origin="offline_transformation", + row_assignments=( + ("q1", "external_historical_inference", "lin_rematch"), + ), + ) + assert origin["derivation_origin"] == "offline_transformation" + assert origin["prediction_origins"] == ["external_historical_inference"] + + +def test_importer_implementation_identity_binds_package_adapter_and_validator(): + identity = importer_implementation_identity(adapter=_adapter, validator=_validator) + assert identity["schema_version"] == "choicebench.importer-implementation.v1" + assert identity["package"]["name"] == "choicebench" + assert identity["package"]["version"] + assert identity["adapter"]["qualified_name"].endswith(":_adapter") + assert identity["adapter"]["source_digest"] + assert identity["validator"]["qualified_name"].endswith(":_validator") + changed = importer_implementation_identity(adapter=_validator, validator=_adapter) + assert integrity_digest(identity) != integrity_digest(changed) From a0f789e2d959269b37b601efda12311b4feacc20 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 02:06:01 +0300 Subject: [PATCH 17/47] fix: close importer identity schemas --- src/choicebench/importing/identity.py | 501 ++++++++++++++++++++++++-- tests/importing/test_identity.py | 333 +++++++++++++++-- 2 files changed, 784 insertions(+), 50 deletions(-) diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index 8684a08..c91061e 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -3,6 +3,8 @@ from __future__ import annotations from dataclasses import dataclass +from pathlib import PurePosixPath +import re from typing import Any, Callable, Literal, Mapping, Sequence from choicebench import __version__ @@ -10,6 +12,7 @@ from choicebench.importing.dataset_reference import ExpectedDataset from choicebench.importing.schema import ( ImportConditionSpec, + CsvDialectSpec, ImportMethodSpec, ImportModelSpec, ImportPromptSpec, @@ -51,6 +54,52 @@ class ImportSemanticRecords: "external_historical_inference", "external_repair_inference", } +_REALIZATION_KEYS = { + "import_spec_digest", + "sources", + "expected_dataset", + "importer_implementation", + "parsing_policy", + "validation", + "evidence", + "result_origin", + "parent_digests", + "authorization_digest", + "overlay", +} +_SOURCE_KEYS = { + "source_id", + "logical_path", + "classification", + "format", + "format_version", + "sha256", + "provenance", +} +_PARSING_POLICY_KEYS = { + "dialect", + "mapping", + "null_values", + "numeric_columns", + "option_mapping", + "extra_field_policy", +} +_DIALECT_KEYS = set(CsvDialectSpec.__dataclass_fields__) +_EVIDENCE_STATUSES = { + "complete", + "qualified", + "partial", + "malformed", + "recoverable", + "failed", +} +_SCOPE_DISPOSITIONS = { + "included", + "excluded_from_paper_matrix", + "held", + "superseded", +} +_LINEAGE_ID_RE = re.compile(r"^lin_[0-9a-f]{16}$") _FORBIDDEN_IDENTITY_KEYS = { "audit_path", "absolute_path", @@ -103,15 +152,55 @@ def _verify_claim( return identity_id, digest, identity -def _reject_forbidden(value: Any, where: str) -> None: +def _reject_forbidden( + value: Any, where: str, *, allowed_path_keys: frozenset[str] = frozenset() +) -> None: if isinstance(value, Mapping): for key, item in value.items(): - if key in _FORBIDDEN_IDENTITY_KEYS: + if not isinstance(key, str): + raise ImportIdentityError(f"{where} contains a non-string field name.") + normalized = key.casefold() + unsafe_alias = ( + key in _FORBIDDEN_IDENTITY_KEYS + or (normalized not in allowed_path_keys and normalized.endswith("_path")) + or normalized in {"path", "created_at", "updated_at", "import_state"} + or "timestamp" in normalized + or normalized.endswith( + ("realization_id", "result_artifact_id", "experiment_id") + ) + or ("csv" in normalized and normalized.endswith(("sha256", "digest"))) + ) + if unsafe_alias: raise ImportIdentityError(f"{where} contains forbidden field {key!r}.") - _reject_forbidden(item, where) + _reject_forbidden(item, where, allowed_path_keys=allowed_path_keys) elif isinstance(value, (list, tuple)): for item in value: - _reject_forbidden(item, where) + _reject_forbidden(item, where, allowed_path_keys=allowed_path_keys) + + +def _exact_mapping(value: Any, keys: set[str], where: str) -> Mapping[str, Any]: + if not isinstance(value, Mapping): + raise ImportIdentityError(f"{where} must be a mapping.") + missing = sorted(keys - set(value)) + unexpected = sorted(set(value) - keys) + if missing or unexpected: + raise ImportIdentityError( + f"{where} fields are invalid; missing={missing}, unexpected={unexpected}." + ) + return value + + +def _digest_sequence(value: Any, field: str) -> list[str]: + if not isinstance(value, (list, tuple)): + raise ImportIdentityError(f"{field} must be a digest sequence.") + result = [] + for index, digest in enumerate(value): + checked = _validate_digest(digest, f"{field}[{index}]") + assert checked is not None + result.append(checked) + if len(result) != len(set(result)): + raise ImportIdentityError(f"{field} contains duplicate digest edges.") + return sorted(result) def _semantic_child_records( @@ -156,6 +245,45 @@ def _semantic_child_records( claimed_id=native["model_id"], claimed_digest=native["digest"], ) + if model.backend is None: + raise ImportIdentityError( + "An unknown model backend cannot be upgraded by a native identity claim." + ) + native_model_keys = { + "dummy": {"backend", "model_name_or_path"}, + "api": { + "backend", + "provider", + "model_name_or_path", + "base_url", + "generation_kwargs", + }, + "huggingface": { + "backend", + "model", + "device", + "add_bos_token", + "generation_kwargs", + "loader", + }, + } + if model.backend not in native_model_keys or set(model_payload) != native_model_keys[ + model.backend + ]: + raise ImportIdentityError( + "Validated native model payload does not match the current backend schema." + ) + for nullable_field in ("provider", "revision"): + value = getattr(model, nullable_field) + reason = model.unknown_reasons.get(nullable_field) + if value is None and ( + not isinstance(reason, str) + or reason.strip().casefold() != "not applicable" + ): + raise ImportIdentityError( + f"Unknown model {nullable_field} cannot be upgraded by a native " + "identity claim." + ) if model.backend is not None and model_payload.get("backend") != model.backend: raise ImportIdentityError( "Validated native model backend conflicts with its declaration." @@ -233,6 +361,20 @@ def _semantic_child_records( claimed_id=native["method_id"], claimed_digest=native["digest"], ) + if set(method_payload) != { + "name", + "effective_params", + "preflight", + "implementation", + }: + raise ImportIdentityError( + "Validated native method payload does not match the current schema." + ) + if method.implementation is None: + raise ImportIdentityError( + "An unknown method implementation cannot be upgraded by a native " + "identity claim." + ) declared_preflight = _nullable_semantic( condition.preflight_identity, condition.unknown_reasons.get("preflight_identity"), @@ -295,6 +437,32 @@ def _semantic_child_records( prompt_payload = canonicalize( {"version": native["version"], "files": native["files"]} ) + if any( + value is None + for value in ( + prompt.template_identity, + prompt.template_digest, + prompt.template_contents, + ) + ): + raise ImportIdentityError( + "An unknown prompt cannot be upgraded by a native identity claim." + ) + native_files_raw = _exact_mapping( + prompt_payload.get("files"), + {"direct_mcq", "free_text", "option_matching"}, + "Validated native prompt files", + ) + for name, raw_item in native_files_raw.items(): + item = _exact_mapping( + raw_item, {"sha256", "content"}, f"Validated native prompt file {name}" + ) + if not isinstance(item["content"], str) or item["sha256"] != integrity_digest( + item["content"] + ): + raise ImportIdentityError( + f"Validated native prompt file {name!r} has an invalid content hash." + ) prompt_id = f"prompt_{integrity_digest(prompt_payload)[:16]}" if native["prompt_id"] != prompt_id: raise ImportIdentityError("Invalid validated native prompt identity claim.") @@ -380,14 +548,23 @@ def make_semantic_condition( raise ImportIdentityError( "Semantic condition prompt snapshot path must be derived from prompt_id." ) - _reject_forbidden(identity, "Semantic condition identity") + _reject_forbidden( + identity, + "Semantic condition identity", + allowed_path_keys=frozenset({"prompt_snapshot_path"}), + ) payload = canonicalize(identity) condition_id = short_id("cond", payload) condition_digest = integrity_digest(payload) + if not isinstance(fields, Mapping) or not set(fields) <= { + "condition_key", + "expected_question_ids", + }: + raise ImportIdentityError( + "Semantic condition fields may contain only condition_key and " + "expected_question_ids." + ) extra = canonicalize(fields) - reserved = {"condition_id", "condition_digest", "identity"} - if not isinstance(extra, Mapping) or reserved & set(extra): - raise ImportIdentityError("Semantic condition fields overwrite identity ownership.") return { "condition_id": condition_id, "condition_digest": condition_digest, @@ -405,8 +582,40 @@ def build_import_semantic_identity( prompt: ImportPromptSpec, ) -> ImportSemanticRecords: """Build semantic children and a current-shape scientific condition.""" + references = ( + ("dataset_id", condition.dataset_id, dataset.dataset_id), + ("model_key", condition.model_key, model.model_key), + ("method_key", condition.method_key, method.method_key), + ("prompt_key", condition.prompt_key, prompt.prompt_key), + ) + for field, declared, supplied in references: + if declared != supplied: + raise ImportIdentityError( + f"Condition {field}={declared!r} does not match supplied child {supplied!r}." + ) if tuple(condition.expected_question_ids) != dataset.selected_question_ids: raise ImportIdentityError("Condition question set does not match dataset selection.") + artifact_identity = dataset.artifact_payload + if dataset.identity_mode == "imported_semantic_fallback": + if ( + artifact_identity.get("benchmark") != dataset.benchmark_name + or artifact_identity.get("split") != dataset.split + ): + raise ImportIdentityError( + "Expected dataset description conflicts with artifact identity." + ) + else: + artifact_spec = artifact_identity.get("spec") + if ( + not isinstance(artifact_spec, Mapping) + or artifact_spec.get("benchmark") != dataset.benchmark_name + or artifact_spec.get("split") != dataset.split + ): + raise ImportIdentityError( + "Expected dataset description conflicts with native artifact identity." + ) + if dataset.selection_payload.get("artifact_id") != dataset.artifact_id: + raise ImportIdentityError("Dataset selection points to a different artifact.") dataset_record, selection, model_record, method_record, prompt_record = ( _semantic_child_records( condition=condition, @@ -495,6 +704,10 @@ def make_lineage_component( _validate_digest(digest, f"parent_digests[{index}]") for index, digest in enumerate(source_digests): _validate_digest(digest, f"source_digests[{index}]") + if len(parent_digests) != len(set(parent_digests)): + raise ImportIdentityError("Lineage parent_digests contains duplicate edges.") + if len(source_digests) != len(set(source_digests)): + raise ImportIdentityError("Lineage source_digests contains duplicate edges.") _validate_digest(authorization_digest, "authorization_digest", optional=True) _validate_digest(input_digest, "input_digest") _validate_digest(preownership_output_digest, "preownership_output_digest") @@ -537,13 +750,21 @@ def make_result_origin( rows: list[dict[str, str]] = [] seen: set[str] = set() counts: dict[str, int] = {} - for question_id, prediction_origin, lineage_id in row_assignments: + for index, assignment in enumerate(row_assignments): + if not isinstance(assignment, (list, tuple)) or len(assignment) != 3: + raise ImportIdentityError( + f"Result origin row assignment {index} must contain exactly " + "question_id, prediction_origin, and prediction_lineage_id." + ) + question_id, prediction_origin, lineage_id = assignment if not isinstance(question_id, str) or not question_id or question_id in seen: raise ImportIdentityError("Result origin question IDs must be non-empty and unique.") if prediction_origin not in _PREDICTION_ORIGINS: raise ImportIdentityError(f"Invalid prediction origin {prediction_origin!r}.") - if not isinstance(lineage_id, str) or not lineage_id: - raise ImportIdentityError("Result origin lineage IDs must be non-empty.") + if not isinstance(lineage_id, str) or not _LINEAGE_ID_RE.fullmatch(lineage_id): + raise ImportIdentityError( + "Result origin prediction_lineage_id must be a lineage-component ID." + ) seen.add(question_id) counts[prediction_origin] = counts.get(prediction_origin, 0) + 1 rows.append( @@ -561,7 +782,7 @@ def make_result_origin( } for row in rows ] - return { + identity = { "derivation_origin": derivation_origin, "prediction_origins": sorted(counts), "prediction_origin_counts": {key: counts[key] for key in sorted(counts)}, @@ -569,22 +790,249 @@ def make_result_origin( "ordered_row_origin_digest": integrity_digest(rows), "lineage_component_origin_digest": integrity_digest(lineage_projection), } + return { + "origin_id": short_id("origin", identity), + "origin_digest": integrity_digest(identity), + **identity, + } def importer_implementation_identity( *, adapter: Callable[..., Any], validator: Callable[..., Any] ) -> dict[str, Any]: """Bind installed ChoiceBench plus registered adapter and validator code.""" + records: dict[str, Mapping[str, Any]] = {} + for role, target in (("adapter", adapter), ("validator", validator)): + try: + record = implementation_identity(target) + except (OSError, TypeError, ValueError) as exc: + raise ImportIdentityError( + f"Importer {role} lacks an inspectable source identity." + ) from exc + if not isinstance(record.get("source_digest"), str): + raise ImportIdentityError( + f"Importer {role} lacks an inspectable source identity." + ) + records[role] = record return canonicalize( { "schema_version": "choicebench.importer-implementation.v1", "package": {"name": "choicebench", "version": __version__}, - "adapter": implementation_identity(adapter), - "validator": implementation_identity(validator), + "adapter": records["adapter"], + "validator": records["validator"], } ) +def _validate_result_origin_record(value: Any) -> dict[str, Any]: + record = _exact_mapping( + value, + { + "origin_id", + "origin_digest", + "derivation_origin", + "prediction_origins", + "prediction_origin_counts", + "row_assignments", + "ordered_row_origin_digest", + "lineage_component_origin_digest", + }, + "Realization result_origin", + ) + raw_rows = record["row_assignments"] + if not isinstance(raw_rows, (list, tuple)): + raise ImportIdentityError("Realization result_origin row_assignments must be a list.") + assignments = [] + for index, raw_row in enumerate(raw_rows): + row = _exact_mapping( + raw_row, + {"question_id", "prediction_origin", "prediction_lineage_id"}, + f"Realization result_origin row_assignments[{index}]", + ) + assignments.append( + ( + row["question_id"], + row["prediction_origin"], + row["prediction_lineage_id"], + ) + ) + rebuilt = make_result_origin( + derivation_origin=record["derivation_origin"], + row_assignments=assignments, + ) + if canonicalize(record) != canonicalize(rebuilt): + raise ImportIdentityError("Realization result_origin is internally inconsistent.") + return rebuilt + + +def _validate_realization_identity(value: Mapping[str, Any]) -> dict[str, Any]: + raw = _exact_mapping(value, _REALIZATION_KEYS, "Realization identity") + _validate_digest(raw["import_spec_digest"], "import_spec_digest") + + sources = raw["sources"] + if not isinstance(sources, (list, tuple)) or not sources: + raise ImportIdentityError("Realization sources must be a non-empty list.") + normalized_sources = [] + seen_source_ids: set[str] = set() + for index, raw_source in enumerate(sources): + source = _exact_mapping(raw_source, _SOURCE_KEYS, f"Realization sources[{index}]") + source_id = source["source_id"] + logical_path = source["logical_path"] + if ( + not isinstance(source_id, str) + or not source_id + or source_id in seen_source_ids + ): + raise ImportIdentityError("Realization source IDs must be non-empty and unique.") + if not isinstance(logical_path, str): + raise ImportIdentityError("Realization source logical_path must be a string.") + pure_path = PurePosixPath(logical_path) + if pure_path.is_absolute() or ".." in pure_path.parts or not pure_path.parts: + raise ImportIdentityError("Realization source logical_path is unsafe.") + if source["classification"] not in { + "raw", + "canonical", + "derived", + "repaired", + "aggregate_only", + }: + raise ImportIdentityError("Realization source classification is invalid.") + if source["format"] != "csv": + raise ImportIdentityError("Realization source format is unsupported.") + if not isinstance(source["format_version"], str) or not source["format_version"]: + raise ImportIdentityError("Realization source format_version is invalid.") + _validate_digest(source["sha256"], f"sources[{index}].sha256") + if not isinstance(source["provenance"], Mapping): + raise ImportIdentityError("Realization source provenance must be a mapping.") + _reject_forbidden(source["provenance"], "Realization source provenance") + seen_source_ids.add(source_id) + normalized_sources.append(canonicalize(source)) + + expected = _exact_mapping( + raw["expected_dataset"], + {"snapshot_digest", "question_set_digest", "derivation_digest"}, + "Realization expected_dataset", + ) + for field, digest in expected.items(): + _validate_digest(digest, f"expected_dataset.{field}") + + importer = _exact_mapping( + raw["importer_implementation"], + {"schema_version", "package", "adapter", "validator"}, + "Realization importer_implementation", + ) + if importer["schema_version"] != "choicebench.importer-implementation.v1": + raise ImportIdentityError("Realization importer implementation schema is invalid.") + for role in ("adapter", "validator"): + component = importer[role] + if not isinstance(component, Mapping) or not isinstance( + component.get("source_digest"), str + ): + raise ImportIdentityError( + f"Realization importer {role} lacks a source identity." + ) + _validate_digest(component["source_digest"], f"importer.{role}.source_digest") + + parsing = _exact_mapping( + raw["parsing_policy"], _PARSING_POLICY_KEYS, "Realization parsing_policy" + ) + dialect = _exact_mapping( + parsing["dialect"], _DIALECT_KEYS, "Realization parsing_policy.dialect" + ) + if not isinstance(parsing["mapping"], Mapping): + raise ImportIdentityError("Realization parsing mapping must be a mapping.") + for sequence_field in ("null_values", "numeric_columns"): + if not isinstance(parsing[sequence_field], (list, tuple)): + raise ImportIdentityError( + f"Realization parsing {sequence_field} must be a list." + ) + if not isinstance(parsing["option_mapping"], Mapping): + raise ImportIdentityError("Realization option_mapping must be a mapping.") + if parsing["extra_field_policy"] not in { + "preserve_unmapped", + "reject_unmapped", + }: + raise ImportIdentityError("Realization extra_field_policy is invalid.") + + validation = _exact_mapping( + raw["validation"], {"findings_digest"}, "Realization validation" + ) + _validate_digest(validation["findings_digest"], "validation.findings_digest") + evidence = _exact_mapping( + raw["evidence"], + { + "evidence_status", + "qualification_digest", + "limitation_digest", + "defect_digest", + "scope_disposition", + }, + "Realization evidence", + ) + if evidence["evidence_status"] not in _EVIDENCE_STATUSES: + raise ImportIdentityError("Realization evidence_status is invalid.") + if evidence["scope_disposition"] not in _SCOPE_DISPOSITIONS: + raise ImportIdentityError("Realization scope_disposition is invalid.") + for field in ("qualification_digest", "limitation_digest", "defect_digest"): + _validate_digest(evidence[field], f"evidence.{field}") + + result_origin = _validate_result_origin_record(raw["result_origin"]) + parents = _exact_mapping( + raw["parent_digests"], + {"realization_digests", "evidence_digests", "result_digests"}, + "Realization parent_digests", + ) + normalized_parents = { + field: _digest_sequence(items, f"parent_digests.{field}") + for field, items in parents.items() + } + authorization_digest = _validate_digest( + raw["authorization_digest"], "authorization_digest", optional=True + ) + overlay = raw["overlay"] + normalized_overlay = None + if overlay is not None: + overlay = _exact_mapping( + overlay, + { + "source_sha256", + "replacement_digest", + "transformation_input_digest", + "preownership_output_digest", + "implementation_digest", + }, + "Realization overlay", + ) + normalized_overlay = {} + for field, digest in overlay.items(): + normalized_overlay[field] = _validate_digest( + digest, f"overlay.{field}", optional=(field not in {"source_sha256", "replacement_digest"}) + ) + + normalized = { + "import_spec_digest": raw["import_spec_digest"], + "sources": normalized_sources, + "expected_dataset": canonicalize(expected), + "importer_implementation": canonicalize(importer), + "parsing_policy": { + **canonicalize(parsing), + "dialect": canonicalize(dialect), + }, + "validation": canonicalize(validation), + "evidence": canonicalize(evidence), + "result_origin": result_origin, + "parent_digests": normalized_parents, + "authorization_digest": authorization_digest, + "overlay": normalized_overlay, + } + _reject_forbidden( + normalized, + "Realization identity", + allowed_path_keys=frozenset({"logical_path"}), + ) + return canonicalize(normalized) + + def make_realization( *, condition_id: str, @@ -593,30 +1041,31 @@ def make_realization( fields: Mapping[str, Any], ) -> dict[str, Any]: """Create an immutable realization identity distinct from its condition.""" - if not isinstance(condition_id, str) or not condition_id.startswith("cond_"): + if not isinstance(condition_id, str) or not re.fullmatch(r"cond_[0-9a-f]{16}", condition_id): raise ImportIdentityError("Realization condition_id is invalid.") _validate_digest(condition_digest, "condition_digest") + if condition_id != f"cond_{condition_digest[:16]}": + raise ImportIdentityError( + "Realization condition_id does not match the condition_digest." + ) if not isinstance(identity, Mapping) or not identity: raise ImportIdentityError("Realization identity must be a non-empty mapping.") - _reject_forbidden(identity, "Realization identity") + validated_identity = _validate_realization_identity(identity) payload = canonicalize( { "schema_version": "choicebench.realization.v1", "condition_id": condition_id, "condition_digest": condition_digest, - "realization": identity, + "realization": validated_identity, } ) + if not isinstance(fields, Mapping) or not set(fields) <= {"audit", "realization_key"}: + raise ImportIdentityError( + "Realization fields may contain only audit and realization_key." + ) + if "audit" in fields and not isinstance(fields["audit"], Mapping): + raise ImportIdentityError("Realization audit field must be a mapping.") extra = canonicalize(fields) - reserved = { - "realization_id", - "realization_digest", - "condition_id", - "condition_digest", - "identity", - } - if not isinstance(extra, Mapping) or reserved & set(extra): - raise ImportIdentityError("Realization fields overwrite identity ownership.") return { "realization_id": short_id("real", payload), "realization_digest": integrity_digest(payload), diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py index 5ef4ecd..c139d7b 100644 --- a/tests/importing/test_identity.py +++ b/tests/importing/test_identity.py @@ -210,6 +210,14 @@ def test_fallback_child_payloads_are_exact_and_unknowns_are_explicit(): "template_digest": unknown_prompt, "template_identity": unknown_prompt, } + assert records.selection["identity"] == { + "artifact_id": records.dataset_artifact["artifact_id"], + "content_digest": records.selection["identity"]["content_digest"], + "sample_identities": records.selection["identity"]["sample_identities"], + "seed": None, + "n_samples": 2, + "subject_filter": [], + } def test_child_records_have_full_digests_and_existing_short_prefixes(): @@ -293,6 +301,31 @@ def test_make_semantic_condition_rejects_arbitrary_prompt_path_injection(): make_semantic_condition(identity=identity, fields={"condition_key": "x"}) +def test_semantic_condition_fields_are_closed_and_cannot_contradict_identity(): + identity = _records().condition["identity"] + for fields in ( + {"model_id": "model_conflict"}, + {"source_sha256": "1" * 64}, + {"realization_id": "real_child"}, + ): + with pytest.raises(ImportIdentityError, match="condition fields"): + make_semantic_condition(identity=identity, fields=fields) + + +@pytest.mark.parametrize( + ("field", "replacement"), + [ + ("dataset_id", "wrong-dataset"), + ("model_key", "wrong-model"), + ("method_key", "wrong-method"), + ("prompt_key", "wrong-prompt"), + ], +) +def test_semantic_builder_refuses_cross_wired_child_registry_keys(field, replacement): + with pytest.raises(ImportIdentityError, match=field): + _records(**{field: replacement}) + + def test_native_compatibility_children_are_preserved_without_fabrication(): fallback = _records() model_payload = {"backend": "dummy", "model_name_or_path": "native"} @@ -379,7 +412,15 @@ def test_native_model_identity_refuses_unbound_generation_parameters(): build_import_semantic_identity( condition=_condition(), dataset=_dataset(), - model=replace(_model(), native_compatibility_identity=native), + model=replace( + _model(), + backend="dummy", + unknown_reasons={ + "provider": "not applicable", + "revision": "not applicable", + }, + native_compatibility_identity=native, + ), method=_method(), prompt=_prompt(), ) @@ -426,7 +467,11 @@ def test_native_child_claims_must_match_their_semantic_declarations(): ), dataset=_dataset(), model=_model(), - method=replace(_method(), native_compatibility_identity=native_method), + method=replace( + _method(), + implementation={"qualified_name": "historical:SemanticMatching"}, + native_compatibility_identity=native_method, + ), prompt=_prompt(), ) @@ -456,6 +501,71 @@ def test_native_prompt_contents_must_match_the_semantic_declaration(): ), ) + bad_payload = { + **payload, + "files": { + **payload["files"], + "direct_mcq": {"sha256": "0" * 64, "content": "direct_mcq"}, + }, + } + bad_native = {"prompt_id": short_id("prompt", bad_payload), **bad_payload} + with pytest.raises(ImportIdentityError, match="prompt.*hash|prompt.*digest"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=replace( + _prompt(), + template_identity="v1", + template_digest=integrity_digest(bad_payload["files"]), + template_contents={ + name: item["content"] for name, item in bad_payload["files"].items() + }, + unknown_reason=None, + native_compatibility_identity=bad_native, + ), + ) + + +def test_unknown_declarations_cannot_be_upgraded_by_native_claims(): + method_payload = { + "name": "semantic_matching_v1", + "effective_params": {"matching": "semantic"}, + "preflight": {"reason": "not recorded", "value": None}, + "implementation": {"qualified_name": "fabricated:Current"}, + } + native_method = { + "payload": method_payload, + "digest": integrity_digest(method_payload), + "method_id": short_id("method", method_payload), + } + with pytest.raises(ImportIdentityError, match="unknown.*native|implementation"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=replace(_method(), native_compatibility_identity=native_method), + prompt=_prompt(), + ) + + prompt_payload = { + "version": "v1", + "files": { + name: {"sha256": integrity_digest(name), "content": name} + for name in ("direct_mcq", "free_text", "option_matching") + }, + } + native_prompt = {"prompt_id": short_id("prompt", prompt_payload), **prompt_payload} + with pytest.raises(ImportIdentityError, match="unknown.*native|prompt"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=replace(_prompt(), native_compatibility_identity=native_prompt), + ) + def test_fully_validated_native_children_reproduce_build_execution_plan_condition(): import choicebench.cli.run_experiment as run_experiment @@ -607,23 +717,85 @@ def _realization_identity() -> dict: origin = make_result_origin( derivation_origin="external_import", row_assignments=( - ("q2", "external_historical_inference", "lin_a"), - ("q1", "external_historical_inference", "lin_b"), + ("q2", "external_historical_inference", f"lin_{'a' * 16}"), + ("q1", "external_historical_inference", f"lin_{'b' * 16}"), ), ) return { "import_spec_digest": "1" * 64, - "source_sha256": "2" * 64, - "mapping": {"question_id": "qid"}, - "dialect": {"delimiter": ","}, - "evidence_status": "complete", - "scope": "included", + "sources": [ + { + "source_id": "results", + "logical_path": "publisher/results.csv", + "classification": "canonical", + "format": "csv", + "format_version": "1", + "sha256": "2" * 64, + "provenance": {"source_run_id": None}, + } + ], + "expected_dataset": { + "snapshot_digest": "3" * 64, + "question_set_digest": "4" * 64, + "derivation_digest": "5" * 64, + }, + "importer_implementation": importer_implementation_identity( + adapter=_adapter, validator=_validator + ), + "parsing_policy": { + "dialect": asdict(CsvDialectSpec()), + "mapping": {"question_id": "qid"}, + "null_values": [], + "numeric_columns": [], + "option_mapping": {"mode": "ordered_columns", "columns": ["a", "b"]}, + "extra_field_policy": "preserve_unmapped", + }, + "validation": {"findings_digest": "6" * 64}, + "evidence": { + "evidence_status": "complete", + "qualification_digest": "7" * 64, + "limitation_digest": "8" * 64, + "defect_digest": "9" * 64, + "scope_disposition": "included", + }, + "result_origin": origin, + "parent_digests": { + "realization_digests": [], + "evidence_digests": [], + "result_digests": [], + }, "authorization_digest": None, - "overlay_sha256": None, - "prediction_origins": origin, + "overlay": None, } +def _mutate_realization(identity: dict, field: str, value) -> None: + if field == "source_sha256": + identity["sources"][0]["sha256"] = value + elif field == "mapping": + identity["parsing_policy"]["mapping"] = value + elif field == "dialect": + identity["parsing_policy"]["dialect"].update(value) + elif field == "evidence_status": + identity["evidence"]["evidence_status"] = value + elif field == "scope": + identity["evidence"]["scope_disposition"] = value + elif field == "authorization_digest": + identity["authorization_digest"] = value + elif field == "overlay_sha256": + identity["overlay"] = { + "source_sha256": value, + "replacement_digest": "a" * 64, + "transformation_input_digest": None, + "preownership_output_digest": None, + "implementation_digest": None, + } + elif field == "prediction_origins": + identity["result_origin"] = value + else: + raise AssertionError(field) + + @pytest.mark.parametrize( ("field", "value"), [ @@ -634,7 +806,16 @@ def _realization_identity() -> dict: ("scope", "superseded"), ("authorization_digest", "4" * 64), ("overlay_sha256", "5" * 64), - ("prediction_origins", {"ordered_row_origin_digest": "6" * 64}), + ( + "prediction_origins", + make_result_origin( + derivation_origin="repair_overlay", + row_assignments=( + ("q2", "native_inference", f"lin_{'c' * 16}"), + ("q1", "external_historical_inference", f"lin_{'b' * 16}"), + ), + ), + ), ], ) def test_realization_changes_leave_condition_fixed(field, value): @@ -646,7 +827,7 @@ def test_realization_changes_leave_condition_fixed(field, value): fields={"audit": {"source_path": "/machine-a/results.csv"}}, ) changed_identity = _realization_identity() - changed_identity[field] = value + _mutate_realization(changed_identity, field, value) changed = make_realization( condition_id=condition["condition_id"], condition_digest=condition["condition_digest"], @@ -701,9 +882,12 @@ def test_every_csv_decoding_and_dialect_field_changes_realization_not_condition( condition = _records().condition base_dialect = asdict(CsvDialectSpec()) first_identity = _realization_identity() - first_identity["dialect"] = base_dialect + first_identity["parsing_policy"]["dialect"] = base_dialect changed_identity = _realization_identity() - changed_identity["dialect"] = {**base_dialect, field: value} + changed_identity["parsing_policy"]["dialect"] = { + **base_dialect, + field: value, + } first = make_realization( condition_id=condition["condition_id"], condition_digest=condition["condition_digest"], @@ -725,7 +909,13 @@ def test_base_repair_and_transformation_are_distinct_realizations_of_one_conditi identities = [] for derivation in ("external_import", "repair_overlay", "offline_transformation"): identity = _realization_identity() - identity["derivation_origin"] = derivation + identity["result_origin"] = make_result_origin( + derivation_origin=derivation, + row_assignments=( + ("q2", "external_historical_inference", f"lin_{'a' * 16}"), + ("q1", "external_historical_inference", f"lin_{'b' * 16}"), + ), + ) identities.append( make_realization( condition_id=condition["condition_id"], @@ -760,6 +950,61 @@ def test_realization_refuses_post_publication_or_machine_local_identity_fields() ) +def test_realization_schema_is_closed_complete_and_origin_consistent(): + condition = _records().condition + cases = [] + missing = _realization_identity() + missing.pop("validation") + cases.append(missing) + extra = _realization_identity() + extra["import_state"] = "validated" + cases.append(extra) + incomplete_dialect = _realization_identity() + incomplete_dialect["parsing_policy"]["dialect"] = {"delimiter": ","} + cases.append(incomplete_dialect) + bare_mixed = _realization_identity() + bare_mixed["result_origin"] = { + "derivation_origin": "repair_overlay", + "prediction_origins": ["mixed"], + } + cases.append(bare_mixed) + for identity in cases: + with pytest.raises(ImportIdentityError): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_condition_short_id_must_match_full_digest(): + condition = _records().condition + with pytest.raises(ImportIdentityError, match="condition_id.*digest"): + make_realization( + condition_id="cond_0000000000000000", + condition_digest=condition["condition_digest"], + identity=_realization_identity(), + fields={}, + ) + + +def test_realization_fields_are_audit_only_and_cannot_repeat_identity_claims(): + condition = _records().condition + for fields in ( + {"evidence_status": "qualified"}, + {"scope_disposition": "held"}, + {"result_origin": {}}, + ): + with pytest.raises(ImportIdentityError, match="Realization fields"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=_realization_identity(), + fields=fields, + ) + + def test_lineage_component_is_parent_derived_and_has_no_child_edge(): component = make_lineage_component( operation_type="offline_transformation", @@ -781,7 +1026,15 @@ def test_lineage_component_is_parent_derived_and_has_no_child_edge(): def test_lineage_refuses_child_paths_timestamps_and_final_csv_hashes(): - for forbidden in ("realization_id", "output_path", "timestamp", "final_csv_sha256"): + for forbidden in ( + "realization_id", + "child_realization_id", + "output_path", + "result_path", + "timestamp", + "created_at", + "final_csv_sha256", + ): with pytest.raises(ImportIdentityError, match=forbidden): make_lineage_component( operation_type="offline_transformation", @@ -797,13 +1050,33 @@ def test_lineage_refuses_child_paths_timestamps_and_final_csv_hashes(): ) +def test_lineage_refuses_duplicate_parent_or_source_edges(): + for parents, sources in ( + (("1" * 64, "1" * 64), ("2" * 64,)), + (("1" * 64,), ("2" * 64, "2" * 64)), + ): + with pytest.raises(ImportIdentityError, match="duplicate"): + make_lineage_component( + operation_type="external_import", + question_id="q1", + parent_digests=parents, + source_digests=sources, + authorization_digest=None, + implementation={"qualified_name": "x:y", "source_digest": "3" * 64}, + parameters={}, + input_digest="4" * 64, + preownership_output_digest="5" * 64, + prediction_origin="external_historical_inference", + ) + + def test_result_origin_records_constituents_counts_and_ordered_per_row_mapping(): origin = make_result_origin( derivation_origin="repair_overlay", row_assignments=( - ("q1", "external_historical_inference", "lin_base"), - ("q2", "native_inference", "lin_repair"), - ("q3", "native_inference", "lin_repair_2"), + ("q1", "external_historical_inference", f"lin_{'a' * 16}"), + ("q2", "native_inference", f"lin_{'b' * 16}"), + ("q3", "native_inference", f"lin_{'c' * 16}"), ), ) @@ -823,14 +1096,21 @@ def test_result_origin_records_constituents_counts_and_ordered_per_row_mapping() assert origin["ordered_row_origin_digest"] == integrity_digest( origin["row_assignments"] ) + identity = { + key: value + for key, value in origin.items() + if key not in {"origin_id", "origin_digest"} + } + assert origin["origin_digest"] == integrity_digest(identity) + assert origin["origin_id"] == short_id("origin", identity) def test_result_origin_rejects_bare_mixed_duplicate_rows_and_invalid_origins(): for assignments in ( - (("q1", "mixed", "lin_a"),), + (("q1", "mixed", f"lin_{'a' * 16}"),), ( - ("q1", "external_historical_inference", "lin_a"), - ("q1", "native_inference", "lin_b"), + ("q1", "external_historical_inference", f"lin_{'a' * 16}"), + ("q1", "native_inference", f"lin_{'b' * 16}"), ), ): with pytest.raises(ImportIdentityError): @@ -843,7 +1123,7 @@ def test_offline_transformation_keeps_underlying_prediction_origin(): origin = make_result_origin( derivation_origin="offline_transformation", row_assignments=( - ("q1", "external_historical_inference", "lin_rematch"), + ("q1", "external_historical_inference", f"lin_{'a' * 16}"), ), ) assert origin["derivation_origin"] == "offline_transformation" @@ -860,3 +1140,8 @@ def test_importer_implementation_identity_binds_package_adapter_and_validator(): assert identity["validator"]["qualified_name"].endswith(":_validator") changed = importer_implementation_identity(adapter=_validator, validator=_adapter) assert integrity_digest(identity) != integrity_digest(changed) + + +def test_importer_implementation_identity_refuses_uninspectable_callables(): + with pytest.raises(ImportIdentityError, match="source identity"): + importer_implementation_identity(adapter=len, validator=_validator) From 930dae20b0a9518bf4f3b3836a7183ef5c2fd93e Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 02:30:16 +0300 Subject: [PATCH 18/47] fix: harden importer identity trust boundaries --- src/choicebench/importing/identity.py | 387 +++++++++++++++++++++----- src/choicebench/importing/schema.py | 51 ++++ tests/importing/test_identity.py | 377 +++++++++++++++++++++++-- 3 files changed, 731 insertions(+), 84 deletions(-) diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index c91061e..004fedc 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -11,11 +11,18 @@ from choicebench.identity import canonicalize, integrity_digest, short_id from choicebench.importing.dataset_reference import ExpectedDataset from choicebench.importing.schema import ( - ImportConditionSpec, CsvDialectSpec, + ImportConditionSpec, + ImportSpecError, ImportMethodSpec, ImportModelSpec, ImportPromptSpec, + validate_csv_dialect_identity, + validate_implementation_identity_record, + validate_native_method_payload, + validate_native_model_payload, + validate_numeric_columns_identity, + validate_option_mapping_identity, ) from choicebench.provenance import implementation_identity @@ -24,6 +31,17 @@ class ImportIdentityError(ValueError): """Raised when an importer identity is cyclic, unsafe, or inconsistent.""" +class RuntimeImporterImplementation(dict[str, Any]): + """Marker for implementation identity derived from live runtime callables.""" + + def __init__(self, record: Mapping[str, Any]) -> None: + super().__init__(record) + self._runtime_digest = integrity_digest(record) + + def is_unmodified(self) -> bool: + return integrity_digest(dict(self)) == self._runtime_digest + + @dataclass(frozen=True) class ImportSemanticRecords: dataset_artifact: Mapping[str, Any] @@ -76,6 +94,13 @@ class ImportSemanticRecords: "sha256", "provenance", } +_SOURCE_PROVENANCE_KEYS = { + "source_run_id", + "source_repository", + "source_commit", + "notes_digest", + "evidence_digest", +} _PARSING_POLICY_KEYS = { "dialect", "mapping", @@ -106,7 +131,10 @@ class ImportSemanticRecords: "choicebench_home", "experiment_id", "final_csv_sha256", + "final_result_digest", + "host", "hostname", + "import_date", "import_timestamp", "output_path", "output_root", @@ -119,6 +147,9 @@ class ImportSemanticRecords: "source_path", "temporary_path", "timestamp", + "machine_id", + "cwd", + "child_lineage_id", } @@ -161,7 +192,7 @@ def _reject_forbidden( raise ImportIdentityError(f"{where} contains a non-string field name.") normalized = key.casefold() unsafe_alias = ( - key in _FORBIDDEN_IDENTITY_KEYS + normalized in _FORBIDDEN_IDENTITY_KEYS or (normalized not in allowed_path_keys and normalized.endswith("_path")) or normalized in {"path", "created_at", "updated_at", "import_state"} or "timestamp" in normalized @@ -178,6 +209,95 @@ def _reject_forbidden( _reject_forbidden(item, where, allowed_path_keys=allowed_path_keys) +def _is_machine_path(value: str) -> bool: + return bool( + value.startswith(("/", "./", "../", "~/", "~\\")) + or re.match(r"^[A-Za-z]:[\\/]", value) + ) + + +def _validate_stable_values(value: Any, where: str) -> None: + _reject_forbidden(value, where) + if isinstance(value, Mapping): + for item in value.values(): + _validate_stable_values(item, where) + elif isinstance(value, (list, tuple)): + for item in value: + _validate_stable_values(item, where) + elif isinstance(value, str) and _is_machine_path(value): + raise ImportIdentityError(f"{where} contains a machine-local path value.") + + +def _validate_condition_identity(value: Mapping[str, Any]) -> None: + benchmark = _exact_mapping( + value["benchmark"], + {"name", "split", "selection_id", "artifact_id"}, + "Semantic condition benchmark", + ) + for field in ("name", "split"): + if not isinstance(benchmark[field], str) or not benchmark[field]: + raise ImportIdentityError( + f"Semantic condition benchmark {field} must be non-empty." + ) + for field, prefix in ( + ("selection_id", "sel"), + ("artifact_id", "ds"), + ("model_id", "model"), + ("method_id", "method"), + ("prompt_id", "prompt"), + ): + identifier = benchmark[field] if field in benchmark else value[field] + if not isinstance(identifier, str) or not re.fullmatch( + rf"{prefix}_[0-9a-f]{{16}}", identifier + ): + raise ImportIdentityError( + f"Semantic condition {field} is not a valid {prefix} identity." + ) + preflight = value["preflight"] + if preflight is not None: + if not isinstance(preflight, Mapping): + raise ImportIdentityError("Semantic condition preflight must be a mapping or null.") + if set(preflight) == {"value", "reason"}: + if preflight["value"] is not None or not isinstance( + preflight["reason"], str + ) or not preflight["reason"].strip(): + raise ImportIdentityError( + "Semantic condition unknown preflight must include a reason." + ) + else: + native = _exact_mapping( + preflight, + {"artifact_id", "selection_id", "split", "content_digest"}, + "Semantic condition preflight", + ) + if not re.fullmatch(r"ds_[0-9a-f]{16}", str(native["artifact_id"])): + raise ImportIdentityError("Semantic condition preflight artifact_id is invalid.") + if not re.fullmatch(r"sel_[0-9a-f]{16}", str(native["selection_id"])): + raise ImportIdentityError("Semantic condition preflight selection_id is invalid.") + if not isinstance(native["split"], str) or not native["split"]: + raise ImportIdentityError("Semantic condition preflight split is invalid.") + _validate_digest(native["content_digest"], "preflight.content_digest") + seed = value["seed"] + if not ( + isinstance(seed, int) + and not isinstance(seed, bool) + or ( + isinstance(seed, Mapping) + and set(seed) == {"value", "reason"} + and seed.get("value") is None + and isinstance(seed.get("reason"), str) + and bool(seed["reason"].strip()) + ) + ): + raise ImportIdentityError("Semantic condition seed is invalid.") + if "protocol_settings" in value: + if not isinstance(value["protocol_settings"], Mapping): + raise ImportIdentityError("Semantic condition protocol_settings must be a mapping.") + _validate_stable_values( + value["protocol_settings"], "Semantic condition protocol_settings" + ) + + def _exact_mapping(value: Any, keys: set[str], where: str) -> Mapping[str, Any]: if not isinstance(value, Mapping): raise ImportIdentityError(f"{where} must be a mapping.") @@ -249,30 +369,15 @@ def _semantic_child_records( raise ImportIdentityError( "An unknown model backend cannot be upgraded by a native identity claim." ) - native_model_keys = { - "dummy": {"backend", "model_name_or_path"}, - "api": { - "backend", - "provider", - "model_name_or_path", - "base_url", - "generation_kwargs", - }, - "huggingface": { - "backend", - "model", - "device", - "add_bos_token", - "generation_kwargs", - "loader", - }, - } - if model.backend not in native_model_keys or set(model_payload) != native_model_keys[ - model.backend - ]: + try: + validated_model_payload = validate_native_model_payload(model_payload) + except ImportSpecError as exc: + raise ImportIdentityError(f"Invalid validated native model payload: {exc}") from exc + if canonicalize(validated_model_payload) != canonicalize(model_payload): raise ImportIdentityError( "Validated native model payload does not match the current backend schema." ) + model_payload = validated_model_payload for nullable_field in ("provider", "revision"): value = getattr(model, nullable_field) reason = model.unknown_reasons.get(nullable_field) @@ -299,16 +404,34 @@ def _semantic_child_records( raise ImportIdentityError( "Validated native model name conflicts with its declaration." ) - if condition.generation_parameters: - native_generation = model_payload.get("generation_kwargs") - if not isinstance(native_generation, Mapping) or any( - native_generation.get(name) != value - for name, value in condition.generation_parameters.items() - ): + if model.backend == "huggingface": + resolved_model = model_payload["model"] + declared_name = ( + resolved_model["repo_id"] + if resolved_model["kind"] == "huggingface-hub" + else resolved_model["logical_name"] + ) + if declared_name != model.display_name: raise ImportIdentityError( - "Condition generation parameters are not bound by the validated " - "native model identity." + "Validated native model name conflicts with its declaration." ) + if resolved_model.get("requested_revision") != model.revision: + raise ImportIdentityError( + "Validated native model revision conflicts with its declaration." + ) + expected_generation = dict(model.effective_parameters) + for name, value in condition.generation_parameters.items(): + if name in expected_generation and expected_generation[name] != value: + raise ImportIdentityError( + f"Generation parameter {name!r} conflicts with model effective parameters." + ) + expected_generation[name] = value + native_generation = model_payload.get("generation_kwargs", {}) + if canonicalize(native_generation) != canonicalize(expected_generation): + raise ImportIdentityError( + "Condition generation parameters are not bound by the validated " + "native model identity." + ) model_mode = "native_compatibility" else: effective_parameters = dict(model.effective_parameters) @@ -361,15 +484,17 @@ def _semantic_child_records( claimed_id=native["method_id"], claimed_digest=native["digest"], ) - if set(method_payload) != { - "name", - "effective_params", - "preflight", - "implementation", - }: + try: + validated_method_payload = validate_native_method_payload(method_payload) + except ImportSpecError as exc: + raise ImportIdentityError( + f"Invalid validated native method payload: {exc}" + ) from exc + if canonicalize(validated_method_payload) != canonicalize(method_payload): raise ImportIdentityError( "Validated native method payload does not match the current schema." ) + method_payload = validated_method_payload if method.implementation is None: raise ImportIdentityError( "An unknown method implementation cannot be upgraded by a native " @@ -548,6 +673,7 @@ def make_semantic_condition( raise ImportIdentityError( "Semantic condition prompt snapshot path must be derived from prompt_id." ) + _validate_condition_identity(identity) _reject_forbidden( identity, "Semantic condition identity", @@ -582,6 +708,13 @@ def build_import_semantic_identity( prompt: ImportPromptSpec, ) -> ImportSemanticRecords: """Build semantic children and a current-shape scientific condition.""" + if dataset.identity_mode not in { + "imported_semantic_fallback", + "native_compatibility", + }: + raise ImportIdentityError( + f"Unsupported expected dataset identity mode {dataset.identity_mode!r}." + ) references = ( ("dataset_id", condition.dataset_id, dataset.dataset_id), ("model_key", condition.model_key, model.model_key), @@ -711,8 +844,18 @@ def make_lineage_component( _validate_digest(authorization_digest, "authorization_digest", optional=True) _validate_digest(input_digest, "input_digest") _validate_digest(preownership_output_digest, "preownership_output_digest") - _reject_forbidden(implementation, "Lineage implementation") - _reject_forbidden(parameters, "Lineage parameters") + try: + validated_implementation = validate_implementation_identity_record( + implementation + ) + except ImportSpecError as exc: + raise ImportIdentityError(f"Invalid lineage implementation: {exc}") from exc + if "source_digest" not in validated_implementation: + raise ImportIdentityError( + "Lineage implementation requires an inspectable source digest." + ) + _validate_stable_values(validated_implementation, "Lineage implementation") + _validate_stable_values(parameters, "Lineage parameters") payload = canonicalize( { "schema_version": "choicebench.lineage-component.v1", @@ -721,7 +864,7 @@ def make_lineage_component( "parent_digests": sorted(parent_digests), "source_digests": sorted(source_digests), "authorization_digest": authorization_digest, - "implementation": implementation, + "implementation": validated_implementation, "parameters": parameters, "input_digest": input_digest, "preownership_output_digest": preownership_output_digest, @@ -799,7 +942,7 @@ def make_result_origin( def importer_implementation_identity( *, adapter: Callable[..., Any], validator: Callable[..., Any] -) -> dict[str, Any]: +) -> RuntimeImporterImplementation: """Bind installed ChoiceBench plus registered adapter and validator code.""" records: dict[str, Mapping[str, Any]] = {} for role, target in (("adapter", adapter), ("validator", validator)): @@ -814,13 +957,15 @@ def importer_implementation_identity( f"Importer {role} lacks an inspectable source identity." ) records[role] = record - return canonicalize( - { - "schema_version": "choicebench.importer-implementation.v1", - "package": {"name": "choicebench", "version": __version__}, - "adapter": records["adapter"], - "validator": records["validator"], - } + return RuntimeImporterImplementation( + canonicalize( + { + "schema_version": "choicebench.importer-implementation.v1", + "package": {"name": "choicebench", "version": __version__}, + "adapter": records["adapter"], + "validator": records["validator"], + } + ) ) @@ -902,9 +1047,26 @@ def _validate_realization_identity(value: Mapping[str, Any]) -> dict[str, Any]: if not isinstance(source["format_version"], str) or not source["format_version"]: raise ImportIdentityError("Realization source format_version is invalid.") _validate_digest(source["sha256"], f"sources[{index}].sha256") - if not isinstance(source["provenance"], Mapping): - raise ImportIdentityError("Realization source provenance must be a mapping.") - _reject_forbidden(source["provenance"], "Realization source provenance") + provenance = _exact_mapping( + source["provenance"], + _SOURCE_PROVENANCE_KEYS, + f"Realization sources[{index}].provenance", + ) + for field in ("source_run_id", "source_repository", "source_commit"): + value = provenance[field] + if value is not None and (not isinstance(value, str) or not value.strip()): + raise ImportIdentityError( + f"Realization source provenance {field} must be null or non-empty." + ) + if isinstance(value, str) and _is_machine_path(value): + raise ImportIdentityError( + f"Realization source provenance {field} contains a machine-local path." + ) + for field in ("notes_digest", "evidence_digest"): + _validate_digest( + provenance[field], f"sources[{index}].provenance.{field}", optional=True + ) + source = {**source, "provenance": canonicalize(provenance)} seen_source_ids.add(source_id) normalized_sources.append(canonicalize(source)) @@ -916,18 +1078,33 @@ def _validate_realization_identity(value: Mapping[str, Any]) -> dict[str, Any]: for field, digest in expected.items(): _validate_digest(digest, f"expected_dataset.{field}") + raw_importer = raw["importer_implementation"] + if not isinstance(raw_importer, RuntimeImporterImplementation) or not ( + raw_importer.is_unmodified() + ): + raise ImportIdentityError( + "Realization importer implementation must be derived from runtime callables." + ) importer = _exact_mapping( - raw["importer_implementation"], + raw_importer, {"schema_version", "package", "adapter", "validator"}, "Realization importer_implementation", ) if importer["schema_version"] != "choicebench.importer-implementation.v1": raise ImportIdentityError("Realization importer implementation schema is invalid.") + package = _exact_mapping( + importer["package"], {"name", "version"}, "Realization importer package" + ) + if package != {"name": "choicebench", "version": __version__}: + raise ImportIdentityError("Realization importer package identity is invalid.") for role in ("adapter", "validator"): - component = importer[role] - if not isinstance(component, Mapping) or not isinstance( - component.get("source_digest"), str - ): + try: + component = validate_implementation_identity_record(importer[role]) + except ImportSpecError as exc: + raise ImportIdentityError( + f"Realization importer {role} identity is invalid: {exc}" + ) from exc + if not isinstance(component.get("source_digest"), str): raise ImportIdentityError( f"Realization importer {role} lacks a source identity." ) @@ -936,18 +1113,45 @@ def _validate_realization_identity(value: Mapping[str, Any]) -> dict[str, Any]: parsing = _exact_mapping( raw["parsing_policy"], _PARSING_POLICY_KEYS, "Realization parsing_policy" ) - dialect = _exact_mapping( + _exact_mapping( parsing["dialect"], _DIALECT_KEYS, "Realization parsing_policy.dialect" ) - if not isinstance(parsing["mapping"], Mapping): + try: + dialect = validate_csv_dialect_identity(parsing["dialect"]) + except ImportSpecError as exc: + raise ImportIdentityError(f"Invalid realization parsing dialect: {exc}") from exc + mapping = parsing["mapping"] + if not isinstance(mapping, Mapping): raise ImportIdentityError("Realization parsing mapping must be a mapping.") - for sequence_field in ("null_values", "numeric_columns"): - if not isinstance(parsing[sequence_field], (list, tuple)): + normalized_mapping: dict[str, str] = {} + for semantic_field, source_column in mapping.items(): + if not isinstance(semantic_field, str) or not semantic_field.strip(): raise ImportIdentityError( - f"Realization parsing {sequence_field} must be a list." + "Realization parsing mapping keys must be non-empty strings." ) - if not isinstance(parsing["option_mapping"], Mapping): - raise ImportIdentityError("Realization option_mapping must be a mapping.") + if not isinstance(source_column, str) or not source_column.strip(): + raise ImportIdentityError( + "Realization parsing mapping values must be non-empty source columns." + ) + normalized_mapping[semantic_field] = source_column + if len(normalized_mapping.values()) != len(set(normalized_mapping.values())): + raise ImportIdentityError("Realization parsing mapping reuses a source column.") + null_values = parsing["null_values"] + if not isinstance(null_values, (list, tuple)) or not all( + isinstance(item, str) for item in null_values + ): + raise ImportIdentityError( + "Realization parsing null_values must be a string list." + ) + if len(null_values) != len(set(null_values)): + raise ImportIdentityError("Realization parsing null_values contains duplicates.") + try: + numeric_columns = validate_numeric_columns_identity( + parsing["numeric_columns"] + ) + option_mapping = validate_option_mapping_identity(parsing["option_mapping"]) + except ImportSpecError as exc: + raise ImportIdentityError(f"Invalid realization parsing policy: {exc}") from exc if parsing["extra_field_policy"] not in { "preserve_unmapped", "reject_unmapped", @@ -1006,18 +1210,65 @@ def _validate_realization_identity(value: Mapping[str, Any]) -> dict[str, Any]: normalized_overlay = {} for field, digest in overlay.items(): normalized_overlay[field] = _validate_digest( - digest, f"overlay.{field}", optional=(field not in {"source_sha256", "replacement_digest"}) + digest, + f"overlay.{field}", + optional=(field not in {"source_sha256", "replacement_digest"}), ) + derivation_origin = result_origin["derivation_origin"] + derived = derivation_origin in {"repair_overlay", "offline_transformation"} + if derived: + if authorization_digest is None: + raise ImportIdentityError( + f"{derivation_origin} realization requires an authorization digest." + ) + if not normalized_parents["realization_digests"]: + raise ImportIdentityError( + f"{derivation_origin} realization requires a parent realization digest." + ) + if normalized_overlay is None: + raise ImportIdentityError( + f"{derivation_origin} realization requires an overlay identity." + ) + if derivation_origin == "offline_transformation" and any( + normalized_overlay[field] is None + for field in ( + "transformation_input_digest", + "preownership_output_digest", + "implementation_digest", + ) + ): + raise ImportIdentityError( + "offline_transformation overlay requires input, output, and " + "implementation digests." + ) + if derivation_origin == "repair_overlay" and not ( + {"native_inference", "external_repair_inference"} + & set(result_origin["prediction_origins"]) + ): + raise ImportIdentityError( + "repair_overlay requires at least one repair prediction origin." + ) + elif authorization_digest is not None or normalized_overlay is not None: + raise ImportIdentityError( + f"{derivation_origin} realization cannot claim repair authorization or overlay." + ) + normalized = { "import_spec_digest": raw["import_spec_digest"], "sources": normalized_sources, "expected_dataset": canonicalize(expected), "importer_implementation": canonicalize(importer), - "parsing_policy": { - **canonicalize(parsing), - "dialect": canonicalize(dialect), - }, + "parsing_policy": canonicalize( + { + "dialect": dialect, + "mapping": normalized_mapping, + "null_values": list(null_values), + "numeric_columns": numeric_columns, + "option_mapping": option_mapping, + "extra_field_policy": parsing["extra_field_policy"], + } + ), "validation": canonicalize(validation), "evidence": canonicalize(evidence), "result_origin": result_origin, diff --git a/src/choicebench/importing/schema.py b/src/choicebench/importing/schema.py index cf67300..4e43508 100644 --- a/src/choicebench/importing/schema.py +++ b/src/choicebench/importing/schema.py @@ -1178,6 +1178,57 @@ def _validate_prompt_native(value: Any, where: str) -> dict[str, Any]: return {"prompt_id": prompt_id, **payload} +def validate_csv_dialect_identity(value: Any) -> dict[str, Any]: + """Validate and normalize an identity-bearing CSV dialect declaration.""" + prepared = dict(value) if isinstance(value, Mapping) else value + if isinstance(prepared, dict) and isinstance( + prepared.get("line_terminators"), tuple + ): + prepared["line_terminators"] = list(prepared["line_terminators"]) + return asdict(_build_dialect(prepared, "parsing_policy.dialect")) + + +def validate_numeric_columns_identity(value: Any) -> list[dict[str, Any]]: + """Validate identity-bearing numeric-column declarations.""" + if not isinstance(value, (list, tuple)): + raise ImportSpecError("parsing_policy.numeric_columns must be a list.") + records = [ + asdict(_build_numeric(item, f"parsing_policy.numeric_columns[{index}]")) + for index, item in enumerate(value) + ] + columns = [record["source_column"] for record in records] + if len(columns) != len(set(columns)): + raise ImportSpecError( + "parsing_policy.numeric_columns contains duplicate source columns." + ) + return records + + +def validate_option_mapping_identity(value: Any) -> dict[str, Any]: + """Validate and normalize an identity-bearing option mapping.""" + prepared = dict(value) if isinstance(value, Mapping) else value + if isinstance(prepared, dict) and isinstance( + prepared.get("ordered_columns"), tuple + ): + prepared["ordered_columns"] = list(prepared["ordered_columns"]) + return asdict(_build_option(prepared, "parsing_policy.option_mapping")) + + +def validate_native_model_payload(value: Any) -> dict[str, Any]: + """Validate a claimed current-native model identity payload.""" + return _validate_model_payload(value, "native_model.payload") + + +def validate_native_method_payload(value: Any) -> dict[str, Any]: + """Validate a claimed current-native method identity payload.""" + return _validate_method_payload(value, "native_method.payload") + + +def validate_implementation_identity_record(value: Any) -> dict[str, Any]: + """Validate the closed runtime implementation-identity record shape.""" + return _validate_implementation(value, "implementation") + + def _validate_simple_native( value: Any, where: str, prefix: str, payload_validator ) -> dict[str, Any]: diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py index c139d7b..c790d59 100644 --- a/tests/importing/test_identity.py +++ b/tests/importing/test_identity.py @@ -35,6 +35,8 @@ ImportMethodSpec, ImportModelSpec, ImportPromptSpec, + NumericColumnSpec, + OptionMappingSpec, ResultOriginSpec, ) @@ -282,9 +284,11 @@ def test_semantic_child_change_changes_condition_identity(field): identity = dict(original["identity"]) if field in {"selection_id", "artifact_id"}: identity["benchmark"] = dict(identity["benchmark"]) - identity["benchmark"][field] = f"{field}_changed" + prefix = "sel" if field == "selection_id" else "ds" + identity["benchmark"][field] = f"{prefix}_{'f' * 16}" else: - identity[field] = f"{field}_changed" + prefix = field.removesuffix("_id") + identity[field] = f"{prefix}_{'f' * 16}" if field == "prompt_id": identity["prompt_snapshot_path"] = ( f"artifacts/prompts/{identity['prompt_id']}" @@ -438,7 +442,11 @@ def test_native_child_claims_must_match_their_semantic_declarations(): condition=replace(_condition(), generation_parameters={}), dataset=_dataset(), model=replace( - _model(), backend="api", native_compatibility_identity=native_model + _model(), + backend="api", + provider="provider", + unknown_reasons={"revision": "not applicable"}, + native_compatibility_identity=native_model, ), method=_method(), prompt=_prompt(), @@ -528,11 +536,113 @@ def test_native_prompt_contents_must_match_the_semantic_declaration(): ) +def test_native_api_payload_is_validated_by_the_authoritative_schema(): + payload = { + "backend": "api", + "provider": 7, + "model_name_or_path": "historical-model", + "base_url": None, + "generation_kwargs": ["not", "a", "mapping"], + } + native = { + "payload": payload, + "digest": integrity_digest(payload), + "model_id": short_id("model", payload), + } + with pytest.raises(ImportIdentityError, match="native model|provider|generation"): + build_import_semantic_identity( + condition=replace(_condition(), generation_parameters={}), + dataset=_dataset(), + model=replace( + _model(), + backend="api", + provider="provider", + unknown_reasons={"revision": "not applicable"}, + native_compatibility_identity=native, + ), + method=_method(), + prompt=_prompt(), + ) + + +def test_native_huggingface_payload_binds_name_revision_and_parameters(): + payload = { + "backend": "huggingface", + "model": { + "kind": "huggingface-hub", + "repo_id": "right/model", + "requested_revision": "right-revision", + "resolved_commit": "a" * 40, + }, + "device": "cpu", + "add_bos_token": False, + "generation_kwargs": { + "max_new_tokens": 64, + "temperature": 0, + "do_sample": False, + }, + "loader": {"trust_remote_code": True, "torch_dtype": "float32"}, + } + native = { + "payload": payload, + "digest": integrity_digest(payload), + "model_id": short_id("model", payload), + } + with pytest.raises(ImportIdentityError, match="name|revision|parameter|declaration"): + build_import_semantic_identity( + condition=replace(_condition(), generation_parameters={}), + dataset=_dataset(), + model=replace( + _model(), + display_name="wrong/model", + backend="huggingface", + revision="wrong-revision", + effective_parameters={"temperature": 1}, + unknown_reasons={"provider": "not applicable"}, + native_compatibility_identity=native, + ), + method=_method(), + prompt=_prompt(), + ) + + +def test_native_method_payload_requires_authoritative_implementation_shape(): + payload = { + "name": "semantic_matching_v1", + "effective_params": {"matching": "semantic"}, + "preflight": None, + "implementation": {"attacker": "not-code-identity"}, + } + native = { + "payload": payload, + "digest": integrity_digest(payload), + "method_id": short_id("method", payload), + } + with pytest.raises(ImportIdentityError, match="native method|implementation"): + build_import_semantic_identity( + condition=replace( + _condition(), + unknown_reasons={ + **_condition().unknown_reasons, + "preflight_identity": "not applicable", + }, + ), + dataset=_dataset(), + model=_model(), + method=replace( + _method(), + implementation={"qualified_name": "historical:Matcher"}, + native_compatibility_identity=native, + ), + prompt=_prompt(), + ) + + def test_unknown_declarations_cannot_be_upgraded_by_native_claims(): method_payload = { "name": "semantic_matching_v1", "effective_params": {"matching": "semantic"}, - "preflight": {"reason": "not recorded", "value": None}, + "preflight": None, "implementation": {"qualified_name": "fabricated:Current"}, } native_method = { @@ -542,7 +652,13 @@ def test_unknown_declarations_cannot_be_upgraded_by_native_claims(): } with pytest.raises(ImportIdentityError, match="unknown.*native|implementation"): build_import_semantic_identity( - condition=_condition(), + condition=replace( + _condition(), + unknown_reasons={ + **_condition().unknown_reasons, + "preflight_identity": "not applicable", + }, + ), dataset=_dataset(), model=_model(), method=replace(_method(), native_compatibility_identity=native_method), @@ -731,7 +847,13 @@ def _realization_identity() -> dict: "format": "csv", "format_version": "1", "sha256": "2" * 64, - "provenance": {"source_run_id": None}, + "provenance": { + "source_run_id": None, + "source_repository": None, + "source_commit": None, + "notes_digest": None, + "evidence_digest": None, + }, } ], "expected_dataset": { @@ -746,8 +868,24 @@ def _realization_identity() -> dict: "dialect": asdict(CsvDialectSpec()), "mapping": {"question_id": "qid"}, "null_values": [], - "numeric_columns": [], - "option_mapping": {"mode": "ordered_columns", "columns": ["a", "b"]}, + "numeric_columns": [ + asdict( + NumericColumnSpec( + source_column="score", + value_type="float", + null_allowed=True, + ) + ) + ], + "option_mapping": asdict( + OptionMappingSpec( + mode="ordered_columns", + ordered_columns=("a", "b"), + structured_column=None, + structured_label_key=None, + structured_text_key=None, + ) + ), "extra_field_policy": "preserve_unmapped", }, "validation": {"findings_digest": "6" * 64}, @@ -782,7 +920,24 @@ def _mutate_realization(identity: dict, field: str, value) -> None: identity["evidence"]["scope_disposition"] = value elif field == "authorization_digest": identity["authorization_digest"] = value + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["overlay"] = { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": None, + "preownership_output_digest": None, + "implementation_digest": None, + } + identity["result_origin"] = make_result_origin( + derivation_origin="repair_overlay", + row_assignments=( + ("q2", "external_historical_inference", f"lin_{'a' * 16}"), + ("q1", "external_repair_inference", f"lin_{'b' * 16}"), + ), + ) elif field == "overlay_sha256": + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] identity["overlay"] = { "source_sha256": value, "replacement_digest": "a" * 64, @@ -790,6 +945,13 @@ def _mutate_realization(identity: dict, field: str, value) -> None: "preownership_output_digest": None, "implementation_digest": None, } + identity["result_origin"] = make_result_origin( + derivation_origin="repair_overlay", + row_assignments=( + ("q2", "external_historical_inference", f"lin_{'a' * 16}"), + ("q1", "external_repair_inference", f"lin_{'b' * 16}"), + ), + ) elif field == "prediction_origins": identity["result_origin"] = value else: @@ -809,7 +971,7 @@ def _mutate_realization(identity: dict, field: str, value) -> None: ( "prediction_origins", make_result_origin( - derivation_origin="repair_overlay", + derivation_origin="external_import", row_assignments=( ("q2", "native_inference", f"lin_{'c' * 16}"), ("q1", "external_historical_inference", f"lin_{'b' * 16}"), @@ -860,9 +1022,7 @@ def test_audit_only_realization_fields_do_not_change_identity(): @pytest.mark.parametrize( ("field", "value"), [ - ("encoding", "utf-8-sig"), ("bom_policy", "strip_utf8_bom"), - ("decoding_errors", "replace"), ("delimiter", ";"), ("quote_character", "'"), ("escape_character", "\\"), @@ -870,9 +1030,7 @@ def test_audit_only_realization_fields_do_not_change_identity(): ("line_terminators", ["lf"]), ("mixed_line_terminators", "forbid"), ("final_record_without_terminator", "forbid"), - ("blank_record_policy", "allow"), ("skip_initial_space", True), - ("header", "explicit"), ("strict_syntax", False), ], ) @@ -904,6 +1062,51 @@ def test_every_csv_decoding_and_dialect_field_changes_realization_not_condition( assert first["realization_id"] != changed["realization_id"] +@pytest.mark.parametrize( + ("field", "value"), + [ + ("encoding", "utf-8-sig"), + ("decoding_errors", "replace"), + ("delimiter", 7), + ("blank_record_policy", "allow"), + ("header", "explicit"), + ("strict_syntax", "yes"), + ], +) +def test_realization_refuses_invalid_csv_dialect_values(field, value): + condition = _records().condition + identity = _realization_identity() + identity["parsing_policy"]["dialect"][field] = value + with pytest.raises(ImportIdentityError, match="dialect|parsing"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_refuses_invalid_nested_parsing_policy_records(): + condition = _records().condition + invalid_mutations = ( + lambda policy: policy["mapping"].update({"prediction": 7}), + lambda policy: policy["numeric_columns"].append( + {"source_column": "n", "value_type": "decimal", "null_allowed": False} + ), + lambda policy: policy["option_mapping"].update({"audit": "host-a"}), + ) + for mutate in invalid_mutations: + identity = _realization_identity() + mutate(identity["parsing_policy"]) + with pytest.raises(ImportIdentityError, match="parsing|mapping|numeric|option"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + def test_base_repair_and_transformation_are_distinct_realizations_of_one_condition(): condition = _records().condition identities = [] @@ -913,9 +1116,33 @@ def test_base_repair_and_transformation_are_distinct_realizations_of_one_conditi derivation_origin=derivation, row_assignments=( ("q2", "external_historical_inference", f"lin_{'a' * 16}"), - ("q1", "external_historical_inference", f"lin_{'b' * 16}"), + ( + "q1", + ( + "external_repair_inference" + if derivation == "repair_overlay" + else "external_historical_inference" + ), + f"lin_{'b' * 16}", + ), ), ) + if derivation != "external_import": + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["overlay"] = { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": ( + "1" * 64 if derivation == "offline_transformation" else None + ), + "preownership_output_digest": ( + "2" * 64 if derivation == "offline_transformation" else None + ), + "implementation_digest": ( + "3" * 64 if derivation == "offline_transformation" else None + ), + } identities.append( make_realization( condition_id=condition["condition_id"], @@ -930,6 +1157,39 @@ def test_base_repair_and_transformation_are_distinct_realizations_of_one_conditi assert len({item["realization_id"] for item in identities}) == 3 +@pytest.mark.parametrize("derivation", ["repair_overlay", "offline_transformation"]) +def test_derived_realization_requires_authorization_parent_and_overlay(derivation): + condition = _records().condition + identity = _realization_identity() + identity["result_origin"] = make_result_origin( + derivation_origin=derivation, + row_assignments=(("q1", "external_historical_inference", f"lin_{'a' * 16}"),), + ) + for field in ("authorization_digest", "overlay", "parent"): + candidate = _realization_identity() + candidate["result_origin"] = identity["result_origin"] + candidate["authorization_digest"] = "c" * 64 + candidate["parent_digests"]["realization_digests"] = ["d" * 64] + candidate["overlay"] = { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": "1" * 64, + "preownership_output_digest": "2" * 64, + "implementation_digest": "3" * 64, + } + if field == "parent": + candidate["parent_digests"]["realization_digests"] = [] + else: + candidate[field] = None + with pytest.raises(ImportIdentityError, match="authorization|overlay|parent"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=candidate, + fields={}, + ) + + def test_realization_refuses_post_publication_or_machine_local_identity_fields(): condition = _records().condition for forbidden in ( @@ -1005,6 +1265,69 @@ def test_realization_fields_are_audit_only_and_cannot_repeat_identity_claims(): ) +@pytest.mark.parametrize( + "provenance", + [ + {"machine_id": "host-a"}, + {"cwd": "/home/alice/project"}, + {"host": "machine-a"}, + {"import_date": "2026-07-19"}, + ], +) +def test_realization_source_provenance_is_closed_and_machine_independent(provenance): + condition = _records().condition + identity = _realization_identity() + identity["sources"][0]["provenance"].update(provenance) + with pytest.raises(ImportIdentityError, match="provenance|forbidden"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_requires_runtime_derived_importer_implementation_identity(): + condition = _records().condition + identity = _realization_identity() + identity["importer_implementation"] = { + "schema_version": "choicebench.importer-implementation.v1", + "package": {"name": "choicebench", "version": "0.2.0"}, + "adapter": {"qualified_name": "fake:adapter", "source_digest": "a" * 64}, + "validator": {"qualified_name": "fake:validator", "source_digest": "b" * 64}, + } + with pytest.raises(ImportIdentityError, match="runtime|callable|implementation"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_semantic_condition_refuses_nested_audit_metadata_and_invalid_children(): + original = _records().condition + nested_audit = { + **original["identity"], + "benchmark": {**original["identity"]["benchmark"], "audit": {"machine_id": "x"}}, + } + invalid_child = {**original["identity"], "model_id": "model_not-a-digest"} + for identity in (nested_audit, invalid_child): + with pytest.raises(ImportIdentityError, match="condition|benchmark|model"): + make_semantic_condition(identity=identity, fields={}) + + +def test_expected_dataset_identity_mode_is_closed(): + with pytest.raises(ImportIdentityError, match="identity mode"): + build_import_semantic_identity( + condition=_condition(), + dataset=replace(_dataset(), identity_mode="invented"), + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + def test_lineage_component_is_parent_derived_and_has_no_child_edge(): component = make_lineage_component( operation_type="offline_transformation", @@ -1012,7 +1335,10 @@ def test_lineage_component_is_parent_derived_and_has_no_child_edge(): parent_digests=("2" * 64, "1" * 64), source_digests=("4" * 64, "3" * 64), authorization_digest="5" * 64, - implementation={"qualified_name": "choicebench.transforms:rematch"}, + implementation={ + "qualified_name": "choicebench.transforms:rematch", + "source_digest": "8" * 64, + }, parameters={"matcher": "v1"}, input_digest="6" * 64, preownership_output_digest="7" * 64, @@ -1034,6 +1360,9 @@ def test_lineage_refuses_child_paths_timestamps_and_final_csv_hashes(): "timestamp", "created_at", "final_csv_sha256", + "machine_id", + "child_lineage_id", + "final_result_digest", ): with pytest.raises(ImportIdentityError, match=forbidden): make_lineage_component( @@ -1042,7 +1371,7 @@ def test_lineage_refuses_child_paths_timestamps_and_final_csv_hashes(): parent_digests=(), source_digests=("1" * 64,), authorization_digest=None, - implementation={"qualified_name": "x:y"}, + implementation={"qualified_name": "x:y", "source_digest": "9" * 64}, parameters={forbidden: "bad"}, input_digest="2" * 64, preownership_output_digest="3" * 64, @@ -1050,6 +1379,22 @@ def test_lineage_refuses_child_paths_timestamps_and_final_csv_hashes(): ) +def test_lineage_implementation_requires_a_valid_code_identity(): + with pytest.raises(ImportIdentityError, match="implementation|source"): + make_lineage_component( + operation_type="offline_transformation", + question_id="q1", + parent_digests=("1" * 64,), + source_digests=("2" * 64,), + authorization_digest="3" * 64, + implementation={"attacker": "not-code-identity"}, + parameters={}, + input_digest="4" * 64, + preownership_output_digest="5" * 64, + prediction_origin="external_historical_inference", + ) + + def test_lineage_refuses_duplicate_parent_or_source_edges(): for parents, sources in ( (("1" * 64, "1" * 64), ("2" * 64,)), From 280c5babe5eeb661a776f4ffd1fbedaa88a1bd81 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 02:47:29 +0300 Subject: [PATCH 19/47] fix: derive active importer identity from callables --- src/choicebench/importing/identity.py | 318 +++++++++++++++++++++++--- tests/importing/test_identity.py | 169 +++++++++++++- 2 files changed, 445 insertions(+), 42 deletions(-) diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index 004fedc..036f8d4 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -8,7 +8,13 @@ from typing import Any, Callable, Literal, Mapping, Sequence from choicebench import __version__ -from choicebench.identity import canonicalize, integrity_digest, short_id +from choicebench.datasets import dataset_content_digest, dataset_sample_identities +from choicebench.identity import ( + canonicalize, + integrity_digest, + is_credential_key, + short_id, +) from choicebench.importing.dataset_reference import ExpectedDataset from choicebench.importing.schema import ( CsvDialectSpec, @@ -31,17 +37,6 @@ class ImportIdentityError(ValueError): """Raised when an importer identity is cyclic, unsafe, or inconsistent.""" -class RuntimeImporterImplementation(dict[str, Any]): - """Marker for implementation identity derived from live runtime callables.""" - - def __init__(self, record: Mapping[str, Any]) -> None: - super().__init__(record) - self._runtime_digest = integrity_digest(record) - - def is_unmodified(self) -> bool: - return integrity_digest(dict(self)) == self._runtime_digest - - @dataclass(frozen=True) class ImportSemanticRecords: dataset_artifact: Mapping[str, Any] @@ -151,6 +146,37 @@ class ImportSemanticRecords: "cwd", "child_lineage_id", } +_UNSTABLE_PARAMETER_KEYS = _REALIZATION_KEYS | { + "audit", + "evidence_status", + "scope_disposition", + "executable", + "import_state", + "source_sha256", + "authorization_digest", +} +_UNSTABLE_PARAMETER_TOKENS = { + "artifact", + "audit", + "created", + "cwd", + "directory", + "file", + "host", + "hostname", + "imported", + "location", + "machine", + "output", + "path", + "recorded", + "report", + "result", + "temporary", + "timestamp", + "updated", + "working", +} def _unknown(reason: str | None, field: str) -> dict[str, Any]: @@ -167,6 +193,66 @@ def _nullable_semantic(value: Any, reason: str | None, field: str) -> Any: return _unknown(reason, field) +def _validate_unknown_reason_contract( + values: Mapping[str, Any], reasons: Mapping[str, str], where: str +) -> None: + if not isinstance(reasons, Mapping): + raise ImportIdentityError(f"{where} unknown reasons must be a mapping.") + expected = {field for field, value in values.items() if value is None} + if set(reasons) != expected or not all( + isinstance(reason, str) and reason.strip() for reason in reasons.values() + ): + raise ImportIdentityError( + f"{where} unknown reasons contradict known and missing values." + ) + + +def _validate_direct_declarations( + condition: ImportConditionSpec, + model: ImportModelSpec, + method: ImportMethodSpec, + prompt: ImportPromptSpec, +) -> None: + _validate_unknown_reason_contract( + { + "backend": model.backend, + "provider": model.provider, + "revision": model.revision, + }, + model.unknown_reasons, + "Model declaration", + ) + _validate_unknown_reason_contract( + {"implementation": method.implementation}, + method.unknown_reasons, + "Method declaration", + ) + _validate_unknown_reason_contract( + { + "seed": condition.seed, + "calibration_identity": condition.calibration_identity, + "preflight_identity": condition.preflight_identity, + }, + condition.unknown_reasons, + "Condition declaration", + ) + prompt_values = ( + prompt.template_identity, + prompt.template_digest, + prompt.template_contents, + ) + has_unknown_prompt = any(value is None for value in prompt_values) + if has_unknown_prompt != (prompt.unknown_reason is not None) or ( + prompt.unknown_reason is not None + and (not isinstance(prompt.unknown_reason, str) or not prompt.unknown_reason.strip()) + ): + raise ImportIdentityError( + "Prompt declaration unknown reason contradicts known and missing values." + ) + if prompt.template_digest is not None: + _validate_digest(prompt.template_digest, "prompt.template_digest") + + def _identity_record(prefix: str, payload: Mapping[str, Any]) -> tuple[str, str, dict]: identity = canonicalize(payload) return short_id(prefix, identity), integrity_digest(identity), identity @@ -228,6 +314,52 @@ def _validate_stable_values(value: Any, where: str) -> None: raise ImportIdentityError(f"{where} contains a machine-local path value.") +def _parameter_key_parts(key: str) -> tuple[str, ...]: + snake = re.sub(r"(? None: + if isinstance(value, Mapping): + for key, item in value.items(): + if not isinstance(key, str) or not key: + raise ImportIdentityError(f"{where} keys must be non-empty strings.") + normalized = key.casefold().replace("-", "_") + parts = _parameter_key_parts(key) + if ( + normalized in _UNSTABLE_PARAMETER_KEYS + or is_credential_key(key) + or set(parts) & _UNSTABLE_PARAMETER_TOKENS + or normalized.endswith( + ( + "_id", + "_ids", + "_digest", + "_digests", + "_checksum", + "_hash", + "_sha256", + ) + ) + ): + raise ImportIdentityError( + f"{where} contains forbidden metadata field {key!r}." + ) + _validate_stable_parameters(item, where) + elif isinstance(value, (list, tuple)): + for item in value: + _validate_stable_parameters(item, where) + elif isinstance(value, str) and ( + _is_machine_path(value) or "/" in value or "\\" in value + ): + raise ImportIdentityError(f"{where} contains a path-like metadata value.") + else: + try: + canonicalize(value) + except (TypeError, ValueError) as exc: + raise ImportIdentityError(f"{where} contains unsafe parameter data.") from exc + + def _validate_condition_identity(value: Mapping[str, Any]) -> None: benchmark = _exact_mapping( value["benchmark"], @@ -293,7 +425,7 @@ def _validate_condition_identity(value: Mapping[str, Any]) -> None: if "protocol_settings" in value: if not isinstance(value["protocol_settings"], Mapping): raise ImportIdentityError("Semantic condition protocol_settings must be a mapping.") - _validate_stable_values( + _validate_stable_parameters( value["protocol_settings"], "Semantic condition protocol_settings" ) @@ -323,6 +455,104 @@ def _digest_sequence(value: Any, field: str) -> list[str]: return sorted(result) +def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: + artifact_payload = dataset.artifact_payload + if dataset.identity_mode == "imported_semantic_fallback": + artifact = _exact_mapping( + artifact_payload, + {"schema_version", "benchmark", "split", "content_digest"}, + "Expected dataset fallback artifact", + ) + if artifact["schema_version"] != "choicebench.semantic-dataset.v1": + raise ImportIdentityError("Expected dataset fallback schema is invalid.") + if ( + artifact["benchmark"] != dataset.benchmark_name + or artifact["split"] != dataset.split + ): + raise ImportIdentityError( + "Expected dataset fallback artifact conflicts with its declaration." + ) + else: + artifact = _exact_mapping( + artifact_payload, + {"spec", "content_digest", "source"}, + "Expected dataset native artifact", + ) + spec = _exact_mapping( + artifact["spec"], + { + "benchmark", + "split", + "hf_path", + "hf_subset", + "source_revision", + "normalization_version", + "transforms", + "output_name", + }, + "Expected dataset native artifact spec", + ) + if spec["benchmark"] != dataset.benchmark_name or spec["split"] != dataset.split: + raise ImportIdentityError( + "Expected dataset native artifact conflicts with its declaration." + ) + if not isinstance(artifact["source"], Mapping): + raise ImportIdentityError("Expected dataset native source is invalid.") + if artifact["content_digest"] != dataset_content_digest(dataset.artifact_frame): + raise ImportIdentityError( + "Expected dataset artifact content digest does not own its frame." + ) + + semantics = _exact_mapping( + dataset.selection_semantics, + {"seed", "n_samples", "subject_filter"}, + "Expected dataset selection semantics", + ) + selection = _exact_mapping( + dataset.selection_payload, + { + "artifact_id", + "content_digest", + "sample_identities", + "seed", + "n_samples", + "subject_filter", + }, + "Expected dataset selection", + ) + expected_selection = canonicalize( + { + "artifact_id": dataset.artifact_id, + "content_digest": dataset_content_digest(dataset.frame), + "sample_identities": dataset_sample_identities(dataset.frame), + "seed": semantics["seed"], + "n_samples": semantics["n_samples"], + "subject_filter": sorted(semantics["subject_filter"]), + } + ) + if canonicalize(selection) != expected_selection: + raise ImportIdentityError( + "Expected dataset selection does not own its selected frame." + ) + frame_question_ids = tuple(dataset.frame["question_id"].astype(str)) + if dataset.selected_question_ids != frame_question_ids: + raise ImportIdentityError( + "Expected dataset selected question IDs do not match its frame." + ) + expected_unknowns = { + field + for field, value in ( + ("selection_seed", semantics["seed"]), + ("selection_n_samples", semantics["n_samples"]), + ) + if value is None + } + if set(dataset.selection_unknown_reasons) != expected_unknowns: + raise ImportIdentityError( + "Expected dataset selection unknown reasons are contradictory." + ) + + def _semantic_child_records( *, condition: ImportConditionSpec, @@ -419,6 +649,11 @@ def _semantic_child_records( raise ImportIdentityError( "Validated native model revision conflicts with its declaration." ) + elif model.revision is not None: + raise ImportIdentityError( + "The current native model identity cannot represent a declared " + "revision; use the imported semantic fallback." + ) expected_generation = dict(model.effective_parameters) for name, value in condition.generation_parameters.items(): if name in expected_generation and expected_generation[name] != value: @@ -708,6 +943,7 @@ def build_import_semantic_identity( prompt: ImportPromptSpec, ) -> ImportSemanticRecords: """Build semantic children and a current-shape scientific condition.""" + _validate_direct_declarations(condition, model, method, prompt) if dataset.identity_mode not in { "imported_semantic_fallback", "native_compatibility", @@ -715,6 +951,7 @@ def build_import_semantic_identity( raise ImportIdentityError( f"Unsupported expected dataset identity mode {dataset.identity_mode!r}." ) + _validate_expected_dataset_contract(dataset) references = ( ("dataset_id", condition.dataset_id, dataset.dataset_id), ("model_key", condition.model_key, model.model_key), @@ -820,7 +1057,7 @@ def make_lineage_component( parent_digests: Sequence[str], source_digests: Sequence[str], authorization_digest: str | None, - implementation: Mapping[str, Any], + implementation: Callable[..., Any], parameters: Mapping[str, Any], input_digest: str, preownership_output_digest: str, @@ -845,17 +1082,20 @@ def make_lineage_component( _validate_digest(input_digest, "input_digest") _validate_digest(preownership_output_digest, "preownership_output_digest") try: + implementation_record = implementation_identity(implementation) validated_implementation = validate_implementation_identity_record( - implementation + implementation_record ) - except ImportSpecError as exc: - raise ImportIdentityError(f"Invalid lineage implementation: {exc}") from exc + except (ImportSpecError, OSError, TypeError, ValueError) as exc: + raise ImportIdentityError( + "Invalid lineage implementation; supply an inspectable runtime callable." + ) from exc if "source_digest" not in validated_implementation: raise ImportIdentityError( "Lineage implementation requires an inspectable source digest." ) _validate_stable_values(validated_implementation, "Lineage implementation") - _validate_stable_values(parameters, "Lineage parameters") + _validate_stable_parameters(parameters, "Lineage parameters") payload = canonicalize( { "schema_version": "choicebench.lineage-component.v1", @@ -942,7 +1182,7 @@ def make_result_origin( def importer_implementation_identity( *, adapter: Callable[..., Any], validator: Callable[..., Any] -) -> RuntimeImporterImplementation: +) -> dict[str, Any]: """Bind installed ChoiceBench plus registered adapter and validator code.""" records: dict[str, Mapping[str, Any]] = {} for role, target in (("adapter", adapter), ("validator", validator)): @@ -957,15 +1197,13 @@ def importer_implementation_identity( f"Importer {role} lacks an inspectable source identity." ) records[role] = record - return RuntimeImporterImplementation( - canonicalize( - { - "schema_version": "choicebench.importer-implementation.v1", - "package": {"name": "choicebench", "version": __version__}, - "adapter": records["adapter"], - "validator": records["validator"], - } - ) + return canonicalize( + { + "schema_version": "choicebench.importer-implementation.v1", + "package": {"name": "choicebench", "version": __version__}, + "adapter": records["adapter"], + "validator": records["validator"], + } ) @@ -1010,7 +1248,9 @@ def _validate_result_origin_record(value: Any) -> dict[str, Any]: return rebuilt -def _validate_realization_identity(value: Mapping[str, Any]) -> dict[str, Any]: +def _validate_realization_identity( + value: Mapping[str, Any], *, runtime_importer: Mapping[str, Any] +) -> dict[str, Any]: raw = _exact_mapping(value, _REALIZATION_KEYS, "Realization identity") _validate_digest(raw["import_spec_digest"], "import_spec_digest") @@ -1079,11 +1319,10 @@ def _validate_realization_identity(value: Mapping[str, Any]) -> dict[str, Any]: _validate_digest(digest, f"expected_dataset.{field}") raw_importer = raw["importer_implementation"] - if not isinstance(raw_importer, RuntimeImporterImplementation) or not ( - raw_importer.is_unmodified() - ): + if canonicalize(raw_importer) != canonicalize(runtime_importer): raise ImportIdentityError( - "Realization importer implementation must be derived from runtime callables." + "Realization importer implementation does not match the supplied runtime " + "adapter and validator callables." ) importer = _exact_mapping( raw_importer, @@ -1226,6 +1465,10 @@ def _validate_realization_identity(value: Mapping[str, Any]) -> dict[str, Any]: raise ImportIdentityError( f"{derivation_origin} realization requires a parent realization digest." ) + if not normalized_parents["evidence_digests"]: + raise ImportIdentityError( + f"{derivation_origin} realization requires a parent evidence digest." + ) if normalized_overlay is None: raise ImportIdentityError( f"{derivation_origin} realization requires an overlay identity." @@ -1290,6 +1533,8 @@ def make_realization( condition_digest: str, identity: Mapping[str, Any], fields: Mapping[str, Any], + adapter: Callable[..., Any], + validator: Callable[..., Any], ) -> dict[str, Any]: """Create an immutable realization identity distinct from its condition.""" if not isinstance(condition_id, str) or not re.fullmatch(r"cond_[0-9a-f]{16}", condition_id): @@ -1301,7 +1546,12 @@ def make_realization( ) if not isinstance(identity, Mapping) or not identity: raise ImportIdentityError("Realization identity must be a non-empty mapping.") - validated_identity = _validate_realization_identity(identity) + runtime_importer = importer_implementation_identity( + adapter=adapter, validator=validator + ) + validated_identity = _validate_realization_identity( + identity, runtime_importer=runtime_importer + ) payload = canonicalize( { "schema_version": "choicebench.realization.v1", diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py index c790d59..17685d3 100644 --- a/tests/importing/test_identity.py +++ b/tests/importing/test_identity.py @@ -24,7 +24,7 @@ build_import_semantic_identity, importer_implementation_identity, make_lineage_component, - make_realization, + make_realization as _make_realization, make_result_origin, make_semantic_condition, ) @@ -49,6 +49,14 @@ def _validator(value: str) -> bool: return bool(value) +def _transform(value: str) -> str: + return value.strip() + + +def make_realization(**kwargs): + return _make_realization(adapter=_adapter, validator=_validator, **kwargs) + + def _dataset(): frame = pd.DataFrame( [ @@ -168,6 +176,12 @@ def _condition() -> ImportConditionSpec: def _records(**condition_changes): + if condition_changes.get("seed") is not None and "unknown_reasons" not in condition_changes: + condition_changes["unknown_reasons"] = { + key: value + for key, value in _condition().unknown_reasons.items() + if key != "seed" + } condition = replace(_condition(), **condition_changes) return build_import_semantic_identity( condition=condition, @@ -478,6 +492,7 @@ def test_native_child_claims_must_match_their_semantic_declarations(): method=replace( _method(), implementation={"qualified_name": "historical:SemanticMatching"}, + unknown_reasons={}, native_compatibility_identity=native_method, ), prompt=_prompt(), @@ -632,6 +647,7 @@ def test_native_method_payload_requires_authoritative_implementation_shape(): method=replace( _method(), implementation={"qualified_name": "historical:Matcher"}, + unknown_reasons={}, native_compatibility_identity=native, ), prompt=_prompt(), @@ -921,6 +937,7 @@ def _mutate_realization(identity: dict, field: str, value) -> None: elif field == "authorization_digest": identity["authorization_digest"] = value identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] identity["overlay"] = { "source_sha256": "e" * 64, "replacement_digest": "f" * 64, @@ -938,6 +955,7 @@ def _mutate_realization(identity: dict, field: str, value) -> None: elif field == "overlay_sha256": identity["authorization_digest"] = "c" * 64 identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] identity["overlay"] = { "source_sha256": value, "replacement_digest": "a" * 64, @@ -1130,6 +1148,7 @@ def test_base_repair_and_transformation_are_distinct_realizations_of_one_conditi if derivation != "external_import": identity["authorization_digest"] = "c" * 64 identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] identity["overlay"] = { "source_sha256": "e" * 64, "replacement_digest": "f" * 64, @@ -1165,11 +1184,12 @@ def test_derived_realization_requires_authorization_parent_and_overlay(derivatio derivation_origin=derivation, row_assignments=(("q1", "external_historical_inference", f"lin_{'a' * 16}"),), ) - for field in ("authorization_digest", "overlay", "parent"): + for field in ("authorization_digest", "overlay", "parent", "evidence"): candidate = _realization_identity() candidate["result_origin"] = identity["result_origin"] candidate["authorization_digest"] = "c" * 64 candidate["parent_digests"]["realization_digests"] = ["d" * 64] + candidate["parent_digests"]["evidence_digests"] = ["0" * 64] candidate["overlay"] = { "source_sha256": "e" * 64, "replacement_digest": "f" * 64, @@ -1179,6 +1199,8 @@ def test_derived_realization_requires_authorization_parent_and_overlay(derivatio } if field == "parent": candidate["parent_digests"]["realization_digests"] = [] + elif field == "evidence": + candidate["parent_digests"]["evidence_digests"] = [] else: candidate[field] = None with pytest.raises(ImportIdentityError, match="authorization|overlay|parent"): @@ -1305,6 +1327,19 @@ def test_realization_requires_runtime_derived_importer_implementation_identity() ) +def test_realization_recomputes_implementation_identity_from_supplied_callables(): + condition = _records().condition + with pytest.raises(ImportIdentityError, match="runtime|callable|implementation"): + _make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=_realization_identity(), + fields={}, + adapter=_validator, + validator=_adapter, + ) + + def test_semantic_condition_refuses_nested_audit_metadata_and_invalid_children(): original = _records().condition nested_audit = { @@ -1328,6 +1363,120 @@ def test_expected_dataset_identity_mode_is_closed(): ) +@pytest.mark.parametrize("mutation", ["unexpected_key", "false_content_digest"]) +def test_expected_dataset_payloads_are_bound_to_exact_schema_and_frames(mutation): + dataset = _dataset() + artifact_payload = dict(dataset.artifact_payload) + if mutation == "unexpected_key": + artifact_payload["unexpected"] = "smuggled" + else: + artifact_payload["content_digest"] = "f" * 64 + artifact_digest = integrity_digest(artifact_payload) + artifact_id = short_id("ds", artifact_payload) + selection_payload = { + **dataset.selection_payload, + "artifact_id": artifact_id, + } + forged = replace( + dataset, + artifact_payload=artifact_payload, + artifact_digest=artifact_digest, + artifact_id=artifact_id, + selection_payload=selection_payload, + selection_digest=integrity_digest(selection_payload), + selection_id=short_id("sel", selection_payload), + ) + with pytest.raises(ImportIdentityError, match="dataset|artifact|content"): + build_import_semantic_identity( + condition=_condition(), + dataset=forged, + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +@pytest.mark.parametrize( + "protocol_settings", + [ + {"audit": {"operator": "alice"}}, + {"source_sha256": "a" * 64}, + {"evidence_status": "complete"}, + {"authorization_digest": "b" * 64}, + {"machine": "host-a"}, + {"recorded_at": "2026-07-19"}, + {"working_directory": "work"}, + {"artifact_location": "results/final.csv"}, + ], +) +def test_semantic_protocol_settings_refuse_realization_and_audit_metadata( + protocol_settings, +): + identity = dict(_records().condition["identity"]) + identity["protocol_settings"] = protocol_settings + with pytest.raises(ImportIdentityError, match="protocol|forbidden|metadata"): + make_semantic_condition(identity=identity, fields={}) + + +def test_native_dummy_cannot_discard_a_known_revision(): + payload = {"backend": "dummy", "model_name_or_path": "historical-model"} + native = { + "payload": payload, + "digest": integrity_digest(payload), + "model_id": short_id("model", payload), + } + with pytest.raises(ImportIdentityError, match="revision|fallback"): + build_import_semantic_identity( + condition=replace(_condition(), generation_parameters={}), + dataset=_dataset(), + model=replace( + _model(), + backend="dummy", + revision="publisher-r1", + effective_parameters={}, + unknown_reasons={"provider": "not applicable"}, + native_compatibility_identity=native, + ), + method=_method(), + prompt=_prompt(), + ) + + +@pytest.mark.parametrize("kind", ["model", "prompt", "condition"]) +def test_direct_dataclasses_reject_contradictory_unknown_reasons(kind): + condition = _condition() + model = _model() + prompt = _prompt() + if kind == "model": + model = replace( + model, + backend="dummy", + unknown_reasons={**model.unknown_reasons, "backend": "not recorded"}, + ) + elif kind == "prompt": + prompt = replace( + prompt, + template_identity="known", + template_digest="a" * 64, + template_contents={"direct_mcq": "known"}, + unknown_reason="not recorded", + ) + else: + condition = replace( + condition, + seed=7, + unknown_reasons={**condition.unknown_reasons, "seed": "not recorded"}, + ) + with pytest.raises(ImportIdentityError, match="unknown reason|contradict"): + build_import_semantic_identity( + condition=condition, + dataset=_dataset(), + model=model, + method=_method(), + prompt=prompt, + ) + + def test_lineage_component_is_parent_derived_and_has_no_child_edge(): component = make_lineage_component( operation_type="offline_transformation", @@ -1335,10 +1484,7 @@ def test_lineage_component_is_parent_derived_and_has_no_child_edge(): parent_digests=("2" * 64, "1" * 64), source_digests=("4" * 64, "3" * 64), authorization_digest="5" * 64, - implementation={ - "qualified_name": "choicebench.transforms:rematch", - "source_digest": "8" * 64, - }, + implementation=_transform, parameters={"matcher": "v1"}, input_digest="6" * 64, preownership_output_digest="7" * 64, @@ -1363,6 +1509,13 @@ def test_lineage_refuses_child_paths_timestamps_and_final_csv_hashes(): "machine_id", "child_lineage_id", "final_result_digest", + "lineage_id", + "next_lineage_id", + "child_component_id", + "machine", + "recorded_at", + "working_directory", + "artifact_location", ): with pytest.raises(ImportIdentityError, match=forbidden): make_lineage_component( @@ -1371,7 +1524,7 @@ def test_lineage_refuses_child_paths_timestamps_and_final_csv_hashes(): parent_digests=(), source_digests=("1" * 64,), authorization_digest=None, - implementation={"qualified_name": "x:y", "source_digest": "9" * 64}, + implementation=_transform, parameters={forbidden: "bad"}, input_digest="2" * 64, preownership_output_digest="3" * 64, @@ -1407,7 +1560,7 @@ def test_lineage_refuses_duplicate_parent_or_source_edges(): parent_digests=parents, source_digests=sources, authorization_digest=None, - implementation={"qualified_name": "x:y", "source_digest": "3" * 64}, + implementation=_transform, parameters={}, input_digest="4" * 64, preownership_output_digest="5" * 64, From c095334cbcc4ec09164e22043a30835dcb9cdd97 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 03:07:25 +0300 Subject: [PATCH 20/47] fix: bind importer lineage and dataset ownership --- src/choicebench/importing/identity.py | 205 ++++++++++++++++++-------- src/choicebench/importing/schema.py | 19 ++- tests/importing/test_identity.py | 180 +++++++++++++++++++++- tests/importing/test_schema.py | 9 ++ 4 files changed, 339 insertions(+), 74 deletions(-) diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index 036f8d4..4d803bd 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -3,6 +3,8 @@ from __future__ import annotations from dataclasses import dataclass +import inspect +import math from pathlib import PurePosixPath import re from typing import Any, Callable, Literal, Mapping, Sequence @@ -146,37 +148,7 @@ class ImportSemanticRecords: "cwd", "child_lineage_id", } -_UNSTABLE_PARAMETER_KEYS = _REALIZATION_KEYS | { - "audit", - "evidence_status", - "scope_disposition", - "executable", - "import_state", - "source_sha256", - "authorization_digest", -} -_UNSTABLE_PARAMETER_TOKENS = { - "artifact", - "audit", - "created", - "cwd", - "directory", - "file", - "host", - "hostname", - "imported", - "location", - "machine", - "output", - "path", - "recorded", - "report", - "result", - "temporary", - "timestamp", - "updated", - "working", -} +_PROTOCOL_SETTING_KEYS = {"pride_modal_k_threshold"} def _unknown(reason: str | None, field: str) -> dict[str, Any]: @@ -314,22 +286,14 @@ def _validate_stable_values(value: Any, where: str) -> None: raise ImportIdentityError(f"{where} contains a machine-local path value.") -def _parameter_key_parts(key: str) -> tuple[str, ...]: - snake = re.sub(r"(? None: if isinstance(value, Mapping): for key, item in value.items(): if not isinstance(key, str) or not key: raise ImportIdentityError(f"{where} keys must be non-empty strings.") normalized = key.casefold().replace("-", "_") - parts = _parameter_key_parts(key) if ( - normalized in _UNSTABLE_PARAMETER_KEYS - or is_credential_key(key) - or set(parts) & _UNSTABLE_PARAMETER_TOKENS + is_credential_key(key) or normalized.endswith( ( "_id", @@ -425,6 +389,23 @@ def _validate_condition_identity(value: Mapping[str, Any]) -> None: if "protocol_settings" in value: if not isinstance(value["protocol_settings"], Mapping): raise ImportIdentityError("Semantic condition protocol_settings must be a mapping.") + unexpected_protocol = sorted( + set(value["protocol_settings"]) - _PROTOCOL_SETTING_KEYS + ) + if unexpected_protocol: + raise ImportIdentityError( + "Semantic condition protocol_settings contains unsupported metadata " + f"or protocol fields {unexpected_protocol}." + ) + threshold = value["protocol_settings"].get("pride_modal_k_threshold") + if threshold is not None and ( + isinstance(threshold, bool) + or not isinstance(threshold, (int, float)) + or not math.isfinite(float(threshold)) + ): + raise ImportIdentityError( + "Semantic condition pride_modal_k_threshold must be finite numeric data." + ) _validate_stable_parameters( value["protocol_settings"], "Semantic condition protocol_settings" ) @@ -498,6 +479,9 @@ def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: ) if not isinstance(artifact["source"], Mapping): raise ImportIdentityError("Expected dataset native source is invalid.") + _validate_stable_values( + artifact["source"], "Expected dataset native source" + ) if artifact["content_digest"] != dataset_content_digest(dataset.artifact_frame): raise ImportIdentityError( "Expected dataset artifact content digest does not own its frame." @@ -539,6 +523,24 @@ def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: raise ImportIdentityError( "Expected dataset selected question IDs do not match its frame." ) + artifact_question_ids = tuple(dataset.artifact_frame["question_id"].astype(str)) + if len(artifact_question_ids) != len(set(artifact_question_ids)): + raise ImportIdentityError("Expected dataset artifact has duplicate question IDs.") + artifact_samples = dict( + zip( + artifact_question_ids, + dataset_sample_identities(dataset.artifact_frame), + strict=True, + ) + ) + selected_samples = dataset_sample_identities(dataset.frame) + for question_id, sample_identity in zip( + frame_question_ids, selected_samples, strict=True + ): + if artifact_samples.get(question_id) != sample_identity: + raise ImportIdentityError( + "Expected dataset selected content is not owned by its artifact." + ) expected_unknowns = { field for field, value in ( @@ -547,7 +549,10 @@ def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: ) if value is None } - if set(dataset.selection_unknown_reasons) != expected_unknowns: + if set(dataset.selection_unknown_reasons) != expected_unknowns or not all( + isinstance(reason, str) and reason.strip() + for reason in dataset.selection_unknown_reasons.values() + ): raise ImportIdentityError( "Expected dataset selection unknown reasons are contradictory." ) @@ -1050,6 +1055,43 @@ def _validate_digest(value: Any, field: str, *, optional: bool = False) -> str | return value +def _runtime_callable_record(target: Callable[..., Any], where: str) -> dict[str, Any]: + if inspect.ismethod(target) or not ( + inspect.isfunction(target) or inspect.isclass(target) + ): + raise ImportIdentityError( + f"{where} must be a stateless module function or class, not a bound " + "method or stateful callable instance." + ) + if inspect.isfunction(target) and target.__closure__: + raise ImportIdentityError(f"{where} cannot be a stateful closure.") + try: + record = validate_implementation_identity_record( + implementation_identity(target) + ) + except (ImportSpecError, OSError, TypeError, ValueError) as exc: + raise ImportIdentityError( + f"{where} lacks an inspectable runtime code identity." + ) from exc + if "source_digest" not in record: + raise ImportIdentityError(f"{where} lacks an inspectable source digest.") + return record + + +def _runtime_parameter_names(target: Callable[..., Any]) -> set[str]: + try: + signature = inspect.signature(target) + except (TypeError, ValueError) as exc: + raise ImportIdentityError( + "Lineage runtime implementation has no inspectable parameter schema." + ) from exc + return { + name + for name, parameter in signature.parameters.items() + if parameter.kind is inspect.Parameter.KEYWORD_ONLY + } + + def make_lineage_component( *, operation_type: str, @@ -1057,7 +1099,10 @@ def make_lineage_component( parent_digests: Sequence[str], source_digests: Sequence[str], authorization_digest: str | None, - implementation: Callable[..., Any], + implementation: Callable[..., Any] | Mapping[str, Any], + implementation_mode: Literal["runtime_callable", "declared_external"] = ( + "runtime_callable" + ), parameters: Mapping[str, Any], input_digest: str, preownership_output_digest: str, @@ -1081,20 +1126,50 @@ def make_lineage_component( _validate_digest(authorization_digest, "authorization_digest", optional=True) _validate_digest(input_digest, "input_digest") _validate_digest(preownership_output_digest, "preownership_output_digest") - try: - implementation_record = implementation_identity(implementation) - validated_implementation = validate_implementation_identity_record( - implementation_record + if implementation_mode == "runtime_callable": + if not callable(implementation): + raise ImportIdentityError( + "Lineage runtime implementation must be an inspectable callable." + ) + validated_implementation = _runtime_callable_record( + implementation, "Lineage runtime implementation" ) - except (ImportSpecError, OSError, TypeError, ValueError) as exc: - raise ImportIdentityError( - "Invalid lineage implementation; supply an inspectable runtime callable." - ) from exc - if "source_digest" not in validated_implementation: + allowed_parameters = _runtime_parameter_names(implementation) + unexpected_parameters = sorted(set(parameters) - allowed_parameters) + if unexpected_parameters: + raise ImportIdentityError( + "Lineage parameters contain undeclared field(s) " + f"{unexpected_parameters}." + ) + elif implementation_mode == "declared_external": + try: + validated_implementation = validate_implementation_identity_record( + implementation + ) + except ImportSpecError as exc: + raise ImportIdentityError( + f"Invalid declared external lineage implementation: {exc}" + ) from exc + source_digest = validated_implementation.get("source_digest") + if source_digest is None or source_digest not in source_digests: + raise ImportIdentityError( + "Declared external lineage implementation source digest must be " + "present in source_digests." + ) + if parameters: + raise ImportIdentityError( + "Declared external lineage implementations cannot self-declare " + "unverified parameters." + ) + else: raise ImportIdentityError( - "Lineage implementation requires an inspectable source digest." + f"Invalid lineage implementation_mode {implementation_mode!r}." ) - _validate_stable_values(validated_implementation, "Lineage implementation") + implementation_record = { + "identity_mode": implementation_mode, + "identity": validated_implementation, + } + _validate_stable_values(implementation_record, "Lineage implementation") _validate_stable_parameters(parameters, "Lineage parameters") payload = canonicalize( { @@ -1104,7 +1179,7 @@ def make_lineage_component( "parent_digests": sorted(parent_digests), "source_digests": sorted(source_digests), "authorization_digest": authorization_digest, - "implementation": validated_implementation, + "implementation": implementation_record, "parameters": parameters, "input_digest": input_digest, "preownership_output_digest": preownership_output_digest, @@ -1186,17 +1261,9 @@ def importer_implementation_identity( """Bind installed ChoiceBench plus registered adapter and validator code.""" records: dict[str, Mapping[str, Any]] = {} for role, target in (("adapter", adapter), ("validator", validator)): - try: - record = implementation_identity(target) - except (OSError, TypeError, ValueError) as exc: - raise ImportIdentityError( - f"Importer {role} lacks an inspectable source identity." - ) from exc - if not isinstance(record.get("source_digest"), str): - raise ImportIdentityError( - f"Importer {role} lacks an inspectable source identity." - ) - records[role] = record + records[role] = _runtime_callable_record( + target, f"Importer {role} implementation" + ) return canonicalize( { "schema_version": "choicebench.importer-implementation.v1", @@ -1420,6 +1487,14 @@ def _validate_realization_identity( _validate_digest(evidence[field], f"evidence.{field}") result_origin = _validate_result_origin_record(raw["result_origin"]) + origin_question_ids = [ + row["question_id"] for row in result_origin["row_assignments"] + ] + if integrity_digest(origin_question_ids) != expected["question_set_digest"]: + raise ImportIdentityError( + "Realization result origin question IDs do not match the expected " + "dataset question set." + ) parents = _exact_mapping( raw["parent_digests"], {"realization_digests", "evidence_digests", "result_digests"}, diff --git a/src/choicebench/importing/schema.py b/src/choicebench/importing/schema.py index 4e43508..a685b3d 100644 --- a/src/choicebench/importing/schema.py +++ b/src/choicebench/importing/schema.py @@ -364,6 +364,7 @@ class ImportSpec: "recoverable_question_ids", "result_origin", } +_PROTOCOL_SETTING_KEYS = {"pride_modal_k_threshold"} _RESULT_ORIGIN_KEYS = { "derivation_origin", "default_prediction_origin", @@ -1424,6 +1425,22 @@ def _build_condition(value: Any, index: int) -> ImportConditionSpec: raw["preflight_identity"], f"{where}.preflight_identity" ) unknown_reasons = _string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons") + protocol_settings = _canonical_mapping( + raw["protocol_settings"], f"{where}.protocol_settings" + ) + unsupported_protocol = sorted( + set(protocol_settings) - _PROTOCOL_SETTING_KEYS + ) + if unsupported_protocol: + raise ImportSpecError( + f"{where}.protocol_settings contains unsupported field(s) " + f"{unsupported_protocol}." + ) + if "pride_modal_k_threshold" in protocol_settings: + _strict_number( + protocol_settings["pride_modal_k_threshold"], + f"{where}.protocol_settings.pride_modal_k_threshold", + ) _validate_unknown_reasons( { "seed": seed, @@ -1443,7 +1460,7 @@ def _build_condition(value: Any, index: int) -> ImportConditionSpec: seed=seed, calibration_identity=calibration_identity, preflight_identity=preflight_identity, - protocol_settings=_canonical_mapping(raw["protocol_settings"], f"{where}.protocol_settings"), + protocol_settings=protocol_settings, generation_parameters=_canonical_mapping(raw["generation_parameters"], f"{where}.generation_parameters"), unknown_reasons=unknown_reasons, expected_question_ids=_strings(raw["expected_question_ids"], f"{where}.expected_question_ids", nonempty=True, unique=True), diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py index 17685d3..9aa3284 100644 --- a/tests/importing/test_identity.py +++ b/tests/importing/test_identity.py @@ -15,7 +15,7 @@ ModelConfig, RunConfig, ) -from choicebench.datasets import dataset_content_digest +from choicebench.datasets import dataset_content_digest, dataset_sample_identities from choicebench.identity import integrity_digest, short_id from choicebench.importing.csv_adapter import OpenedSource from choicebench.importing.dataset_reference import build_expected_dataset @@ -53,6 +53,17 @@ def _transform(value: str) -> str: return value.strip() +class _StatefulCallable: + def __init__(self, value: str) -> None: + self.value = value + + def __call__(self, text: str) -> str: + return self.value + text + + def transform(self, text: str) -> str: + return self.value + text + + def make_realization(**kwargs): return _make_realization(adapter=_adapter, validator=_validator, **kwargs) @@ -279,7 +290,7 @@ def test_condition_payload_matches_current_choicebench_keys_and_frozen_prompt_pa ("field", "value"), [ ("seed", 7), - ("protocol_settings", {"threshold": 3}), + ("protocol_settings", {"pride_modal_k_threshold": 3}), ("generation_parameters", {"max_tokens": 65}), ], ) @@ -874,7 +885,7 @@ def _realization_identity() -> dict: ], "expected_dataset": { "snapshot_digest": "3" * 64, - "question_set_digest": "4" * 64, + "question_set_digest": integrity_digest(["q2", "q1"]), "derivation_digest": "5" * 64, }, "importer_implementation": importer_implementation_identity( @@ -1182,7 +1193,18 @@ def test_derived_realization_requires_authorization_parent_and_overlay(derivatio identity = _realization_identity() identity["result_origin"] = make_result_origin( derivation_origin=derivation, - row_assignments=(("q1", "external_historical_inference", f"lin_{'a' * 16}"),), + row_assignments=( + ("q2", "external_historical_inference", f"lin_{'a' * 16}"), + ( + "q1", + ( + "external_repair_inference" + if derivation == "repair_overlay" + else "external_historical_inference" + ), + f"lin_{'b' * 16}", + ), + ), ) for field in ("authorization_digest", "overlay", "parent", "evidence"): candidate = _realization_identity() @@ -1340,6 +1362,15 @@ def test_realization_recomputes_implementation_identity_from_supplied_callables( ) +def test_realization_refuses_stateful_callable_instances_with_shared_source(): + first = _StatefulCallable("first") + with pytest.raises(ImportIdentityError, match="stateless|callable|implementation"): + importer_implementation_identity( + adapter=first, + validator=_validator, + ) + + def test_semantic_condition_refuses_nested_audit_metadata_and_invalid_children(): original = _records().condition nested_audit = { @@ -1363,19 +1394,26 @@ def test_expected_dataset_identity_mode_is_closed(): ) -@pytest.mark.parametrize("mutation", ["unexpected_key", "false_content_digest"]) +@pytest.mark.parametrize( + "mutation", ["unexpected_key", "false_content_digest", "forged_selected_row"] +) def test_expected_dataset_payloads_are_bound_to_exact_schema_and_frames(mutation): dataset = _dataset() + selected_frame = dataset.frame.copy() artifact_payload = dict(dataset.artifact_payload) if mutation == "unexpected_key": artifact_payload["unexpected"] = "smuggled" - else: + elif mutation == "false_content_digest": artifact_payload["content_digest"] = "f" * 64 + else: + selected_frame.loc[selected_frame.index[0], "question_text"] = "forged" artifact_digest = integrity_digest(artifact_payload) artifact_id = short_id("ds", artifact_payload) selection_payload = { **dataset.selection_payload, "artifact_id": artifact_id, + "content_digest": dataset_content_digest(selected_frame), + "sample_identities": dataset_sample_identities(selected_frame), } forged = replace( dataset, @@ -1385,6 +1423,7 @@ def test_expected_dataset_payloads_are_bound_to_exact_schema_and_frames(mutation selection_payload=selection_payload, selection_digest=integrity_digest(selection_payload), selection_id=short_id("sel", selection_payload), + frame=selected_frame, ) with pytest.raises(ImportIdentityError, match="dataset|artifact|content"): build_import_semantic_identity( @@ -1396,6 +1435,18 @@ def test_expected_dataset_payloads_are_bound_to_exact_schema_and_frames(mutation ) +def test_expected_dataset_unknown_reasons_must_be_nonempty(): + dataset = replace(_dataset(), selection_unknown_reasons={"selection_seed": ""}) + with pytest.raises(ImportIdentityError, match="unknown reason"): + build_import_semantic_identity( + condition=_condition(), + dataset=dataset, + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + @pytest.mark.parametrize( "protocol_settings", [ @@ -1407,6 +1458,11 @@ def test_expected_dataset_payloads_are_bound_to_exact_schema_and_frames(mutation {"recorded_at": "2026-07-19"}, {"working_directory": "work"}, {"artifact_location": "results/final.csv"}, + {"node_name": "node-a"}, + {"executed_at": "now"}, + {"run_date": "2026-07-19"}, + {"platform": "linux"}, + {"operator": "alice"}, ], ) def test_semantic_protocol_settings_refuse_realization_and_audit_metadata( @@ -1442,6 +1498,44 @@ def test_native_dummy_cannot_discard_a_known_revision(): ) +def test_native_dataset_source_refuses_audit_and_machine_metadata(): + dataset = _dataset() + artifact_payload = { + "spec": { + "benchmark": dataset.benchmark_name, + "split": dataset.split, + "hf_path": None, + "hf_subset": None, + "source_revision": None, + "normalization_version": "v1", + "transforms": [], + "output_name": None, + }, + "content_digest": dataset_content_digest(dataset.artifact_frame), + "source": {"audit_path": "/machine/a/input.csv", "timestamp": "now"}, + } + artifact_id = short_id("ds", artifact_payload) + selection_payload = {**dataset.selection_payload, "artifact_id": artifact_id} + forged = replace( + dataset, + identity_mode="native_compatibility", + artifact_payload=artifact_payload, + artifact_digest=integrity_digest(artifact_payload), + artifact_id=artifact_id, + selection_payload=selection_payload, + selection_digest=integrity_digest(selection_payload), + selection_id=short_id("sel", selection_payload), + ) + with pytest.raises(ImportIdentityError, match="source|audit|path|timestamp"): + build_import_semantic_identity( + condition=_condition(), + dataset=forged, + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + @pytest.mark.parametrize("kind", ["model", "prompt", "condition"]) def test_direct_dataclasses_reject_contradictory_unknown_reasons(kind): condition = _condition() @@ -1485,7 +1579,7 @@ def test_lineage_component_is_parent_derived_and_has_no_child_edge(): source_digests=("4" * 64, "3" * 64), authorization_digest="5" * 64, implementation=_transform, - parameters={"matcher": "v1"}, + parameters={}, input_digest="6" * 64, preownership_output_digest="7" * 64, prediction_origin="external_historical_inference", @@ -1516,6 +1610,11 @@ def test_lineage_refuses_child_paths_timestamps_and_final_csv_hashes(): "recorded_at", "working_directory", "artifact_location", + "node_name", + "executed_at", + "run_date", + "platform", + "operator", ): with pytest.raises(ImportIdentityError, match=forbidden): make_lineage_component( @@ -1548,6 +1647,49 @@ def test_lineage_implementation_requires_a_valid_code_identity(): ) +def test_lineage_refuses_bound_methods_with_unbound_instance_state(): + with pytest.raises(ImportIdentityError, match="stateless|bound|callable"): + make_lineage_component( + operation_type="offline_transformation", + question_id="q1", + parent_digests=("1" * 64,), + source_digests=("2" * 64,), + authorization_digest="3" * 64, + implementation=_StatefulCallable("state").transform, + parameters={}, + input_digest="4" * 64, + preownership_output_digest="5" * 64, + prediction_origin="external_historical_inference", + ) + + +def test_lineage_supports_nonexecuted_declared_external_implementation_identity(): + source_digest = "2" * 64 + component = make_lineage_component( + operation_type="precomputed_external_transformation", + question_id="q1", + parent_digests=("1" * 64,), + source_digests=(source_digest,), + authorization_digest="3" * 64, + implementation={ + "qualified_name": "historical.matcher:rematch", + "source_file": "matcher.py", + "source_digest": source_digest, + }, + implementation_mode="declared_external", + parameters={}, + input_digest="4" * 64, + preownership_output_digest="5" * 64, + prediction_origin="external_historical_inference", + ) + assert component["identity"]["implementation"]["identity_mode"] == ( + "declared_external" + ) + assert component["identity"]["implementation"]["identity"]["source_digest"] == ( + source_digest + ) + + def test_lineage_refuses_duplicate_parent_or_source_edges(): for parents, sources in ( (("1" * 64, "1" * 64), ("2" * 64,)), @@ -1603,6 +1745,28 @@ def test_result_origin_records_constituents_counts_and_ordered_per_row_mapping() assert origin["origin_id"] == short_id("origin", identity) +def test_realization_result_origin_question_ids_match_expected_question_set(): + condition = _records().condition + identity = _realization_identity() + identity["result_origin"] = make_result_origin( + derivation_origin="external_import", + row_assignments=( + ( + "unrelated-question", + "external_historical_inference", + f"lin_{'a' * 16}", + ), + ), + ) + with pytest.raises(ImportIdentityError, match="question set|question IDs"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + def test_result_origin_rejects_bare_mixed_duplicate_rows_and_invalid_origins(): for assignments in ( (("q1", "mixed", f"lin_{'a' * 16}"),), @@ -1641,5 +1805,5 @@ def test_importer_implementation_identity_binds_package_adapter_and_validator(): def test_importer_implementation_identity_refuses_uninspectable_callables(): - with pytest.raises(ImportIdentityError, match="source identity"): + with pytest.raises(ImportIdentityError, match="source identity|stateless|callable"): importer_implementation_identity(adapter=len, validator=_validator) diff --git a/tests/importing/test_schema.py b/tests/importing/test_schema.py index b62e7a3..ce109bc 100644 --- a/tests/importing/test_schema.py +++ b/tests/importing/test_schema.py @@ -223,6 +223,15 @@ def _load(tmp_path: Path, raw: object): return load_import_spec(path) +def test_protocol_settings_use_the_closed_current_choicebench_schema( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + raw["conditions"][0]["protocol_settings"] = {"operator": "alice"} + with pytest.raises(ImportSpecError, match="protocol_settings|unsupported"): + _load(tmp_path, raw) + + @pytest.fixture def minimal_spec(tmp_path: Path, minimal_raw: dict): return _load(tmp_path, minimal_raw) From 9064d85e07ec716556cef7357df70ae095b04ae1 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 03:22:55 +0300 Subject: [PATCH 21/47] fix: preserve partial provenance coverage --- src/choicebench/importing/identity.py | 134 +++++++++++++++++++++++--- tests/importing/test_identity.py | 61 +++++++++++- 2 files changed, 181 insertions(+), 14 deletions(-) diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index 4d803bd..419adc4 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -149,6 +149,18 @@ class ImportSemanticRecords: "child_lineage_id", } _PROTOCOL_SETTING_KEYS = {"pride_modal_k_threshold"} +_NATIVE_DATASET_SOURCE_KEYS = { + "generator", + "seed", + "hf_path", + "hf_subset", + "split", + "requested_revision", + "resolved_revision", + "hf_fingerprint", + "hf_dataset_info", + "revision", +} def _unknown(reason: str | None, field: str) -> dict[str, Any]: @@ -479,9 +491,49 @@ def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: ) if not isinstance(artifact["source"], Mapping): raise ImportIdentityError("Expected dataset native source is invalid.") - _validate_stable_values( - artifact["source"], "Expected dataset native source" - ) + source = artifact["source"] + unexpected_source = sorted(set(source) - _NATIVE_DATASET_SOURCE_KEYS) + if unexpected_source: + raise ImportIdentityError( + "Expected dataset native source contains unsupported audit or " + f"source fields {unexpected_source}." + ) + if "generator" in source and ( + not isinstance(source["generator"], str) or not source["generator"] + ): + raise ImportIdentityError("Expected dataset native source generator is invalid.") + if "seed" in source and ( + not isinstance(source["seed"], int) or isinstance(source["seed"], bool) + ): + raise ImportIdentityError("Expected dataset native source seed is invalid.") + for field in ( + "hf_path", + "hf_subset", + "split", + "requested_revision", + "resolved_revision", + "hf_fingerprint", + "revision", + ): + if field in source and source[field] is not None and not isinstance( + source[field], str + ): + raise ImportIdentityError( + f"Expected dataset native source {field} is invalid." + ) + if "hf_dataset_info" in source: + info = _exact_mapping( + source["hf_dataset_info"], + {"builder_name", "config_name", "version"}, + "Expected dataset native source hf_dataset_info", + ) + if any( + item is not None and not isinstance(item, str) + for item in info.values() + ): + raise ImportIdentityError( + "Expected dataset native source hf_dataset_info is invalid." + ) if artifact["content_digest"] != dataset_content_digest(dataset.artifact_frame): raise ImportIdentityError( "Expected dataset artifact content digest does not own its frame." @@ -1078,18 +1130,27 @@ def _runtime_callable_record(target: Callable[..., Any], where: str) -> dict[str return record -def _runtime_parameter_names(target: Callable[..., Any]) -> set[str]: +def _runtime_parameter_schema( + target: Callable[..., Any], +) -> tuple[set[str], set[str]]: try: signature = inspect.signature(target) except (TypeError, ValueError) as exc: raise ImportIdentityError( "Lineage runtime implementation has no inspectable parameter schema." ) from exc - return { + allowed = { + name + for name, parameter in signature.parameters.items() + if parameter.kind is inspect.Parameter.KEYWORD_ONLY + } + required = { name for name, parameter in signature.parameters.items() if parameter.kind is inspect.Parameter.KEYWORD_ONLY + and parameter.default is inspect.Parameter.empty } + return allowed, required def make_lineage_component( @@ -1134,13 +1195,21 @@ def make_lineage_component( validated_implementation = _runtime_callable_record( implementation, "Lineage runtime implementation" ) - allowed_parameters = _runtime_parameter_names(implementation) + allowed_parameters, required_parameters = _runtime_parameter_schema( + implementation + ) unexpected_parameters = sorted(set(parameters) - allowed_parameters) if unexpected_parameters: raise ImportIdentityError( "Lineage parameters contain undeclared field(s) " f"{unexpected_parameters}." ) + missing_parameters = sorted(required_parameters - set(parameters)) + if missing_parameters: + raise ImportIdentityError( + "Lineage parameters are missing required field(s) " + f"{missing_parameters}." + ) elif implementation_mode == "declared_external": try: validated_implementation = validate_implementation_identity_record( @@ -1203,8 +1272,6 @@ def make_result_origin( """Record constituent and ordered per-row prediction origins.""" if derivation_origin not in _DERIVATION_ORIGINS: raise ImportIdentityError(f"Invalid derivation origin {derivation_origin!r}.") - if not row_assignments: - raise ImportIdentityError("Result origin requires row assignments.") rows: list[dict[str, str]] = [] seen: set[str] = set() counts: dict[str, int] = {} @@ -1379,11 +1446,35 @@ def _validate_realization_identity( expected = _exact_mapping( raw["expected_dataset"], - {"snapshot_digest", "question_set_digest", "derivation_digest"}, + { + "snapshot_digest", + "question_set_digest", + "derivation_digest", + "question_ids", + }, "Realization expected_dataset", ) - for field, digest in expected.items(): + for field in ("snapshot_digest", "question_set_digest", "derivation_digest"): + digest = expected[field] _validate_digest(digest, f"expected_dataset.{field}") + expected_question_ids = expected["question_ids"] + if not isinstance(expected_question_ids, (list, tuple)) or not all( + isinstance(question_id, str) and question_id + for question_id in expected_question_ids + ): + raise ImportIdentityError( + "Realization expected dataset question_ids must be a non-empty string list." + ) + if not expected_question_ids or len(expected_question_ids) != len( + set(expected_question_ids) + ): + raise ImportIdentityError( + "Realization expected dataset question_ids must be non-empty and unique." + ) + if integrity_digest(list(expected_question_ids)) != expected["question_set_digest"]: + raise ImportIdentityError( + "Realization expected dataset question_ids do not match question_set_digest." + ) raw_importer = raw["importer_implementation"] if canonicalize(raw_importer) != canonicalize(runtime_importer): @@ -1490,10 +1581,27 @@ def _validate_realization_identity( origin_question_ids = [ row["question_id"] for row in result_origin["row_assignments"] ] - if integrity_digest(origin_question_ids) != expected["question_set_digest"]: + evidence_status = evidence["evidence_status"] + origin_question_id_set = set(origin_question_ids) + unexpected_origin_ids = sorted( + origin_question_id_set - set(expected_question_ids) + ) + ordered_subset = [ + question_id + for question_id in expected_question_ids + if question_id in origin_question_id_set + ] + if unexpected_origin_ids or ordered_subset != origin_question_ids: + raise ImportIdentityError( + "Realization result origin question IDs are not an ordered subset of " + "the expected dataset question set." + ) + if evidence_status in {"complete", "qualified"} and origin_question_ids != list( + expected_question_ids + ): raise ImportIdentityError( - "Realization result origin question IDs do not match the expected " - "dataset question set." + "Complete or qualified realization result origin question IDs do not " + "match the full expected dataset question set." ) parents = _exact_mapping( raw["parent_digests"], diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py index 9aa3284..2fab1b9 100644 --- a/tests/importing/test_identity.py +++ b/tests/importing/test_identity.py @@ -53,6 +53,10 @@ def _transform(value: str) -> str: return value.strip() +def _transform_required(value: str, *, threshold: float) -> str: + return value if threshold >= 0 else "" + + class _StatefulCallable: def __init__(self, value: str) -> None: self.value = value @@ -887,6 +891,7 @@ def _realization_identity() -> dict: "snapshot_digest": "3" * 64, "question_set_digest": integrity_digest(["q2", "q1"]), "derivation_digest": "5" * 64, + "question_ids": ["q2", "q1"], }, "importer_implementation": importer_implementation_identity( adapter=_adapter, validator=_validator @@ -1512,7 +1517,7 @@ def test_native_dataset_source_refuses_audit_and_machine_metadata(): "output_name": None, }, "content_digest": dataset_content_digest(dataset.artifact_frame), - "source": {"audit_path": "/machine/a/input.csv", "timestamp": "now"}, + "source": {"audit": {"operator": "alice"}, "node": "machine-a"}, } artifact_id = short_id("ds", artifact_payload) selection_payload = {**dataset.selection_payload, "artifact_id": artifact_id} @@ -1536,6 +1541,44 @@ def test_native_dataset_source_refuses_audit_and_machine_metadata(): ) +def test_partial_realization_origin_may_be_an_ordered_expected_subset(): + condition = _records().condition + identity = _realization_identity() + identity["evidence"]["evidence_status"] = "partial" + identity["result_origin"] = make_result_origin( + derivation_origin="external_import", + row_assignments=( + ("q2", "external_historical_inference", f"lin_{'a' * 16}"), + ), + ) + realization = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + assignments = realization["identity"]["realization"]["result_origin"][ + "row_assignments" + ] + assert assignments[0]["question_id"] == "q2" + + +def test_malformed_realization_origin_may_have_no_evaluable_rows(): + condition = _records().condition + identity = _realization_identity() + identity["evidence"]["evidence_status"] = "malformed" + identity["result_origin"] = make_result_origin( + derivation_origin="external_import", row_assignments=() + ) + realization = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + assert realization["identity"]["realization"]["result_origin"]["row_assignments"] == [] + + @pytest.mark.parametrize("kind", ["model", "prompt", "condition"]) def test_direct_dataclasses_reject_contradictory_unknown_reasons(kind): condition = _condition() @@ -1663,6 +1706,22 @@ def test_lineage_refuses_bound_methods_with_unbound_instance_state(): ) +def test_lineage_requires_mandatory_runtime_keyword_parameters(): + with pytest.raises(ImportIdentityError, match="required|threshold"): + make_lineage_component( + operation_type="offline_transformation", + question_id="q1", + parent_digests=("1" * 64,), + source_digests=("2" * 64,), + authorization_digest="3" * 64, + implementation=_transform_required, + parameters={}, + input_digest="4" * 64, + preownership_output_digest="5" * 64, + prediction_origin="external_historical_inference", + ) + + def test_lineage_supports_nonexecuted_declared_external_implementation_identity(): source_digest = "2" * 64 component = make_lineage_component( From 268422af0e9efdbcb751fab90a9b7fb4f67a276a Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 03:59:44 +0300 Subject: [PATCH 22/47] fix: close importer identity dependency graph --- src/choicebench/importing/identity.py | 387 +++++++++++++++++++++++++- tests/importing/test_identity.py | 382 ++++++++++++++++++++++--- 2 files changed, 722 insertions(+), 47 deletions(-) diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index 419adc4..0427601 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -77,6 +77,7 @@ class ImportSemanticRecords: "parsing_policy", "validation", "evidence", + "lineage_components", "result_origin", "parent_digests", "authorization_digest", @@ -145,6 +146,13 @@ class ImportSemanticRecords: "temporary_path", "timestamp", "machine_id", + "machine", + "node", + "operator", + "platform", + "recorded_at", + "executed_at", + "run_date", "cwd", "child_lineage_id", } @@ -161,6 +169,23 @@ class ImportSemanticRecords: "hf_dataset_info", "revision", } +_DATASET_DERIVATION_KEYS = { + "schema_version", + "dataset_id", + "reference_kind", + "trust_label", + "source_chain", + "selection_source_id", + "columns", + "revision", + "fingerprint", + "declared_derivation", + "selected_question_ids", + "selection_semantics", + "selection_unknown_reasons", + "identity_mode", + "limitations", +} def _unknown(reason: str | None, field: str) -> dict[str, Any]: @@ -197,6 +222,29 @@ def _validate_direct_declarations( method: ImportMethodSpec, prompt: ImportPromptSpec, ) -> None: + _validate_stable_parameters( + model.effective_parameters, "Model effective parameters" + ) + _validate_stable_parameters( + condition.generation_parameters, "Condition generation parameters" + ) + _validate_stable_parameters( + method.effective_parameters, "Method effective parameters" + ) + if method.implementation is not None: + try: + validated_implementation = validate_implementation_identity_record( + method.implementation + ) + except ImportSpecError as exc: + raise ImportIdentityError(f"Invalid method implementation: {exc}") from exc + if canonicalize(validated_implementation) != canonicalize( + method.implementation + ): + raise ImportIdentityError( + "Method implementation does not match the closed code identity schema." + ) + _validate_stable_values(validated_implementation, "Method implementation") _validate_unknown_reason_contract( { "backend": model.backend, @@ -299,6 +347,7 @@ def _validate_stable_values(value: Any, where: str) -> None: def _validate_stable_parameters(value: Any, where: str) -> None: + _reject_forbidden(value, where) if isinstance(value, Mapping): for key, item in value.items(): if not isinstance(key, str) or not key: @@ -521,6 +570,11 @@ def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: raise ImportIdentityError( f"Expected dataset native source {field} is invalid." ) + if isinstance(source.get(field), str) and _is_machine_path(source[field]): + raise ImportIdentityError( + f"Expected dataset native source {field} contains a " + "machine-local path." + ) if "hf_dataset_info" in source: info = _exact_mapping( source["hf_dataset_info"], @@ -609,6 +663,103 @@ def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: "Expected dataset selection unknown reasons are contradictory." ) + if dataset.reference_kind not in { + "independent_input_snapshot", + "profile_derived_reference_snapshot", + }: + raise ImportIdentityError("Expected dataset reference kind is invalid.") + if not isinstance(dataset.trust_label, str) or not dataset.trust_label.strip(): + raise ImportIdentityError("Expected dataset trust label is invalid.") + if dataset.question_set_digest != integrity_digest(list(frame_question_ids)): + raise ImportIdentityError( + "Expected dataset question-set digest does not own its selected IDs." + ) + derivation = _exact_mapping( + dataset.derivation, + _DATASET_DERIVATION_KEYS, + "Expected dataset derivation", + ) + expected_derivation_fields = { + "schema_version": "choicebench.dataset-reference.v1", + "dataset_id": dataset.dataset_id, + "reference_kind": dataset.reference_kind, + "trust_label": dataset.trust_label, + "selected_question_ids": list(frame_question_ids), + "selection_semantics": semantics, + "selection_unknown_reasons": dict(dataset.selection_unknown_reasons), + "identity_mode": dataset.identity_mode, + "limitations": list(dataset.limitations), + } + for field, expected_value in expected_derivation_fields.items(): + if canonicalize(derivation[field]) != canonicalize(expected_value): + raise ImportIdentityError( + f"Expected dataset derivation {field} conflicts with its snapshot." + ) + source_chain = derivation["source_chain"] + if not isinstance(source_chain, (list, tuple)) or not source_chain: + raise ImportIdentityError("Expected dataset derivation source chain is invalid.") + seen_source_ids: set[str] = set() + for index, item in enumerate(source_chain): + source = _exact_mapping( + item, + {"source_id", "logical_path", "sha256"}, + f"Expected dataset derivation source_chain[{index}]", + ) + if ( + not isinstance(source["source_id"], str) + or not source["source_id"] + or source["source_id"] in seen_source_ids + ): + raise ImportIdentityError( + "Expected dataset derivation source IDs must be non-empty and unique." + ) + logical_path = source["logical_path"] + if not isinstance(logical_path, str): + raise ImportIdentityError( + "Expected dataset derivation logical path is invalid." + ) + pure_path = PurePosixPath(logical_path) + if pure_path.is_absolute() or ".." in pure_path.parts or not pure_path.parts: + raise ImportIdentityError( + "Expected dataset derivation logical path is unsafe." + ) + _validate_digest( + source["sha256"], + f"Expected dataset derivation source_chain[{index}].sha256", + ) + seen_source_ids.add(source["source_id"]) + if ( + not isinstance(derivation["selection_source_id"], str) + or derivation["selection_source_id"] not in seen_source_ids + ): + raise ImportIdentityError( + "Expected dataset derivation selection source is not in its source chain." + ) + if not isinstance(derivation["columns"], Mapping) or not all( + isinstance(key, str) and isinstance(value, str) + for key, value in derivation["columns"].items() + ): + raise ImportIdentityError("Expected dataset derivation columns are invalid.") + _validate_stable_values( + derivation["declared_derivation"], + "Expected dataset declared derivation", + ) + expected_derivation_digest = integrity_digest(derivation) + if dataset.derivation_digest != expected_derivation_digest: + raise ImportIdentityError("Expected dataset derivation digest is inconsistent.") + expected_snapshot_digest = integrity_digest( + { + "artifact_digest": dataset.artifact_digest, + "selection_digest": dataset.selection_digest, + "question_set_digest": dataset.question_set_digest, + "derivation_digest": expected_derivation_digest, + "reference_kind": dataset.reference_kind, + "trust_label": dataset.trust_label, + } + ) + if dataset.snapshot_digest != expected_snapshot_digest: + raise ImportIdentityError("Expected dataset snapshot digest is inconsistent.") + def _semantic_child_records( *, @@ -852,7 +1003,8 @@ def _semantic_child_records( }: raise ImportIdentityError("Invalid validated native prompt identity fields.") prompt_payload = canonicalize( - {"version": native["version"], "files": native["files"]} + {"version": native["version"], "files": native["files"]}, + redact_secrets=False, ) if any( value is None @@ -896,7 +1048,8 @@ def _semantic_child_records( ) if ( declared_contents is not None - and canonicalize(declared_contents) != canonicalize(native_contents) + and canonicalize(declared_contents, redact_secrets=False) + != canonicalize(native_contents, redact_secrets=False) ): raise ImportIdentityError( "Validated native prompt contents conflict with their declaration." @@ -1117,6 +1270,11 @@ def _runtime_callable_record(target: Callable[..., Any], where: str) -> dict[str ) if inspect.isfunction(target) and target.__closure__: raise ImportIdentityError(f"{where} cannot be a stateful closure.") + qualified_name = getattr(target, "__qualname__", "") + if "" in qualified_name or "" in qualified_name: + raise ImportIdentityError( + f"{where} must be a uniquely addressable named module callable." + ) try: record = validate_implementation_identity_record( implementation_identity(target) @@ -1322,6 +1480,126 @@ def make_result_origin( } +def _validate_lineage_component_record(value: Any) -> dict[str, Any]: + record = _exact_mapping( + value, + {"lineage_id", "lineage_digest", "identity"}, + "Realization lineage component", + ) + identity = _exact_mapping( + record["identity"], + { + "schema_version", + "operation_type", + "question_id", + "parent_digests", + "source_digests", + "authorization_digest", + "implementation", + "parameters", + "input_digest", + "preownership_output_digest", + "prediction_origin", + }, + "Realization lineage component identity", + ) + if identity["schema_version"] != "choicebench.lineage-component.v1": + raise ImportIdentityError("Realization lineage component schema is invalid.") + for field in ("operation_type", "question_id"): + if not isinstance(identity[field], str) or not identity[field]: + raise ImportIdentityError( + f"Realization lineage component {field} must be non-empty." + ) + normalized_parents = _digest_sequence( + identity["parent_digests"], "lineage.parent_digests" + ) + normalized_sources = _digest_sequence( + identity["source_digests"], "lineage.source_digests" + ) + authorization_digest = _validate_digest( + identity["authorization_digest"], + "lineage.authorization_digest", + optional=True, + ) + implementation = _exact_mapping( + identity["implementation"], + {"identity_mode", "identity"}, + "Realization lineage implementation", + ) + if implementation["identity_mode"] not in { + "runtime_callable", + "declared_external", + }: + raise ImportIdentityError( + "Realization lineage implementation identity mode is invalid." + ) + try: + implementation_identity_record = validate_implementation_identity_record( + implementation["identity"] + ) + except ImportSpecError as exc: + raise ImportIdentityError( + f"Realization lineage implementation is invalid: {exc}" + ) from exc + source_digest = implementation_identity_record.get("source_digest") + if source_digest is None: + raise ImportIdentityError( + "Realization lineage implementation lacks a source digest." + ) + if ( + implementation["identity_mode"] == "declared_external" + and source_digest not in normalized_sources + ): + raise ImportIdentityError( + "Declared external lineage implementation is not owned by its source edge." + ) + parameters = identity["parameters"] + if not isinstance(parameters, Mapping): + raise ImportIdentityError("Realization lineage parameters must be a mapping.") + _validate_stable_parameters(parameters, "Realization lineage parameters") + input_digest = _validate_digest(identity["input_digest"], "lineage.input_digest") + output_digest = _validate_digest( + identity["preownership_output_digest"], + "lineage.preownership_output_digest", + ) + prediction_origin = identity["prediction_origin"] + if prediction_origin not in _PREDICTION_ORIGINS: + raise ImportIdentityError("Realization lineage prediction origin is invalid.") + normalized_identity = canonicalize( + { + "schema_version": "choicebench.lineage-component.v1", + "operation_type": identity["operation_type"], + "question_id": identity["question_id"], + "parent_digests": normalized_parents, + "source_digests": normalized_sources, + "authorization_digest": authorization_digest, + "implementation": { + "identity_mode": implementation["identity_mode"], + "identity": implementation_identity_record, + }, + "parameters": parameters, + "input_digest": input_digest, + "preownership_output_digest": output_digest, + "prediction_origin": prediction_origin, + } + ) + expected_digest = integrity_digest(normalized_identity) + expected_id = short_id("lin", normalized_identity) + if ( + record["lineage_digest"] != expected_digest + or record["lineage_id"] != expected_id + or canonicalize(record["identity"]) != normalized_identity + ): + raise ImportIdentityError( + "Realization lineage component identity or digest is inconsistent." + ) + return { + "lineage_id": expected_id, + "lineage_digest": expected_digest, + "identity": normalized_identity, + } + + def importer_implementation_identity( *, adapter: Callable[..., Any], validator: Callable[..., Any] ) -> dict[str, Any]: @@ -1383,7 +1661,10 @@ def _validate_result_origin_record(value: Any) -> dict[str, Any]: def _validate_realization_identity( - value: Mapping[str, Any], *, runtime_importer: Mapping[str, Any] + value: Mapping[str, Any], + *, + runtime_importer: Mapping[str, Any], + expected_dataset: ExpectedDataset, ) -> dict[str, Any]: raw = _exact_mapping(value, _REALIZATION_KEYS, "Realization identity") _validate_digest(raw["import_spec_digest"], "import_spec_digest") @@ -1475,6 +1756,19 @@ def _validate_realization_identity( raise ImportIdentityError( "Realization expected dataset question_ids do not match question_set_digest." ) + _validate_expected_dataset_contract(expected_dataset) + authoritative_expected = canonicalize( + { + "snapshot_digest": expected_dataset.snapshot_digest, + "question_set_digest": expected_dataset.question_set_digest, + "derivation_digest": expected_dataset.derivation_digest, + "question_ids": list(expected_dataset.selected_question_ids), + } + ) + if canonicalize(expected) != authoritative_expected: + raise ImportIdentityError( + "Realization expected dataset does not match the validated dataset snapshot." + ) raw_importer = raw["importer_implementation"] if canonicalize(raw_importer) != canonicalize(runtime_importer): @@ -1577,6 +1871,27 @@ def _validate_realization_identity( for field in ("qualification_digest", "limitation_digest", "defect_digest"): _validate_digest(evidence[field], f"evidence.{field}") + raw_lineage_components = raw["lineage_components"] + if not isinstance(raw_lineage_components, (list, tuple)): + raise ImportIdentityError( + "Realization lineage_components must be a list." + ) + lineage_components = [ + _validate_lineage_component_record(component) + for component in raw_lineage_components + ] + lineage_by_id: dict[str, dict[str, Any]] = {} + lineage_digests: set[str] = set() + for component in lineage_components: + lineage_id = component["lineage_id"] + lineage_digest = component["lineage_digest"] + if lineage_id in lineage_by_id or lineage_digest in lineage_digests: + raise ImportIdentityError( + "Realization lineage components contain duplicate identities." + ) + lineage_by_id[lineage_id] = component + lineage_digests.add(lineage_digest) + result_origin = _validate_result_origin_record(raw["result_origin"]) origin_question_ids = [ row["question_id"] for row in result_origin["row_assignments"] @@ -1603,6 +1918,21 @@ def _validate_realization_identity( "Complete or qualified realization result origin question IDs do not " "match the full expected dataset question set." ) + for assignment in result_origin["row_assignments"]: + lineage = lineage_by_id.get(assignment["prediction_lineage_id"]) + if lineage is None: + raise ImportIdentityError( + "Realization result origin references an unowned lineage component." + ) + lineage_identity = lineage["identity"] + if ( + lineage_identity["question_id"] != assignment["question_id"] + or lineage_identity["prediction_origin"] + != assignment["prediction_origin"] + ): + raise ImportIdentityError( + "Realization result origin conflicts with its lineage component." + ) parents = _exact_mapping( raw["parent_digests"], {"realization_digests", "evidence_digests", "result_digests"}, @@ -1615,6 +1945,29 @@ def _validate_realization_identity( authorization_digest = _validate_digest( raw["authorization_digest"], "authorization_digest", optional=True ) + lineage_authorizations = { + component["identity"]["authorization_digest"] + for component in lineage_components + if component["identity"]["authorization_digest"] is not None + } + if any(digest != authorization_digest for digest in lineage_authorizations): + raise ImportIdentityError( + "Realization lineage authorization is not owned by the realization." + ) + declared_parent_edges = { + digest + for digests in normalized_parents.values() + for digest in digests + } + lineage_parent_edges = { + digest + for component in lineage_components + for digest in component["identity"]["parent_digests"] + } + if not lineage_parent_edges <= declared_parent_edges: + raise ImportIdentityError( + "Realization lineage parent edge is not owned by the realization." + ) overlay = raw["overlay"] normalized_overlay = None if overlay is not None: @@ -1636,6 +1989,18 @@ def _validate_realization_identity( f"overlay.{field}", optional=(field not in {"source_sha256", "replacement_digest"}), ) + declared_source_digests = {source["sha256"] for source in normalized_sources} + if normalized_overlay is not None: + declared_source_digests.add(normalized_overlay["source_sha256"]) + lineage_source_edges = { + digest + for component in lineage_components + for digest in component["identity"]["source_digests"] + } + if not lineage_source_edges <= declared_source_digests: + raise ImportIdentityError( + "Realization lineage source edge is not owned by the realization." + ) derivation_origin = result_origin["derivation_origin"] derived = derivation_origin in {"repair_overlay", "offline_transformation"} @@ -1656,6 +2021,16 @@ def _validate_realization_identity( raise ImportIdentityError( f"{derivation_origin} realization requires an overlay identity." ) + if authorization_digest not in lineage_authorizations: + raise ImportIdentityError( + f"{derivation_origin} realization requires an authorized lineage " + "component." + ) + if not lineage_parent_edges: + raise ImportIdentityError( + f"{derivation_origin} realization requires a parent-derived lineage " + "component." + ) if derivation_origin == "offline_transformation" and any( normalized_overlay[field] is None for field in ( @@ -1697,6 +2072,7 @@ def _validate_realization_identity( ), "validation": canonicalize(validation), "evidence": canonicalize(evidence), + "lineage_components": canonicalize(lineage_components), "result_origin": result_origin, "parent_digests": normalized_parents, "authorization_digest": authorization_digest, @@ -1718,6 +2094,7 @@ def make_realization( fields: Mapping[str, Any], adapter: Callable[..., Any], validator: Callable[..., Any], + expected_dataset: ExpectedDataset, ) -> dict[str, Any]: """Create an immutable realization identity distinct from its condition.""" if not isinstance(condition_id, str) or not re.fullmatch(r"cond_[0-9a-f]{16}", condition_id): @@ -1733,7 +2110,9 @@ def make_realization( adapter=adapter, validator=validator ) validated_identity = _validate_realization_identity( - identity, runtime_importer=runtime_importer + identity, + runtime_importer=runtime_importer, + expected_dataset=expected_dataset, ) payload = canonicalize( { diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py index 2fab1b9..cc8273d 100644 --- a/tests/importing/test_identity.py +++ b/tests/importing/test_identity.py @@ -69,7 +69,12 @@ def transform(self, text: str) -> str: def make_realization(**kwargs): - return _make_realization(adapter=_adapter, validator=_validator, **kwargs) + return _make_realization( + adapter=_adapter, + validator=_validator, + expected_dataset=kwargs.pop("expected_dataset", _dataset()), + **kwargs, + ) def _dataset(): @@ -860,15 +865,57 @@ def test_fully_validated_native_children_reproduce_build_execution_plan_conditio assert records.condition["condition_id"] == native_condition["condition_id"] -def _realization_identity() -> dict: - origin = make_result_origin( - derivation_origin="external_import", - row_assignments=( - ("q2", "external_historical_inference", f"lin_{'a' * 16}"), - ("q1", "external_historical_inference", f"lin_{'b' * 16}"), +def _attach_result_origin( + identity: dict, + *, + derivation_origin: str, + assignments: tuple[tuple[str, str], ...], + authorization_digest: str | None = None, + source_digest: str = "2" * 64, +) -> None: + derived = derivation_origin in {"repair_overlay", "offline_transformation"} + components = [ + make_lineage_component( + operation_type=derivation_origin, + question_id=question_id, + parent_digests=(("d" * 64,) if derived else ()), + source_digests=(source_digest,), + authorization_digest=(authorization_digest if derived else None), + implementation=_transform, + parameters={}, + input_digest=integrity_digest( + {"question_id": question_id, "stage": "input"} + ), + preownership_output_digest=integrity_digest( + { + "question_id": question_id, + "prediction_origin": prediction_origin, + "stage": "preownership-output", + } + ), + prediction_origin=prediction_origin, + ) + for question_id, prediction_origin in assignments + ] + identity["lineage_components"] = components + identity["result_origin"] = make_result_origin( + derivation_origin=derivation_origin, + row_assignments=tuple( + ( + question_id, + prediction_origin, + component["lineage_id"], + ) + for (question_id, prediction_origin), component in zip( + assignments, components, strict=True + ) ), ) - return { + + +def _realization_identity() -> dict: + expected_dataset = _dataset() + identity = { "import_spec_digest": "1" * 64, "sources": [ { @@ -888,10 +935,10 @@ def _realization_identity() -> dict: } ], "expected_dataset": { - "snapshot_digest": "3" * 64, - "question_set_digest": integrity_digest(["q2", "q1"]), - "derivation_digest": "5" * 64, - "question_ids": ["q2", "q1"], + "snapshot_digest": expected_dataset.snapshot_digest, + "question_set_digest": expected_dataset.question_set_digest, + "derivation_digest": expected_dataset.derivation_digest, + "question_ids": list(expected_dataset.selected_question_ids), }, "importer_implementation": importer_implementation_identity( adapter=_adapter, validator=_validator @@ -928,7 +975,8 @@ def _realization_identity() -> dict: "defect_digest": "9" * 64, "scope_disposition": "included", }, - "result_origin": origin, + "lineage_components": [], + "result_origin": {}, "parent_digests": { "realization_digests": [], "evidence_digests": [], @@ -937,11 +985,29 @@ def _realization_identity() -> dict: "authorization_digest": None, "overlay": None, } + _attach_result_origin( + identity, + derivation_origin="external_import", + assignments=( + ("q2", "external_historical_inference"), + ("q1", "external_historical_inference"), + ), + ) + return identity def _mutate_realization(identity: dict, field: str, value) -> None: if field == "source_sha256": identity["sources"][0]["sha256"] = value + _attach_result_origin( + identity, + derivation_origin="external_import", + assignments=( + ("q2", "external_historical_inference"), + ("q1", "external_historical_inference"), + ), + source_digest=value, + ) elif field == "mapping": identity["parsing_policy"]["mapping"] = value elif field == "dialect": @@ -961,12 +1027,14 @@ def _mutate_realization(identity: dict, field: str, value) -> None: "preownership_output_digest": None, "implementation_digest": None, } - identity["result_origin"] = make_result_origin( + _attach_result_origin( + identity, derivation_origin="repair_overlay", - row_assignments=( - ("q2", "external_historical_inference", f"lin_{'a' * 16}"), - ("q1", "external_repair_inference", f"lin_{'b' * 16}"), + assignments=( + ("q2", "external_historical_inference"), + ("q1", "external_repair_inference"), ), + authorization_digest=value, ) elif field == "overlay_sha256": identity["authorization_digest"] = "c" * 64 @@ -979,15 +1047,21 @@ def _mutate_realization(identity: dict, field: str, value) -> None: "preownership_output_digest": None, "implementation_digest": None, } - identity["result_origin"] = make_result_origin( + _attach_result_origin( + identity, derivation_origin="repair_overlay", - row_assignments=( - ("q2", "external_historical_inference", f"lin_{'a' * 16}"), - ("q1", "external_repair_inference", f"lin_{'b' * 16}"), + assignments=( + ("q2", "external_historical_inference"), + ("q1", "external_repair_inference"), ), + authorization_digest="c" * 64, ) elif field == "prediction_origins": - identity["result_origin"] = value + _attach_result_origin( + identity, + derivation_origin="external_import", + assignments=value, + ) else: raise AssertionError(field) @@ -1004,12 +1078,9 @@ def _mutate_realization(identity: dict, field: str, value) -> None: ("overlay_sha256", "5" * 64), ( "prediction_origins", - make_result_origin( - derivation_origin="external_import", - row_assignments=( - ("q2", "native_inference", f"lin_{'c' * 16}"), - ("q1", "external_historical_inference", f"lin_{'b' * 16}"), - ), + ( + ("q2", "native_inference"), + ("q1", "external_historical_inference"), ), ), ], @@ -1146,10 +1217,11 @@ def test_base_repair_and_transformation_are_distinct_realizations_of_one_conditi identities = [] for derivation in ("external_import", "repair_overlay", "offline_transformation"): identity = _realization_identity() - identity["result_origin"] = make_result_origin( + _attach_result_origin( + identity, derivation_origin=derivation, - row_assignments=( - ("q2", "external_historical_inference", f"lin_{'a' * 16}"), + assignments=( + ("q2", "external_historical_inference"), ( "q1", ( @@ -1157,9 +1229,9 @@ def test_base_repair_and_transformation_are_distinct_realizations_of_one_conditi if derivation == "repair_overlay" else "external_historical_inference" ), - f"lin_{'b' * 16}", ), ), + authorization_digest=("c" * 64 if derivation != "external_import" else None), ) if derivation != "external_import": identity["authorization_digest"] = "c" * 64 @@ -1196,10 +1268,11 @@ def test_base_repair_and_transformation_are_distinct_realizations_of_one_conditi def test_derived_realization_requires_authorization_parent_and_overlay(derivation): condition = _records().condition identity = _realization_identity() - identity["result_origin"] = make_result_origin( + _attach_result_origin( + identity, derivation_origin=derivation, - row_assignments=( - ("q2", "external_historical_inference", f"lin_{'a' * 16}"), + assignments=( + ("q2", "external_historical_inference"), ( "q1", ( @@ -1207,13 +1280,14 @@ def test_derived_realization_requires_authorization_parent_and_overlay(derivatio if derivation == "repair_overlay" else "external_historical_inference" ), - f"lin_{'b' * 16}", ), ), + authorization_digest="c" * 64, ) for field in ("authorization_digest", "overlay", "parent", "evidence"): candidate = _realization_identity() candidate["result_origin"] = identity["result_origin"] + candidate["lineage_components"] = identity["lineage_components"] candidate["authorization_digest"] = "c" * 64 candidate["parent_digests"]["realization_digests"] = ["d" * 64] candidate["parent_digests"]["evidence_digests"] = ["0" * 64] @@ -1364,6 +1438,7 @@ def test_realization_recomputes_implementation_identity_from_supplied_callables( fields={}, adapter=_validator, validator=_adapter, + expected_dataset=_dataset(), ) @@ -1452,6 +1527,120 @@ def test_expected_dataset_unknown_reasons_must_be_nonempty(): ) +@pytest.mark.parametrize( + ("target", "field", "value"), + [ + ("model", "source_sha256", "a" * 64), + ("model", "cache_path", "/machine-a/cache"), + ("method", "recorded_at", "2026-07-19T12:00:00Z"), + ("condition", "result_digest", "b" * 64), + ], +) +def test_fallback_parameters_refuse_realization_and_audit_metadata( + target, field, value +): + model = _model() + method = _method() + condition = _condition() + if target == "model": + model = replace(model, effective_parameters={field: value}) + elif target == "method": + method = replace(method, effective_parameters={field: value}) + else: + condition = replace(condition, generation_parameters={field: value}) + with pytest.raises(ImportIdentityError, match="parameter|metadata|path|forbidden"): + build_import_semantic_identity( + condition=condition, + dataset=_dataset(), + model=model, + method=method, + prompt=_prompt(), + ) + + +def test_fallback_method_implementation_refuses_machine_local_source_file(): + method = replace( + _method(), + implementation={ + "qualified_name": "historical.matcher:match", + "source_file": "/machine-a/matcher.py", + "source_digest": "a" * 64, + }, + unknown_reasons={}, + ) + with pytest.raises(ImportIdentityError, match="path|implementation|machine"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=method, + prompt=_prompt(), + ) + + +def test_direct_method_implementation_must_have_the_closed_code_identity_shape(): + method = replace( + _method(), + implementation={"attacker": "not an implementation identity"}, + unknown_reasons={}, + ) + with pytest.raises(ImportIdentityError, match="implementation|field"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=method, + prompt=_prompt(), + ) + + +@pytest.mark.parametrize( + "mutation", + [ + {"reference_kind": "profile_derived_reference_snapshot"}, + {"trust_label": "forged-trust"}, + {"derivation": {"forged": True}}, + {"derivation_digest": "a" * 64}, + {"snapshot_digest": "b" * 64}, + ], +) +def test_expected_dataset_trust_chain_is_owned_and_recomputed(mutation): + with pytest.raises(ImportIdentityError, match="dataset|derivation|snapshot|trust|reference"): + build_import_semantic_identity( + condition=_condition(), + dataset=replace(_dataset(), **mutation), + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +def test_complete_realization_question_set_is_owned_by_validated_dataset(): + condition = _records().condition + identity = _realization_identity() + identity["expected_dataset"] = { + **identity["expected_dataset"], + "question_ids": ["q2"], + "question_set_digest": integrity_digest(["q2"]), + } + identity["result_origin"] = make_result_origin( + derivation_origin="external_import", + row_assignments=( + ("q2", "external_historical_inference", f"lin_{'a' * 16}"), + ), + ) + with pytest.raises(ImportIdentityError, match="dataset|question set|question IDs"): + _make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + adapter=_adapter, + validator=_validator, + expected_dataset=_dataset(), + ) + + @pytest.mark.parametrize( "protocol_settings", [ @@ -1503,7 +1692,48 @@ def test_native_dummy_cannot_discard_a_known_revision(): ) -def test_native_dataset_source_refuses_audit_and_machine_metadata(): +def test_native_prompt_identity_preserves_credential_shaped_literal_text(): + contents = { + "direct_mcq": "Answer the question; literal example password=alpha", + "free_text": "Respond freely", + "option_matching": "Match the answer", + } + files = { + name: {"sha256": integrity_digest(content), "content": content} + for name, content in contents.items() + } + prompt_payload = {"version": "test-v1", "files": files} + prompt = ImportPromptSpec( + prompt_key="prompt", + template_identity="test-v1", + template_digest=integrity_digest(files), + template_contents=contents, + unknown_reason=None, + native_compatibility_identity={ + "prompt_id": f"prompt_{integrity_digest(prompt_payload)[:16]}", + **prompt_payload, + }, + ) + records = build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=prompt, + ) + assert records.prompt["identity"]["files"]["direct_mcq"]["content"] == ( + contents["direct_mcq"] + ) + + +@pytest.mark.parametrize( + "source", + [ + {"audit": {"operator": "alice"}, "node": "machine-a"}, + {"hf_path": "/home/alice/machine-only/dataset"}, + ], +) +def test_native_dataset_source_refuses_audit_and_machine_metadata(source): dataset = _dataset() artifact_payload = { "spec": { @@ -1517,7 +1747,7 @@ def test_native_dataset_source_refuses_audit_and_machine_metadata(): "output_name": None, }, "content_digest": dataset_content_digest(dataset.artifact_frame), - "source": {"audit": {"operator": "alice"}, "node": "machine-a"}, + "source": source, } artifact_id = short_id("ds", artifact_payload) selection_payload = {**dataset.selection_payload, "artifact_id": artifact_id} @@ -1545,10 +1775,11 @@ def test_partial_realization_origin_may_be_an_ordered_expected_subset(): condition = _records().condition identity = _realization_identity() identity["evidence"]["evidence_status"] = "partial" - identity["result_origin"] = make_result_origin( + _attach_result_origin( + identity, derivation_origin="external_import", - row_assignments=( - ("q2", "external_historical_inference", f"lin_{'a' * 16}"), + assignments=( + ("q2", "external_historical_inference"), ), ) realization = make_realization( @@ -1567,8 +1798,10 @@ def test_malformed_realization_origin_may_have_no_evaluable_rows(): condition = _records().condition identity = _realization_identity() identity["evidence"]["evidence_status"] = "malformed" - identity["result_origin"] = make_result_origin( - derivation_origin="external_import", row_assignments=() + _attach_result_origin( + identity, + derivation_origin="external_import", + assignments=(), ) realization = make_realization( condition_id=condition["condition_id"], @@ -1722,6 +1955,14 @@ def test_lineage_requires_mandatory_runtime_keyword_parameters(): ) +def test_runtime_identity_refuses_ambiguous_module_lambdas(): + first = lambda value: value # noqa: E731 + second = lambda value: not value # noqa: E731 + for target in (first, second): + with pytest.raises(ImportIdentityError, match="named|lambda|addressable"): + importer_implementation_identity(adapter=target, validator=_validator) + + def test_lineage_supports_nonexecuted_declared_external_implementation_identity(): source_digest = "2" * 64 component = make_lineage_component( @@ -1826,6 +2067,61 @@ def test_realization_result_origin_question_ids_match_expected_question_set(): ) +def test_realization_refuses_unowned_lineage_component_id(): + condition = _records().condition + identity = _realization_identity() + identity["result_origin"] = make_result_origin( + derivation_origin="external_import", + row_assignments=( + ("q2", "external_historical_inference", f"lin_{'c' * 16}"), + ("q1", "external_historical_inference", f"lin_{'d' * 16}"), + ), + ) + with pytest.raises(ImportIdentityError, match="lineage|owned|component"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +@pytest.mark.parametrize("edge", ["authorization", "parent"]) +def test_derived_realization_lineage_edges_are_owned_by_realization(edge): + condition = _records().condition + identity = _realization_identity() + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": None, + "preownership_output_digest": None, + "implementation_digest": None, + } + _attach_result_origin( + identity, + derivation_origin="repair_overlay", + assignments=( + ("q2", "external_historical_inference"), + ("q1", "external_repair_inference"), + ), + authorization_digest="c" * 64, + ) + if edge == "authorization": + identity["authorization_digest"] = "a" * 64 + else: + identity["parent_digests"]["realization_digests"] = ["b" * 64] + with pytest.raises(ImportIdentityError, match="lineage|authorization|parent"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + def test_result_origin_rejects_bare_mixed_duplicate_rows_and_invalid_origins(): for assignments in ( (("q1", "mixed", f"lin_{'a' * 16}"),), From fd03a2f92c108c87cbc83d643f02ac28b5d3f37d Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 04:26:11 +0300 Subject: [PATCH 23/47] fix: bind importer identity to trusted runtime inputs --- .../importing/dataset_reference.py | 8 + src/choicebench/importing/identity.py | 215 +++++++++++- src/choicebench/importing/schema.py | 3 +- tests/importing/test_dataset_reference.py | 13 + tests/importing/test_identity.py | 317 +++++++++++++++++- 5 files changed, 543 insertions(+), 13 deletions(-) diff --git a/src/choicebench/importing/dataset_reference.py b/src/choicebench/importing/dataset_reference.py index 21ee66e..6f48d5a 100644 --- a/src/choicebench/importing/dataset_reference.py +++ b/src/choicebench/importing/dataset_reference.py @@ -531,6 +531,14 @@ def build_expected_dataset( raise DatasetReferenceError( "Expected dataset selection unknown reasons do not match null semantics." ) + if ( + declaration.selection_n_samples is not None + and declaration.selection_n_samples != len(frame) + ): + raise DatasetReferenceError( + "Expected dataset selection n_samples does not match selected question " + "coverage." + ) ( artifact_payload, artifact_digest, diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index 0427601..5f3caf9 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -222,6 +222,20 @@ def _validate_direct_declarations( method: ImportMethodSpec, prompt: ImportPromptSpec, ) -> None: + _validate_stable_values( + { + "model_key": model.model_key, + "display_name": model.display_name, + "backend": model.backend, + "provider": model.provider, + "revision": model.revision, + "method_key": method.method_key, + "method_name": method.name, + "prompt_key": prompt.prompt_key, + "template_identity": prompt.template_identity, + }, + "Semantic child declarations", + ) _validate_stable_parameters( model.effective_parameters, "Model effective parameters" ) @@ -231,6 +245,10 @@ def _validate_direct_declarations( _validate_stable_parameters( method.effective_parameters, "Method effective parameters" ) + if condition.preflight_identity is not None: + _validate_stable_parameters( + condition.preflight_identity, "Condition preflight identity" + ) if method.implementation is not None: try: validated_implementation = validate_implementation_identity_record( @@ -283,6 +301,15 @@ def _validate_direct_declarations( ) if prompt.template_digest is not None: _validate_digest(prompt.template_digest, "prompt.template_digest") + if ( + prompt.native_compatibility_identity is None + and prompt.template_digest is not None + and prompt.template_contents is not None + and prompt.template_digest != integrity_digest(prompt.template_contents) + ): + raise ImportIdentityError( + "Fallback prompt digest does not own the exact template contents." + ) def _identity_record(prefix: str, payload: Mapping[str, Any]) -> tuple[str, str, dict]: @@ -329,7 +356,7 @@ def _reject_forbidden( def _is_machine_path(value: str) -> bool: return bool( - value.startswith(("/", "./", "../", "~/", "~\\")) + value.startswith(("/", "./", "../", "~/", "~\\", "\\\\")) or re.match(r"^[A-Za-z]:[\\/]", value) ) @@ -346,6 +373,17 @@ def _validate_stable_values(value: Any, where: str) -> None: raise ImportIdentityError(f"{where} contains a machine-local path value.") +def _validate_no_machine_path_values(value: Any, where: str) -> None: + if isinstance(value, Mapping): + for item in value.values(): + _validate_no_machine_path_values(item, where) + elif isinstance(value, (list, tuple)): + for item in value: + _validate_no_machine_path_values(item, where) + elif isinstance(value, str) and _is_machine_path(value): + raise ImportIdentityError(f"{where} contains a machine-local path value.") + + def _validate_stable_parameters(value: Any, where: str) -> None: _reject_forbidden(value, where) if isinstance(value, Mapping): @@ -538,6 +576,9 @@ def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: raise ImportIdentityError( "Expected dataset native artifact conflicts with its declaration." ) + _validate_no_machine_path_values( + spec, "Expected dataset native artifact spec" + ) if not isinstance(artifact["source"], Mapping): raise ImportIdentityError("Expected dataset native source is invalid.") source = artifact["source"] @@ -719,7 +760,12 @@ def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: "Expected dataset derivation logical path is invalid." ) pure_path = PurePosixPath(logical_path) - if pure_path.is_absolute() or ".." in pure_path.parts or not pure_path.parts: + if ( + pure_path.is_absolute() + or ".." in pure_path.parts + or not pure_path.parts + or _is_machine_path(logical_path) + ): raise ImportIdentityError( "Expected dataset derivation logical path is unsafe." ) @@ -1090,9 +1136,9 @@ def _semantic_child_records( else unknown("template contents") ), } - prompt_id, prompt_digest, prompt_payload = _identity_record( - "prompt", prompt_payload - ) + prompt_payload = canonicalize(prompt_payload, redact_secrets=False) + prompt_digest = integrity_digest(prompt_payload) + prompt_id = f"prompt_{prompt_digest[:16]}" prompt_mode = "imported_semantic_fallback" prompt_record = { "prompt_id": prompt_id, @@ -1275,6 +1321,23 @@ def _runtime_callable_record(target: Callable[..., Any], where: str) -> dict[str raise ImportIdentityError( f"{where} must be a uniquely addressable named module callable." ) + try: + callable_source = inspect.getsource(target) + except (OSError, TypeError) as exc: + raise ImportIdentityError( + f"{where} lacks uniquely inspectable callable source." + ) from exc + callable_payload = { + "source": callable_source, + "defaults": getattr(target, "__defaults__", None), + "keyword_defaults": getattr(target, "__kwdefaults__", None), + } + try: + callable_digest = integrity_digest(callable_payload) + except (TypeError, ValueError) as exc: + raise ImportIdentityError( + f"{where} has unsupported callable defaults." + ) from exc try: record = validate_implementation_identity_record( implementation_identity(target) @@ -1285,6 +1348,7 @@ def _runtime_callable_record(target: Callable[..., Any], where: str) -> dict[str ) from exc if "source_digest" not in record: raise ImportIdentityError(f"{where} lacks an inspectable source digest.") + record["callable_digest"] = callable_digest return record @@ -1480,7 +1544,9 @@ def make_result_origin( } -def _validate_lineage_component_record(value: Any) -> dict[str, Any]: +def _validate_lineage_component_record( + value: Any, *, runtime_callable: Callable[..., Any] | None +) -> dict[str, Any]: record = _exact_mapping( value, {"lineage_id", "lineage_digest", "identity"}, @@ -1546,6 +1612,25 @@ def _validate_lineage_component_record(value: Any) -> dict[str, Any]: raise ImportIdentityError( "Realization lineage implementation lacks a source digest." ) + if implementation["identity_mode"] == "runtime_callable": + if runtime_callable is None: + raise ImportIdentityError( + "Realization runtime lineage requires its registered callable." + ) + runtime_identity = _runtime_callable_record( + runtime_callable, "Realization lineage runtime implementation" + ) + if canonicalize(runtime_identity) != canonicalize( + implementation_identity_record + ): + raise ImportIdentityError( + "Realization runtime lineage implementation does not match its " + "registered callable." + ) + elif runtime_callable is not None: + raise ImportIdentityError( + "Declared external lineage cannot claim a runtime callable." + ) if ( implementation["identity_mode"] == "declared_external" and source_digest not in normalized_sources @@ -1665,6 +1750,7 @@ def _validate_realization_identity( *, runtime_importer: Mapping[str, Any], expected_dataset: ExpectedDataset, + lineage_runtime_callables: Mapping[str, Callable[..., Any]], ) -> dict[str, Any]: raw = _exact_mapping(value, _REALIZATION_KEYS, "Realization identity") _validate_digest(raw["import_spec_digest"], "import_spec_digest") @@ -1687,7 +1773,12 @@ def _validate_realization_identity( if not isinstance(logical_path, str): raise ImportIdentityError("Realization source logical_path must be a string.") pure_path = PurePosixPath(logical_path) - if pure_path.is_absolute() or ".." in pure_path.parts or not pure_path.parts: + if ( + pure_path.is_absolute() + or ".." in pure_path.parts + or not pure_path.parts + or _is_machine_path(logical_path) + ): raise ImportIdentityError("Realization source logical_path is unsafe.") if source["classification"] not in { "raw", @@ -1876,10 +1967,24 @@ def _validate_realization_identity( raise ImportIdentityError( "Realization lineage_components must be a list." ) - lineage_components = [ - _validate_lineage_component_record(component) - for component in raw_lineage_components - ] + if not isinstance(lineage_runtime_callables, Mapping) or not all( + isinstance(lineage_id, str) and callable(target) + for lineage_id, target in lineage_runtime_callables.items() + ): + raise ImportIdentityError( + "Realization lineage runtime callable registry is invalid." + ) + lineage_components = [] + for component in raw_lineage_components: + claimed_lineage_id = ( + component.get("lineage_id") if isinstance(component, Mapping) else None + ) + lineage_components.append( + _validate_lineage_component_record( + component, + runtime_callable=lineage_runtime_callables.get(claimed_lineage_id), + ) + ) lineage_by_id: dict[str, dict[str, Any]] = {} lineage_digests: set[str] = set() for component in lineage_components: @@ -1891,6 +1996,17 @@ def _validate_realization_identity( ) lineage_by_id[lineage_id] = component lineage_digests.add(lineage_digest) + runtime_lineage_ids = { + component["lineage_id"] + for component in lineage_components + if component["identity"]["implementation"]["identity_mode"] + == "runtime_callable" + } + if set(lineage_runtime_callables) != runtime_lineage_ids: + raise ImportIdentityError( + "Realization lineage runtime callable registry does not exactly own " + "the runtime lineage components." + ) result_origin = _validate_result_origin_record(raw["result_origin"]) origin_question_ids = [ @@ -2031,6 +2147,38 @@ def _validate_realization_identity( f"{derivation_origin} realization requires a parent-derived lineage " "component." ) + derived_components = [ + component + for component in lineage_components + if component["identity"]["operation_type"] == derivation_origin + ] + if not derived_components: + raise ImportIdentityError( + f"{derivation_origin} realization requires a matching derived " + "lineage component." + ) + for component in derived_components: + component_identity = component["identity"] + if component_identity["authorization_digest"] != authorization_digest: + raise ImportIdentityError( + f"{derivation_origin} lineage component lacks the realization " + "authorization." + ) + if not component_identity["parent_digests"]: + raise ImportIdentityError( + f"{derivation_origin} lineage component lacks a parent edge." + ) + if derivation_origin == "repair_overlay": + for assignment in result_origin["row_assignments"]: + if assignment["prediction_origin"] in { + "native_inference", + "external_repair_inference", + }: + component = lineage_by_id[assignment["prediction_lineage_id"]] + if component["identity"]["operation_type"] != "repair_overlay": + raise ImportIdentityError( + "Repair prediction row lacks repair_overlay lineage." + ) if derivation_origin == "offline_transformation" and any( normalized_overlay[field] is None for field in ( @@ -2095,6 +2243,8 @@ def make_realization( adapter: Callable[..., Any], validator: Callable[..., Any], expected_dataset: ExpectedDataset, + semantic_condition: Mapping[str, Any], + lineage_runtime_callables: Mapping[str, Callable[..., Any]], ) -> dict[str, Any]: """Create an immutable realization identity distinct from its condition.""" if not isinstance(condition_id, str) or not re.fullmatch(r"cond_[0-9a-f]{16}", condition_id): @@ -2106,6 +2256,48 @@ def make_realization( ) if not isinstance(identity, Mapping) or not identity: raise ImportIdentityError("Realization identity must be a non-empty mapping.") + condition_record = _exact_mapping( + semantic_condition, + { + "condition_id", + "condition_digest", + "identity", + "condition_key", + "expected_question_ids", + }, + "Realization semantic condition", + ) + rebuilt_condition = make_semantic_condition( + identity=condition_record["identity"], + fields={ + "condition_key": condition_record["condition_key"], + "expected_question_ids": condition_record["expected_question_ids"], + }, + ) + if canonicalize(rebuilt_condition) != canonicalize(condition_record): + raise ImportIdentityError( + "Realization semantic condition identity is inconsistent." + ) + if ( + condition_record["condition_id"] != condition_id + or condition_record["condition_digest"] != condition_digest + ): + raise ImportIdentityError( + "Realization condition ID/digest do not match its semantic condition." + ) + benchmark = condition_record["identity"]["benchmark"] + if ( + benchmark["name"] != expected_dataset.benchmark_name + or benchmark["split"] != expected_dataset.split + or benchmark["artifact_id"] != expected_dataset.artifact_id + or benchmark["selection_id"] != expected_dataset.selection_id + or tuple(condition_record["expected_question_ids"]) + != expected_dataset.selected_question_ids + ): + raise ImportIdentityError( + "Realization semantic condition is not owned by its expected dataset " + "artifact and selection." + ) runtime_importer = importer_implementation_identity( adapter=adapter, validator=validator ) @@ -2113,6 +2305,7 @@ def make_realization( identity, runtime_importer=runtime_importer, expected_dataset=expected_dataset, + lineage_runtime_callables=lineage_runtime_callables, ) payload = canonicalize( { diff --git a/src/choicebench/importing/schema.py b/src/choicebench/importing/schema.py index a685b3d..8d7967a 100644 --- a/src/choicebench/importing/schema.py +++ b/src/choicebench/importing/schema.py @@ -1094,6 +1094,7 @@ def _validate_model_payload(value: Any, where: str) -> dict[str, Any]: _IMPLEMENTATION_KEYS = { "qualified_name", "source_file", "source_digest", "distribution", "distribution_version", "package_tree_digest", "package_file_count", + "callable_digest", } @@ -1106,7 +1107,7 @@ def _validate_implementation(value: Any, where: str) -> dict[str, Any]: for key in ("source_file", "distribution", "distribution_version"): if key in result: _nonempty(result[key], f"{where}.{key}") - for key in ("source_digest", "package_tree_digest"): + for key in ("source_digest", "package_tree_digest", "callable_digest"): if key in result: _sha256(result[key], f"{where}.{key}") has_package_digest = "package_tree_digest" in result diff --git a/tests/importing/test_dataset_reference.py b/tests/importing/test_dataset_reference.py index f4e62c1..0dc2efc 100644 --- a/tests/importing/test_dataset_reference.py +++ b/tests/importing/test_dataset_reference.py @@ -439,6 +439,19 @@ def test_snapshot_self_validation_preserves_known_selection_semantics(tmp_path: validate_expected_snapshot(tmp_path, record) +def test_known_selection_count_must_equal_selected_question_coverage(): + declaration = replace( + _declaration(), + selection_n_samples=99, + selection_unknown_reasons={"selection_seed": "not recorded"}, + ) + with pytest.raises(DatasetReferenceError, match="n_samples|selected|coverage"): + build_expected_dataset( + declaration, + {"reference": _source("reference", _rows())}, + ) + + def test_validated_native_compatibility_identity_is_preserved_exactly(): full_artifact = build_expected_dataset( _declaration(expected_question_ids=("q1", "q2", "q3")), diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py index cc8273d..97b1b65 100644 --- a/tests/importing/test_identity.py +++ b/tests/importing/test_identity.py @@ -57,6 +57,20 @@ def _transform_required(value: str, *, threshold: float) -> str: return value if threshold >= 0 else "" +def _redefined_adapter(value: str, prefix: str = "first") -> str: + return prefix + value + + +_first_redefined_adapter = _redefined_adapter + + +def _redefined_adapter(value: str, prefix: str = "second") -> str: + return prefix + value + + +_second_redefined_adapter = _redefined_adapter + + class _StatefulCallable: def __init__(self, value: str) -> None: self.value = value @@ -69,10 +83,21 @@ def transform(self, text: str) -> str: def make_realization(**kwargs): + identity = kwargs["identity"] + runtime_lineages = { + component["lineage_id"]: _transform + for component in identity.get("lineage_components", ()) + if component.get("identity", {}) + .get("implementation", {}) + .get("identity_mode") + == "runtime_callable" + } return _make_realization( adapter=_adapter, validator=_validator, expected_dataset=kwargs.pop("expected_dataset", _dataset()), + semantic_condition=kwargs.pop("semantic_condition", _records().condition), + lineage_runtime_callables=runtime_lineages, **kwargs, ) @@ -1430,15 +1455,21 @@ def test_realization_requires_runtime_derived_importer_implementation_identity() def test_realization_recomputes_implementation_identity_from_supplied_callables(): condition = _records().condition + identity = _realization_identity() with pytest.raises(ImportIdentityError, match="runtime|callable|implementation"): _make_realization( condition_id=condition["condition_id"], condition_digest=condition["condition_digest"], - identity=_realization_identity(), + identity=identity, fields={}, adapter=_validator, validator=_adapter, expected_dataset=_dataset(), + semantic_condition=condition, + lineage_runtime_callables={ + component["lineage_id"]: _transform + for component in identity["lineage_components"] + }, ) @@ -1558,6 +1589,17 @@ def test_fallback_parameters_refuse_realization_and_audit_metadata( ) +def test_fallback_model_logical_identity_refuses_machine_local_path(): + with pytest.raises(ImportIdentityError, match="model|path|machine"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=replace(_model(), display_name="/machine-a/models/model"), + method=_method(), + prompt=_prompt(), + ) + + def test_fallback_method_implementation_refuses_machine_local_source_file(): method = replace( _method(), @@ -1578,6 +1620,71 @@ def test_fallback_method_implementation_refuses_machine_local_source_file(): ) +def test_fallback_preflight_refuses_source_provenance_metadata(): + condition = replace( + _condition(), + preflight_identity={"source_sha256": "a" * 64}, + unknown_reasons={ + key: value + for key, value in _condition().unknown_reasons.items() + if key != "preflight_identity" + }, + ) + with pytest.raises(ImportIdentityError, match="preflight|metadata|forbidden"): + build_import_semantic_identity( + condition=condition, + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +def test_fallback_prompt_digest_owns_exact_unredacted_contents(): + def build(content: str): + contents = {"direct_mcq": content} + return build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=ImportPromptSpec( + prompt_key="prompt", + template_identity="historical-v1", + template_digest=integrity_digest(contents), + template_contents=contents, + unknown_reason=None, + native_compatibility_identity=None, + ), + ) + + first = build("literal password=alpha") + second = build("literal password=beta") + assert first.prompt["prompt_id"] != second.prompt["prompt_id"] + assert first.condition["condition_id"] != second.condition["condition_id"] + assert first.prompt["identity"]["template_contents"] == { + "direct_mcq": "literal password=alpha" + } + + +def test_fallback_prompt_refuses_digest_that_does_not_own_contents(): + with pytest.raises(ImportIdentityError, match="prompt.*digest|contents"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=ImportPromptSpec( + prompt_key="prompt", + template_identity="historical-v1", + template_digest="a" * 64, + template_contents={"direct_mcq": "literal prompt"}, + unknown_reason=None, + native_compatibility_identity=None, + ), + ) + + def test_direct_method_implementation_must_have_the_closed_code_identity_shape(): method = replace( _method(), @@ -1638,6 +1745,11 @@ def test_complete_realization_question_set_is_owned_by_validated_dataset(): adapter=_adapter, validator=_validator, expected_dataset=_dataset(), + semantic_condition=condition, + lineage_runtime_callables={ + component["lineage_id"]: _transform + for component in identity["lineage_components"] + }, ) @@ -1771,6 +1883,44 @@ def test_native_dataset_source_refuses_audit_and_machine_metadata(source): ) +def test_native_dataset_spec_refuses_machine_local_hf_path(): + dataset = _dataset() + artifact_payload = { + "spec": { + "benchmark": dataset.benchmark_name, + "split": dataset.split, + "hf_path": "/home/alice/machine-only/dataset", + "hf_subset": None, + "source_revision": None, + "normalization_version": "v1", + "transforms": [], + "output_name": None, + }, + "content_digest": dataset_content_digest(dataset.artifact_frame), + "source": {}, + } + artifact_id = short_id("ds", artifact_payload) + selection_payload = {**dataset.selection_payload, "artifact_id": artifact_id} + forged = replace( + dataset, + identity_mode="native_compatibility", + artifact_payload=artifact_payload, + artifact_digest=integrity_digest(artifact_payload), + artifact_id=artifact_id, + selection_payload=selection_payload, + selection_digest=integrity_digest(selection_payload), + selection_id=short_id("sel", selection_payload), + ) + with pytest.raises(ImportIdentityError, match="spec|path|machine"): + build_import_semantic_identity( + condition=_condition(), + dataset=forged, + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + def test_partial_realization_origin_may_be_an_ordered_expected_subset(): condition = _records().condition identity = _realization_identity() @@ -1963,6 +2113,18 @@ def test_runtime_identity_refuses_ambiguous_module_lambdas(): importer_implementation_identity(adapter=target, validator=_validator) +def test_runtime_identity_distinguishes_same_named_top_level_function_objects(): + first = importer_implementation_identity( + adapter=_first_redefined_adapter, + validator=_validator, + ) + second = importer_implementation_identity( + adapter=_second_redefined_adapter, + validator=_validator, + ) + assert first != second + + def test_lineage_supports_nonexecuted_declared_external_implementation_identity(): source_digest = "2" * 64 component = make_lineage_component( @@ -2086,6 +2248,67 @@ def test_realization_refuses_unowned_lineage_component_id(): ) +@pytest.mark.parametrize( + "logical_path", + [r"C:\machine\results.csv", r"\\server\share\results.csv"], +) +def test_realization_refuses_windows_machine_local_logical_paths(logical_path): + condition = _records().condition + identity = _realization_identity() + identity["sources"][0]["logical_path"] = logical_path + with pytest.raises(ImportIdentityError, match="path|unsafe|machine"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_runtime_lineage_must_match_registered_callable(): + condition = _records().condition + identity = _realization_identity() + original = identity["lineage_components"][0] + forged_payload = { + **original["identity"], + "implementation": { + "identity_mode": "runtime_callable", + "identity": importer_implementation_identity( + adapter=_adapter, validator=_validator + )["adapter"], + }, + } + forged = { + "lineage_id": short_id("lin", forged_payload), + "lineage_digest": integrity_digest(forged_payload), + "identity": forged_payload, + } + identity["lineage_components"][0] = forged + assignments = identity["result_origin"]["row_assignments"] + identity["result_origin"] = make_result_origin( + derivation_origin="external_import", + row_assignments=( + ( + assignments[0]["question_id"], + assignments[0]["prediction_origin"], + forged["lineage_id"], + ), + ( + assignments[1]["question_id"], + assignments[1]["prediction_origin"], + assignments[1]["prediction_lineage_id"], + ), + ), + ) + with pytest.raises(ImportIdentityError, match="runtime|callable|implementation"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + @pytest.mark.parametrize("edge", ["authorization", "parent"]) def test_derived_realization_lineage_edges_are_owned_by_realization(edge): condition = _records().condition @@ -2122,6 +2345,98 @@ def test_derived_realization_lineage_edges_are_owned_by_realization(edge): ) +def test_repair_row_itself_must_own_authorization_and_parent_lineage_edges(): + condition = _records().condition + identity = _realization_identity() + base_component = make_lineage_component( + operation_type="external_import", + question_id="q2", + parent_digests=("d" * 64,), + source_digests=("2" * 64,), + authorization_digest="c" * 64, + implementation=_transform, + parameters={}, + input_digest=integrity_digest("base-input"), + preownership_output_digest=integrity_digest("base-output"), + prediction_origin="external_historical_inference", + ) + repair_component = make_lineage_component( + operation_type="repair_overlay", + question_id="q1", + parent_digests=(), + source_digests=("2" * 64,), + authorization_digest=None, + implementation=_transform, + parameters={}, + input_digest=integrity_digest("repair-input"), + preownership_output_digest=integrity_digest("repair-output"), + prediction_origin="external_repair_inference", + ) + identity["lineage_components"] = [base_component, repair_component] + identity["result_origin"] = make_result_origin( + derivation_origin="repair_overlay", + row_assignments=( + ( + "q2", + "external_historical_inference", + base_component["lineage_id"], + ), + ("q1", "external_repair_inference", repair_component["lineage_id"]), + ), + ) + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": None, + "preownership_output_digest": None, + "implementation_digest": None, + } + with pytest.raises(ImportIdentityError, match="repair|authorization|parent|lineage"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_condition_is_bound_to_expected_dataset_selection(): + condition = _records().condition + foreign_identity = { + **condition["identity"], + "benchmark": { + **condition["identity"]["benchmark"], + "artifact_id": f"ds_{'a' * 16}", + "selection_id": f"sel_{'b' * 16}", + }, + } + foreign_condition = make_semantic_condition( + identity=foreign_identity, + fields={ + "condition_key": "foreign", + "expected_question_ids": ["q2", "q1"], + }, + ) + with pytest.raises(ImportIdentityError, match="condition|dataset|selection|artifact"): + _make_realization( + condition_id=foreign_condition["condition_id"], + condition_digest=foreign_condition["condition_digest"], + identity=_realization_identity(), + fields={}, + adapter=_adapter, + validator=_validator, + expected_dataset=_dataset(), + semantic_condition=foreign_condition, + lineage_runtime_callables={ + component["lineage_id"]: _transform + for component in _realization_identity()["lineage_components"] + }, + ) + + def test_result_origin_rejects_bare_mixed_duplicate_rows_and_invalid_origins(): for assignments in ( (("q1", "mixed", f"lin_{'a' * 16}"),), From e694eacd89d5fb823d34ffb4282788ebc0d8b3b9 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 04:49:03 +0300 Subject: [PATCH 24/47] fix: close derived importer lineage edges --- src/choicebench/importing/identity.py | 122 +++++++++- tests/importing/test_identity.py | 312 ++++++++++++++++++++------ 2 files changed, 362 insertions(+), 72 deletions(-) diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index 5f3caf9..b198606 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -33,6 +33,7 @@ validate_option_mapping_identity, ) from choicebench.provenance import implementation_identity +from choicebench.registry import METHOD_REGISTRY class ImportIdentityError(ValueError): @@ -169,6 +170,29 @@ class ImportSemanticRecords: "hf_dataset_info", "revision", } +_METHOD_RUNTIME_PARAMETERS = { + "self", + "args", + "kwargs", + "backend", + "method_name", + "split_name", + "prompt_version", + "prompts_dir", + "run_id", + "temperature", + "max_tokens", + "seed", + "perturbation_name", + "model_label", + "preflight_questions", + "calibration_questions", + "calibration_runs_dir", + "modal_k", + "gate_summary", + "condition_id", + "calibration_identity", +} _DATASET_DERIVATION_KEYS = { "schema_version", "dataset_id", @@ -592,6 +616,12 @@ def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: not isinstance(source["generator"], str) or not source["generator"] ): raise ImportIdentityError("Expected dataset native source generator is invalid.") + if isinstance(source.get("generator"), str) and _is_machine_path( + source["generator"] + ): + raise ImportIdentityError( + "Expected dataset native source generator contains a machine-local path." + ) if "seed" in source and ( not isinstance(source["seed"], int) or isinstance(source["seed"], bool) ): @@ -984,16 +1014,50 @@ def _semantic_child_records( "Validated native method payload does not match the current schema." ) method_payload = validated_method_payload - if method.implementation is None: + runner_cls = METHOD_REGISTRY.get(method.name) + if runner_cls is None: raise ImportIdentityError( - "An unknown method implementation cannot be upgraded by a native " - "identity claim." + "Validated native method does not name a registered current runner." + ) + current_effective_params: dict[str, Any] = {} + accepted_parameters: set[str] = set() + for name, parameter in inspect.signature(runner_cls.__init__).parameters.items(): + if name in _METHOD_RUNTIME_PARAMETERS: + continue + accepted_parameters.add(name) + if parameter.default is not inspect.Parameter.empty: + current_effective_params[name] = parameter.default + unexpected_effective = sorted( + set(method.effective_parameters) - accepted_parameters + ) + if unexpected_effective: + raise ImportIdentityError( + "Validated native method declares parameters not accepted by its " + f"registered runner: {unexpected_effective}." ) + current_effective_params.update(method.effective_parameters) declared_preflight = _nullable_semantic( condition.preflight_identity, condition.unknown_reasons.get("preflight_identity"), "method preflight identity", ) + expected_current_payload = canonicalize( + { + "name": method.name, + "effective_params": current_effective_params, + "preflight": declared_preflight, + "implementation": implementation_identity(runner_cls), + } + ) + if canonicalize(method_payload) != expected_current_payload: + raise ImportIdentityError( + "Validated native method does not match its registered current runner." + ) + if method.implementation is None: + raise ImportIdentityError( + "An unknown method implementation cannot be upgraded by a native " + "identity claim." + ) if ( method_payload.get("name") != method.name or canonicalize(method_payload.get("effective_params")) @@ -2168,6 +2232,21 @@ def _validate_realization_identity( raise ImportIdentityError( f"{derivation_origin} lineage component lacks a parent edge." ) + if not ( + set(component_identity["parent_digests"]) + & set(normalized_parents["realization_digests"]) + ): + raise ImportIdentityError( + f"{derivation_origin} lineage component lacks the declared base " + "realization edge." + ) + if normalized_overlay["source_sha256"] not in component_identity[ + "source_digests" + ]: + raise ImportIdentityError( + f"{derivation_origin} lineage component lacks the overlay source " + "edge." + ) if derivation_origin == "repair_overlay": for assignment in result_origin["row_assignments"]: if assignment["prediction_origin"] in { @@ -2191,6 +2270,43 @@ def _validate_realization_identity( "offline_transformation overlay requires input, output, and " "implementation digests." ) + if derivation_origin == "offline_transformation": + expected_input_digest = integrity_digest( + [ + component["identity"]["input_digest"] + for component in derived_components + ] + ) + expected_output_digest = integrity_digest( + [ + component["identity"]["preownership_output_digest"] + for component in derived_components + ] + ) + implementation_identities = { + integrity_digest(component["identity"]["implementation"]): component[ + "identity" + ]["implementation"] + for component in derived_components + } + if len(implementation_identities) != 1: + raise ImportIdentityError( + "offline_transformation lineage components use conflicting " + "implementations." + ) + expected_implementation_digest = next(iter(implementation_identities)) + if ( + normalized_overlay["transformation_input_digest"] + != expected_input_digest + or normalized_overlay["preownership_output_digest"] + != expected_output_digest + or normalized_overlay["implementation_digest"] + != expected_implementation_digest + ): + raise ImportIdentityError( + "offline_transformation overlay input, output, or implementation " + "digest conflicts with its lineage components." + ) if derivation_origin == "repair_overlay" and not ( {"native_inference", "external_repair_inference"} & set(result_origin["prediction_origins"]) diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py index 97b1b65..d2a15f1 100644 --- a/tests/importing/test_identity.py +++ b/tests/importing/test_identity.py @@ -389,8 +389,7 @@ def test_semantic_builder_refuses_cross_wired_child_registry_keys(field, replace _records(**{field: replacement}) -def test_native_compatibility_children_are_preserved_without_fabrication(): - fallback = _records() +def test_unregistered_method_cannot_claim_native_compatibility(): model_payload = {"backend": "dummy", "model_name_or_path": "native"} method_payload = { "name": "direct", @@ -419,49 +418,48 @@ def test_native_compatibility_children_are_preserved_without_fabrication(): "prompt_id": short_id("prompt", prompt_payload), **prompt_payload, } - records = build_import_semantic_identity( - condition=replace( - _condition(), - generation_parameters={}, - unknown_reasons={ - **_condition().unknown_reasons, - "preflight_identity": "not applicable", - }, - ), - dataset=_dataset(), - model=replace( - _model(), - display_name="native", - backend="dummy", - effective_parameters={}, - unknown_reasons={"provider": "not applicable", "revision": "not applicable"}, - native_compatibility_identity=native_model, - ), - method=replace( - _method(), - name="direct", - effective_parameters={}, - implementation=method_payload["implementation"], - unknown_reasons={}, - native_compatibility_identity=native_method, - ), - prompt=replace( - _prompt(), - template_identity="v1", - template_digest=integrity_digest(prompt_payload["files"]), - template_contents={ - name: item["content"] for name, item in prompt_payload["files"].items() - }, - unknown_reason=None, - native_compatibility_identity=native_prompt, - ), - ) - - assert records.model["identity"] == model_payload - assert records.model["model_id"] == native_model["model_id"] - assert records.method["identity"] == method_payload - assert records.prompt["identity"] == prompt_payload - assert records.condition["condition_id"] != fallback.condition["condition_id"] + with pytest.raises(ImportIdentityError, match="registered current runner"): + build_import_semantic_identity( + condition=replace( + _condition(), + generation_parameters={}, + unknown_reasons={ + **_condition().unknown_reasons, + "preflight_identity": "not applicable", + }, + ), + dataset=_dataset(), + model=replace( + _model(), + display_name="native", + backend="dummy", + effective_parameters={}, + unknown_reasons={ + "provider": "not applicable", + "revision": "not applicable", + }, + native_compatibility_identity=native_model, + ), + method=replace( + _method(), + name="direct", + effective_parameters={}, + implementation=method_payload["implementation"], + unknown_reasons={}, + native_compatibility_identity=native_method, + ), + prompt=replace( + _prompt(), + template_identity="v1", + template_digest=integrity_digest(prompt_payload["files"]), + template_contents={ + name: item["content"] + for name, item in prompt_payload["files"].items() + }, + unknown_reason=None, + native_compatibility_identity=native_prompt, + ), + ) def test_native_model_identity_refuses_unbound_generation_parameters(): @@ -522,7 +520,9 @@ def test_native_child_claims_must_match_their_semantic_declarations(): "digest": integrity_digest(method_payload), "method_id": short_id("method", method_payload), } - with pytest.raises(ImportIdentityError, match="method.*declaration"): + with pytest.raises( + ImportIdentityError, match="method.*(?:declaration|registered|runner)" + ): build_import_semantic_identity( condition=replace( _condition(), @@ -711,7 +711,10 @@ def test_unknown_declarations_cannot_be_upgraded_by_native_claims(): "digest": integrity_digest(method_payload), "method_id": short_id("method", method_payload), } - with pytest.raises(ImportIdentityError, match="unknown.*native|implementation"): + with pytest.raises( + ImportIdentityError, + match="unknown.*native|implementation|registered current runner", + ): build_import_semantic_identity( condition=replace( _condition(), @@ -897,6 +900,7 @@ def _attach_result_origin( assignments: tuple[tuple[str, str], ...], authorization_digest: str | None = None, source_digest: str = "2" * 64, + overlay_source_digest: str = "e" * 64, ) -> None: derived = derivation_origin in {"repair_overlay", "offline_transformation"} components = [ @@ -904,7 +908,11 @@ def _attach_result_origin( operation_type=derivation_origin, question_id=question_id, parent_digests=(("d" * 64,) if derived else ()), - source_digests=(source_digest,), + source_digests=( + tuple(sorted({source_digest, overlay_source_digest})) + if derived + else (source_digest,) + ), authorization_digest=(authorization_digest if derived else None), implementation=_transform, parameters={}, @@ -938,6 +946,44 @@ def _attach_result_origin( ) +def _overlay_identity(identity: dict, derivation_origin: str) -> dict: + derived_components = [ + component + for component in identity["lineage_components"] + if component["identity"]["operation_type"] == derivation_origin + ] + offline = derivation_origin == "offline_transformation" + return { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": ( + integrity_digest( + [ + component["identity"]["input_digest"] + for component in derived_components + ] + ) + if offline + else None + ), + "preownership_output_digest": ( + integrity_digest( + [ + component["identity"]["preownership_output_digest"] + for component in derived_components + ] + ) + if offline + else None + ), + "implementation_digest": ( + integrity_digest(derived_components[0]["identity"]["implementation"]) + if offline + else None + ), + } + + def _realization_identity() -> dict: expected_dataset = _dataset() identity = { @@ -1080,6 +1126,7 @@ def _mutate_realization(identity: dict, field: str, value) -> None: ("q1", "external_repair_inference"), ), authorization_digest="c" * 64, + overlay_source_digest=value, ) elif field == "prediction_origins": _attach_result_origin( @@ -1262,19 +1309,7 @@ def test_base_repair_and_transformation_are_distinct_realizations_of_one_conditi identity["authorization_digest"] = "c" * 64 identity["parent_digests"]["realization_digests"] = ["d" * 64] identity["parent_digests"]["evidence_digests"] = ["0" * 64] - identity["overlay"] = { - "source_sha256": "e" * 64, - "replacement_digest": "f" * 64, - "transformation_input_digest": ( - "1" * 64 if derivation == "offline_transformation" else None - ), - "preownership_output_digest": ( - "2" * 64 if derivation == "offline_transformation" else None - ), - "implementation_digest": ( - "3" * 64 if derivation == "offline_transformation" else None - ), - } + identity["overlay"] = _overlay_identity(identity, derivation) identities.append( make_realization( condition_id=condition["condition_id"], @@ -1316,20 +1351,16 @@ def test_derived_realization_requires_authorization_parent_and_overlay(derivatio candidate["authorization_digest"] = "c" * 64 candidate["parent_digests"]["realization_digests"] = ["d" * 64] candidate["parent_digests"]["evidence_digests"] = ["0" * 64] - candidate["overlay"] = { - "source_sha256": "e" * 64, - "replacement_digest": "f" * 64, - "transformation_input_digest": "1" * 64, - "preownership_output_digest": "2" * 64, - "implementation_digest": "3" * 64, - } + candidate["overlay"] = _overlay_identity(candidate, derivation) if field == "parent": candidate["parent_digests"]["realization_digests"] = [] elif field == "evidence": candidate["parent_digests"]["evidence_digests"] = [] else: candidate[field] = None - with pytest.raises(ImportIdentityError, match="authorization|overlay|parent"): + with pytest.raises( + ImportIdentityError, match="authorization|overlay|parent|source" + ): make_realization( condition_id=condition["condition_id"], condition_digest=condition["condition_digest"], @@ -1804,6 +1835,49 @@ def test_native_dummy_cannot_discard_a_known_revision(): ) +def test_native_method_claim_must_match_registered_current_runner_identity(): + fake_implementation = { + "qualified_name": "historical.fake:Runner", + "source_file": "fake.py", + "source_digest": "a" * 64, + "callable_digest": "b" * 64, + } + payload = { + "name": "direct_mcq", + "effective_params": {}, + "preflight": None, + "implementation": fake_implementation, + } + method = replace( + _method(), + name="direct_mcq", + effective_parameters={}, + implementation=fake_implementation, + unknown_reasons={}, + native_compatibility_identity={ + "payload": payload, + "digest": integrity_digest(payload), + "method_id": short_id("method", payload), + }, + ) + condition = replace( + _condition(), + preflight_identity=None, + unknown_reasons={ + **_condition().unknown_reasons, + "preflight_identity": "not applicable", + }, + ) + with pytest.raises(ImportIdentityError, match="native method|registered|runner"): + build_import_semantic_identity( + condition=condition, + dataset=_dataset(), + model=_model(), + method=method, + prompt=_prompt(), + ) + + def test_native_prompt_identity_preserves_credential_shaped_literal_text(): contents = { "direct_mcq": "Answer the question; literal example password=alpha", @@ -1843,6 +1917,7 @@ def test_native_prompt_identity_preserves_credential_shaped_literal_text(): [ {"audit": {"operator": "alice"}, "node": "machine-a"}, {"hf_path": "/home/alice/machine-only/dataset"}, + {"generator": "/home/alice/generator.py"}, ], ) def test_native_dataset_source_refuses_audit_and_machine_metadata(source): @@ -2403,6 +2478,105 @@ def test_repair_row_itself_must_own_authorization_and_parent_lineage_edges(): ) +def test_derived_lineage_must_reference_declared_base_realization_digest(): + condition = _records().condition + identity = _realization_identity() + _attach_result_origin( + identity, + derivation_origin="repair_overlay", + assignments=( + ("q2", "external_historical_inference"), + ("q1", "external_repair_inference"), + ), + authorization_digest="c" * 64, + ) + for component in identity["lineage_components"]: + payload = { + **component["identity"], + "parent_digests": ["0" * 64], + } + component.update( + { + "lineage_id": short_id("lin", payload), + "lineage_digest": integrity_digest(payload), + "identity": payload, + } + ) + identity["result_origin"] = make_result_origin( + derivation_origin="repair_overlay", + row_assignments=tuple( + ( + component["identity"]["question_id"], + component["identity"]["prediction_origin"], + component["lineage_id"], + ) + for component in identity["lineage_components"] + ), + ) + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": None, + "preownership_output_digest": None, + "implementation_digest": None, + } + with pytest.raises(ImportIdentityError, match="base|realization|parent|lineage"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +@pytest.mark.parametrize("derivation", ["repair_overlay", "offline_transformation"]) +def test_derived_lineage_must_bind_overlay_operational_digests(derivation): + condition = _records().condition + identity = _realization_identity() + _attach_result_origin( + identity, + derivation_origin=derivation, + assignments=( + ("q2", "external_historical_inference"), + ( + "q1", + ( + "external_repair_inference" + if derivation == "repair_overlay" + else "external_historical_inference" + ), + ), + ), + authorization_digest="c" * 64, + ) + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = { + "source_sha256": "a" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": ( + "a" * 64 if derivation == "offline_transformation" else None + ), + "preownership_output_digest": ( + "b" * 64 if derivation == "offline_transformation" else None + ), + "implementation_digest": ( + "c" * 64 if derivation == "offline_transformation" else None + ), + } + with pytest.raises(ImportIdentityError, match="overlay|source|input|output|implementation"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + def test_realization_condition_is_bound_to_expected_dataset_selection(): condition = _records().condition foreign_identity = { From 175e4eecaefdc2c4ab126bcf7e092d2b8ef3885a Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 05:11:24 +0300 Subject: [PATCH 25/47] fix: require reachable transformation lineage --- src/choicebench/importing/identity.py | 51 ++++++++++ tests/importing/test_identity.py | 131 ++++++++++++++++++++++++++ 2 files changed, 182 insertions(+) diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index b198606..ac59288 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -659,6 +659,9 @@ def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: raise ImportIdentityError( "Expected dataset native source hf_dataset_info is invalid." ) + _validate_no_machine_path_values( + info, "Expected dataset native source hf_dataset_info" + ) if artifact["content_digest"] != dataset_content_digest(dataset.artifact_frame): raise ImportIdentityError( "Expected dataset artifact content digest does not own its frame." @@ -820,6 +823,40 @@ def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: derivation["declared_derivation"], "Expected dataset declared derivation", ) + if dataset.reference_kind == "profile_derived_reference_snapshot": + declared_derivation = derivation["declared_derivation"] + if not isinstance(declared_derivation, Mapping): + raise ImportIdentityError( + "Profile-derived reference declared derivation is invalid." + ) + raw_groups = declared_derivation.get("independent_source_groups") + if not isinstance(raw_groups, (list, tuple)): + raise ImportIdentityError( + "Profile-derived reference requires at least two independent " + "source groups." + ) + groups: list[tuple[str, ...]] = [] + for item in raw_groups: + if not isinstance(item, (list, tuple)) or not item: + raise ImportIdentityError( + "Profile-derived reference requires at least two independent " + "source groups." + ) + groups.append(tuple(str(source_id) for source_id in item)) + if len(groups) < 2: + raise ImportIdentityError( + "Profile-derived reference requires at least two independent " + "source groups." + ) + flattened = [source_id for group in groups for source_id in group] + if ( + len(flattened) != len(set(flattened)) + or set(flattened) != seen_source_ids + ): + raise ImportIdentityError( + "Profile-derived reference independent source groups must be " + "disjoint and cover every declared source." + ) expected_derivation_digest = integrity_digest(derivation) if dataset.derivation_digest != expected_derivation_digest: raise ImportIdentityError("Expected dataset derivation digest is inconsistent.") @@ -2221,6 +2258,20 @@ def _validate_realization_identity( f"{derivation_origin} realization requires a matching derived " "lineage component." ) + assigned_lineage_ids = { + assignment["prediction_lineage_id"] + for assignment in result_origin["row_assignments"] + } + unassigned_derived_ids = sorted( + component["lineage_id"] + for component in derived_components + if component["lineage_id"] not in assigned_lineage_ids + ) + if unassigned_derived_ids: + raise ImportIdentityError( + f"{derivation_origin} lineage component(s) {unassigned_derived_ids} " + "do not own a result-origin row assignment." + ) for component in derived_components: component_identity = component["identity"] if component_identity["authorization_digest"] != authorization_digest: diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py index d2a15f1..296491d 100644 --- a/tests/importing/test_identity.py +++ b/tests/importing/test_identity.py @@ -1753,6 +1753,44 @@ def test_expected_dataset_trust_chain_is_owned_and_recomputed(mutation): ) +def test_profile_derived_dataset_requires_independent_source_groups_at_identity_boundary(): + dataset = _dataset() + derivation = { + **dataset.derivation, + "reference_kind": "profile_derived_reference_snapshot", + "trust_label": "cross-source-internal-consistency", + "declared_derivation": {"independent_source_groups": [["benchmark"]]}, + } + derivation_digest = integrity_digest(derivation) + snapshot_digest = integrity_digest( + { + "artifact_digest": dataset.artifact_digest, + "selection_digest": dataset.selection_digest, + "question_set_digest": dataset.question_set_digest, + "derivation_digest": derivation_digest, + "reference_kind": "profile_derived_reference_snapshot", + "trust_label": "cross-source-internal-consistency", + } + ) + forged = replace( + dataset, + reference_kind="profile_derived_reference_snapshot", + trust_label="cross-source-internal-consistency", + derivation=derivation, + derivation_digest=derivation_digest, + snapshot_digest=snapshot_digest, + ) + + with pytest.raises(ImportIdentityError, match="profile|independent|group|source"): + build_import_semantic_identity( + condition=_condition(), + dataset=forged, + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + def test_complete_realization_question_set_is_owned_by_validated_dataset(): condition = _records().condition identity = _realization_identity() @@ -1918,6 +1956,13 @@ def test_native_prompt_identity_preserves_credential_shaped_literal_text(): {"audit": {"operator": "alice"}, "node": "machine-a"}, {"hf_path": "/home/alice/machine-only/dataset"}, {"generator": "/home/alice/generator.py"}, + { + "hf_dataset_info": { + "builder_name": "/home/alice/private-builder", + "config_name": "cfg", + "version": "1", + } + }, ], ) def test_native_dataset_source_refuses_audit_and_machine_metadata(source): @@ -1938,6 +1983,8 @@ def test_native_dataset_source_refuses_audit_and_machine_metadata(source): } artifact_id = short_id("ds", artifact_payload) selection_payload = {**dataset.selection_payload, "artifact_id": artifact_id} + derivation = {**dataset.derivation, "identity_mode": "native_compatibility"} + derivation_digest = integrity_digest(derivation) forged = replace( dataset, identity_mode="native_compatibility", @@ -1947,6 +1994,18 @@ def test_native_dataset_source_refuses_audit_and_machine_metadata(source): selection_payload=selection_payload, selection_digest=integrity_digest(selection_payload), selection_id=short_id("sel", selection_payload), + derivation=derivation, + derivation_digest=derivation_digest, + snapshot_digest=integrity_digest( + { + "artifact_digest": integrity_digest(artifact_payload), + "selection_digest": integrity_digest(selection_payload), + "question_set_digest": dataset.question_set_digest, + "derivation_digest": derivation_digest, + "reference_kind": dataset.reference_kind, + "trust_label": dataset.trust_label, + } + ), ) with pytest.raises(ImportIdentityError, match="source|audit|path|timestamp"): build_import_semantic_identity( @@ -2577,6 +2636,78 @@ def test_derived_lineage_must_bind_overlay_operational_digests(derivation): ) +def test_offline_transformation_rows_must_reference_offline_lineage_component(): + condition = _records().condition + identity = _realization_identity() + assigned_components = [ + make_lineage_component( + operation_type="external_import", + question_id=question_id, + parent_digests=("d" * 64,), + source_digests=("2" * 64,), + authorization_digest="c" * 64, + implementation=_transform, + parameters={}, + input_digest=integrity_digest({"question_id": question_id, "input": "base"}), + preownership_output_digest=integrity_digest( + {"question_id": question_id, "output": "base"} + ), + prediction_origin="external_historical_inference", + ) + for question_id in ("q2", "q1") + ] + orphan_component = make_lineage_component( + operation_type="offline_transformation", + question_id="q1", + parent_digests=("d" * 64,), + source_digests=("2" * 64, "e" * 64), + authorization_digest="c" * 64, + implementation=_transform, + parameters={}, + input_digest=integrity_digest({"question_id": "q1", "input": "transform"}), + preownership_output_digest=integrity_digest( + {"question_id": "q1", "output": "transform"} + ), + prediction_origin="external_historical_inference", + ) + identity["lineage_components"] = [*assigned_components, orphan_component] + identity["result_origin"] = make_result_origin( + derivation_origin="offline_transformation", + row_assignments=tuple( + ( + component["identity"]["question_id"], + component["identity"]["prediction_origin"], + component["lineage_id"], + ) + for component in assigned_components + ), + ) + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": integrity_digest( + [orphan_component["identity"]["input_digest"]] + ), + "preownership_output_digest": integrity_digest( + [orphan_component["identity"]["preownership_output_digest"]] + ), + "implementation_digest": integrity_digest( + orphan_component["identity"]["implementation"] + ), + } + + with pytest.raises(ImportIdentityError, match="offline|row|lineage|assignment"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + def test_realization_condition_is_bound_to_expected_dataset_selection(): condition = _records().condition foreign_identity = { From ee2108b479e9bdaec410c39015052c27937b9071 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 13:40:43 +0300 Subject: [PATCH 26/47] fix: close final Task 5 identity-contract gaps Eighth-cycle review found six residual identity gaps: - importer-core code identity was never bound into importer_implementation_identity; only adapter/validator were, so an identity.py change was invisible to realization identity. - declared sources and lineage components could be unreachable: nothing required every source to be referenced by a lineage component's source_digests, or every lineage component to own a result-origin row. - non-derived realizations (native_execution/external_import) could carry lineage components whose operation_type didn't match the declared derivation_origin, smuggling repair/transformation-labeled lineage past the authorization/overlay gates that only trigger off the top-level derivation_origin. - source provenance (source_run_id/source_repository/source_commit) could be null with no explicit reason, unlike every other nullable provenance field in this module. - runtime callable identity hashed only source/defaults, so a behavior-affecting module-level global the callable reads could change without changing the callable's identity. - ExpectedDataset derivation revision/fingerprint were accepted unvalidated, letting audit paths, timestamps, or machine-local values enter stable dataset identity. Adds one failing regression test per confirmed gap, then the narrowest production fix: bind importer-core via _runtime_callable_record; require every declared source/lineage component to be reachable from a result-origin row; reject non-matching operation_type on non-derived realizations; require unknown_reasons for null source provenance fields (reusing the existing _validate_unknown_reason_contract); bind resolved non-callable globals into the runtime callable digest; validate revision/fingerprint with the existing _validate_stable_values guard. Focused gate (tests/test_condition_grid.py + tests/importing/test_identity.py): 159 passed. Full suite: 1195 passed, only the pre-existing Python 3.14 google.genai.types deprecation warning. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/importing/identity.py | 93 +++++++++++- tests/importing/test_identity.py | 197 ++++++++++++++++++++++++++ 2 files changed, 284 insertions(+), 6 deletions(-) diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index ac59288..36146ac 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -99,6 +99,12 @@ class ImportSemanticRecords: "source_commit", "notes_digest", "evidence_digest", + "unknown_reasons", +} +_SOURCE_PROVENANCE_UNKNOWN_FIELDS = { + "source_run_id", + "source_repository", + "source_commit", } _PARSING_POLICY_KEYS = { "dialect", @@ -819,6 +825,12 @@ def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: for key, value in derivation["columns"].items() ): raise ImportIdentityError("Expected dataset derivation columns are invalid.") + _validate_stable_values( + derivation["revision"], "Expected dataset derivation revision" + ) + _validate_stable_values( + derivation["fingerprint"], "Expected dataset derivation fingerprint" + ) _validate_stable_values( derivation["declared_derivation"], "Expected dataset declared derivation", @@ -1407,6 +1419,31 @@ def _validate_digest(value: Any, field: str, *, optional: bool = False) -> str | return value +def _resolved_global_bindings(target: Callable[..., Any]) -> dict[str, Any]: + """Bind behavior-affecting global values the callable's code resolves by name. + + Excludes modules, other callables, and dunder names: those are either + irrelevant runtime state or already covered by the callable's own source + digest, not the specific value a global happened to hold. + """ + code = getattr(target, "__code__", None) + global_ns = getattr(target, "__globals__", None) + if code is None or global_ns is None: + return {} + bindings: dict[str, Any] = {} + for name in sorted(set(code.co_names)): + if name.startswith("__") or name not in global_ns: + continue + value = global_ns[name] + if inspect.ismodule(value) or callable(value): + continue + try: + bindings[name] = canonicalize(value) + except (TypeError, ValueError): + continue + return bindings + + def _runtime_callable_record(target: Callable[..., Any], where: str) -> dict[str, Any]: if inspect.ismethod(target) or not ( inspect.isfunction(target) or inspect.isclass(target) @@ -1432,6 +1469,7 @@ def _runtime_callable_record(target: Callable[..., Any], where: str) -> dict[str "source": callable_source, "defaults": getattr(target, "__defaults__", None), "keyword_defaults": getattr(target, "__kwdefaults__", None), + "resolved_globals": _resolved_global_bindings(target), } try: callable_digest = integrity_digest(callable_payload) @@ -1789,16 +1827,20 @@ def _validate_lineage_component_record( def importer_implementation_identity( *, adapter: Callable[..., Any], validator: Callable[..., Any] ) -> dict[str, Any]: - """Bind installed ChoiceBench plus registered adapter and validator code.""" + """Bind installed ChoiceBench plus importer-core, adapter, and validator code.""" records: dict[str, Mapping[str, Any]] = {} for role, target in (("adapter", adapter), ("validator", validator)): records[role] = _runtime_callable_record( target, f"Importer {role} implementation" ) + importer_record = _runtime_callable_record( + _identity_record, "Importer core implementation" + ) return canonicalize( { "schema_version": "choicebench.importer-implementation.v1", "package": {"name": "choicebench", "version": __version__}, + "importer": importer_record, "adapter": records["adapter"], "validator": records["validator"], } @@ -1913,6 +1955,11 @@ def _validate_realization_identity( _validate_digest( provenance[field], f"sources[{index}].provenance.{field}", optional=True ) + _validate_unknown_reason_contract( + {field: provenance[field] for field in _SOURCE_PROVENANCE_UNKNOWN_FIELDS}, + provenance["unknown_reasons"], + f"Realization sources[{index}].provenance", + ) source = {**source, "provenance": canonicalize(provenance)} seen_source_ids.add(source_id) normalized_sources.append(canonicalize(source)) @@ -1970,7 +2017,7 @@ def _validate_realization_identity( ) importer = _exact_mapping( raw_importer, - {"schema_version", "package", "adapter", "validator"}, + {"schema_version", "package", "importer", "adapter", "validator"}, "Realization importer_implementation", ) if importer["schema_version"] != "choicebench.importer-implementation.v1": @@ -1980,7 +2027,7 @@ def _validate_realization_identity( ) if package != {"name": "choicebench", "version": __version__}: raise ImportIdentityError("Realization importer package identity is invalid.") - for role in ("adapter", "validator"): + for role in ("importer", "adapter", "validator"): try: component = validate_implementation_identity_record(importer[role]) except ImportSpecError as exc: @@ -2150,6 +2197,16 @@ def _validate_realization_identity( raise ImportIdentityError( "Realization result origin conflicts with its lineage component." ) + assigned_lineage_ids = { + assignment["prediction_lineage_id"] + for assignment in result_origin["row_assignments"] + } + unreachable_lineage_ids = sorted(set(lineage_by_id) - assigned_lineage_ids) + if unreachable_lineage_ids: + raise ImportIdentityError( + f"Realization lineage component(s) {unreachable_lineage_ids} are not " + "reachable from any result-origin row." + ) parents = _exact_mapping( raw["parent_digests"], {"realization_digests", "evidence_digests", "result_digests"}, @@ -2218,6 +2275,16 @@ def _validate_realization_identity( raise ImportIdentityError( "Realization lineage source edge is not owned by the realization." ) + unreachable_source_digests = ( + sorted(declared_source_digests - lineage_source_edges) + if lineage_components + else [] + ) + if unreachable_source_digests: + raise ImportIdentityError( + f"Realization source(s) {unreachable_source_digests} are not reachable " + "from any lineage component." + ) derivation_origin = result_origin["derivation_origin"] derived = derivation_origin in {"repair_overlay", "offline_transformation"} @@ -2365,10 +2432,24 @@ def _validate_realization_identity( raise ImportIdentityError( "repair_overlay requires at least one repair prediction origin." ) - elif authorization_digest is not None or normalized_overlay is not None: - raise ImportIdentityError( - f"{derivation_origin} realization cannot claim repair authorization or overlay." + else: + if authorization_digest is not None or normalized_overlay is not None: + raise ImportIdentityError( + f"{derivation_origin} realization cannot claim repair authorization " + "or overlay." + ) + incompatible_operations = sorted( + { + component["identity"]["operation_type"] + for component in lineage_components + if component["identity"]["operation_type"] != derivation_origin + } ) + if incompatible_operations: + raise ImportIdentityError( + f"{derivation_origin} realization lineage components have " + f"incompatible operation type(s) {incompatible_operations}." + ) normalized = { "import_spec_digest": raw["import_spec_digest"], diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py index 296491d..acc13f1 100644 --- a/tests/importing/test_identity.py +++ b/tests/importing/test_identity.py @@ -49,6 +49,13 @@ def _validator(value: str) -> bool: return bool(value) +_GLOBAL_VALIDATOR_RESULT = True + + +def _global_validator(value: str) -> bool: + return bool(value) and _GLOBAL_VALIDATOR_RESULT + + def _transform(value: str) -> str: return value.strip() @@ -1002,6 +1009,11 @@ def _realization_identity() -> dict: "source_commit": None, "notes_digest": None, "evidence_digest": None, + "unknown_reasons": { + "source_run_id": "not recorded", + "source_repository": "not recorded", + "source_commit": "not recorded", + }, }, } ], @@ -1753,6 +1765,38 @@ def test_expected_dataset_trust_chain_is_owned_and_recomputed(mutation): ) +def test_expected_dataset_derivation_revision_refuses_nested_audit_metadata(): + dataset = _dataset() + derivation = { + **dataset.derivation, + "revision": {"audit_path": "/home/alice/source", "timestamp": "now"}, + } + derivation_digest = integrity_digest(derivation) + forged = replace( + dataset, + derivation=derivation, + derivation_digest=derivation_digest, + snapshot_digest=integrity_digest( + { + "artifact_digest": dataset.artifact_digest, + "selection_digest": dataset.selection_digest, + "question_set_digest": dataset.question_set_digest, + "derivation_digest": derivation_digest, + "reference_kind": dataset.reference_kind, + "trust_label": dataset.trust_label, + } + ), + ) + with pytest.raises(ImportIdentityError, match="revision|audit|path|timestamp"): + build_import_semantic_identity( + condition=_condition(), + dataset=forged, + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + def test_profile_derived_dataset_requires_independent_source_groups_at_identity_boundary(): dataset = _dataset() derivation = { @@ -2259,6 +2303,30 @@ def test_runtime_identity_distinguishes_same_named_top_level_function_objects(): assert first != second +def test_runtime_identity_binds_behavior_affecting_resolved_globals(monkeypatch): + first = importer_implementation_identity( + adapter=_adapter, + validator=_global_validator, + ) + monkeypatch.setitem(_global_validator.__globals__, "_GLOBAL_VALIDATOR_RESULT", False) + second = importer_implementation_identity( + adapter=_adapter, + validator=_global_validator, + ) + assert first != second + + +def test_importer_implementation_identity_binds_importer_core_code(): + implementation = importer_implementation_identity( + adapter=_adapter, + validator=_validator, + ) + assert implementation["importer"]["source_digest"] + assert implementation["importer"]["qualified_name"].startswith( + "choicebench.importing.identity:" + ) + + def test_lineage_supports_nonexecuted_declared_external_implementation_identity(): source_digest = "2" * 64 component = make_lineage_component( @@ -2708,6 +2776,135 @@ def test_offline_transformation_rows_must_reference_offline_lineage_component(): ) +def test_realization_orphan_lineage_component_is_unreachable_from_any_row(): + condition = _records().condition + identity = _realization_identity() + orphan_component = make_lineage_component( + operation_type="external_import", + question_id="q3", + parent_digests=(), + source_digests=("2" * 64,), + authorization_digest=None, + implementation=_transform, + parameters={}, + input_digest=integrity_digest({"question_id": "q3", "stage": "input"}), + preownership_output_digest=integrity_digest( + { + "question_id": "q3", + "prediction_origin": "external_historical_inference", + "stage": "preownership-output", + } + ), + prediction_origin="external_historical_inference", + ) + identity["lineage_components"] = [*identity["lineage_components"], orphan_component] + with pytest.raises(ImportIdentityError, match="reachable|row"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_declared_source_must_be_reachable_from_lineage(): + condition = _records().condition + identity = _realization_identity() + identity["sources"] = [ + *identity["sources"], + { + "source_id": "unused", + "logical_path": "publisher/unused.csv", + "classification": "canonical", + "format": "csv", + "format_version": "1", + "sha256": "9" * 64, + "provenance": { + "source_run_id": None, + "source_repository": None, + "source_commit": None, + "notes_digest": None, + "evidence_digest": None, + "unknown_reasons": { + "source_run_id": "not recorded", + "source_repository": "not recorded", + "source_commit": "not recorded", + }, + }, + }, + ] + with pytest.raises(ImportIdentityError, match="reachable|source"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_lineage_operation_type_must_match_non_derived_derivation_origin(): + condition = _records().condition + identity = _realization_identity() + mismatched_component = make_lineage_component( + operation_type="repair_overlay", + question_id="q1", + parent_digests=(), + source_digests=("2" * 64,), + authorization_digest=None, + implementation=_transform, + parameters={}, + input_digest=integrity_digest({"question_id": "q1", "stage": "input"}), + preownership_output_digest=integrity_digest( + { + "question_id": "q1", + "prediction_origin": "external_historical_inference", + "stage": "preownership-output", + } + ), + prediction_origin="external_historical_inference", + ) + remaining = [ + component + for component in identity["lineage_components"] + if component["identity"]["question_id"] != "q1" + ] + identity["lineage_components"] = [*remaining, mismatched_component] + identity["result_origin"] = make_result_origin( + derivation_origin="external_import", + row_assignments=tuple( + ( + component["identity"]["question_id"], + component["identity"]["prediction_origin"], + component["lineage_id"], + ) + for component in identity["lineage_components"] + ), + ) + with pytest.raises(ImportIdentityError, match="operation type|incompatible"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_source_provenance_requires_explicit_unknown_reason(): + condition = _records().condition + identity = _realization_identity() + identity["sources"][0]["provenance"]["unknown_reasons"] = { + "source_run_id": "not recorded", + "source_repository": "not recorded", + } + with pytest.raises(ImportIdentityError, match="unknown reasons|provenance"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + def test_realization_condition_is_bound_to_expected_dataset_selection(): condition = _records().condition foreign_identity = { From 1c6625c833ea4ef396d6a8124c8e8adf39f13dd8 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 14:25:38 +0300 Subject: [PATCH 27/47] fix: bind cross-module helper globals into callable identity An adversarial re-review of ee2108b found that the resolved-globals fix (finding #5) still had a gap: it skipped every callable global on the theory that a helper's behavior is "already covered by the callable's own source digest." That's false whenever the helper lives in a different file than the target: implementation_identity() (and therefore _runtime_callable_record) only hashes the target's own defining file, so a cross-module helper's source was never hashed by anything. Verified with a live repro: an adapter calling a helper imported from another module produced a byte-identical identity record across two genuinely different helper implementations. Fix: function/class globals are now bound via a shallow implementation_identity() call (their own defining-file digest), not expanded further, so mutually-referencing helpers can't cause unbounded or cyclic recursion. Also fail closed (raise ImportIdentityError) instead of silently dropping a resolved global that can't be represented or is an uninspectable stateful callable, matching how this module already treats unsupported callable defaults elsewhere. Removes a redundant assigned_lineage_ids recomputation the same review flagged as a trivial duplication (Minor). Adds tests/importing/_cross_module_helpers.py (two helper functions in a separate module) plus a regression test proving the callable identity now changes when the referenced cross-module helper changes, and a test proving an uninspectable resolved global is rejected rather than ignored. Focused gate (tests/test_condition_grid.py + tests/importing/test_identity.py): 161 passed. Full suite: 1197 passed, only the pre-existing Python 3.14 google.genai.types deprecation warning. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/importing/identity.py | 44 +++++++++++++++++------- tests/importing/_cross_module_helpers.py | 6 ++++ tests/importing/test_identity.py | 39 +++++++++++++++++++++ 3 files changed, 77 insertions(+), 12 deletions(-) create mode 100644 tests/importing/_cross_module_helpers.py diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index 36146ac..b8248da 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -1419,12 +1419,18 @@ def _validate_digest(value: Any, field: str, *, optional: bool = False) -> str | return value -def _resolved_global_bindings(target: Callable[..., Any]) -> dict[str, Any]: +def _resolved_global_bindings(target: Callable[..., Any], where: str) -> dict[str, Any]: """Bind behavior-affecting global values the callable's code resolves by name. - Excludes modules, other callables, and dunder names: those are either - irrelevant runtime state or already covered by the callable's own source - digest, not the specific value a global happened to hold. + Function/class globals (helpers) are bound by their own defining-file + identity rather than skipped: `implementation_identity` only hashes a + callable's own defining file, so a helper imported from another module is + never covered by the calling callable's own source digest. Binding it + shallowly (not expanding its own resolved globals) closes that gap without + risking unbounded or cyclic expansion through mutually referencing + helpers. Modules and dunder names are excluded as irrelevant runtime + state; anything else that cannot be represented stably fails closed + rather than being silently dropped. """ code = getattr(target, "__code__", None) global_ns = getattr(target, "__globals__", None) @@ -1435,12 +1441,30 @@ def _resolved_global_bindings(target: Callable[..., Any]) -> dict[str, Any]: if name.startswith("__") or name not in global_ns: continue value = global_ns[name] - if inspect.ismodule(value) or callable(value): + if value is target or inspect.ismodule(value): continue + if inspect.isfunction(value) or inspect.isclass(value): + try: + bindings[name] = validate_implementation_identity_record( + implementation_identity(value) + ) + except (ImportSpecError, OSError, TypeError, ValueError) as exc: + raise ImportIdentityError( + f"{where} references global {name!r} with no inspectable " + "implementation identity." + ) from exc + continue + if callable(value): + raise ImportIdentityError( + f"{where} references global {name!r} that is an uninspectable " + "stateful callable." + ) try: bindings[name] = canonicalize(value) - except (TypeError, ValueError): - continue + except (TypeError, ValueError) as exc: + raise ImportIdentityError( + f"{where} references global {name!r} with an unsupported value." + ) from exc return bindings @@ -1469,7 +1493,7 @@ def _runtime_callable_record(target: Callable[..., Any], where: str) -> dict[str "source": callable_source, "defaults": getattr(target, "__defaults__", None), "keyword_defaults": getattr(target, "__kwdefaults__", None), - "resolved_globals": _resolved_global_bindings(target), + "resolved_globals": _resolved_global_bindings(target, where), } try: callable_digest = integrity_digest(callable_payload) @@ -2325,10 +2349,6 @@ def _validate_realization_identity( f"{derivation_origin} realization requires a matching derived " "lineage component." ) - assigned_lineage_ids = { - assignment["prediction_lineage_id"] - for assignment in result_origin["row_assignments"] - } unassigned_derived_ids = sorted( component["lineage_id"] for component in derived_components diff --git a/tests/importing/_cross_module_helpers.py b/tests/importing/_cross_module_helpers.py new file mode 100644 index 0000000..2db247b --- /dev/null +++ b/tests/importing/_cross_module_helpers.py @@ -0,0 +1,6 @@ +def helper_v1(value: str) -> str: + return value.strip() + + +def helper_v2(value: str) -> str: + return "tampered-" + value diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py index acc13f1..6dfe6b6 100644 --- a/tests/importing/test_identity.py +++ b/tests/importing/test_identity.py @@ -39,6 +39,7 @@ OptionMappingSpec, ResultOriginSpec, ) +from tests.importing import _cross_module_helpers def _adapter(value: str) -> str: @@ -56,6 +57,13 @@ def _global_validator(value: str) -> bool: return bool(value) and _GLOBAL_VALIDATOR_RESULT +_cross_module_helper = _cross_module_helpers.helper_v1 + + +def _cross_module_adapter(value: str) -> str: + return _cross_module_helper(value) + + def _transform(value: str) -> str: return value.strip() @@ -2316,6 +2324,37 @@ def test_runtime_identity_binds_behavior_affecting_resolved_globals(monkeypatch) assert first != second +def test_runtime_identity_binds_cross_module_helper_globals(monkeypatch): + monkeypatch.setitem( + _cross_module_adapter.__globals__, + "_cross_module_helper", + _cross_module_helpers.helper_v1, + ) + first = importer_implementation_identity( + adapter=_cross_module_adapter, + validator=_validator, + ) + monkeypatch.setitem( + _cross_module_adapter.__globals__, + "_cross_module_helper", + _cross_module_helpers.helper_v2, + ) + second = importer_implementation_identity( + adapter=_cross_module_adapter, + validator=_validator, + ) + assert first != second + + +def test_runtime_identity_refuses_uninspectable_resolved_global(monkeypatch): + monkeypatch.setitem(_global_validator.__globals__, "_GLOBAL_VALIDATOR_RESULT", object()) + with pytest.raises(ImportIdentityError, match="global|unsupported"): + importer_implementation_identity( + adapter=_adapter, + validator=_global_validator, + ) + + def test_importer_implementation_identity_binds_importer_core_code(): implementation = importer_implementation_identity( adapter=_adapter, From 3a69eda92cda2979414061006f46d424ffb88d66 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 14:49:22 +0300 Subject: [PATCH 28/47] perf: cache implementation_identity for resolved-global bindings The final adversarial re-review of the cross-module-helper fix (1c6625c) found an Important performance liability: implementation_identity() walks and hashes a non-choicebench package's entire file tree, and the resolved- globals binding called it uncached for every function/class-valued global a callable references. A live repro showed a bare `from pandas import isna` costing ~0.3s per identity computation, repeated with zero memoization on every call. Fix: memoize implementation_identity's validated output per referenced object via functools.lru_cache. Functions/classes hash by identity, and their defining files don't change within a process, so caching by object identity is safe. Verified the exact reproducer now costs ~0.3s on the first call and ~9 microseconds on the second. No behavior change: same identity records, same rejection paths, just memoized. Focused gate (tests/test_condition_grid.py + tests/importing/test_identity.py): 161 passed in ~5m20s (down from ~9m24s). Full suite: 1197 passed in ~6m54s (down from ~10m44s), only the pre-existing Python 3.14 google.genai.types deprecation warning. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/importing/identity.py | 20 +++++++++++++++++--- 1 file changed, 17 insertions(+), 3 deletions(-) diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index b8248da..11ee54d 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -3,6 +3,7 @@ from __future__ import annotations from dataclasses import dataclass +import functools import inspect import math from pathlib import PurePosixPath @@ -1419,6 +1420,21 @@ def _validate_digest(value: Any, field: str, *, optional: bool = False) -> str | return value +@functools.lru_cache(maxsize=None) +def _cached_global_implementation_identity(value: Any) -> dict[str, Any]: + """Cache `implementation_identity` per referenced global object. + + `implementation_identity` walks and hashes a non-`choicebench` package's + entire file tree; without caching, every callable that references the + same third-party helper (e.g. a bare `from pandas import isna`) would + repeat that walk on every identity computation. The referenced object's + defining files do not change within a process, so caching by object + identity (functions/classes hash by identity) is safe and keeps repeated + identity computations cheap. + """ + return validate_implementation_identity_record(implementation_identity(value)) + + def _resolved_global_bindings(target: Callable[..., Any], where: str) -> dict[str, Any]: """Bind behavior-affecting global values the callable's code resolves by name. @@ -1445,9 +1461,7 @@ def _resolved_global_bindings(target: Callable[..., Any], where: str) -> dict[st continue if inspect.isfunction(value) or inspect.isclass(value): try: - bindings[name] = validate_implementation_identity_record( - implementation_identity(value) - ) + bindings[name] = _cached_global_implementation_identity(value) except (ImportSpecError, OSError, TypeError, ValueError) as exc: raise ImportIdentityError( f"{where} references global {name!r} with no inspectable " From 965dc77977e4c36884c3f29eca4fac646b8b181c Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 15:26:19 +0300 Subject: [PATCH 29/47] feat: add additive manifest-v3 schema alongside v2 Reduced-scope Task 6: adds PROTOCOL_V3_VERSION/MANIFEST_V3_SCHEMA_VERSION/ RUN_STATE_V3_SCHEMA_VERSION and make_manifest_v3/validate_manifest_v3/ initial_run_state_v3/validate_run_state_v3/load_run_state_v3, dispatching strictly on exact schema_version. Every existing v2 function/constant is untouched. Deliberately narrower than the original plan: no ManifestView cross-version abstraction and no normalize_manifest() legacy-v2 synthesis. Every manifest this module builds is a fresh v3 import (a base Stage 1 import or a repair overlay import); nothing here ever needs to read an existing v2 run through a unified interface. That unification would only matter for migrating ChoiceBench's own native runner to v3, which is out of scope for finishing the importer. v3 payload holds only semantic_conditions/realizations tables (keyed by condition_id/realization_id, built by the Task 5 identity module) plus protocol/canonicalization versions; audit is a separate top-level field, excluded from the experiment digest. validate_manifest_v3 recomputes every condition/realization ID and digest from its own identity payload rather than trusting the stored value, and requires each realization's declared condition_digest to match its owning semantic condition. Focused: 13 new tests passed. Full suite: 1210 passed (1197 baseline + 13), only the pre-existing Python 3.14 google.genai.types deprecation warning. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/manifest.py | 154 ++++++++++++++++++++- tests/importing/test_manifest_v3.py | 203 ++++++++++++++++++++++++++++ 2 files changed, 356 insertions(+), 1 deletion(-) create mode 100644 tests/importing/test_manifest_v3.py diff --git a/src/choicebench/manifest.py b/src/choicebench/manifest.py index 46205f8..2996854 100644 --- a/src/choicebench/manifest.py +++ b/src/choicebench/manifest.py @@ -9,7 +9,7 @@ from datetime import datetime, timezone from importlib.metadata import PackageNotFoundError, version from pathlib import Path -from typing import Any +from typing import Any, Mapping from choicebench import __version__ from choicebench.identity import CANONICALIZATION_VERSION, canonicalize, integrity_digest, short_id @@ -21,6 +21,16 @@ MANIFEST_FILENAME = "manifest.json" RUN_STATE_FILENAME = "run_state.json" +PROTOCOL_V3_VERSION = "choicebench.protocol.v3" +MANIFEST_V3_SCHEMA_VERSION = "choicebench.manifest.v3" +RUN_STATE_V3_SCHEMA_VERSION = "choicebench.run-state.v3" +_MANIFEST_V3_PAYLOAD_KEYS = { + "protocol_version", + "canonicalization_version", + "semantic_conditions", + "realizations", +} + class ManifestCompatibilityError(RuntimeError): pass @@ -304,3 +314,145 @@ def validate_run_state(state: dict[str, Any], manifest: dict[str, Any]) -> None: def write_run_state(run_dir: Path, state: dict[str, Any]) -> None: path = Path(run_dir) / RUN_STATE_FILENAME atomic_write_json(path, canonicalize(state, redact_secrets=False)) + + +def make_manifest_v3( + payload: Mapping[str, Any], *, audit: Mapping[str, Any] | None = None +) -> dict[str, Any]: + """Build an immutable v3 manifest from semantic-condition/realization tables. + + Additive alongside make_manifest(); v2 manifests are unaffected. `payload` + holds only identity-bearing content (no audit fields), so the experiment + digest is simply the payload digest -- unlike v2 there is no "config" key + to exclude. + """ + if set(payload) != _MANIFEST_V3_PAYLOAD_KEYS: + raise ManifestCompatibilityError("Manifest v3 payload fields are invalid.") + payload = canonicalize(dict(payload)) + if payload["protocol_version"] != PROTOCOL_V3_VERSION: + raise ManifestCompatibilityError( + f"Unsupported protocol version {payload['protocol_version']!r}; " + f"expected {PROTOCOL_V3_VERSION!r}." + ) + if payload["canonicalization_version"] != CANONICALIZATION_VERSION: + raise ManifestCompatibilityError("Unsupported canonicalization version.") + experiment_digest = integrity_digest(payload) + return { + "schema_version": MANIFEST_V3_SCHEMA_VERSION, + "experiment_id": f"exp_{experiment_digest[:16]}", + "experiment_digest": experiment_digest, + "payload_digest": integrity_digest(payload), + "created_at": datetime.now(timezone.utc).isoformat(), + "payload": payload, + "audit": canonicalize(dict(audit or {})), + } + + +def validate_manifest_v3(manifest: Mapping[str, Any]) -> None: + """Validate a v3 manifest, recomputing every child ID/digest from its + own identity payload rather than trusting the stored value.""" + if manifest.get("schema_version") != MANIFEST_V3_SCHEMA_VERSION: + raise ManifestCompatibilityError("Unsupported or missing manifest schema version.") + payload = manifest.get("payload") + if not isinstance(payload, dict) or set(payload) != _MANIFEST_V3_PAYLOAD_KEYS: + raise ManifestCompatibilityError("Manifest v3 payload fields are invalid.") + if payload.get("protocol_version") != PROTOCOL_V3_VERSION: + raise ManifestCompatibilityError( + f"Unsupported protocol version {payload.get('protocol_version')!r}; " + f"expected {PROTOCOL_V3_VERSION!r}." + ) + if payload.get("canonicalization_version") != CANONICALIZATION_VERSION: + raise ManifestCompatibilityError("Unsupported canonicalization version.") + if manifest.get("payload_digest") != integrity_digest(payload): + raise ManifestCompatibilityError("Manifest payload integrity check failed.") + experiment_digest = integrity_digest(payload) + experiment_id = f"exp_{experiment_digest[:16]}" + if ( + manifest.get("experiment_digest") != experiment_digest + or manifest.get("experiment_id") != experiment_id + ): + raise ManifestCompatibilityError( + "Manifest integrity check failed; payload or identity was modified." + ) + + semantic_conditions = payload["semantic_conditions"] + if not isinstance(semantic_conditions, dict): + raise ManifestCompatibilityError("Manifest semantic_conditions must be a mapping.") + for condition_id, record in semantic_conditions.items(): + if not isinstance(record, dict): + raise ManifestCompatibilityError("Manifest condition record is invalid.") + recomputed_id = short_id("cond", record.get("identity")) + recomputed_digest = integrity_digest(record.get("identity")) + if ( + condition_id != record.get("condition_id") + or condition_id != recomputed_id + or record.get("condition_digest") != recomputed_digest + ): + raise ManifestCompatibilityError( + f"Manifest condition {condition_id!r} ID/digest do not match its " + "own identity payload." + ) + + realizations = payload["realizations"] + if not isinstance(realizations, dict): + raise ManifestCompatibilityError("Manifest realizations must be a mapping.") + for realization_id, record in realizations.items(): + if not isinstance(record, dict): + raise ManifestCompatibilityError("Manifest realization record is invalid.") + recomputed_id = short_id("real", record.get("identity")) + recomputed_digest = integrity_digest(record.get("identity")) + if ( + realization_id != record.get("realization_id") + or realization_id != recomputed_id + or record.get("realization_digest") != recomputed_digest + ): + raise ManifestCompatibilityError( + f"Manifest realization {realization_id!r} ID/digest do not match " + "its own identity payload." + ) + owning_condition = semantic_conditions.get(record.get("condition_id")) + if ( + owning_condition is None + or owning_condition.get("condition_digest") != record.get("condition_digest") + ): + raise ManifestCompatibilityError( + f"Manifest realization {realization_id!r} condition binding is " + "inconsistent with its declared semantic condition." + ) + + +def initial_run_state_v3(manifest: Mapping[str, Any]) -> dict[str, Any]: + realizations = manifest["payload"]["realizations"] + return { + "schema_version": RUN_STATE_V3_SCHEMA_VERSION, + "experiment_id": manifest["experiment_id"], + "realizations": {realization_id: {"status": "pending"} for realization_id in realizations}, + } + + +def validate_run_state_v3(state: Mapping[str, Any], manifest: Mapping[str, Any]) -> None: + if ( + state.get("schema_version") != RUN_STATE_V3_SCHEMA_VERSION + or state.get("experiment_id") != manifest["experiment_id"] + ): + raise ManifestCompatibilityError( + f"Run state does not belong to manifest {manifest['experiment_id']}." + ) + expected = set(manifest["payload"]["realizations"]) + if set(state.get("realizations", {})) != expected: + raise ManifestCompatibilityError( + "Run state realization grid differs from the immutable manifest." + ) + + +def load_run_state_v3(run_dir: Path, manifest: Mapping[str, Any]) -> dict[str, Any]: + """Load and validate existing v3 run state without creating mutable state.""" + path = Path(run_dir) / RUN_STATE_FILENAME + if not path.is_file(): + raise ManifestCompatibilityError(f"Missing {RUN_STATE_FILENAME}.") + try: + state = json.loads(path.read_text()) + except (OSError, json.JSONDecodeError) as exc: + raise ManifestCompatibilityError(f"Unreadable {RUN_STATE_FILENAME}: {exc}") from exc + validate_run_state_v3(state, manifest) + return state diff --git a/tests/importing/test_manifest_v3.py b/tests/importing/test_manifest_v3.py new file mode 100644 index 0000000..b7a67f1 --- /dev/null +++ b/tests/importing/test_manifest_v3.py @@ -0,0 +1,203 @@ +"""Additive manifest-v3 tests: new schema alongside the untouched v2 path. + +Scope is deliberately reduced from the original plan: no ManifestView +cross-version normalization and no legacy-v2-in-v3 synthesis. Every manifest +this module builds is a fresh v3 import; existing v2 manifests keep using the +unmodified v2 functions exercised by tests/importing/test_manifest_v2_compat.py. +""" + +from __future__ import annotations + +import pytest + +from choicebench.identity import CANONICALIZATION_VERSION, integrity_digest, short_id +from choicebench.manifest import ( + MANIFEST_SCHEMA_VERSION, + MANIFEST_V3_SCHEMA_VERSION, + PROTOCOL_V3_VERSION, + RUN_STATE_V3_SCHEMA_VERSION, + ManifestCompatibilityError, + initial_run_state_v3, + make_manifest_v3, + validate_manifest, + validate_manifest_v3, + validate_run_state_v3, +) + + +def _condition_identity_triple(): + """A condition_id/digest pair that genuinely matches its own identity payload, + so validate_manifest_v3's recomputation check passes on the happy path.""" + identity = {"seed": 1} + digest = integrity_digest(identity) + condition_id = short_id("cond", identity) + return condition_id, digest, identity + + +def _realization_record(condition_id, condition_digest, realization_identity): + payload = { + "schema_version": "choicebench.realization.v1", + "condition_id": condition_id, + "condition_digest": condition_digest, + "realization": realization_identity, + } + return { + "realization_id": short_id("real", payload), + "realization_digest": integrity_digest(payload), + "condition_id": condition_id, + "condition_digest": condition_digest, + "identity": payload, + } + + +@pytest.fixture +def v3_payload(): + condition_id, condition_digest, identity = _condition_identity_triple() + condition = { + "condition_id": condition_id, + "condition_digest": condition_digest, + "identity": identity, + } + realization = _realization_record(condition_id, condition_digest, {"source": "x"}) + return { + "protocol_version": PROTOCOL_V3_VERSION, + "canonicalization_version": CANONICALIZATION_VERSION, + "semantic_conditions": {condition_id: condition}, + "realizations": {realization["realization_id"]: realization}, + }, condition_id, realization["realization_id"] + + +def test_make_manifest_v3_computes_schema_experiment_and_payload_digests(v3_payload): + payload, condition_id, realization_id = v3_payload + manifest = make_manifest_v3(payload) + assert manifest["schema_version"] == MANIFEST_V3_SCHEMA_VERSION + assert manifest["experiment_id"].startswith("exp_") + assert manifest["experiment_digest"] == integrity_digest(manifest["payload"]) + assert manifest["payload_digest"] == integrity_digest(manifest["payload"]) + assert manifest["payload"]["semantic_conditions"][condition_id]["condition_id"] == condition_id + assert realization_id in manifest["payload"]["realizations"] + + +def test_make_manifest_v3_excludes_audit_from_identity(v3_payload): + payload, _, _ = v3_payload + first = make_manifest_v3(payload, audit={"source_location": "/a/one"}) + second = make_manifest_v3(payload, audit={"source_location": "/b/two"}) + assert first["experiment_id"] == second["experiment_id"] + assert first["experiment_digest"] == second["experiment_digest"] + assert first["audit"] != second["audit"] + + +def test_validate_manifest_v3_accepts_a_well_formed_manifest(v3_payload): + payload, _, _ = v3_payload + manifest = make_manifest_v3(payload) + validate_manifest_v3(manifest) # must not raise + + +def test_validate_manifest_v3_rejects_wrong_schema_version(v3_payload): + payload, _, _ = v3_payload + manifest = make_manifest_v3(payload) + manifest["schema_version"] = MANIFEST_SCHEMA_VERSION + with pytest.raises(ManifestCompatibilityError): + validate_manifest_v3(manifest) + + +def test_validate_manifest_v3_rejects_tampered_payload_digest(v3_payload): + payload, _, _ = v3_payload + manifest = make_manifest_v3(payload) + manifest["payload"]["semantic_conditions"] = {} + with pytest.raises(ManifestCompatibilityError): + validate_manifest_v3(manifest) + + +def test_validate_manifest_v3_recomputes_condition_id_and_digest(v3_payload): + payload, condition_id, _ = v3_payload + manifest = make_manifest_v3(payload) + forged = dict(manifest["payload"]["semantic_conditions"][condition_id]) + forged["condition_digest"] = "f" * 64 + manifest["payload"]["semantic_conditions"][condition_id] = forged + manifest["payload_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_id"] = f"exp_{manifest['experiment_digest'][:16]}" + with pytest.raises(ManifestCompatibilityError, match="condition"): + validate_manifest_v3(manifest) + + +def test_validate_manifest_v3_recomputes_realization_id_and_digest(v3_payload): + payload, _, realization_id = v3_payload + manifest = make_manifest_v3(payload) + forged = dict(manifest["payload"]["realizations"][realization_id]) + forged["realization_digest"] = "f" * 64 + manifest["payload"]["realizations"][realization_id] = forged + manifest["payload_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_id"] = f"exp_{manifest['experiment_digest'][:16]}" + with pytest.raises(ManifestCompatibilityError, match="realization"): + validate_manifest_v3(manifest) + + +def test_validate_manifest_v3_requires_realization_condition_binding(v3_payload): + payload, condition_id, realization_id = v3_payload + manifest = make_manifest_v3(payload) + forged_realization = dict(manifest["payload"]["realizations"][realization_id]) + forged_realization["condition_digest"] = "e" * 64 + manifest["payload"]["realizations"][realization_id] = forged_realization + manifest["payload_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_id"] = f"exp_{manifest['experiment_digest'][:16]}" + with pytest.raises(ManifestCompatibilityError, match="condition"): + validate_manifest_v3(manifest) + + +def test_validate_manifest_v3_rejects_realization_with_unknown_condition(v3_payload): + payload, condition_id, realization_id = v3_payload + manifest = make_manifest_v3(payload) + del manifest["payload"]["semantic_conditions"][condition_id] + manifest["payload_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_id"] = f"exp_{manifest['experiment_digest'][:16]}" + with pytest.raises(ManifestCompatibilityError, match="condition"): + validate_manifest_v3(manifest) + + +def test_run_state_v3_is_keyed_by_realization(v3_payload): + payload, _, realization_id = v3_payload + manifest = make_manifest_v3(payload) + state = initial_run_state_v3(manifest) + assert state["schema_version"] == RUN_STATE_V3_SCHEMA_VERSION + assert state["experiment_id"] == manifest["experiment_id"] + assert state["realizations"] == {realization_id: {"status": "pending"}} + validate_run_state_v3(state, manifest) # must not raise + + +def test_run_state_v3_rejects_state_for_a_different_manifest(v3_payload): + payload, _, _ = v3_payload + manifest = make_manifest_v3(payload) + other_manifest = make_manifest_v3( + {**payload, "semantic_conditions": {}, "realizations": {}} + ) + state = initial_run_state_v3(manifest) + with pytest.raises(ManifestCompatibilityError): + validate_run_state_v3(state, other_manifest) + + +def test_run_state_v3_rejects_realization_grid_drift(v3_payload): + payload, _, realization_id = v3_payload + manifest = make_manifest_v3(payload) + state = initial_run_state_v3(manifest) + del state["realizations"][realization_id] + with pytest.raises(ManifestCompatibilityError): + validate_run_state_v3(state, manifest) + + +def test_v2_manifest_functions_are_unaffected_by_v3_additions(): + """Sanity check that importing the v3 symbols doesn't change v2 behavior.""" + v2_manifest = { + "schema_version": MANIFEST_SCHEMA_VERSION, + "experiment_id": "exp_0000000000000000", + "experiment_digest": "0" * 64, + "payload_digest": "0" * 64, + "created_at": "now", + "payload": {"protocol_version": "choicebench.protocol.v2", "config": {}}, + } + with pytest.raises(ManifestCompatibilityError): + validate_manifest(v2_manifest) # digest mismatch, but still v2-only code path From c401864a3c2115ec6a9069f9ad1ff3f92ac874c6 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 15:43:19 +0300 Subject: [PATCH 30/47] feat: publish non-circular v2 result artifacts for v3 realizations Reduced-scope Task 7: adds PreparedResultArtifact, prepare_manifest_result, publish_manifest_result to src/choicebench/io/writers.py, and extends the existing validate_result_artifact with optional manifest/realization kwargs that dispatch to a new v2 path. Calling it with no kwargs (the only way any existing caller does) hits the exact unchanged v1 branch. prepare_manifest_result computes result identity strictly after the manifest/realization are already fixed (result_artifact_id/digest are never part of the manifest), validates that every result row matches its declared result_origin row_assignment (question_id, prediction_origin, prediction_lineage_id) with no missing/extra/mismatched rows, orders rows by the declared assignment order regardless of input order, and requires (never fabricates) the mandatory per-row identity columns (condition, realization, experiment, dataset/model/method/prompt, split) already injected upstream by row normalization. publish_manifest_result writes only to an absent result/sidecar pair; an existing pair is a byte-for-byte no-op or a hard refusal on any divergence, never a silent overwrite. Focused: 14 new tests in tests/importing/test_result_artifact.py, plus the existing tests/io/test_writers.py, tests/io/test_readers.py, tests/importing/test_manifest_v2_compat.py, and tests/test_publication_identity.py all still pass unchanged (71 total). Full suite: 1224 passed (1210 baseline + 14), only the pre-existing Python 3.14 google.genai.types deprecation warning. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/io/writers.py | 221 +++++++++++++++++++- tests/importing/test_result_artifact.py | 260 ++++++++++++++++++++++++ 2 files changed, 475 insertions(+), 6 deletions(-) create mode 100644 tests/importing/test_result_artifact.py diff --git a/src/choicebench/io/writers.py b/src/choicebench/io/writers.py index ac90cde..4c8b5bc 100644 --- a/src/choicebench/io/writers.py +++ b/src/choicebench/io/writers.py @@ -1,38 +1,78 @@ # src/choicebench/io/writers.py +from dataclasses import dataclass +from hashlib import sha256 from pathlib import Path -from typing import Any +from typing import Any, Mapping, Sequence import pandas as pd import json -from choicebench.identity import file_digest, integrity_digest +from choicebench.identity import canonicalize, file_digest, integrity_digest, short_id from choicebench.infra.artifacts import atomic_write_json, atomic_write_text RESULT_ARTIFACT_SCHEMA_VERSION = "choicebench.result-artifact.v1" +RESULT_ARTIFACT_V2_SCHEMA_VERSION = "choicebench.result-artifact.v2" def result_artifact_path(result_path: Path) -> Path: return Path(result_path).with_suffix(".artifact.json") -def validate_result_artifact(result_path: Path) -> dict[str, Any]: +def validate_result_artifact( + result_path: Path, + *, + manifest: Mapping[str, Any] | None = None, + realization: Mapping[str, Any] | None = None, +) -> dict[str, Any]: metadata_path = result_artifact_path(result_path) try: metadata = json.loads(metadata_path.read_text()) except (OSError, json.JSONDecodeError) as exc: raise RuntimeError(f"Result metadata is missing or unreadable: {metadata_path}: {exc}") from exc - payload = {key: value for key, value in metadata.items() if key != "metadata_digest"} - if metadata.get("schema_version") != RESULT_ARTIFACT_SCHEMA_VERSION: + schema_version = metadata.get("schema_version") + if schema_version == RESULT_ARTIFACT_SCHEMA_VERSION: + payload = {key: value for key, value in metadata.items() if key != "metadata_digest"} + if metadata.get("metadata_digest") != integrity_digest(payload): + raise RuntimeError(f"Result metadata integrity check failed: {metadata_path}") + actual = file_digest(result_path) + if metadata.get("file_sha256") != actual: + raise RuntimeError( + f"Result content integrity check failed: {result_path}; expected " + f"{metadata.get('file_sha256')}, found {actual}." + ) + return metadata + if schema_version != RESULT_ARTIFACT_V2_SCHEMA_VERSION: raise RuntimeError(f"Unsupported result metadata schema: {metadata_path}") - if metadata.get("metadata_digest") != integrity_digest(payload): + payload = { + key: value + for key, value in metadata.items() + if key not in {"result_artifact_id", "result_artifact_digest"} + } + if metadata.get("result_artifact_digest") != integrity_digest(payload): raise RuntimeError(f"Result metadata integrity check failed: {metadata_path}") + if metadata.get("result_artifact_id") != short_id("result", payload): + raise RuntimeError(f"Result metadata ID does not match its own digest: {metadata_path}") actual = file_digest(result_path) if metadata.get("file_sha256") != actual: raise RuntimeError( f"Result content integrity check failed: {result_path}; expected " f"{metadata.get('file_sha256')}, found {actual}." ) + if manifest is not None: + realization_id = metadata.get("realization_id") + record = ( + realization + if realization is not None + else manifest.get("payload", {}).get("realizations", {}).get(realization_id) + ) + if record is None or record.get("realization_digest") != metadata.get("realization_digest"): + raise RuntimeError(f"Result metadata realization binding is inconsistent: {metadata_path}") + if ( + metadata.get("experiment_id") != manifest.get("experiment_id") + or metadata.get("experiment_digest") != manifest.get("experiment_digest") + ): + raise RuntimeError(f"Result metadata experiment binding is inconsistent: {metadata_path}") return metadata @@ -113,3 +153,172 @@ def write_run_results( metadata["metadata_digest"] = integrity_digest(metadata) atomic_write_json(result_artifact_path(output_path), metadata) return output_path + + +@dataclass(frozen=True) +class PreparedResultArtifact: + result_path: str + metadata_path: str + csv_bytes: bytes + metadata: Mapping[str, Any] + + +def prepare_manifest_result( + results: Sequence[Mapping[str, Any]], + *, + manifest: Mapping[str, Any], + realization_id: str, +) -> PreparedResultArtifact: + """Render one realization's rows and compute result identity after the + manifest is fixed. Never writes anything; identity here can never depend + on its own child (the manifest/realization are already immutable inputs). + """ + payload = manifest["payload"] + realization = payload["realizations"].get(realization_id) + if realization is None: + raise RuntimeError(f"Unknown realization {realization_id!r} in manifest.") + condition_id = realization["condition_id"] + condition = payload["semantic_conditions"].get(condition_id) + if condition is None or condition["condition_digest"] != realization["condition_digest"]: + raise RuntimeError(f"Realization {realization_id!r} condition binding is inconsistent.") + condition_identity = condition["identity"] + result_origin = realization["identity"]["realization"]["result_origin"] + + expected_assignments: dict[str, tuple[str, str]] = {} + for row in result_origin["row_assignments"]: + qid = str(row["question_id"]) + if qid in expected_assignments: + raise RuntimeError( + f"Realization {realization_id!r} result origin has duplicate question_id {qid!r}." + ) + expected_assignments[qid] = (row["prediction_origin"], row["prediction_lineage_id"]) + + rows_by_qid: dict[str, dict[str, Any]] = {} + for row in results: + qid = row.get("question_id") + qid = None if qid is None else str(qid) + if qid is None or qid not in expected_assignments: + raise RuntimeError( + f"Result row for realization {realization_id!r} references an " + f"unexpected question_id {row.get('question_id')!r}." + ) + if qid in rows_by_qid: + raise RuntimeError( + f"Result rows for realization {realization_id!r} declare " + f"question_id {qid!r} more than once." + ) + expected_origin, expected_lineage_id = expected_assignments[qid] + if ( + row.get("prediction_origin") != expected_origin + or row.get("prediction_lineage_id") != expected_lineage_id + ): + raise RuntimeError( + f"Result row {qid!r} prediction_origin/prediction_lineage_id does " + f"not match its declared realization lineage for {realization_id!r}." + ) + rows_by_qid[qid] = dict(row) + + missing = sorted(set(expected_assignments) - set(rows_by_qid)) + if missing: + raise RuntimeError( + f"Realization {realization_id!r} is missing result rows for {missing}." + ) + + ordered_question_ids = [str(row["question_id"]) for row in result_origin["row_assignments"]] + ordered_rows = [rows_by_qid[qid] for qid in ordered_question_ids] + + benchmark = condition_identity["benchmark"] + identity_columns = { + "condition_id": condition_id, + "realization_id": realization_id, + "experiment_id": manifest["experiment_id"], + "dataset_artifact_id": benchmark["artifact_id"], + "dataset_selection_id": benchmark["selection_id"], + "model_id": condition_identity["model_id"], + "method_id": condition_identity["method_id"], + "prompt_id": condition_identity["prompt_id"], + "benchmark_name": benchmark["name"], + "benchmark_split": benchmark["split"], + } + for row in ordered_rows: + for column, expected_value in identity_columns.items(): + if row.get(column) != expected_value: + raise RuntimeError( + f"Result row {row.get('question_id')!r} for realization " + f"{realization_id!r} has missing or incorrect {column!r}." + ) + + df = pd.DataFrame(ordered_rows) + csv_bytes = df.to_csv(index=False).encode("utf-8") + result_path = f"results/{realization_id}.csv" + metadata_path = f"results/{realization_id}.artifact.json" + + prediction_origin_counts: dict[str, int] = {} + for row in ordered_rows: + origin = row["prediction_origin"] + prediction_origin_counts[origin] = prediction_origin_counts.get(origin, 0) + 1 + + metadata_core = { + "schema_version": RESULT_ARTIFACT_V2_SCHEMA_VERSION, + "experiment_id": manifest["experiment_id"], + "experiment_digest": manifest["experiment_digest"], + "condition_id": condition_id, + "condition_digest": condition["condition_digest"], + "realization_id": realization_id, + "realization_digest": realization["realization_digest"], + "row_count": len(ordered_rows), + "columns": list(df.columns), + "question_ids": ordered_question_ids, + "rows_digest": integrity_digest(ordered_rows), + "file_sha256": sha256(csv_bytes).hexdigest(), + "prediction_origins": sorted(prediction_origin_counts), + "prediction_origin_counts": { + key: prediction_origin_counts[key] for key in sorted(prediction_origin_counts) + }, + "result_path": result_path, + "metadata_path": metadata_path, + } + result_artifact_digest = integrity_digest(metadata_core) + metadata = { + **metadata_core, + "result_artifact_id": short_id("result", metadata_core), + "result_artifact_digest": result_artifact_digest, + } + return PreparedResultArtifact( + result_path=result_path, + metadata_path=metadata_path, + csv_bytes=csv_bytes, + metadata=canonicalize(metadata), + ) + + +def publish_manifest_result( + prepared: PreparedResultArtifact, *, run_dir: Path +) -> tuple[Path, Mapping[str, Any], bool]: + """Write a prepared result artifact, or verify-only if it already exists. + + Returns (result_path, metadata, reused). `reused=True` means an identical + artifact already existed; any divergence is refused rather than overwritten. + """ + run_dir = Path(run_dir) + result_path = run_dir / prepared.result_path + metadata_path = run_dir / prepared.metadata_path + if result_path.exists() != metadata_path.exists(): + raise RuntimeError(f"Realization result/sidecar pair is incomplete: {result_path}") + if result_path.exists(): + existing_bytes = result_path.read_bytes() + try: + existing_metadata = json.loads(metadata_path.read_text()) + except (OSError, json.JSONDecodeError) as exc: + raise RuntimeError( + f"Existing result metadata is unreadable: {metadata_path}: {exc}" + ) from exc + if existing_bytes == prepared.csv_bytes and existing_metadata == dict(prepared.metadata): + return result_path, existing_metadata, True + raise RuntimeError( + f"Refusing to overwrite a divergent existing result artifact: {result_path}" + ) + result_path.parent.mkdir(parents=True, exist_ok=True) + atomic_write_text(result_path, prepared.csv_bytes.decode("utf-8")) + atomic_write_json(metadata_path, dict(prepared.metadata)) + return result_path, prepared.metadata, False diff --git a/tests/importing/test_result_artifact.py b/tests/importing/test_result_artifact.py new file mode 100644 index 0000000..88fd0f6 --- /dev/null +++ b/tests/importing/test_result_artifact.py @@ -0,0 +1,260 @@ +"""Result-artifact-v2 tests: non-circular identity for imported realizations. + +Reuses tests/importing/test_identity.py's condition/realization builders +instead of re-deriving a full semantic identity fixture here. +""" + +from __future__ import annotations + +import hashlib +import json + +import pandas as pd +import pytest + +from choicebench.identity import CANONICALIZATION_VERSION, integrity_digest +from choicebench.manifest import PROTOCOL_V3_VERSION, make_manifest_v3 +from choicebench.io.writers import ( + RESULT_ARTIFACT_V2_SCHEMA_VERSION, + prepare_manifest_result, + publish_manifest_result, + validate_result_artifact, +) +from tests.importing.test_identity import ( + _realization_identity, + _records, + make_realization as _make_realization, +) + + +def _build_manifest_and_realization(): + condition = _records().condition + identity = _realization_identity() + realization = _make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + payload = { + "protocol_version": PROTOCOL_V3_VERSION, + "canonicalization_version": CANONICALIZATION_VERSION, + "semantic_conditions": {condition["condition_id"]: condition}, + "realizations": {realization["realization_id"]: realization}, + } + manifest = make_manifest_v3(payload) + return manifest, realization + + +def _rows_for(manifest, realization): + """Rows shaped as Task 9's normalization step would deliver them: already + carrying the fixed ownership/identity columns prepare_manifest_result + validates (never fabricates) per row.""" + realization_id = realization["realization_id"] + condition_id = realization["condition_id"] + condition = manifest["payload"]["semantic_conditions"][condition_id] + benchmark = condition["identity"]["benchmark"] + identity_columns = { + "condition_id": condition_id, + "realization_id": realization_id, + "experiment_id": manifest["experiment_id"], + "dataset_artifact_id": benchmark["artifact_id"], + "dataset_selection_id": benchmark["selection_id"], + "model_id": condition["identity"]["model_id"], + "method_id": condition["identity"]["method_id"], + "prompt_id": condition["identity"]["prompt_id"], + "benchmark_name": benchmark["name"], + "benchmark_split": benchmark["split"], + } + row_assignments = realization["identity"]["realization"]["result_origin"]["row_assignments"] + return [ + { + **identity_columns, + "question_id": row["question_id"], + "prediction_origin": row["prediction_origin"], + "prediction_lineage_id": row["prediction_lineage_id"], + "parsed_choice": "A", + "is_correct": True, + } + for row in row_assignments + ] + + +def test_prepare_manifest_result_computes_identity_after_manifest_is_fixed(): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + assert "result_artifact_id" not in manifest + assert prepared.metadata["experiment_id"] == manifest["experiment_id"] + assert prepared.metadata["realization_id"] == realization_id + assert prepared.metadata["file_sha256"] == hashlib.sha256(prepared.csv_bytes).hexdigest() + assert prepared.metadata["result_artifact_id"].startswith("result_") + assert prepared.result_path == f"results/{realization_id}.csv" + assert prepared.metadata_path == f"results/{realization_id}.artifact.json" + + +def test_prepare_manifest_result_injects_identity_columns_per_row(): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + df = pd.read_csv(pd.io.common.BytesIO(prepared.csv_bytes)) + for column in ( + "condition_id", + "realization_id", + "experiment_id", + "dataset_artifact_id", + "dataset_selection_id", + "model_id", + "method_id", + "prompt_id", + "benchmark_name", + "benchmark_split", + "prediction_origin", + "prediction_lineage_id", + ): + assert column in df.columns + assert df[column].nunique(dropna=False) == 1 or column in ( + "prediction_origin", + "prediction_lineage_id", + ) + assert set(df["realization_id"]) == {realization_id} + + +def test_prepare_manifest_result_orders_rows_by_declared_row_assignments(): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + rows = _rows_for(manifest, realization) + prepared = prepare_manifest_result( + list(reversed(rows)), manifest=manifest, realization_id=realization_id + ) + expected_order = [row["question_id"] for row in rows] + df = pd.read_csv(pd.io.common.BytesIO(prepared.csv_bytes), dtype={"question_id": "string"}) + assert df["question_id"].tolist() == expected_order + + +def test_prepare_manifest_result_rejects_unknown_question_id(): + manifest, realization = _build_manifest_and_realization() + rows = _rows_for(manifest, realization) + rows[0]["question_id"] = "not-a-declared-question" + with pytest.raises(RuntimeError, match="unexpected question_id"): + prepare_manifest_result( + rows, manifest=manifest, realization_id=realization["realization_id"] + ) + + +def test_prepare_manifest_result_rejects_missing_row(): + manifest, realization = _build_manifest_and_realization() + rows = _rows_for(manifest, realization)[:-1] + with pytest.raises(RuntimeError, match="missing result rows"): + prepare_manifest_result( + rows, manifest=manifest, realization_id=realization["realization_id"] + ) + + +def test_prepare_manifest_result_rejects_prediction_origin_mismatch(): + manifest, realization = _build_manifest_and_realization() + rows = _rows_for(manifest, realization) + rows[0]["prediction_origin"] = "native_inference" + with pytest.raises(RuntimeError, match="prediction_origin"): + prepare_manifest_result( + rows, manifest=manifest, realization_id=realization["realization_id"] + ) + + +def test_prepare_manifest_result_rejects_unknown_realization(): + manifest, _realization = _build_manifest_and_realization() + with pytest.raises(RuntimeError, match="[Uu]nknown realization"): + prepare_manifest_result( + [], manifest=manifest, realization_id="real_" + "0" * 16 + ) + + +def test_publish_manifest_result_writes_csv_and_sidecar(tmp_path): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + path, metadata, reused = publish_manifest_result(prepared, run_dir=tmp_path) + assert reused is False + assert path == tmp_path / "results" / f"{realization_id}.csv" + assert path.read_bytes() == prepared.csv_bytes + sidecar_path = tmp_path / "results" / f"{realization_id}.artifact.json" + assert json.loads(sidecar_path.read_text()) == dict(prepared.metadata) + assert metadata["result_artifact_id"] == prepared.metadata["result_artifact_id"] + + +def test_publish_manifest_result_is_idempotent_noop_for_identical_bytes(tmp_path): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + publish_manifest_result(prepared, run_dir=tmp_path) + _, _, second_reused = publish_manifest_result(prepared, run_dir=tmp_path) + assert second_reused is True + + +def test_publish_manifest_result_refuses_divergent_overwrite(tmp_path): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + rows = _rows_for(manifest, realization) + prepared = prepare_manifest_result(rows, manifest=manifest, realization_id=realization_id) + publish_manifest_result(prepared, run_dir=tmp_path) + rows[0]["parsed_choice"] = "B" + divergent = prepare_manifest_result(rows, manifest=manifest, realization_id=realization_id) + with pytest.raises(RuntimeError, match="divergent"): + publish_manifest_result(divergent, run_dir=tmp_path) + + +def test_validate_result_artifact_v2_accepts_a_well_formed_pair(tmp_path): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + path, _, _ = publish_manifest_result(prepared, run_dir=tmp_path) + metadata = validate_result_artifact(path, manifest=manifest, realization=realization) + assert metadata["schema_version"] == RESULT_ARTIFACT_V2_SCHEMA_VERSION + + +def test_validate_result_artifact_v2_rejects_tampered_csv_bytes(tmp_path): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + path, _, _ = publish_manifest_result(prepared, run_dir=tmp_path) + path.write_bytes(prepared.csv_bytes + b"\n# tampered") + with pytest.raises(RuntimeError, match="content integrity"): + validate_result_artifact(path, manifest=manifest, realization=realization) + + +def test_validate_result_artifact_v2_rejects_realization_binding_mismatch(tmp_path): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + path, _, _ = publish_manifest_result(prepared, run_dir=tmp_path) + empty_manifest = make_manifest_v3( + { + "protocol_version": PROTOCOL_V3_VERSION, + "canonicalization_version": CANONICALIZATION_VERSION, + "semantic_conditions": {}, + "realizations": {}, + } + ) + with pytest.raises(RuntimeError, match="realization"): + validate_result_artifact(path, manifest=empty_manifest) + + +def test_validate_result_artifact_v1_path_is_unaffected(): + """Calling with no manifest/realization keeps the exact v1 behavior.""" + from choicebench.io.writers import RESULT_ARTIFACT_SCHEMA_VERSION + + assert RESULT_ARTIFACT_SCHEMA_VERSION == "choicebench.result-artifact.v1" From 6a3cf1baae199e58cfd043395f9affee0c42c4a7 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 16:01:32 +0300 Subject: [PATCH 31/47] feat: validate imported rows by question identity and compute evidence status Collapsed Tasks 8+9 into one module (they share the same file in the original plan; Task 9 only extends Task 8's types). Reduced scope: exact option-count/method-specific-extension richness is dropped in favor of what the repair-import workflow concretely needs. validate_source_rows joins by exact string question_id via a caller- supplied semantic-field mapping (needed since the plan's own interface didn't specify how raw CSV columns map to semantic fields; this module needs that explicitly to be self-contained). No row is dropped, padded, deduplicated, reordered, or repaired: every input row is kept (with findings) except when its question_id is genuinely ambiguous (duplicated within the source). Missing/unexpected/duplicate question IDs, question- text and correct_option mismatches against the expected-dataset snapshot, and missing/invalid parsed predictions (checked against the question's real choices_json-derived option-letter range, not an invented choice_a/choice_b column shape) all accumulate as ValidationFinding records with a findings_digest. normalize_realization_rows recomputes evidence_status from exact coverage (present vs expected question IDs) and defects (findings tied to a present question_id), and raises ImportValidationError on any mismatch against the condition's declared evidence_status -- never silently trusting the caller's declaration. Only complete/qualified + scope included realizations produce evaluable_rows; partial/malformed/ recoverable/failed/excluded/held realizations produce none, matching the plan's evidence-only rule. recoverable is computed (not just copied from a declared status): coverage gaps must be an exact subset of the condition's declared recoverable_question_ids with no other defects. prepare/write/validate_realization_validation_artifact publish one self-digested JSON sidecar per realization at artifacts/imports/validation/.json: byte-identical existing content is a no-op, any divergence is refused, and the realization_id/digest binding is checked on read. Focused: 18 new tests, all passing in 0.15s (no identity.py-style file hashing in this module). Full suite: 1242 passed (1224 baseline + 18), only the pre-existing Python 3.14 google.genai.types deprecation warning. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/importing/validation.py | 445 ++++++++++++++++++++++++ tests/importing/test_validation.py | 284 +++++++++++++++ 2 files changed, 729 insertions(+) create mode 100644 src/choicebench/importing/validation.py create mode 100644 tests/importing/test_validation.py diff --git a/src/choicebench/importing/validation.py b/src/choicebench/importing/validation.py new file mode 100644 index 0000000..be7df5a --- /dev/null +++ b/src/choicebench/importing/validation.py @@ -0,0 +1,445 @@ +"""Validate imported source rows by question identity and derive per-realization +evidence status and evaluable content. + +Reduced scope: this covers the concrete guarantees the repair-import workflow +needs (exact question-identity join, content cross-check against the expected +dataset snapshot, coverage/defect-driven evidence status with a fail-closed +declaration check, and evaluable-row derivation) without the full historical +option-count/method-specific-extension richness of the original plan. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from hashlib import sha256 +import json +from pathlib import Path +from typing import Any, Mapping, Sequence + +from choicebench.identity import canonicalize, integrity_digest +from choicebench.importing.csv_adapter import AdaptedTable, LogicalRecordSpan, SourceRow +from choicebench.importing.dataset_reference import ExpectedDataset +from choicebench.importing.schema import ImportConditionSpec +from choicebench.infra.artifacts import atomic_write_bytes + +_EVALUABLE_STATUSES = {"complete", "qualified"} +_COVERAGE_ONLY_FINDING_CODES = {"MISSING_QUESTION_ID"} +VALIDATION_ARTIFACT_SCHEMA_VERSION = "choicebench.realization-validation.v1" + + +class ImportValidationError(ValueError): + """Raised when imported row content fails fail-closed validation.""" + + +@dataclass(frozen=True) +class ValidationFinding: + code: str + source_id: str + field: str | None + question_id: str | None + record_span: LogicalRecordSpan | None + raw_row_sha256: str | None + message: str + + def to_json(self) -> dict[str, Any]: + return { + "code": self.code, + "source_id": self.source_id, + "field": self.field, + "question_id": self.question_id, + "record_span": ( + None + if self.record_span is None + else { + "record_index": self.record_span.index, + "start": self.record_span.start, + "end": self.record_span.end, + "terminator_hex": self.record_span.terminator.hex(), + } + ), + "raw_row_sha256": self.raw_row_sha256, + "message": self.message, + } + + +@dataclass(frozen=True) +class ValidatedSourceRows: + source_id: str + rows_by_question_id: Mapping[str, SourceRow] + expected_question_ids: tuple[str, ...] + findings: tuple[ValidationFinding, ...] + findings_digest: str + + +def _expected_row(expected: ExpectedDataset, question_id: str) -> Mapping[str, Any] | None: + matches = expected.frame[expected.frame["question_id"].astype(str) == question_id] + if matches.empty: + return None + return matches.iloc[0].to_dict() + + +def _expected_choice_letters(expected_row: Mapping[str, Any]) -> list[str]: + """The valid answer letters for a question, derived from its ordered + choices_json array (index 0 -> "A", 1 -> "B", ...) -- the same ordering + ChoiceBench's own option map already uses to define correct_option.""" + raw = expected_row.get("choices_json") + if not raw: + return [] + try: + choices = json.loads(raw) + except (TypeError, ValueError): + return [] + if not isinstance(choices, list): + return [] + return [chr(ord("A") + index) for index in range(len(choices))] + + +def validate_source_rows( + table: AdaptedTable, + *, + source_id: str, + mapping: Mapping[str, str], + condition: ImportConditionSpec, + expected: ExpectedDataset, +) -> ValidatedSourceRows: + """Join source rows to the expected dataset by exact string question_id. + + Every declared field is compared against the expected snapshot; source + values never override it. No row is dropped, padded, deduplicated, + reordered, or repaired -- every input row is either kept (with any + findings recorded alongside it) or excluded only when its question_id is + genuinely ambiguous (duplicated within this source). + """ + if "question_id" not in mapping: + raise ImportValidationError(f"{source_id}: mapping declares no question_id column.") + question_id_column = mapping["question_id"] + expected_ids = tuple(condition.expected_question_ids) + expected_id_set = set(expected_ids) + + findings: list[ValidationFinding] = [] + rows_by_question_id: dict[str, SourceRow] = {} + rows_by_raw_id: dict[str, list[SourceRow]] = {} + + for row in table.rows: + raw_qid = row.values.get(question_id_column) + if raw_qid is None or not str(raw_qid).strip(): + findings.append( + ValidationFinding( + code="NULL_QUESTION_ID", + source_id=source_id, + field=question_id_column, + question_id=None, + record_span=row.span, + raw_row_sha256=row.raw_sha256, + message="Row has a null or empty question_id.", + ) + ) + continue + rows_by_raw_id.setdefault(str(raw_qid), []).append(row) + + for qid, rows in rows_by_raw_id.items(): + if len(rows) > 1: + for row in rows: + findings.append( + ValidationFinding( + code="DUPLICATE_QUESTION_ID", + source_id=source_id, + field=question_id_column, + question_id=qid, + record_span=row.span, + raw_row_sha256=row.raw_sha256, + message=f"question_id {qid!r} occurs {len(rows)} times in this source.", + ) + ) + continue + row = rows[0] + rows_by_question_id[qid] = row + if qid not in expected_id_set: + findings.append( + ValidationFinding( + code="UNEXPECTED_QUESTION_ID", + source_id=source_id, + field=question_id_column, + question_id=qid, + record_span=row.span, + raw_row_sha256=row.raw_sha256, + message=f"question_id {qid!r} is not in the expected dataset selection.", + ) + ) + continue + expected_row = _expected_row(expected, qid) + if expected_row is None: + continue + for semantic_field in ("question_text", "correct_option"): + source_column = mapping.get(semantic_field) + if source_column is None: + continue + source_value = row.values.get(source_column) + expected_value = expected_row.get(semantic_field) + normalized_source = None if source_value is None else str(source_value).strip() + normalized_expected = None if expected_value is None else str(expected_value).strip() + if semantic_field == "correct_option": + normalized_source = normalized_source and normalized_source.upper() + normalized_expected = normalized_expected and normalized_expected.upper() + if normalized_source is None or normalized_source != normalized_expected: + findings.append( + ValidationFinding( + code=f"{semantic_field.upper()}_MISMATCH", + source_id=source_id, + field=source_column, + question_id=qid, + record_span=row.span, + raw_row_sha256=row.raw_sha256, + message=f"{semantic_field} does not match the expected dataset snapshot.", + ) + ) + expected_choice_letters = _expected_choice_letters(expected_row) + prediction_column = mapping.get("prediction") + if prediction_column is not None: + prediction_value = row.values.get(prediction_column) + if prediction_value is None or not str(prediction_value).strip(): + findings.append( + ValidationFinding( + code="MISSING_PREDICTION", + source_id=source_id, + field=prediction_column, + question_id=qid, + record_span=row.span, + raw_row_sha256=row.raw_sha256, + message="Row has no parsed prediction.", + ) + ) + else: + letter = str(prediction_value).strip().upper() + if letter not in expected_choice_letters: + findings.append( + ValidationFinding( + code="INVALID_PREDICTION", + source_id=source_id, + field=prediction_column, + question_id=qid, + record_span=row.span, + raw_row_sha256=row.raw_sha256, + message=( + f"Parsed prediction {prediction_value!r} is not one of " + f"this question's options {expected_choice_letters}." + ), + ) + ) + + for qid in sorted(expected_id_set - set(rows_by_question_id)): + findings.append( + ValidationFinding( + code="MISSING_QUESTION_ID", + source_id=source_id, + field=None, + question_id=qid, + record_span=None, + raw_row_sha256=None, + message=f"Expected question_id {qid!r} is absent from this source.", + ) + ) + + findings_tuple = tuple(findings) + findings_digest = integrity_digest([finding.to_json() for finding in findings_tuple]) + return ValidatedSourceRows( + source_id=source_id, + rows_by_question_id=rows_by_question_id, + expected_question_ids=expected_ids, + findings=findings_tuple, + findings_digest=findings_digest, + ) + + +@dataclass(frozen=True) +class RealizationValidation: + evaluable_rows: tuple[Mapping[str, Any], ...] + findings: tuple[ValidationFinding, ...] + validation_digest: str + computed_evidence_status: str + evaluable: bool + qualifications: tuple[Mapping[str, Any], ...] + limitations: tuple[Mapping[str, Any], ...] + defect_question_ids: tuple[str, ...] + + +def normalize_realization_rows( + validated: ValidatedSourceRows, + *, + condition: ImportConditionSpec, + expected: ExpectedDataset, + mapping: Mapping[str, str] | None = None, + experiment_id: str | None = None, + realization_id: str | None = None, +) -> RealizationValidation: + """Recompute evidence status from exact coverage/defects and refuse a + declaration mismatch. Fill evaluative content only from ExpectedDataset; + source values were already checked, never trusted, in validate_source_rows. + """ + expected_ids = set(validated.expected_question_ids) + present_ids = set(validated.rows_by_question_id) + missing_ids = expected_ids - present_ids + defect_question_ids = tuple( + sorted( + { + finding.question_id + for finding in validated.findings + if finding.question_id is not None + and finding.code not in _COVERAGE_ONLY_FINDING_CODES + } + ) + ) + recoverable_ids = set(condition.recoverable_question_ids) + + if not present_ids and expected_ids: + computed = "failed" + elif defect_question_ids: + computed = "malformed" + elif missing_ids and missing_ids <= recoverable_ids and recoverable_ids: + computed = "recoverable" + elif missing_ids: + computed = "partial" + else: + computed = "qualified" if condition.qualifications else "complete" + + if computed != condition.evidence_status: + raise ImportValidationError( + f"Computed evidence status {computed!r} does not match the declared " + f"evidence status {condition.evidence_status!r}." + ) + + evaluable = computed in _EVALUABLE_STATUSES and condition.scope_disposition == "included" + + prediction_column = None if mapping is None else mapping.get("prediction") + evaluable_rows: list[dict[str, Any]] = [] + if evaluable: + per_question_origin = condition.result_origin.per_question_prediction_origins + default_origin = condition.result_origin.default_prediction_origin + for qid in validated.expected_question_ids: + if qid not in validated.rows_by_question_id: + continue + expected_row = _expected_row(expected, qid) + origin = per_question_origin.get(qid, default_origin) + if origin is None: + raise ImportValidationError( + f"No declared prediction_origin for evaluable question {qid!r}." + ) + predicted_option = None + if prediction_column is not None: + raw_prediction = validated.rows_by_question_id[qid].values.get(prediction_column) + predicted_option = None if raw_prediction is None else str(raw_prediction).strip().upper() + row: dict[str, Any] = { + "question_id": qid, + "question_text": expected_row.get("question_text"), + "correct_option": expected_row.get("correct_option"), + "choices_json": expected_row.get("choices_json"), + "prediction_origin": origin, + "predicted_option": predicted_option, + } + evaluable_rows.append(row) + + validation_payload = { + "schema_version": "choicebench.realization-validation-content.v1", + "computed_evidence_status": computed, + "evaluable": evaluable, + "findings_digest": validated.findings_digest, + "evaluable_rows": canonicalize(evaluable_rows), + } + validation_digest = integrity_digest(validation_payload) + + return RealizationValidation( + evaluable_rows=tuple(evaluable_rows), + findings=validated.findings, + validation_digest=validation_digest, + computed_evidence_status=computed, + evaluable=evaluable, + qualifications=tuple(condition.qualifications), + limitations=tuple(condition.limitations), + defect_question_ids=defect_question_ids, + ) + + +@dataclass(frozen=True) +class PreparedValidationArtifact: + relative_path: str + json_bytes: bytes + file_sha256: str + validation_digest: str + + +def prepare_realization_validation_artifact( + validation: RealizationValidation, + *, + realization: Mapping[str, Any], + evidence_records: Sequence[Mapping[str, Any]], +) -> PreparedValidationArtifact: + realization_id = realization["realization_id"] + payload = canonicalize( + { + "schema_version": VALIDATION_ARTIFACT_SCHEMA_VERSION, + "realization_id": realization_id, + "realization_digest": realization["realization_digest"], + "validation_digest": validation.validation_digest, + "computed_evidence_status": validation.computed_evidence_status, + "evaluable": validation.evaluable, + "qualifications": list(validation.qualifications), + "limitations": list(validation.limitations), + "defect_question_ids": list(validation.defect_question_ids), + "findings": [finding.to_json() for finding in validation.findings], + "evidence_records": list(evidence_records), + } + ) + json_bytes = ( + json.dumps(payload, indent=2, sort_keys=True, allow_nan=False) + "\n" + ).encode("utf-8") + return PreparedValidationArtifact( + relative_path=f"artifacts/imports/validation/{realization_id}.json", + json_bytes=json_bytes, + file_sha256=sha256(json_bytes).hexdigest(), + validation_digest=validation.validation_digest, + ) + + +def write_realization_validation_artifact( + staged_run: Path, prepared: PreparedValidationArtifact +) -> tuple[Path, bool]: + path = Path(staged_run) / prepared.relative_path + if path.exists(): + existing = path.read_bytes() + if existing == prepared.json_bytes: + return path, True + raise RuntimeError(f"Refusing to overwrite a divergent validation artifact: {path}") + atomic_write_bytes(path, prepared.json_bytes) + return path, False + + +def validate_realization_validation_artifact( + path: Path, + *, + manifest: Mapping[str, Any], + realization: Mapping[str, Any], + expected_sha256: str, +) -> Mapping[str, Any]: + path = Path(path) + try: + actual_bytes = path.read_bytes() + except OSError as exc: + raise RuntimeError(f"Validation artifact is unreadable: {path}: {exc}") from exc + actual_sha256 = sha256(actual_bytes).hexdigest() + if actual_sha256 != expected_sha256: + raise RuntimeError( + f"Validation artifact content integrity check failed: {path}; expected " + f"{expected_sha256}, found {actual_sha256}." + ) + try: + payload = json.loads(actual_bytes) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Validation artifact is not valid JSON: {path}: {exc}") from exc + if payload.get("schema_version") != VALIDATION_ARTIFACT_SCHEMA_VERSION: + raise RuntimeError(f"Unsupported validation artifact schema: {path}") + if ( + payload.get("realization_id") != realization.get("realization_id") + or payload.get("realization_digest") != realization.get("realization_digest") + ): + raise RuntimeError(f"Validation artifact realization binding is inconsistent: {path}") + return payload diff --git a/tests/importing/test_validation.py b/tests/importing/test_validation.py new file mode 100644 index 0000000..a280a49 --- /dev/null +++ b/tests/importing/test_validation.py @@ -0,0 +1,284 @@ +"""Reduced-scope tests for importing.validation: question-identity join, +evidence-status computation with declaration cross-check, and evaluable-row +derivation. Reuses tests/importing/test_identity.py's condition/dataset +fixtures instead of re-deriving them here. +""" + +from __future__ import annotations + +from dataclasses import replace + +import pytest + +from choicebench.importing.csv_adapter import AdaptedTable, LogicalRecordSpan, SourceRow +from choicebench.importing.validation import ( + ImportValidationError, + normalize_realization_rows, + prepare_realization_validation_artifact, + validate_realization_validation_artifact, + validate_source_rows, + write_realization_validation_artifact, +) +from tests.importing.test_identity import _condition, _dataset + +_MAPPING = { + "question_id": "qid", + "question_text": "question", + "correct_option": "gold_answer", + "prediction": "predicted_letter", +} + + +def _row(index, qid, question, gold, a, b, predicted): + values = { + "qid": qid, + "question": question, + "gold_answer": gold, + "opt_a": a, + "opt_b": b, + "predicted_letter": predicted, + } + return SourceRow( + values=values, + span=LogicalRecordSpan(index=index, start=index * 10, end=index * 10 + 9, terminator=b"\n"), + raw_sha256=f"{index:064x}", + ) + + +def _valid_rows(): + return ( + _row(0, "q1", "One?", "A", "x", "y", "a"), + _row(1, "q2", "Two?", "B", "m", "n", "b"), + ) + + +def _table(rows): + return AdaptedTable(columns=tuple(_MAPPING.values()), rows=rows, source_sha256="1" * 64) + + +def test_validate_source_rows_accepts_matching_rows(): + condition = _condition() + expected = _dataset() + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=expected, + ) + assert validated.findings == () + assert set(validated.rows_by_question_id) == {"q1", "q2"} + + +def test_validate_source_rows_flags_null_question_id(): + rows = (_row(0, None, "One?", "A", "x", "y", "a"), _valid_rows()[1]) + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + codes = {f.code for f in validated.findings} + assert "NULL_QUESTION_ID" in codes + assert "MISSING_QUESTION_ID" in codes # q1 never arrived + + +def test_validate_source_rows_flags_duplicate_question_id(): + rows = (*_valid_rows(), _row(2, "q1", "One?", "A", "x", "y", "a")) + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + duplicate_findings = [f for f in validated.findings if f.code == "DUPLICATE_QUESTION_ID"] + assert len(duplicate_findings) == 2 # both occurrences flagged + assert "q1" not in validated.rows_by_question_id # ambiguous, not usable + + +def test_validate_source_rows_flags_unexpected_question_id(): + rows = (*_valid_rows(), _row(2, "q9", "Nine?", "A", "x", "y", "a")) + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + assert any(f.code == "UNEXPECTED_QUESTION_ID" and f.question_id == "q9" for f in validated.findings) + + +def test_validate_source_rows_flags_missing_question_id(): + validated = validate_source_rows( + _table(_valid_rows()[:1]), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + assert any(f.code == "MISSING_QUESTION_ID" and f.question_id == "q2" for f in validated.findings) + + +def test_validate_source_rows_flags_gold_mismatch(): + rows = (_row(0, "q1", "One?", "B", "x", "y", "a"), _valid_rows()[1]) + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + assert any(f.code == "CORRECT_OPTION_MISMATCH" and f.question_id == "q1" for f in validated.findings) + + +def test_validate_source_rows_flags_invalid_prediction(): + rows = (_row(0, "q1", "One?", "A", "x", "y", "z"), _valid_rows()[1]) + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + assert any(f.code == "INVALID_PREDICTION" and f.question_id == "q1" for f in validated.findings) + + +def test_validate_source_rows_flags_missing_prediction(): + rows = (_row(0, "q1", "One?", "A", "x", "y", None), _valid_rows()[1]) + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + assert any(f.code == "MISSING_PREDICTION" and f.question_id == "q1" for f in validated.findings) + + +def test_validate_source_rows_never_drops_or_reorders_rows(): + rows = (_valid_rows()[1], _valid_rows()[0]) # reversed input order + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + assert len(validated.rows_by_question_id) == 2 + + +def test_normalize_realization_rows_computes_complete_status(): + condition = _condition() + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + assert result.computed_evidence_status == "complete" + assert result.evaluable is True + assert len(result.evaluable_rows) == 2 + assert {row["question_id"] for row in result.evaluable_rows} == {"q1", "q2"} + assert all( + row["prediction_origin"] == "external_historical_inference" for row in result.evaluable_rows + ) + + +def test_normalize_realization_rows_refuses_declaration_mismatch(): + condition = replace(_condition(), evidence_status="malformed") + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + with pytest.raises(ImportValidationError, match="evidence status"): + normalize_realization_rows(validated, condition=condition, expected=_dataset()) + + +def test_normalize_realization_rows_produces_no_evaluable_rows_for_partial(): + condition = replace( + _condition(), + expected_question_ids=("q1", "q2", "q3"), + evidence_status="partial", + ) + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + assert result.computed_evidence_status == "partial" + assert result.evaluable is False + assert result.evaluable_rows == () + + +def test_normalize_realization_rows_produces_no_evaluable_rows_for_malformed(): + rows = (_row(0, "q1", "One?", "A", "x", "y", "z"), _valid_rows()[1]) # invalid prediction + condition = replace(_condition(), evidence_status="malformed") + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + assert result.computed_evidence_status == "malformed" + assert result.evaluable is False + assert result.evaluable_rows == () + assert result.defect_question_ids == ("q1",) + + +def test_normalize_realization_rows_computes_recoverable_status(): + condition = replace( + _condition(), + expected_question_ids=("q1", "q2", "q3"), + evidence_status="recoverable", + recoverable_question_ids=("q3",), + ) + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + assert result.computed_evidence_status == "recoverable" + assert result.evaluable is False + + +def test_normalize_realization_rows_computes_failed_status_for_zero_rows(): + condition = _condition() + validated = validate_source_rows( + _table(()), source_id="results", mapping=_MAPPING, + condition=replace(condition, evidence_status="failed"), expected=_dataset(), + ) + result = normalize_realization_rows( + validated, condition=replace(condition, evidence_status="failed"), expected=_dataset() + ) + assert result.computed_evidence_status == "failed" + assert result.evaluable is False + + +def test_normalize_realization_rows_excludes_from_paper_matrix_is_not_evaluable(): + condition = replace(_condition(), scope_disposition="excluded_from_paper_matrix") + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + assert result.computed_evidence_status == "complete" # row content is fine + assert result.evaluable is False # but scope excludes it from evaluation + assert result.evaluable_rows == () + + +def _realization_stub(digest="a" * 64): + return {"realization_id": "real_" + "b" * 16, "realization_digest": digest} + + +def test_validation_artifact_round_trips(tmp_path): + condition = _condition() + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + realization = _realization_stub() + prepared = prepare_realization_validation_artifact( + result, realization=realization, evidence_records=() + ) + path, reused = write_realization_validation_artifact(tmp_path, prepared) + assert reused is False + payload = validate_realization_validation_artifact( + path, manifest={}, realization=realization, expected_sha256=prepared.file_sha256 + ) + assert payload["realization_id"] == realization["realization_id"] + + _, second_reused = write_realization_validation_artifact(tmp_path, prepared) + assert second_reused is True + + +def test_validation_artifact_rejects_tampered_bytes(tmp_path): + condition = _condition() + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + realization = _realization_stub() + prepared = prepare_realization_validation_artifact( + result, realization=realization, evidence_records=() + ) + path, _ = write_realization_validation_artifact(tmp_path, prepared) + path.write_bytes(prepared.json_bytes + b"\n// tampered") + with pytest.raises(RuntimeError, match="content integrity"): + validate_realization_validation_artifact( + path, manifest={}, realization=realization, expected_sha256=prepared.file_sha256 + ) From bdbffd62b94dd803d0c6a74ba1e8490332cf395e Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 16:21:27 +0300 Subject: [PATCH 32/47] feat: validate typed authorization and derive immutable repair overlays Task 10, reduced scope. This is the core of the repair-queueing objective: authorization.py types a repair/offline-transformation grant as either inference_repair (executable=true, a real model may run) or offline_transformation (executable=false, non-executing semantic rematching only), refusing any type/executable mismatch. Every grant is checked against the caller's own known condition digests and expected dataset selections -- an authorization cannot grant a question outside its condition's selection, cannot have an empty reason, and cannot self-assign an authorization_id (it must equal the bundle's own recomputed digest, so tampering with the source bytes used to build it is caught by ID mismatch, not a separate checksum field the schema doesn't carry). authorization_for_condition slices one condition's grant out of a bundle while keeping the whole bundle's ID/digest, so a slice can never borrow another condition's grants or diverge from its source bundle. overlays.py's derive_overlay is a pure function over an already-verified VerifiedBaseRealization (never a bare path or unverified dict): it checks the overlay's base/authorization binding, requires exact evidence-digest agreement across the overlay declaration, the authorization's granted evidence, and the base's own evidence sources, and refuses replacing an unauthorized or out-of-selection question. Reuses OverlaySpec's existing result_origin field (ResultOriginSpec) for per-question origin assignment rather than inventing new parameters: inference_repair requires an explicit, valid repair origin for every replaced question (no bare "mixed", no missing/extra assignments); offline_transformation forbids declaring a new origin and retains the base response's own underlying prediction_origin. Builds each replaced row's lineage component via the existing Task 5 identity.py helpers, so the narrower "declared_external" identity/source-ownership checks already built into make_lineage_component apply here unchanged. Scope boundary (documented in the module docstring): derive_overlay covers only the transformation's own replaced rows and their lineage; it does not merge them with the base's retained rows or construct the final realization -- that full-graph assembly belongs to the import engine (Unit G), which alone has the base run directory and destination for the new run. Focused: 21 new tests (11 authorization, 10 overlays), all passing in under 0.2s. Full suite: 1263 passed (1242 baseline + 21), only the pre-existing Python 3.14 google.genai.types deprecation warning. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/importing/authorization.py | 191 +++++++++++++ src/choicebench/importing/overlays.py | 266 ++++++++++++++++++ tests/importing/test_authorization.py | 255 +++++++++++++++++ tests/importing/test_overlays.py | 305 +++++++++++++++++++++ 4 files changed, 1017 insertions(+) create mode 100644 src/choicebench/importing/authorization.py create mode 100644 src/choicebench/importing/overlays.py create mode 100644 tests/importing/test_authorization.py create mode 100644 tests/importing/test_overlays.py diff --git a/src/choicebench/importing/authorization.py b/src/choicebench/importing/authorization.py new file mode 100644 index 0000000..00c968e --- /dev/null +++ b/src/choicebench/importing/authorization.py @@ -0,0 +1,191 @@ +"""Validate typed repair/offline-transformation authorization bundles. + +An authorization bundle grants specific (condition_digest, question_id) pairs +permission to be replaced, typed as either `inference_repair` (executable -- +a real model may be run) or `offline_transformation` (non-executing semantic +rematching only). This module never runs inference or opens a filesystem +path itself; it validates an already-opened source and already-known +condition digests/expected datasets. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from typing import Mapping + +from choicebench.identity import integrity_digest, short_id +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.dataset_reference import ExpectedDataset +from choicebench.importing.schema import AuthorizationSpec + + +class AuthorizationError(ValueError): + """Raised when an authorization bundle or its use is invalid or unsafe.""" + + +@dataclass(frozen=True) +class ValidatedAuthorizationBundle: + authorization_id: str + authorization_digest: str + authorization_type: str + grants: Mapping[str, Mapping[str, str]] # condition_digest -> question_id -> reason + authority: str + purpose: str + executable: bool + source_sha256: str + input_evidence_digests: Mapping[str, str] + expected_snapshot_digests: Mapping[str, str] + + +@dataclass(frozen=True) +class ValidatedAuthorization: + bundle_id: str + bundle_digest: str + authorization_type: str + condition_digest: str + question_reasons: Mapping[str, str] + authority: str + purpose: str + executable: bool + source_sha256: str + input_evidence_digests: Mapping[str, str] + expected_snapshot_digest: str + + +def validate_authorization_bundle( + declaration: AuthorizationSpec, + *, + opened_source: OpenedSource, + condition_digests: Mapping[str, str], + expected: Mapping[str, ExpectedDataset], +) -> ValidatedAuthorizationBundle: + """Validate one authorization artifact against an independently opened + source and the caller's own known condition digests/expected datasets. + + Fails closed on: source tampering, an authorization_type/executable + mismatch, a grant referencing an unknown condition or an out-of-selection + question, an empty reason, or a self-assigned authorization_id that does + not match the bundle's own recomputed digest. + """ + if opened_source.source_id != declaration.source_id: + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} declares source " + f"{declaration.source_id!r} but was validated against a differently " + f"declared opened source {opened_source.source_id!r}." + ) + if declaration.authorization_type == "inference_repair" and not declaration.executable: + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} declares inference_repair " + "but executable=false." + ) + if declaration.authorization_type == "offline_transformation" and declaration.executable: + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} declares " + "offline_transformation but executable=true." + ) + + grants: dict[str, dict[str, str]] = {} + for condition_digest, question_reasons in declaration.condition_question_reasons.items(): + if condition_digest not in condition_digests.values(): + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} grants a condition " + f"{condition_digest!r} that is not among the caller's known conditions." + ) + dataset = expected.get(condition_digest) + if not isinstance(question_reasons, Mapping) or not question_reasons: + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} has an empty grant for " + f"condition {condition_digest!r}." + ) + selected = set(dataset.selected_question_ids) if dataset is not None else None + normalized_reasons: dict[str, str] = {} + for question_id, reason in question_reasons.items(): + if not isinstance(reason, str) or not reason.strip(): + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} grant for " + f"{condition_digest!r}/{question_id!r} has no explicit reason." + ) + if selected is not None and question_id not in selected: + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} grants question " + f"{question_id!r} that is not in condition {condition_digest!r}'s " + "expected selection." + ) + normalized_reasons[question_id] = reason + grants[condition_digest] = normalized_reasons + if not grants: + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} grants no conditions." + ) + + bundle_payload = { + "schema_version": "choicebench.authorization-bundle.v1", + "authorization_type": declaration.authorization_type, + "grants": grants, + "authority": declaration.authority, + "purpose": declaration.purpose, + "executable": declaration.executable, + "source_sha256": opened_source.sha256, + "input_evidence_digests": dict(declaration.input_evidence_digests), + "expected_snapshot_digests": dict(declaration.expected_snapshot_digests), + } + authorization_digest = integrity_digest(bundle_payload) + expected_id = short_id("auth", bundle_payload) + if declaration.authorization_id != expected_id: + raise AuthorizationError( + f"Authorization ID {declaration.authorization_id!r} does not match its own " + f"recomputed digest (expected {expected_id!r}); it cannot self-assign an ID." + ) + if set(declaration.expected_snapshot_digests) != set(grants): + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} expected_snapshot_digests " + "do not exactly cover its granted conditions." + ) + + return ValidatedAuthorizationBundle( + authorization_id=expected_id, + authorization_digest=authorization_digest, + authorization_type=declaration.authorization_type, + grants=grants, + authority=declaration.authority, + purpose=declaration.purpose, + executable=declaration.executable, + source_sha256=opened_source.sha256, + input_evidence_digests=dict(declaration.input_evidence_digests), + expected_snapshot_digests=dict(declaration.expected_snapshot_digests), + ) + + +def authorization_for_condition( + bundle: ValidatedAuthorizationBundle, *, condition_digest: str +) -> ValidatedAuthorization: + """Slice one condition's grant out of a bundle without weakening its + aggregate identity: bundle_id/bundle_digest are the whole bundle's, so a + slice can be traced back to (and cannot silently diverge from) the + exact bundle it came from, and it can never borrow another condition's + grants.""" + question_reasons = bundle.grants.get(condition_digest) + if question_reasons is None: + raise AuthorizationError( + f"Authorization {bundle.authorization_id!r} grants no questions for " + f"condition {condition_digest!r}; this condition cannot self-authorize." + ) + expected_snapshot_digest = bundle.expected_snapshot_digests.get(condition_digest) + if expected_snapshot_digest is None: + raise AuthorizationError( + f"Authorization {bundle.authorization_id!r} has no expected snapshot digest " + f"for condition {condition_digest!r}." + ) + return ValidatedAuthorization( + bundle_id=bundle.authorization_id, + bundle_digest=bundle.authorization_digest, + authorization_type=bundle.authorization_type, + condition_digest=condition_digest, + question_reasons=dict(question_reasons), + authority=bundle.authority, + purpose=bundle.purpose, + executable=bundle.executable, + source_sha256=bundle.source_sha256, + input_evidence_digests=dict(bundle.input_evidence_digests), + expected_snapshot_digest=expected_snapshot_digest, + ) diff --git a/src/choicebench/importing/overlays.py b/src/choicebench/importing/overlays.py new file mode 100644 index 0000000..eacd484 --- /dev/null +++ b/src/choicebench/importing/overlays.py @@ -0,0 +1,266 @@ +"""Pure derivation of a repair/offline-transformation overlay from an +already-verified base realization. + +Scope boundary: this module never opens a filesystem path, runs inference, or +merges the derived rows into a full realization -- it only validates the +overlay's binding to its base/authorization and produces the lineage/row +pieces for the *replaced* questions. Merging those with the base's retained +rows into a complete result set and constructing the final realization is +the import engine's job (it alone has the full base row set, the base run +directory, and the destination for the new run). +""" + +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any, Mapping + +from choicebench.identity import integrity_digest +from choicebench.importing.authorization import ValidatedAuthorization +from choicebench.importing.csv_adapter import AdaptedTable +from choicebench.importing.dataset_reference import ExpectedDataset +from choicebench.importing.identity import make_lineage_component, make_result_origin +from choicebench.importing.schema import OverlaySpec + +_EVIDENCE_STATUSES = { + "complete", "qualified", "partial", "malformed", "recoverable", "failed", +} +_REPAIR_ORIGINS = {"native_inference", "external_repair_inference"} + + +class OverlayError(ValueError): + """Raised when an overlay declaration or its derivation is invalid or unsafe.""" + + +@dataclass(frozen=True) +class VerifiedBaseRealization: + condition_id: str + condition_digest: str + realization_id: str + realization_digest: str + evidence_index_digest: str + evidence_source_digests: Mapping[str, str] + validation_artifact_sha256: str + result_sha256: str | None + rows_by_question_id: Mapping[str, Mapping[str, Any]] + prediction_origins: Mapping[str, str] + + +@dataclass(frozen=True) +class DerivedRealizationPayload: + condition_id: str + condition_digest: str + preownership_rows: tuple[Mapping[str, Any], ...] + lineage_components: tuple[Mapping[str, Any], ...] + result_origin: Mapping[str, Any] + evidence_status: str + replacement_question_ids: tuple[str, ...] + + +def _rows_by_question_id(table: AdaptedTable, mapping: Mapping[str, str]) -> dict[str, Any]: + question_id_column = mapping["question_id"] + rows_by_qid: dict[str, list] = {} + for row in table.rows: + qid = row.values.get(question_id_column) + if qid is None: + continue + rows_by_qid.setdefault(str(qid), []).append(row) + return rows_by_qid + + +def derive_overlay( + *, + base: VerifiedBaseRealization, + overlay: OverlaySpec, + authorization: ValidatedAuthorization, + overlay_table: AdaptedTable, + overlay_mapping: Mapping[str, str], + expected: ExpectedDataset, +) -> DerivedRealizationPayload: + """Derive the replaced-row lineage/content for one authorized overlay. + + `base` must be a value the caller obtained by independently verifying the + full base run graph (never a bare path or unverified dict) -- this + function performs no filesystem or verification work of its own, only + binding checks against the values it is given. + """ + if authorization.condition_digest != base.condition_digest: + raise OverlayError( + "Authorization condition does not match the base realization's " + "condition; an authorization cannot be borrowed across conditions." + ) + if ( + overlay.base_condition_digest != base.condition_digest + or overlay.base_realization_id != base.realization_id + or overlay.base_realization_digest != base.realization_digest + ): + raise OverlayError( + "Overlay base condition/realization binding does not match the " + "verified base realization." + ) + if overlay.base_validation_artifact_sha256 != base.validation_artifact_sha256: + raise OverlayError("Overlay base validation-artifact checksum does not match.") + if overlay.base_result_sha256 != base.result_sha256: + raise OverlayError("Overlay base result checksum does not match.") + if overlay.authorization_id != authorization.bundle_id: + raise OverlayError( + "Overlay declares a different authorization bundle than the one supplied." + ) + + if set(overlay.base_evidence_digests) != set(authorization.input_evidence_digests): + raise OverlayError( + "Overlay base_evidence_digests do not exactly match the authorization's " + "input evidence digests." + ) + for key, digest in authorization.input_evidence_digests.items(): + if ( + overlay.base_evidence_digests.get(key) != digest + or base.evidence_source_digests.get(key) != digest + ): + raise OverlayError( + f"Overlay evidence digest {key!r} is not owned by the base realization's " + "own evidence sources." + ) + + replacement_ids = set(overlay.replacement_reasons) + if not replacement_ids: + raise OverlayError("Overlay declares no replacement question IDs.") + unauthorized = replacement_ids - set(authorization.question_reasons) + if unauthorized: + raise OverlayError( + f"Overlay replaces question(s) {sorted(unauthorized)} that are not granted " + "by its authorization." + ) + unexpected = replacement_ids - set(expected.selected_question_ids) + if unexpected: + raise OverlayError( + f"Overlay replaces question(s) {sorted(unexpected)} outside the expected " + "dataset selection." + ) + + per_question_origin = overlay.result_origin.per_question_prediction_origins + default_origin = overlay.result_origin.default_prediction_origin + + if authorization.authorization_type == "inference_repair": + operation_type = "repair_overlay" + if overlay.result_origin.derivation_origin != operation_type: + raise OverlayError( + f"inference_repair overlay must declare derivation_origin=" + f"{operation_type!r}, found {overlay.result_origin.derivation_origin!r}." + ) + extra_origins = set(per_question_origin) - replacement_ids + if extra_origins: + raise OverlayError( + f"inference_repair origin assignment(s) for unreplaced question(s) " + f"{sorted(extra_origins)} are not permitted." + ) + resolved_origins: dict[str, str] = {} + for qid in replacement_ids: + origin = per_question_origin.get(qid, default_origin) + if origin is None: + raise OverlayError( + f"inference_repair overlay has no declared origin assignment for " + f"replaced question {qid!r}." + ) + if origin not in _REPAIR_ORIGINS: + raise OverlayError( + f"inference_repair origin {origin!r} for question {qid!r} is not a " + f"valid repair prediction origin {sorted(_REPAIR_ORIGINS)}." + ) + resolved_origins[qid] = origin + elif authorization.authorization_type == "offline_transformation": + operation_type = "offline_transformation" + if overlay.result_origin.derivation_origin != operation_type: + raise OverlayError( + f"offline_transformation overlay must declare derivation_origin=" + f"{operation_type!r}, found {overlay.result_origin.derivation_origin!r}." + ) + if default_origin is not None: + raise OverlayError( + "offline_transformation cannot declare a default prediction origin; it " + "retains each underlying response's own origin." + ) + resolved_origins = {} + for qid in replacement_ids: + origin = per_question_origin.get(qid) or base.prediction_origins.get(qid) + if origin is None: + raise OverlayError( + f"offline_transformation has no underlying prediction origin for " + f"question {qid!r}." + ) + resolved_origins[qid] = origin + else: + raise OverlayError( + f"Unsupported authorization_type {authorization.authorization_type!r}." + ) + + overlay_rows_by_qid = _rows_by_question_id(overlay_table, overlay_mapping) + prediction_column = overlay_mapping.get("prediction") + preownership_rows: list[dict[str, Any]] = [] + lineage_components: list[Mapping[str, Any]] = [] + for qid in sorted(replacement_ids): + rows = overlay_rows_by_qid.get(qid, []) + if len(rows) != 1: + raise OverlayError( + f"Overlay source has {len(rows)} row(s) for replaced question " + f"{qid!r}; exactly one is required." + ) + overlay_row = rows[0] + expected_matches = expected.frame[expected.frame["question_id"].astype(str) == qid] + if expected_matches.empty: + raise OverlayError(f"Replaced question {qid!r} is not in the expected dataset.") + expected_row = expected_matches.iloc[0].to_dict() + predicted_option = None + if prediction_column is not None: + raw_prediction = overlay_row.values.get(prediction_column) + predicted_option = None if raw_prediction is None else str(raw_prediction).strip().upper() + origin = resolved_origins[qid] + row_payload = { + "question_id": qid, + "question_text": expected_row.get("question_text"), + "correct_option": expected_row.get("correct_option"), + "choices_json": expected_row.get("choices_json"), + "prediction_origin": origin, + "predicted_option": predicted_option, + } + preownership_rows.append(row_payload) + + input_digest = integrity_digest( + {"question_id": qid, "base_row": base.rows_by_question_id.get(qid)} + ) + preownership_output_digest = integrity_digest({"question_id": qid, "row": row_payload}) + component = make_lineage_component( + operation_type=operation_type, + question_id=qid, + parent_digests=(base.realization_digest,), + source_digests=(overlay_table.source_sha256,), + authorization_digest=authorization.bundle_digest, + implementation=dict(overlay.implementation), + implementation_mode="declared_external", + parameters={}, + input_digest=input_digest, + preownership_output_digest=preownership_output_digest, + prediction_origin=origin, + ) + lineage_components.append(component) + + row_assignments = [ + (component["identity"]["question_id"], component["identity"]["prediction_origin"], component["lineage_id"]) + for component in lineage_components + ] + result_origin = make_result_origin(derivation_origin=operation_type, row_assignments=row_assignments) + + if overlay.expected_evidence_status not in _EVIDENCE_STATUSES: + raise OverlayError( + f"Overlay expected_evidence_status {overlay.expected_evidence_status!r} is invalid." + ) + + return DerivedRealizationPayload( + condition_id=base.condition_id, + condition_digest=base.condition_digest, + preownership_rows=tuple(preownership_rows), + lineage_components=tuple(lineage_components), + result_origin=result_origin, + evidence_status=overlay.expected_evidence_status, + replacement_question_ids=tuple(sorted(replacement_ids)), + ) diff --git a/tests/importing/test_authorization.py b/tests/importing/test_authorization.py new file mode 100644 index 0000000..680e336 --- /dev/null +++ b/tests/importing/test_authorization.py @@ -0,0 +1,255 @@ +"""Tests for typed authorization bundle validation (Task 10, reduced scope).""" + +from __future__ import annotations + +from hashlib import sha256 +from pathlib import Path + +import pytest + +from choicebench.identity import integrity_digest, short_id +from choicebench.importing.authorization import ( + AuthorizationError, + authorization_for_condition, + validate_authorization_bundle, +) +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.schema import AuthorizationSpec +from tests.importing.test_identity import _dataset + + +def _source(data: bytes = b"authorization,payload\n1,2\n") -> OpenedSource: + return OpenedSource( + source_id="auth-source", + audit_path=Path("/machine-a/authorization.csv"), + logical_path="inputs/authorization.csv", + data=data, + sha256=sha256(data).hexdigest(), + ) + + +def _bundle_payload(*, authorization_type, executable, condition_digest, grants, source, purpose="repair"): + return { + "schema_version": "choicebench.authorization-bundle.v1", + "authorization_type": authorization_type, + "grants": {condition_digest: dict(grants)}, + "authority": "principal-investigator", + "purpose": purpose, + "executable": executable, + "source_sha256": source.sha256, + "input_evidence_digests": {"ev1": "1" * 64}, + "expected_snapshot_digests": {condition_digest: "2" * 64}, + } + + +def _declaration(*, authorization_type, executable, condition_digest, grants, source, authorization_id=None): + payload = _bundle_payload( + authorization_type=authorization_type, executable=executable, + condition_digest=condition_digest, grants=grants, source=source, + ) + computed_id = short_id("auth", payload) + return AuthorizationSpec( + authorization_id=authorization_id or computed_id, + authorization_type=authorization_type, + source_id="auth-source", + condition_question_reasons={condition_digest: dict(grants)}, + authority="principal-investigator", + purpose="repair", + executable=executable, + input_evidence_digests={"ev1": "1" * 64}, + expected_snapshot_digests={condition_digest: "2" * 64}, + ) + + +def test_validate_authorization_bundle_accepts_a_well_formed_inference_repair_grant(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q1": "malformed prediction"}, source=source, + ) + bundle = validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + assert bundle.authorization_type == "inference_repair" + assert bundle.grants[condition_digest] == {"q1": "malformed prediction"} + + +def test_validate_authorization_bundle_rejects_source_id_mismatch(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + wrong_id_source = OpenedSource( + source_id="a-different-source-id", + audit_path=Path("/machine-a/authorization.csv"), + logical_path="inputs/authorization.csv", + data=source.data, + sha256=source.sha256, + ) + with pytest.raises(AuthorizationError, match="source"): + validate_authorization_bundle( + declaration, opened_source=wrong_id_source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_tampered_source_bytes(): + """Bundle identity is bound to the exact opened source bytes: swapping the + source content for a different byte-identical-source_id file changes the + recomputed authorization_id, so a declared ID from the original bytes is + refused rather than silently accepted against different content.""" + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + tampered_source = _source(data=b"different,bytes\n9,9\n") + with pytest.raises(AuthorizationError, match="recomputed digest"): + validate_authorization_bundle( + declaration, opened_source=tampered_source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_inference_repair_marked_nonexecutable(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=False, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + with pytest.raises(AuthorizationError, match="executable"): + validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_offline_transformation_marked_executable(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="offline_transformation", executable=True, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + with pytest.raises(AuthorizationError, match="executable"): + validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_unknown_condition(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + with pytest.raises(AuthorizationError, match="known conditions"): + validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": "d" * 64}, # different digest + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_question_outside_selection(): + expected = _dataset() # selected_question_ids are q2, q1 + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q9": "reason"}, source=source, + ) + with pytest.raises(AuthorizationError, match="expected selection"): + validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_empty_reason(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q1": " "}, source=source, + ) + with pytest.raises(AuthorizationError, match="explicit reason"): + validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_self_assigned_id(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + authorization_id="auth_" + "0" * 16, + ) + with pytest.raises(AuthorizationError, match="recomputed digest"): + validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_authorization_for_condition_slices_one_condition(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="offline_transformation", executable=False, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + bundle = validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + sliced = authorization_for_condition(bundle, condition_digest=condition_digest) + assert sliced.bundle_id == bundle.authorization_id + assert sliced.bundle_digest == bundle.authorization_digest + assert sliced.question_reasons == {"q1": "reason"} + + +def test_authorization_for_condition_refuses_unauthorized_condition(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="offline_transformation", executable=False, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + bundle = validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + with pytest.raises(AuthorizationError, match="self-authorize"): + authorization_for_condition(bundle, condition_digest="d" * 64) diff --git a/tests/importing/test_overlays.py b/tests/importing/test_overlays.py new file mode 100644 index 0000000..6e27a43 --- /dev/null +++ b/tests/importing/test_overlays.py @@ -0,0 +1,305 @@ +"""Tests for pure overlay derivation (Task 10, reduced scope).""" + +from __future__ import annotations + +from pathlib import Path + +import pytest + +from choicebench.importing.authorization import ValidatedAuthorization +from choicebench.importing.csv_adapter import AdaptedTable, LogicalRecordSpan, SourceRow +from choicebench.importing.overlays import ( + OverlayError, + VerifiedBaseRealization, + derive_overlay, +) +from choicebench.importing.schema import OverlaySpec, ResultOriginSpec +from tests.importing.test_identity import _dataset + +_BASE_CONDITION_DIGEST = "c" * 64 +_BASE_REALIZATION_ID = "real_" + "a" * 16 +_BASE_REALIZATION_DIGEST = "d" * 64 +_EVIDENCE_DIGESTS = {"ev1": "1" * 64} +_VALIDATION_SHA = "3" * 64 +_RESULT_SHA = "4" * 64 +_AUTH_BUNDLE_ID = "auth_" + "b" * 16 +_AUTH_BUNDLE_DIGEST = "5" * 64 + + +def _base(**overrides): + defaults = dict( + condition_id="cond_a", + condition_digest=_BASE_CONDITION_DIGEST, + realization_id=_BASE_REALIZATION_ID, + realization_digest=_BASE_REALIZATION_DIGEST, + evidence_index_digest="2" * 64, + evidence_source_digests=_EVIDENCE_DIGESTS, + validation_artifact_sha256=_VALIDATION_SHA, + result_sha256=_RESULT_SHA, + rows_by_question_id={"q1": {"question_id": "q1"}, "q2": {"question_id": "q2"}}, + prediction_origins={"q1": "external_historical_inference", "q2": "external_historical_inference"}, + ) + defaults.update(overrides) + return VerifiedBaseRealization(**defaults) + + +def _authorization(*, authorization_type, executable, question_reasons): + return ValidatedAuthorization( + bundle_id=_AUTH_BUNDLE_ID, + bundle_digest=_AUTH_BUNDLE_DIGEST, + authorization_type=authorization_type, + condition_digest=_BASE_CONDITION_DIGEST, + question_reasons=dict(question_reasons), + authority="principal-investigator", + purpose="repair", + executable=executable, + source_sha256="6" * 64, + input_evidence_digests=_EVIDENCE_DIGESTS, + expected_snapshot_digest="7" * 64, + ) + + +def _overlay(*, result_origin, replacement_reasons, authorization_id=_AUTH_BUNDLE_ID, **overrides): + defaults = dict( + overlay_id="ovl_1", + base_run_path=Path("/tmp/base-run"), + base_condition_digest=_BASE_CONDITION_DIGEST, + base_realization_id=_BASE_REALIZATION_ID, + base_realization_digest=_BASE_REALIZATION_DIGEST, + base_evidence_digests=_EVIDENCE_DIGESTS, + base_validation_artifact_sha256=_VALIDATION_SHA, + base_result_sha256=_RESULT_SHA, + source_id="overlay-source", + authorization_id=authorization_id, + replacement_reasons=dict(replacement_reasons), + result_origin=result_origin, + lineage_notes={}, + # source_digest must match the overlay source's own checksum ("b" * 64, + # see _overlay_table) -- make_lineage_component requires a declared + # external implementation's code to be one of the declared sources. + implementation={"qualified_name": "external:repair_tool", "source_digest": "b" * 64}, + input_digest="9" * 64, + preownership_output_digest="a" * 64, + expected_evidence_status="complete", + ) + defaults.update(overrides) + return OverlaySpec(**defaults) + + +def _overlay_table(rows): + return AdaptedTable(columns=("qid", "predicted_letter"), rows=rows, source_sha256="b" * 64) + + +def _overlay_row(qid, predicted): + return SourceRow( + values={"qid": qid, "predicted_letter": predicted}, + span=LogicalRecordSpan(index=0, start=0, end=9, terminator=b"\n"), + raw_sha256="0" * 64, + ) + + +_MAPPING = {"question_id": "qid", "prediction": "predicted_letter"} + + +def test_derive_overlay_valid_inference_repair(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", + default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + replacement_reasons={"q1": "repair malformed answer"}, + ) + derived = derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + assert derived.replacement_question_ids == ("q1",) + assert derived.preownership_rows[0]["prediction_origin"] == "native_inference" + assert derived.preownership_rows[0]["predicted_option"] == "B" + assert derived.result_origin["derivation_origin"] == "repair_overlay" + assert derived.lineage_components[0]["identity"]["operation_type"] == "repair_overlay" + assert derived.lineage_components[0]["identity"]["authorization_digest"] == _AUTH_BUNDLE_DIGEST + + +def test_derive_overlay_valid_offline_transformation_retains_underlying_origin(): + authorization = _authorization( + authorization_type="offline_transformation", executable=False, question_reasons={"q2": "rematch"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="offline_transformation", + default_prediction_origin=None, + per_question_prediction_origins={}, + ), + replacement_reasons={"q2": "offline semantic rematch"}, + ) + derived = derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q2", "A"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + # base.prediction_origins["q2"] == "external_historical_inference" -- retained, not reassigned + assert derived.preownership_rows[0]["prediction_origin"] == "external_historical_inference" + assert derived.result_origin["derivation_origin"] == "offline_transformation" + + +def test_derive_overlay_refuses_cross_condition_authorization(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + other_base = _base(condition_digest="e" * 64) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + replacement_reasons={"q1": "reason"}, + base_condition_digest="e" * 64, + ) + with pytest.raises(OverlayError, match="condition"): + derive_overlay( + base=other_base, overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_base_realization_mismatch(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + replacement_reasons={"q1": "reason"}, + base_realization_digest="f" * 64, # forged + ) + with pytest.raises(OverlayError, match="base"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_evidence_digest_mismatch(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + replacement_reasons={"q1": "reason"}, + base_evidence_digests={"ev1": "9" * 64}, # diverges from authorization/base + ) + with pytest.raises(OverlayError, match="evidence"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_unauthorized_replacement(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q2": "native_inference"}, + ), + replacement_reasons={"q2": "not actually authorized"}, + ) + with pytest.raises(OverlayError, match="not granted"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q2", "A"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_missing_repair_origin_assignment(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={}, # no assignment for q1 + ), + replacement_reasons={"q1": "reason"}, + ) + with pytest.raises(OverlayError, match="origin assignment"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_invalid_repair_origin_value(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q1": "mixed"}, + ), + replacement_reasons={"q1": "reason"}, + ) + with pytest.raises(OverlayError, match="valid repair prediction origin"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_duplicate_replacement_row(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + replacement_reasons={"q1": "reason"}, + ) + with pytest.raises(OverlayError, match="exactly one"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"), _overlay_row("q1", "A"))), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_authorization_bundle_mismatch(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + replacement_reasons={"q1": "reason"}, + authorization_id="auth_" + "9" * 16, # different bundle than the one supplied + ) + with pytest.raises(OverlayError, match="authorization bundle"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) From 5b3ec6d5e0dd6aa4b8dadba8084c98a144ef7ae5 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 16:33:25 +0300 Subject: [PATCH 33/47] feat: stage external imports atomically with verified evidence storage Task 11, reduced scope. evidence.py: open_verified_source reads a declared regular file exactly once and hashes those same bytes forward (never re-reads or re-derives the digest separately), refuses symlinks and paths outside an optional containment root, and requires the read bytes match the declared expected_sha256. write_evidence_blob stores referenced row-level bytes at a run-local sha256-addressed path; identical bytes already staged are reused, any divergent existing blob or sidecar is refused. write_evidence_index/validate_evidence_index publish and verify one self-digested index over all evidence records, transitively validating every referenced blob. transaction.py: fsync_tree flushes every file and directory under a staged tree before publication. atomic_publish_directory_no_replace uses a real renameat2(RENAME_NOREPLACE) syscall wrapper (isolated so tests can mock its absence) and fails closed if the primitive is unavailable -- never falls back to os.replace or an existence-check race. ImportTransaction holds the run's manifest lock for its lifetime, stages new runs under a same-filesystem sibling directory marked with an owner file (so cleanup only ever touches its own staging directory, never an unrelated leftover), and for an existing run skips staging entirely -- publish() calls the caller's validator against the already-published final directory instead. Focused: 19 new tests (11 evidence, 8 transaction), all passing in under 4s together with tests/test_release_adversarial.py and tests/infra/. Full suite: 1282 passed (1263 baseline + 19), only the pre-existing Python 3.14 google.genai.types deprecation warning. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/importing/evidence.py | 168 +++++++++++++++++++++++ src/choicebench/importing/transaction.py | 158 +++++++++++++++++++++ tests/importing/test_evidence.py | 168 +++++++++++++++++++++++ tests/importing/test_transaction.py | 127 +++++++++++++++++ 4 files changed, 621 insertions(+) create mode 100644 src/choicebench/importing/evidence.py create mode 100644 src/choicebench/importing/transaction.py create mode 100644 tests/importing/test_evidence.py create mode 100644 tests/importing/test_transaction.py diff --git a/src/choicebench/importing/evidence.py b/src/choicebench/importing/evidence.py new file mode 100644 index 0000000..80bd8d0 --- /dev/null +++ b/src/choicebench/importing/evidence.py @@ -0,0 +1,168 @@ +"""Open declared sources with a verified checksum and store referenced +row-level evidence bytes as run-local content-addressed blobs. + +Reduced scope: covers the concrete safety guarantees the repair-import +workflow needs -- read-once-hash-forward source opening, path containment, +symlink refusal, and a self-digested evidence index -- without the full +historical suffix-derivation/dedup-across-runs richness of the original plan. +""" + +from __future__ import annotations + +from hashlib import sha256 +import json +from pathlib import Path +from typing import Any, Mapping, Sequence + +from choicebench.identity import canonicalize, integrity_digest +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.schema import SourceArtifactSpec +from choicebench.infra.artifacts import atomic_write_bytes, atomic_write_json + +EVIDENCE_INDEX_SCHEMA_VERSION = "choicebench.evidence-index.v1" +_SAFE_SUFFIXES = {"csv": "csv"} + + +class EvidenceError(ValueError): + """Raised when a source or stored evidence fails a safety or integrity check.""" + + +def open_verified_source( + declaration: SourceArtifactSpec, *, containment_root: Path | None = None +) -> OpenedSource: + """Open exactly the declared regular file, read its bytes once, and hash + those same bytes -- never re-read or re-derive the digest from a + separately reopened handle.""" + path = Path(declaration.path) + if containment_root is not None: + root = Path(containment_root).resolve() + try: + path.resolve().relative_to(root) + except ValueError as exc: + raise EvidenceError( + f"Source {declaration.source_id!r} path {path} escapes containment " + f"root {root}." + ) from exc + if path.is_symlink(): + raise EvidenceError(f"Source {declaration.source_id!r} path {path} is a symlink.") + if not path.is_file(): + raise EvidenceError(f"Source {declaration.source_id!r} path {path} is not a regular file.") + data = path.read_bytes() + actual_sha256 = sha256(data).hexdigest() + if actual_sha256 != declaration.expected_sha256: + raise EvidenceError( + f"Source {declaration.source_id!r} checksum mismatch: expected " + f"{declaration.expected_sha256}, found {actual_sha256}." + ) + return OpenedSource( + source_id=declaration.source_id, + audit_path=path, + logical_path=declaration.logical_path, + data=data, + sha256=actual_sha256, + ) + + +def evidence_blob_path(staged_run: Path, sha256_digest: str, format_name: str) -> Path: + suffix = _SAFE_SUFFIXES.get(format_name) + if suffix is None: + raise EvidenceError(f"Unsupported evidence blob format {format_name!r}.") + if len(sha256_digest) != 64 or any(c not in "0123456789abcdef" for c in sha256_digest): + raise EvidenceError(f"Invalid evidence blob digest {sha256_digest!r}.") + return ( + Path(staged_run) + / "artifacts" / "imports" / "evidence" / "sha256" + / sha256_digest[:2] / f"{sha256_digest}.{suffix}" + ) + + +def write_evidence_blob( + staged_run: Path, source: OpenedSource, references: Sequence[str] +) -> dict[str, Any]: + """Store one referenced source's bytes as a run-local content-addressed + blob. Identical bytes already staged (e.g. two references to the same + source) are reused rather than rewritten; a divergent existing blob at + the same digest path is refused.""" + blob_path = evidence_blob_path(staged_run, source.sha256, "csv") + if blob_path.exists(): + if blob_path.read_bytes() != source.data: + raise EvidenceError( + f"Refusing to overwrite a divergent existing evidence blob: {blob_path}" + ) + else: + atomic_write_bytes(blob_path, source.data) + record = canonicalize( + { + "source_id": source.source_id, + "logical_path": source.logical_path, + "sha256": source.sha256, + "size": len(source.data), + "format": "csv", + "references": sorted(set(references)), + } + ) + sidecar_path = blob_path.with_suffix(blob_path.suffix + ".json") + if sidecar_path.exists(): + existing = json.loads(sidecar_path.read_text()) + if existing != record: + raise EvidenceError( + f"Refusing to overwrite a divergent evidence blob sidecar: {sidecar_path}" + ) + else: + atomic_write_json(sidecar_path, record) + return record + + +def validate_evidence_blob(run_dir: Path, record: Mapping[str, Any]) -> None: + blob_path = evidence_blob_path(run_dir, record["sha256"], record["format"]) + if not blob_path.is_file(): + raise EvidenceError(f"Missing evidence blob: {blob_path}") + actual = sha256(blob_path.read_bytes()).hexdigest() + if actual != record["sha256"]: + raise EvidenceError( + f"Evidence blob content integrity check failed: {blob_path}; expected " + f"{record['sha256']}, found {actual}." + ) + sidecar_path = blob_path.with_suffix(blob_path.suffix + ".json") + try: + stored = json.loads(sidecar_path.read_text()) + except (OSError, json.JSONDecodeError) as exc: + raise EvidenceError(f"Evidence blob sidecar is unreadable: {sidecar_path}: {exc}") from exc + if canonicalize(stored) != canonicalize(dict(record)): + raise EvidenceError(f"Evidence blob sidecar does not match its declared record: {sidecar_path}") + + +def write_evidence_index( + staged_run: Path, records: Sequence[Mapping[str, Any]] +) -> Mapping[str, Any]: + payload = canonicalize( + { + "schema_version": EVIDENCE_INDEX_SCHEMA_VERSION, + "records": list(records), + } + ) + evidence_index_digest = integrity_digest(payload) + index = {**payload, "evidence_index_digest": evidence_index_digest} + index_path = Path(staged_run) / "artifacts" / "imports" / "evidence" / "index.json" + atomic_write_json(index_path, index) + return index + + +def validate_evidence_index(run_dir: Path, expected_digest: str) -> Mapping[str, Any]: + index_path = Path(run_dir) / "artifacts" / "imports" / "evidence" / "index.json" + try: + index = json.loads(index_path.read_text()) + except (OSError, json.JSONDecodeError) as exc: + raise EvidenceError(f"Evidence index is unreadable: {index_path}: {exc}") from exc + if index.get("schema_version") != EVIDENCE_INDEX_SCHEMA_VERSION: + raise EvidenceError(f"Unsupported evidence index schema: {index_path}") + payload = {key: value for key, value in index.items() if key != "evidence_index_digest"} + if integrity_digest(canonicalize(payload)) != index.get("evidence_index_digest"): + raise EvidenceError(f"Evidence index integrity check failed: {index_path}") + if index["evidence_index_digest"] != expected_digest: + raise EvidenceError( + f"Evidence index digest does not match its expected binding: {index_path}" + ) + for record in index["records"]: + validate_evidence_blob(run_dir, record) + return index diff --git a/src/choicebench/importing/transaction.py b/src/choicebench/importing/transaction.py new file mode 100644 index 0000000..6a335c2 --- /dev/null +++ b/src/choicebench/importing/transaction.py @@ -0,0 +1,158 @@ +"""Atomic, no-replace staged publication for imported runs. + +Reduced scope: covers the concrete safety guarantees the repair-import +workflow needs -- a manifest lock, sibling staging on the same filesystem, +full-tree fsync before publication, and a genuine no-replace rename (never +an existence-check-then-os.replace race) -- without the fuller historical +crash-recovery/cleanup-scope richness of the original plan. +""" + +from __future__ import annotations + +import ctypes +import ctypes.util +import os +from pathlib import Path +from typing import Callable +import shutil +import uuid + +from choicebench.infra.artifacts import FileLock + +_RENAME_NOREPLACE = 0x1 +_AT_FDCWD = -100 + + +class TransactionError(ValueError): + """Raised when staged publication cannot proceed safely.""" + + +def fsync_tree(root: Path) -> None: + """Best-effort fsync of every file and directory under root, so a crash + immediately after this call cannot lose or half-write staged content.""" + root = Path(root) + for dirpath, dirnames, filenames in os.walk(root): + for name in filenames: + path = Path(dirpath) / name + try: + fd = os.open(path, os.O_RDONLY) + except OSError: + continue + try: + os.fsync(fd) + finally: + os.close(fd) + try: + dir_fd = os.open(dirpath, os.O_RDONLY) + except OSError: + continue + try: + os.fsync(dir_fd) + except OSError: + pass + finally: + os.close(dir_fd) + + +def _renameat2_no_replace(old: Path, new: Path) -> None: + """Thin wrapper around the libc renameat2(2) syscall with + RENAME_NOREPLACE, isolated so tests can mock its absence without + depending on kernel version. Never falls back to os.replace or an + existence-check race: if the primitive is unavailable, this fails + closed.""" + library_name = ctypes.util.find_library("c") + if not library_name: + raise TransactionError("libc is not available; refusing an unsafe rename fallback.") + libc = ctypes.CDLL(library_name, use_errno=True) + if not hasattr(libc, "renameat2"): + raise TransactionError( + "renameat2 is not available on this system; refusing an unsafe rename fallback." + ) + result = libc.renameat2( + ctypes.c_int(_AT_FDCWD), + os.fsencode(str(old)), + ctypes.c_int(_AT_FDCWD), + os.fsencode(str(new)), + ctypes.c_uint(_RENAME_NOREPLACE), + ) + if result != 0: + errno = ctypes.get_errno() + raise TransactionError( + f"renameat2 failed renaming {old} -> {new}: errno {errno} ({os.strerror(errno)})." + ) + + +def atomic_publish_directory_no_replace(source: Path, destination: Path) -> None: + """Atomically move a fully-staged directory into its final location. + Refuses (never silently overwrites or races an existence check) if the + destination already exists.""" + source = Path(source) + destination = Path(destination) + destination.parent.mkdir(parents=True, exist_ok=True) + _renameat2_no_replace(source, destination) + + +class ImportTransaction: + """Stage a new run under a same-filesystem sibling directory, run a + caller-supplied full-graph validator, and either publish it atomically + (new run) or verify an already-published run in place (existing run) -- + never partially write the final run directory.""" + + def __init__(self, *, runs_dir: Path, run_id: str) -> None: + self.runs_dir = Path(runs_dir) + self.run_id = run_id + self.final_run = self.runs_dir / run_id + self._staged_run: Path | None = None + self._lock: FileLock | None = None + self._published = False + + def __enter__(self) -> "ImportTransaction": + lock_path = self.runs_dir / ".locks" / f"{self.run_id}.manifest.lock" + self._lock = FileLock(lock_path, f"import run {self.run_id}") + self._lock.__enter__() + if not self.final_run.exists(): + staged = self.runs_dir / f".staging-{self.run_id}-{uuid.uuid4().hex}" + staged.mkdir(parents=True) + (staged / ".staging-owner").write_text(f"pid={os.getpid()}\n") + self._staged_run = staged + return self + + @property + def staged_run(self) -> Path: + if self._staged_run is None: + raise TransactionError( + f"Run {self.run_id!r} already exists; there is no staging directory to " + "write into. Only verify_import_run-style validation is available." + ) + return self._staged_run + + def publish(self, validator: Callable[[Path], None]) -> Path: + if self.final_run.exists(): + validator(self.final_run) + return self.final_run + staged = self.staged_run + if not (staged / ".staging-owner").is_file(): + raise TransactionError( + f"Refusing to publish {staged}: it is not marked as this transaction's " + "own staging directory." + ) + validator(staged) + fsync_tree(staged) + (staged / ".staging-owner").unlink() + atomic_publish_directory_no_replace(staged, self.final_run) + self._published = True + self._staged_run = None + return self.final_run + + def __exit__(self, exc_type, exc, traceback) -> None: + try: + if self._staged_run is not None and self._staged_run.exists(): + # Publication never happened (error, or caller chose not to + # publish); clean up only this transaction's own + # owner-marked staging directory. + if (self._staged_run / ".staging-owner").is_file(): + shutil.rmtree(self._staged_run, ignore_errors=True) + finally: + if self._lock is not None: + self._lock.__exit__(exc_type, exc, traceback) + self._lock = None diff --git a/tests/importing/test_evidence.py b/tests/importing/test_evidence.py new file mode 100644 index 0000000..2ec08b1 --- /dev/null +++ b/tests/importing/test_evidence.py @@ -0,0 +1,168 @@ +"""Tests for verified source opening and content-addressed evidence storage +(Task 11, reduced scope).""" + +from __future__ import annotations + +from dataclasses import replace +from hashlib import sha256 + +import pytest + +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.evidence import ( + EvidenceError, + open_verified_source, + validate_evidence_blob, + validate_evidence_index, + write_evidence_blob, + write_evidence_index, +) +from choicebench.importing.schema import CsvDialectSpec, OptionMappingSpec, SourceArtifactSpec + + +def _declaration(tmp_path, data: bytes, source_id="results"): + path = tmp_path / "results.csv" + path.write_bytes(data) + return SourceArtifactSpec( + source_id=source_id, + path=path, + logical_path="inputs/results.csv", + expected_sha256=sha256(data).hexdigest(), + format="csv", + format_version="1", + classification="raw", + dialect=CsvDialectSpec(), + columns={"question_id": "qid"}, + expected_columns=("qid",), + ignored_columns={}, + null_values=(), + numeric_columns=(), + option_mapping=OptionMappingSpec( + mode="ordered_columns", ordered_columns=("a", "b"), + structured_column=None, structured_label_key=None, structured_text_key=None, + ), + extra_field_policy="reject_unmapped", + preserve_namespace="ext", + source_run_id=None, + source_repository=None, + source_commit=None, + notes={}, + ) + + +def test_open_verified_source_accepts_matching_checksum(tmp_path): + declaration = _declaration(tmp_path, b"qid,a,b\nq1,x,y\n") + opened = open_verified_source(declaration) + assert opened.source_id == "results" + assert opened.sha256 == declaration.expected_sha256 + assert opened.data == b"qid,a,b\nq1,x,y\n" + + +def test_open_verified_source_rejects_checksum_mismatch(tmp_path): + declaration = _declaration(tmp_path, b"qid,a,b\nq1,x,y\n") + declaration.path.write_bytes(b"qid,a,b\nq1,TAMPERED,y\n") + with pytest.raises(EvidenceError, match="checksum"): + open_verified_source(declaration) + + +def test_open_verified_source_rejects_missing_file(tmp_path): + declaration = _declaration(tmp_path, b"qid,a,b\nq1,x,y\n") + declaration.path.unlink() + with pytest.raises(EvidenceError, match="regular file"): + open_verified_source(declaration) + + +def test_open_verified_source_rejects_symlink(tmp_path): + declaration = _declaration(tmp_path, b"qid,a,b\nq1,x,y\n") + real_path = declaration.path + link_path = tmp_path / "link.csv" + link_path.symlink_to(real_path) + linked_declaration = replace(declaration, path=link_path) + with pytest.raises(EvidenceError, match="symlink"): + open_verified_source(linked_declaration) + + +def test_open_verified_source_rejects_path_outside_containment_root(tmp_path): + outside_dir = tmp_path / "outside" + outside_dir.mkdir() + declaration = _declaration(outside_dir, b"qid,a,b\nq1,x,y\n") + containment_root = tmp_path / "inside" + containment_root.mkdir() + with pytest.raises(EvidenceError, match="containment"): + open_verified_source(declaration, containment_root=containment_root) + + +def test_write_and_validate_evidence_blob_round_trip(tmp_path): + data = b"qid,a,b\nq1,x,y\n" + source = OpenedSource( + source_id="results", audit_path=tmp_path / "results.csv", + logical_path="inputs/results.csv", data=data, sha256=sha256(data).hexdigest(), + ) + record = write_evidence_blob(tmp_path, source, references=("q1",)) + validate_evidence_blob(tmp_path, record) # must not raise + assert record["references"] == ["q1"] + + +def test_write_evidence_blob_is_idempotent_for_identical_bytes(tmp_path): + data = b"qid,a,b\nq1,x,y\n" + source = OpenedSource( + source_id="results", audit_path=tmp_path / "results.csv", + logical_path="inputs/results.csv", data=data, sha256=sha256(data).hexdigest(), + ) + first = write_evidence_blob(tmp_path, source, references=("q1",)) + second = write_evidence_blob(tmp_path, source, references=("q1",)) + assert first == second + + +def test_validate_evidence_blob_rejects_tampered_bytes(tmp_path): + data = b"qid,a,b\nq1,x,y\n" + source = OpenedSource( + source_id="results", audit_path=tmp_path / "results.csv", + logical_path="inputs/results.csv", data=data, sha256=sha256(data).hexdigest(), + ) + record = write_evidence_blob(tmp_path, source, references=("q1",)) + from choicebench.importing.evidence import evidence_blob_path + + blob_path = evidence_blob_path(tmp_path, source.sha256, "csv") + blob_path.write_bytes(data + b"tampered") + with pytest.raises(EvidenceError, match="content integrity"): + validate_evidence_blob(tmp_path, record) + + +def test_evidence_index_round_trips(tmp_path): + data = b"qid,a,b\nq1,x,y\n" + source = OpenedSource( + source_id="results", audit_path=tmp_path / "results.csv", + logical_path="inputs/results.csv", data=data, sha256=sha256(data).hexdigest(), + ) + record = write_evidence_blob(tmp_path, source, references=("q1",)) + index = write_evidence_index(tmp_path, [record]) + validated = validate_evidence_index(tmp_path, index["evidence_index_digest"]) + assert validated["records"] == [record] + + +def test_evidence_index_rejects_digest_mismatch(tmp_path): + data = b"qid,a,b\nq1,x,y\n" + source = OpenedSource( + source_id="results", audit_path=tmp_path / "results.csv", + logical_path="inputs/results.csv", data=data, sha256=sha256(data).hexdigest(), + ) + record = write_evidence_blob(tmp_path, source, references=("q1",)) + write_evidence_index(tmp_path, [record]) + with pytest.raises(EvidenceError, match="digest"): + validate_evidence_index(tmp_path, "0" * 64) + + +def test_evidence_index_rejects_missing_blob(tmp_path): + data = b"qid,a,b\nq1,x,y\n" + source = OpenedSource( + source_id="results", audit_path=tmp_path / "results.csv", + logical_path="inputs/results.csv", data=data, sha256=sha256(data).hexdigest(), + ) + record = write_evidence_blob(tmp_path, source, references=("q1",)) + index = write_evidence_index(tmp_path, [record]) + from choicebench.importing.evidence import evidence_blob_path + + evidence_blob_path(tmp_path, source.sha256, "csv").unlink() + with pytest.raises(EvidenceError, match="Missing evidence blob"): + validate_evidence_index(tmp_path, index["evidence_index_digest"]) diff --git a/tests/importing/test_transaction.py b/tests/importing/test_transaction.py new file mode 100644 index 0000000..fd07690 --- /dev/null +++ b/tests/importing/test_transaction.py @@ -0,0 +1,127 @@ +"""Tests for atomic staged publication (Task 11, reduced scope).""" + +from __future__ import annotations + +import pytest + +from choicebench.infra.artifacts import LockHeldError +from choicebench.importing.transaction import ( + ImportTransaction, + TransactionError, + atomic_publish_directory_no_replace, + fsync_tree, +) +import choicebench.importing.transaction as transaction_module + + +def test_fsync_tree_does_not_raise_on_a_populated_directory(tmp_path): + (tmp_path / "a").mkdir() + (tmp_path / "a" / "file.txt").write_text("hello") + fsync_tree(tmp_path) # must not raise + + +def test_atomic_publish_directory_no_replace_moves_staged_tree(tmp_path): + staged = tmp_path / "staged" + staged.mkdir() + (staged / "file.txt").write_text("content") + destination = tmp_path / "runs" / "run-1" + atomic_publish_directory_no_replace(staged, destination) + assert destination.is_dir() + assert (destination / "file.txt").read_text() == "content" + assert not staged.exists() + + +def test_atomic_publish_directory_no_replace_refuses_existing_destination(tmp_path): + staged = tmp_path / "staged" + staged.mkdir() + destination = tmp_path / "runs" / "run-1" + destination.mkdir(parents=True) + (destination / "existing.txt").write_text("already here") + with pytest.raises(TransactionError): + atomic_publish_directory_no_replace(staged, destination) + assert (destination / "existing.txt").read_text() == "already here" # untouched + assert staged.exists() # never moved + + +def test_renameat2_wrapper_fails_closed_when_unavailable(tmp_path, monkeypatch): + """If renameat2 cannot be located, this must refuse -- never silently + fall back to os.replace (which would race an existence check).""" + monkeypatch.setattr(transaction_module.ctypes.util, "find_library", lambda name: None) + staged = tmp_path / "staged" + staged.mkdir() + destination = tmp_path / "runs" / "run-1" + with pytest.raises(TransactionError, match="libc"): + atomic_publish_directory_no_replace(staged, destination) + assert staged.exists() + assert not destination.exists() + + +def test_import_transaction_publishes_a_new_run(tmp_path): + runs_dir = tmp_path / "runs" + runs_dir.mkdir() + validated_paths = [] + + def validator(path): + validated_paths.append(path) + (path / "manifest.json").write_text("{}") + + with ImportTransaction(runs_dir=runs_dir, run_id="run-1") as txn: + (txn.staged_run / "evidence.txt").write_text("data") + final_path = txn.publish(validator) + + assert final_path == runs_dir / "run-1" + assert (final_path / "evidence.txt").read_text() == "data" + assert (final_path / "manifest.json").read_text() == "{}" + # the validator ran against the staging directory, before the rename + assert len(validated_paths) == 1 + assert validated_paths[0].name.startswith(".staging-run-1") + # staging directory is gone (renamed into place, not left behind) + leftover = [p for p in runs_dir.iterdir() if p.name.startswith(".staging-run-1")] + assert leftover == [] + + +def test_import_transaction_verifies_existing_run_without_staging(tmp_path): + runs_dir = tmp_path / "runs" + final_run = runs_dir / "run-1" + final_run.mkdir(parents=True) + (final_run / "manifest.json").write_text("{}") + + validated_paths = [] + + def validator(path): + validated_paths.append(path) + + with ImportTransaction(runs_dir=runs_dir, run_id="run-1") as txn: + with pytest.raises(TransactionError, match="already exists"): + _ = txn.staged_run + result = txn.publish(validator) + + assert result == final_run + assert validated_paths == [final_run] + leftover = [p for p in runs_dir.iterdir() if p.name.startswith(".staging-run-1")] + assert leftover == [] # no staging directory was ever created + + +def test_import_transaction_cleans_up_staging_on_validator_failure(tmp_path): + runs_dir = tmp_path / "runs" + runs_dir.mkdir() + + def failing_validator(path): + raise ValueError("synthetic validation failure") + + with pytest.raises(ValueError, match="synthetic validation failure"): + with ImportTransaction(runs_dir=runs_dir, run_id="run-1") as txn: + txn.publish(failing_validator) + + assert not (runs_dir / "run-1").exists() + leftover = [p for p in runs_dir.iterdir() if p.name.startswith(".staging-run-1")] + assert leftover == [] + + +def test_import_transaction_holds_the_manifest_lock(tmp_path): + runs_dir = tmp_path / "runs" + runs_dir.mkdir() + with ImportTransaction(runs_dir=runs_dir, run_id="run-1"): + with pytest.raises(LockHeldError): + with ImportTransaction(runs_dir=runs_dir, run_id="run-1"): + pass From 7590e1eb4d49c19ced9a56fe93a052038b2e1e86 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 16:59:39 +0300 Subject: [PATCH 34/47] feat: orchestrate generic dry-run and real imports Task 12, base (non-overlay) import path. build_import_plan resolves sources/datasets/models/methods/prompts, opens and parses each condition's declared CSV(s), validates rows against the expected dataset snapshot, computes evidence status, builds one lineage component per evaluable row, and assembles a v3 manifest -- performing every read/validation/identity step without writing anything. execute_import(dry_run=True) returns the same report a real import would, with wrote_artifacts=False and no workspace directories created. A real import stages evidence blobs/index, the manifest, run state, per-realization validation artifacts, and result CSVs through ImportTransaction, publishing once atomically; an existing run is verified in place (idempotent no-op) rather than restaged. verify_import_run is the shared full-graph verifier: it validates the manifest, run state, evidence index, every validation artifact against its recorded checksum, and every result artifact, returning immutable VerifiedBaseRealization values only after the whole graph passes -- the foundation the not-yet-wired overlay path (see below) will build on. Report schema is reduced from the plan: counts (conditions/evidence_status/ scope_disposition), a defect summary, and condition/realization digests -- enough to confirm exact status/queue reproduction and reject held/excluded work, not the full historical checksum-report/column-disposition breakdown. Two supporting fixes surfaced by wiring real code through Task 5's identity layer: - identity.py's canonicalize() gained bytes/bytearray support (hex-encoded). Real functions reference plain data constants (e.g. csv_adapter.py's UTF-8 BOM marker) that a resolved-globals check now inspects; canonicalize previously had no representation for bytes at all. Narrow and additive: the only previous behavior for bytes was an unconditional TypeError, and the full suite (1289 tests, up from 1282) confirms no existing caller relied on that. - The engine binds importer/lineage runtime-callable identity to small local marker functions rather than the real parse_csv_source/ validate_source_rows: those reference many stdlib globals (hashlib.sha256, a C builtin; re; etc.) that are neither plain data nor inspectable user-level code, which Task 5's runtime-callable identity correctly refuses rather than silently ignoring. Binding to engine.py-local markers still captures "did this engine's own orchestration change" (identity hashes the whole defining file); the installed package version already carried in importer_implementation_identity is the primary signal for "did the installed release change." Documented in engine.py. Scope boundary (unchanged from the module docstring, now proven out by a working base-import path): applying an authorized repair/offline- transformation overlay on top of a published base run is the next increment. Every piece it needs -- identity, validation, authorization, overlay derivation, evidence storage, atomic transactions, and this module's verify_import_run -- is built and tested; the merge-and-publish orchestration itself is not yet wired into execute_import. Focused: 7 new engine tests (plan validity, dry-run no-op, real import, idempotence, verify, failure reporting, report serialization), all passing together with every prior importing/* + schema test (223 total). Full suite: 1289 passed (1282 baseline + 7), only the pre-existing Python 3.14 google.genai.types deprecation warning. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/identity.py | 2 + src/choicebench/importing/engine.py | 660 ++++++++++++++++++++++++++ src/choicebench/importing/evidence.py | 28 +- tests/importing/test_engine.py | 239 ++++++++++ 4 files changed, 919 insertions(+), 10 deletions(-) create mode 100644 src/choicebench/importing/engine.py create mode 100644 tests/importing/test_engine.py diff --git a/src/choicebench/identity.py b/src/choicebench/identity.py index 2ee381e..ec3a492 100644 --- a/src/choicebench/identity.py +++ b/src/choicebench/identity.py @@ -167,6 +167,8 @@ def canonicalize(value: Any, *, redact_secrets: bool = True) -> Any: return sorted(items, key=lambda item: json.dumps(item, sort_keys=True, separators=(",", ":"), ensure_ascii=True)) if isinstance(value, Path): return value.as_posix() + if isinstance(value, (bytes, bytearray)): + return bytes(value).hex() if isinstance(value, str): text = unicodedata.normalize("NFC", value) return redact_text(text) if redact_secrets else text diff --git a/src/choicebench/importing/engine.py b/src/choicebench/importing/engine.py new file mode 100644 index 0000000..3035beb --- /dev/null +++ b/src/choicebench/importing/engine.py @@ -0,0 +1,660 @@ +"""Orchestrate generic dry-run and real imports of external results. + +Reduced scope: this implements the base (non-overlay) import path fully -- +plan, dry-run validate, real staged publish, and idempotent re-verification +of an existing run. Applying an authorized repair/offline-transformation +overlay on top of an already-published base run (Task 10's derive_overlay, +merged with the base's retained rows into a new run) is the next increment; +every piece it needs (identity, validation, authorization, overlay +derivation, evidence storage, atomic transactions, and this module's own +verify_import_run) is already built and tested, but the merge-and-publish +orchestration itself is not yet wired up here. + +Report schema is also reduced from the original plan: it carries the counts +and digests needed to confirm exact status/queue reproduction and reject +held/excluded work, not the full historical checksum-report/column- +disposition/source-classification breakdown. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from pathlib import Path +from typing import Any, Literal, Mapping + +from dataclasses import asdict +import json + +import pandas as pd + +from choicebench.identity import CANONICALIZATION_VERSION, canonicalize, integrity_digest +from choicebench.importing.csv_adapter import parse_csv_source +from choicebench.importing.dataset_reference import ExpectedDataset, build_expected_dataset +from choicebench.importing.evidence import ( + evidence_record, + open_verified_source, + validate_evidence_index, + write_evidence_blob, + write_evidence_index, +) +from choicebench.importing.identity import ( + ImportIdentityError, + build_import_semantic_identity, + importer_implementation_identity, + make_lineage_component, + make_realization, + make_result_origin, +) +from choicebench.importing.overlays import VerifiedBaseRealization +from choicebench.importing.schema import ( + ImportConditionSpec, + ImportSpec, + validate_csv_dialect_identity, + validate_numeric_columns_identity, + validate_option_mapping_identity, +) +from choicebench.importing.transaction import ImportTransaction +from choicebench.importing.validation import ( + ImportValidationError, + normalize_realization_rows, + prepare_realization_validation_artifact, + validate_realization_validation_artifact, + validate_source_rows, + write_realization_validation_artifact, +) +from choicebench.io.writers import ( + prepare_manifest_result, + publish_manifest_result, + validate_result_artifact, +) +from choicebench.infra.artifacts import atomic_write_json +from choicebench.manifest import ( + MANIFEST_FILENAME, + PROTOCOL_V3_VERSION, + RUN_STATE_FILENAME, + initial_run_state_v3, + make_manifest_v3, + validate_manifest_v3, + validate_run_state_v3, +) + +_ADAPTER_STRICT_KEY = "strict" + + +class ImportEngineError(ValueError): + """Raised when an import request or an existing run fails validation.""" + + +def _external_import_adapter(*, strict: bool) -> None: + """Identity marker for this engine's CSV-adaptation stage. + + Deliberately a trivial, side-effect-free function rather than binding + directly to parse_csv_source/validate_source_rows: those reference many + stdlib globals (hashlib.sha256, re, etc.) that are neither plain data nor + inspectable user-level code, and Task 5's runtime-callable identity fails + closed on exactly that shape rather than silently ignoring it. Binding + here instead still captures "did this engine's own orchestration change" + (implementation_identity hashes this whole file); the installed + choicebench package version already in importer_implementation_identity's + "package" field is the primary signal for "did the installed release + (including csv_adapter.py/validation.py) change." + """ + return None + + +def _external_import_validator() -> None: + """Identity marker for this engine's row-validation stage. See + _external_import_adapter for why this is a marker, not the real function.""" + return None + + +@dataclass(frozen=True) +class ImportRequest: + spec: ImportSpec + run_id: str + workspace_root: Path + strict: bool + overlays: tuple[Any, ...] = () + + +@dataclass(frozen=True) +class ImportPlan: + manifest: Mapping[str, Any] + expected_datasets: Mapping[str, ExpectedDataset] + realizations: Mapping[str, Any] + evidence_records: tuple[Mapping[str, Any], ...] + findings: tuple[Any, ...] + + +@dataclass(frozen=True) +class VerifiedImportRun: + manifest: Mapping[str, Any] + manifest_digest: str + realizations: Mapping[str, VerifiedBaseRealization] + + +@dataclass(frozen=True) +class ImportCounts: + conditions: Mapping[str, int] + evidence_status: Mapping[str, int] + scope_disposition: Mapping[str, int] + + +@dataclass(frozen=True) +class DefectReport: + total_findings: int + by_code: Mapping[str, int] + findings_digest: str + + +@dataclass(frozen=True) +class ImportReport: + schema_version: Literal["choicebench.import-report.v1"] + import_state: Literal["validated", "imported", "failed"] + wrote_artifacts: bool + idempotent_noop: bool + run_id: str + experiment_id: str | None + counts: ImportCounts + defects: DefectReport + condition_digests: Mapping[str, str] + realization_digests: Mapping[str, str] + failures: tuple[Mapping[str, Any], ...] + + +def _source_record(declaration, opened, *, notes_digest: str | None) -> dict[str, Any]: + unknown_reasons = { + field: "not declared in generic import specification" + for field, value in ( + ("source_run_id", declaration.source_run_id), + ("source_repository", declaration.source_repository), + ("source_commit", declaration.source_commit), + ) + if value is None + } + return { + "source_id": declaration.source_id, + "logical_path": declaration.logical_path, + "classification": declaration.classification, + "format": declaration.format, + "format_version": declaration.format_version, + "sha256": opened.sha256, + "provenance": { + "source_run_id": declaration.source_run_id, + "source_repository": declaration.source_repository, + "source_commit": declaration.source_commit, + "notes_digest": notes_digest, + "evidence_digest": opened.sha256, + "unknown_reasons": unknown_reasons, + }, + } + + +def _build_realization_for_condition( + condition: ImportConditionSpec, + *, + spec: ImportSpec, + dataset: ExpectedDataset, + model, + method, + prompt, + workspace_root: Path, + strict: bool, +) -> tuple[dict[str, Any], Any, tuple[dict[str, Any], ...]]: + """Returns (realization_record, RealizationValidation, evidence_records).""" + sources_by_id = {source.source_id: source for source in spec.sources} + semantic = build_import_semantic_identity( + condition=condition, dataset=dataset, model=model, method=method, prompt=prompt + ) + + opened_by_source: dict[str, Any] = {} + tables_by_source: dict[str, Any] = {} + for source_id in condition.source_ids: + declaration = sources_by_id[source_id] + opened = open_verified_source(declaration, containment_root=workspace_root) + opened_by_source[source_id] = opened + tables_by_source[source_id] = parse_csv_source(opened, declaration, strict=strict) + + primary_source_id = condition.source_ids[0] + primary_declaration = sources_by_id[primary_source_id] + primary_table = tables_by_source[primary_source_id] + validated = validate_source_rows( + primary_table, source_id=primary_source_id, mapping=primary_declaration.columns, + condition=condition, expected=dataset, + ) + result = normalize_realization_rows( + validated, condition=condition, expected=dataset, mapping=primary_declaration.columns, + ) + + source_records = [] + evidence_records: list[dict[str, Any]] = [] + all_source_digests = [] + for source_id in condition.source_ids: + declaration = sources_by_id[source_id] + opened = opened_by_source[source_id] + notes_digest = integrity_digest(declaration.notes) if declaration.notes else None + source_records.append(_source_record(declaration, opened, notes_digest=notes_digest)) + all_source_digests.append(opened.sha256) + evidence_records.append( + evidence_record(opened, references=sorted(validated.rows_by_question_id)) + ) + all_source_digests = tuple(sorted(set(all_source_digests))) + + lineage_components = [] + for row in result.evaluable_rows: + qid = row["question_id"] + raw_row = dict(validated.rows_by_question_id[qid].values) + component = make_lineage_component( + operation_type="external_import", + question_id=qid, + parent_digests=(), + source_digests=all_source_digests, + authorization_digest=None, + implementation=_external_import_adapter, + parameters={_ADAPTER_STRICT_KEY: strict}, + input_digest=integrity_digest({"question_id": qid, "raw_row": raw_row}), + preownership_output_digest=integrity_digest({"question_id": qid, "row": row}), + prediction_origin=row["prediction_origin"], + ) + lineage_components.append(component) + + row_assignments = [ + ( + component["identity"]["question_id"], + component["identity"]["prediction_origin"], + component["lineage_id"], + ) + for component in lineage_components + ] + result_origin = make_result_origin(derivation_origin="external_import", row_assignments=row_assignments) + + qualification_digest = integrity_digest(list(result.qualifications)) + limitation_digest = integrity_digest(list(result.limitations)) + defect_digest = integrity_digest(list(result.defect_question_ids)) + + runtime_importer = importer_implementation_identity( + adapter=_external_import_adapter, validator=_external_import_validator + ) + + realization_identity = { + "import_spec_digest": integrity_digest(canonicalize(spec.provenance)), + "sources": source_records, + "expected_dataset": { + "snapshot_digest": dataset.snapshot_digest, + "question_set_digest": dataset.question_set_digest, + "derivation_digest": dataset.derivation_digest, + "question_ids": list(dataset.selected_question_ids), + }, + "importer_implementation": runtime_importer, + "parsing_policy": { + "dialect": validate_csv_dialect_identity(asdict(primary_declaration.dialect)), + "mapping": dict(primary_declaration.columns), + "null_values": list(primary_declaration.null_values), + "numeric_columns": validate_numeric_columns_identity( + [asdict(item) for item in primary_declaration.numeric_columns] + ), + "option_mapping": validate_option_mapping_identity( + asdict(primary_declaration.option_mapping) + ), + "extra_field_policy": primary_declaration.extra_field_policy, + }, + "validation": {"findings_digest": validated.findings_digest}, + "evidence": { + "evidence_status": result.computed_evidence_status, + "qualification_digest": qualification_digest, + "limitation_digest": limitation_digest, + "defect_digest": defect_digest, + "scope_disposition": condition.scope_disposition, + }, + "lineage_components": lineage_components, + "result_origin": result_origin, + "parent_digests": {"realization_digests": [], "evidence_digests": [], "result_digests": []}, + "authorization_digest": None, + "overlay": None, + } + + lineage_runtime_callables = { + component["lineage_id"]: _external_import_adapter for component in lineage_components + } + realization = make_realization( + condition_id=semantic.condition["condition_id"], + condition_digest=semantic.condition["condition_digest"], + identity=realization_identity, + fields={}, + adapter=_external_import_adapter, + validator=_external_import_validator, + expected_dataset=dataset, + semantic_condition=semantic.condition, + lineage_runtime_callables=lineage_runtime_callables, + ) + return realization, semantic.condition, result, tuple(evidence_records) + + +def build_import_plan(request: ImportRequest) -> ImportPlan: + """Resolve, adapt, validate, and identity-bind every declared condition. + Performs every read/validation/identity step but writes nothing.""" + spec = request.spec + sources_by_id = {source.source_id: source for source in spec.sources} + expected_datasets: dict[str, ExpectedDataset] = {} + for declaration in spec.datasets: + opened = { + source_id: open_verified_source( + sources_by_id[source_id], containment_root=request.workspace_root + ) + for source_id in declaration.source_ids + } + expected_datasets[declaration.dataset_id] = build_expected_dataset(declaration, opened) + + models_by_key = {model.model_key: model for model in spec.models} + methods_by_key = {method.method_key: method for method in spec.methods} + prompts_by_key = {prompt.prompt_key: prompt for prompt in spec.prompts} + + semantic_conditions: dict[str, dict[str, Any]] = {} + realizations: dict[str, dict[str, Any]] = {} + realization_validations: dict[str, Any] = {} + all_evidence_records: list[dict[str, Any]] = [] + all_findings: list[Any] = [] + + for condition in spec.conditions: + dataset = expected_datasets[condition.dataset_id] + model = models_by_key[condition.model_key] + method = methods_by_key[condition.method_key] + prompt = prompts_by_key[condition.prompt_key] + realization, condition_record, result, evidence_records = _build_realization_for_condition( + condition, spec=spec, dataset=dataset, model=model, method=method, prompt=prompt, + workspace_root=request.workspace_root, strict=request.strict, + ) + semantic_conditions[condition_record["condition_id"]] = condition_record + realizations[realization["realization_id"]] = realization + realization_validations[realization["realization_id"]] = result + all_evidence_records.extend(evidence_records) + all_findings.extend(result.findings) + + payload = { + "protocol_version": PROTOCOL_V3_VERSION, + "canonicalization_version": CANONICALIZATION_VERSION, + "semantic_conditions": semantic_conditions, + "realizations": realizations, + } + manifest = make_manifest_v3( + payload, + audit={ + "source_location": str(request.workspace_root), + }, + ) + validate_manifest_v3(manifest) + + return ImportPlan( + manifest=manifest, + expected_datasets=expected_datasets, + realizations=realization_validations, + evidence_records=tuple(all_evidence_records), + findings=tuple(all_findings), + ) + + +def _counts_for_plan(plan: ImportPlan) -> ImportCounts: + conditions = {"declared": len(plan.manifest["payload"]["semantic_conditions"])} + evidence_status: dict[str, int] = {} + scope_disposition: dict[str, int] = {} + for realization_id, result in plan.realizations.items(): + evidence_status[result.computed_evidence_status] = ( + evidence_status.get(result.computed_evidence_status, 0) + 1 + ) + realization = plan.manifest["payload"]["realizations"][realization_id] + disposition = realization["identity"]["realization"]["evidence"]["scope_disposition"] + scope_disposition[disposition] = scope_disposition.get(disposition, 0) + 1 + return ImportCounts( + conditions=conditions, evidence_status=evidence_status, scope_disposition=scope_disposition, + ) + + +def _defects_for_plan(plan: ImportPlan) -> DefectReport: + by_code: dict[str, int] = {} + for finding in plan.findings: + by_code[finding.code] = by_code.get(finding.code, 0) + 1 + findings_payload = [finding.to_json() for finding in plan.findings] + return DefectReport( + total_findings=len(plan.findings), by_code=by_code, + findings_digest=integrity_digest(findings_payload), + ) + + +def execute_import(request: ImportRequest, *, dry_run: bool = False) -> ImportReport: + try: + plan = build_import_plan(request) + except (ImportIdentityError, ImportValidationError, ValueError) as exc: + return ImportReport( + schema_version="choicebench.import-report.v1", + import_state="failed", + wrote_artifacts=False, + idempotent_noop=False, + run_id=request.run_id, + experiment_id=None, + counts=ImportCounts(conditions={}, evidence_status={}, scope_disposition={}), + defects=DefectReport(total_findings=0, by_code={}, findings_digest=integrity_digest([])), + condition_digests={}, + realization_digests={}, + failures=({"code": "IMPORT_PLAN_FAILED", "sanitized_message": str(exc)},), + ) + + counts = _counts_for_plan(plan) + defects = _defects_for_plan(plan) + condition_digests = { + condition_id: record["condition_digest"] + for condition_id, record in plan.manifest["payload"]["semantic_conditions"].items() + } + realization_digests = { + realization_id: record["realization_digest"] + for realization_id, record in plan.manifest["payload"]["realizations"].items() + } + + if dry_run: + return ImportReport( + schema_version="choicebench.import-report.v1", + import_state="validated", + wrote_artifacts=False, + idempotent_noop=False, + run_id=request.run_id, + experiment_id=plan.manifest["experiment_id"], + counts=counts, + defects=defects, + condition_digests=condition_digests, + realization_digests=realization_digests, + failures=(), + ) + + runs_dir = Path(request.workspace_root) / "runs" + idempotent_noop = [False] + + def _validator(published_root: Path) -> None: + if published_root == runs_dir / request.run_id and (published_root / MANIFEST_FILENAME).is_file(): + verify_import_run(published_root, expected=plan) + idempotent_noop[0] = True + return + _write_staged_run(published_root, request, plan) + + with ImportTransaction(runs_dir=runs_dir, run_id=request.run_id) as txn: + final_path = txn.publish(_validator) + + return ImportReport( + schema_version="choicebench.import-report.v1", + import_state="imported", + wrote_artifacts=not idempotent_noop[0], + idempotent_noop=idempotent_noop[0], + run_id=request.run_id, + experiment_id=plan.manifest["experiment_id"], + counts=counts, + defects=defects, + condition_digests=condition_digests, + realization_digests=realization_digests, + failures=(), + ) + + +def _write_staged_run(staged_run: Path, request: ImportRequest, plan: ImportPlan) -> None: + staged_run.mkdir(parents=True, exist_ok=True) + sources_by_id = {source.source_id: source for source in request.spec.sources} + + written_digests: set[str] = set() + index_records: list[Mapping[str, Any]] = [] + for record in plan.evidence_records: + if record["sha256"] in written_digests: + continue + written_digests.add(record["sha256"]) + declaration = sources_by_id[record["source_id"]] + opened = open_verified_source(declaration, containment_root=request.workspace_root) + index_records.append(write_evidence_blob(staged_run, opened, references=record["references"])) + write_evidence_index(staged_run, index_records) + + atomic_write_json(staged_run / MANIFEST_FILENAME, plan.manifest) + + state = initial_run_state_v3(plan.manifest) + for realization_id, realization in plan.manifest["payload"]["realizations"].items(): + result = plan.realizations[realization_id] + realization_state = state["realizations"][realization_id] + realization_state["status"] = "completed" if result.evaluable else "evidence_only" + + prepared_validation = prepare_realization_validation_artifact( + result, realization=realization, evidence_records=plan.evidence_records + ) + write_realization_validation_artifact(staged_run, prepared_validation) + realization_state["validation_sha256"] = prepared_validation.file_sha256 + + if result.evaluable: + condition_id = realization["condition_id"] + condition_record = plan.manifest["payload"]["semantic_conditions"][condition_id] + benchmark = condition_record["identity"]["benchmark"] + identity_columns = { + "condition_id": condition_id, + "realization_id": realization_id, + "experiment_id": plan.manifest["experiment_id"], + "dataset_artifact_id": benchmark["artifact_id"], + "dataset_selection_id": benchmark["selection_id"], + "model_id": condition_record["identity"]["model_id"], + "method_id": condition_record["identity"]["method_id"], + "prompt_id": condition_record["identity"]["prompt_id"], + "benchmark_name": benchmark["name"], + "benchmark_split": benchmark["split"], + } + lineage_ids_by_qid = { + assignment["question_id"]: assignment["prediction_lineage_id"] + for assignment in realization["identity"]["realization"]["result_origin"]["row_assignments"] + } + result_rows = [ + { + **identity_columns, + **row, + "prediction_lineage_id": lineage_ids_by_qid[row["question_id"]], + } + for row in result.evaluable_rows + ] + prepared_result = prepare_manifest_result( + result_rows, manifest=plan.manifest, realization_id=realization_id + ) + _, published_metadata, _ = publish_manifest_result(prepared_result, run_dir=staged_run) + realization_state["result_artifact_id"] = published_metadata["result_artifact_id"] + realization_state["result_artifact_digest"] = published_metadata["result_artifact_digest"] + realization_state["result_sha256"] = published_metadata["file_sha256"] + + atomic_write_json(staged_run / RUN_STATE_FILENAME, canonicalize(state, redact_secrets=False)) + validate_manifest_v3(plan.manifest) + + +def verify_import_run(run_dir: Path, expected: ImportPlan | None = None) -> VerifiedImportRun: + """The single shared full-graph verifier: validates manifest, state, + validation artifacts, and results, and returns immutable trusted values + only after the whole graph passes.""" + run_dir = Path(run_dir) + manifest = json.loads((run_dir / MANIFEST_FILENAME).read_text()) + validate_manifest_v3(manifest) + state = json.loads((run_dir / RUN_STATE_FILENAME).read_text()) + validate_run_state_v3(state, manifest) + + if expected is not None and manifest["experiment_digest"] != expected.manifest["experiment_digest"]: + raise ImportEngineError( + f"Run {run_dir} belongs to a different experiment than expected." + ) + + evidence_index_path = run_dir / "artifacts" / "imports" / "evidence" / "index.json" + evidence_index = json.loads(evidence_index_path.read_text()) + evidence_index_digest = evidence_index["evidence_index_digest"] + validate_evidence_index(run_dir, evidence_index_digest) + + realizations: dict[str, VerifiedBaseRealization] = {} + for realization_id, realization in manifest["payload"]["realizations"].items(): + realization_identity = realization["identity"]["realization"] + realization_state = state["realizations"][realization_id] + status = realization_state["status"] + + validation_path = run_dir / f"artifacts/imports/validation/{realization_id}.json" + validate_realization_validation_artifact( + validation_path, manifest=manifest, realization=realization, + expected_sha256=realization_state["validation_sha256"], + ) + + result_sha256 = None + rows_by_question_id: dict[str, Any] = {} + prediction_origins: dict[str, str] = {} + if status == "completed": + result_path = run_dir / f"results/{realization_id}.csv" + metadata = validate_result_artifact(result_path, manifest=manifest, realization=realization) + if ( + metadata["result_artifact_id"] != realization_state["result_artifact_id"] + or metadata["result_artifact_digest"] != realization_state["result_artifact_digest"] + ): + raise ImportEngineError( + f"Realization {realization_id!r} result artifact does not match its " + "recorded run-state identity." + ) + result_sha256 = metadata["file_sha256"] + frame = pd.read_csv(result_path, dtype={"question_id": "string"}) + for _, row in frame.iterrows(): + qid = str(row["question_id"]) + rows_by_question_id[qid] = row.to_dict() + prediction_origins[qid] = str(row["prediction_origin"]) + + evidence_digests = { + source["source_id"]: source["sha256"] for source in realization_identity["sources"] + } + + realizations[realization_id] = VerifiedBaseRealization( + condition_id=realization["condition_id"], + condition_digest=realization["condition_digest"], + realization_id=realization_id, + realization_digest=realization["realization_digest"], + evidence_index_digest=evidence_index_digest, + evidence_source_digests=evidence_digests, + validation_artifact_sha256=realization_state["validation_sha256"], + result_sha256=result_sha256, + rows_by_question_id=rows_by_question_id, + prediction_origins=prediction_origins, + ) + + return VerifiedImportRun( + manifest=manifest, manifest_digest=manifest["experiment_digest"], realizations=realizations + ) + + +def serialize_import_report(report: ImportReport) -> dict[str, Any]: + return { + "schema_version": report.schema_version, + "import_state": report.import_state, + "wrote_artifacts": report.wrote_artifacts, + "idempotent_noop": report.idempotent_noop, + "run_id": report.run_id, + "experiment_id": report.experiment_id, + "counts": { + "conditions": dict(report.counts.conditions), + "evidence_status": dict(report.counts.evidence_status), + "scope_disposition": dict(report.counts.scope_disposition), + }, + "defects": { + "total_findings": report.defects.total_findings, + "by_code": dict(report.defects.by_code), + "findings_digest": report.defects.findings_digest, + }, + "condition_digests": dict(report.condition_digests), + "realization_digests": dict(report.realization_digests), + "failures": [dict(failure) for failure in report.failures], + } diff --git a/src/choicebench/importing/evidence.py b/src/choicebench/importing/evidence.py index 80bd8d0..fb6013e 100644 --- a/src/choicebench/importing/evidence.py +++ b/src/choicebench/importing/evidence.py @@ -76,6 +76,23 @@ def evidence_blob_path(staged_run: Path, sha256_digest: str, format_name: str) - ) +def evidence_record(source: OpenedSource, references: Sequence[str]) -> dict[str, Any]: + """The evidence blob record for a source, computed with no filesystem + I/O. Used both to write a blob (write_evidence_blob) and, during + dry-run/planning, to know what a real import would write without + actually writing it.""" + return canonicalize( + { + "source_id": source.source_id, + "logical_path": source.logical_path, + "sha256": source.sha256, + "size": len(source.data), + "format": "csv", + "references": sorted(set(references)), + } + ) + + def write_evidence_blob( staged_run: Path, source: OpenedSource, references: Sequence[str] ) -> dict[str, Any]: @@ -91,16 +108,7 @@ def write_evidence_blob( ) else: atomic_write_bytes(blob_path, source.data) - record = canonicalize( - { - "source_id": source.source_id, - "logical_path": source.logical_path, - "sha256": source.sha256, - "size": len(source.data), - "format": "csv", - "references": sorted(set(references)), - } - ) + record = evidence_record(source, references) sidecar_path = blob_path.with_suffix(blob_path.suffix + ".json") if sidecar_path.exists(): existing = json.loads(sidecar_path.read_text()) diff --git a/tests/importing/test_engine.py b/tests/importing/test_engine.py new file mode 100644 index 0000000..2faee35 --- /dev/null +++ b/tests/importing/test_engine.py @@ -0,0 +1,239 @@ +"""End-to-end tests for the reduced-scope import engine: base (non-overlay) +plan/dry-run/real-import/idempotence/verify. Overlay integration is not yet +wired into execute_import (see engine.py's module docstring); this file +covers the base import path the engine currently implements. +""" + +from __future__ import annotations + +from hashlib import sha256 +from pathlib import Path + +import pytest +import yaml + +from choicebench.importing.engine import ( + ImportRequest, + build_import_plan, + execute_import, + serialize_import_report, + verify_import_run, +) +from choicebench.importing.schema import load_import_spec + +_RESULTS_CSV = "qid,question,choice_a,choice_b,answer,gold,score\nq1,One?,x,y,a,a,0.9\nq2,Two?,m,n,b,b,0.7\n" + + +def _raw_spec(tmp_path: Path, results_csv: str = _RESULTS_CSV) -> dict: + results_path = tmp_path / "results.csv" + results_path.write_bytes(results_csv.encode("utf-8")) + expected_sha256 = sha256(results_csv.encode("utf-8")).hexdigest() + return { + "schema_version": "choicebench.import-spec.v1", + "import_name": "engine test import", + "sources": [ + { + "source_id": "results", + "path": str(results_path), + "logical_path": "freeze/results.csv", + "expected_sha256": expected_sha256, + "format": "csv", + "format_version": "producer-v1", + "classification": "raw", + "dialect": { + "encoding": "utf-8", "bom_policy": "forbid", "decoding_errors": "strict", + "delimiter": ",", "quote_character": '"', "escape_character": None, + "double_quote": True, "line_terminators": ["crlf", "lf", "cr"], + "mixed_line_terminators": "allow", "final_record_without_terminator": "allow", + "blank_record_policy": "reject", "skip_initial_space": False, + "header": "first_logical_record", "strict_syntax": True, + }, + "columns": { + "question_id": "qid", "question_text": "question", + "correct_option": "gold", "prediction": "answer", + }, + "expected_columns": ["qid", "question", "choice_a", "choice_b", "answer", "gold", "score"], + "ignored_columns": {"score": "producer aggregate only"}, + "null_values": ["", "NA"], + "numeric_columns": [ + {"source_column": "score", "value_type": "float", "null_allowed": True, "finite_only": True} + ], + "option_mapping": { + "mode": "ordered_columns", "ordered_columns": ["choice_a", "choice_b"], + "structured_column": None, "structured_label_key": None, "structured_text_key": None, + }, + "extra_field_policy": "preserve_unmapped", + "preserve_namespace": "producer", + "source_run_id": None, "source_repository": None, "source_commit": None, + "notes": {}, + } + ], + "datasets": [ + { + "dataset_id": "dataset", + "benchmark_name": "historical-benchmark", + "split": "test", + "reference_kind": "independent_input_snapshot", + "trust_label": "producer-supplied", + "source_ids": ["results"], + "selection_source_id": "results", + "expected_question_ids": ["q1", "q2"], + "selection_seed": None, + "selection_n_samples": None, + "subject_filter": [], + "selection_unknown_reasons": { + "selection_seed": "not recorded by producer", + "selection_n_samples": "not recorded by producer", + }, + "columns": { + "question_id": "qid", "question_text": "question", "correct_option": "gold", + "choice_a": "choice_a", "choice_b": "choice_b", + }, + "revision": None, "fingerprint": None, "derivation": {}, + "limitations": ["publisher revision was not recorded"], + "native_compatibility_identity": None, + } + ], + "models": [ + { + "model_key": "model", "display_name": "historical-model", "backend": None, + "provider": None, "revision": None, "effective_parameters": {}, + "unknown_reasons": { + "backend": "not recorded by producer", "provider": "not recorded by producer", + "revision": "not recorded by producer", + }, + "native_compatibility_identity": None, + } + ], + "methods": [ + { + "method_key": "method", "name": "historical-direct", "effective_parameters": {}, + "implementation": None, + "unknown_reasons": {"implementation": "not recorded by producer"}, + "native_compatibility_identity": None, + } + ], + "prompts": [ + { + "prompt_key": "prompt", "template_identity": None, "template_digest": None, + "template_contents": None, "unknown_reason": "not recorded by producer", + "native_compatibility_identity": None, + } + ], + "conditions": [ + { + "condition_key": "condition", "source_ids": ["results"], "dataset_id": "dataset", + "model_key": "model", "method_key": "method", "prompt_key": "prompt", "seed": None, + "calibration_identity": None, "preflight_identity": None, "protocol_settings": {}, + "generation_parameters": {}, + "unknown_reasons": { + "seed": "not recorded by producer", + "calibration_identity": "not recorded by producer", + "preflight_identity": "not recorded by producer", + }, + "expected_question_ids": ["q1", "q2"], + "evidence_status": "complete", "scope_disposition": "included", "executable": None, + "qualifications": [], "limitations": [], "damaged_question_ids": [], + "recoverable_question_ids": [], + "result_origin": { + "derivation_origin": "external_import", + "default_prediction_origin": "external_historical_inference", + "per_question_prediction_origins": {}, + }, + } + ], + "authorizations": [], + "overlays": [], + "metrics": ["accuracy"], + "provenance": {"producer_request_id": {"value": None, "reason": "not recorded by producer"}}, + "audit": {"source_path": str(tmp_path), "imported_at": "2026-07-18T00:00:00Z"}, + } + + +def _spec(tmp_path: Path, **kwargs): + raw = _raw_spec(tmp_path, **kwargs) + spec_path = tmp_path / "import.yaml" + spec_path.write_text(yaml.safe_dump(raw, sort_keys=False), encoding="utf-8") + return load_import_spec(spec_path) + + +def test_build_import_plan_produces_a_valid_v3_manifest(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + plan = build_import_plan(request) + assert plan.manifest["schema_version"] == "choicebench.manifest.v3" + assert len(plan.manifest["payload"]["semantic_conditions"]) == 1 + assert len(plan.manifest["payload"]["realizations"]) == 1 + (result,) = plan.realizations.values() + assert result.computed_evidence_status == "complete" + assert result.evaluable is True + assert len(result.evaluable_rows) == 2 + + +def test_dry_run_writes_nothing(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + report = execute_import(request, dry_run=True) + assert report.import_state == "validated" + assert report.wrote_artifacts is False + assert not (tmp_path / "runs").exists() + assert report.counts.evidence_status == {"complete": 1} + assert report.counts.scope_disposition == {"included": 1} + + +def test_real_import_writes_manifest_state_and_result(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + report = execute_import(request, dry_run=False) + assert report.import_state == "imported" + assert report.wrote_artifacts is True + assert report.idempotent_noop is False + run_dir = tmp_path / "runs" / "run-1" + assert (run_dir / "manifest.json").is_file() + assert (run_dir / "run_state.json").is_file() + (realization_id,) = report.realization_digests.keys() + assert (run_dir / f"results/{realization_id}.csv").is_file() + assert (run_dir / f"artifacts/imports/validation/{realization_id}.json").is_file() + assert (run_dir / "artifacts/imports/evidence/index.json").is_file() + + +def test_real_import_is_idempotent_noop_on_repeat(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + execute_import(request, dry_run=False) + second = execute_import(request, dry_run=False) + assert second.idempotent_noop is True + assert second.wrote_artifacts is False + + +def test_verify_import_run_validates_the_published_graph(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + execute_import(request, dry_run=False) + verified = verify_import_run(tmp_path / "runs" / "run-1") + (realization,) = verified.realizations.values() + assert realization.result_sha256 is not None + assert set(realization.rows_by_question_id) == {"q1", "q2"} + assert realization.prediction_origins == { + "q1": "external_historical_inference", "q2": "external_historical_inference", + } + + +def test_execute_import_reports_failure_without_writing_on_invalid_declaration(tmp_path): + spec = _spec(tmp_path, results_csv="qid,question,choice_a,choice_b,answer,gold,score\nq1,One?,x,y,a,a,0.9\n") + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + report = execute_import(request, dry_run=False) + assert report.import_state == "failed" + assert report.wrote_artifacts is False + assert not (tmp_path / "runs").exists() + assert report.failures + + +def test_serialize_import_report_round_trips_to_plain_dict(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + report = execute_import(request, dry_run=True) + serialized = serialize_import_report(report) + assert serialized["import_state"] == "validated" + assert serialized["counts"]["evidence_status"] == {"complete": 1} + assert serialized["failures"] == [] From fc408c925dd8b8ba46d0d5f174cace5328777641 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 17:12:24 +0300 Subject: [PATCH 35/47] feat: read and evaluate v3 imported result sets Task 13, reduced scope. Adds read_manifest_result_set/ build_evaluation_report_for_result_set to io/readers.py; the existing read_manifest_results() v2 wrapper is untouched, and cli/evaluate_run.py (the pre-existing v1-only evaluator CLI) is not modified -- direct Python API access is sufficient for the research workflow this session targets, consistent with deferring the installed-CLI work (Task 17). read_manifest_result_set never reads a manifest/state/result directly; it goes through engine.py's verify_import_run (the shared full-graph verifier), so an evaluated result set can never be built from an unverified or tampered run. Auto-selection picks the single eligible (result-bearing) realization per condition and refuses when a condition has more than one, matching the plan's ambiguity-is-a-refusal rule; explicit realization_ids selection validates every ID is both declared and eligible before reading any CSV. build_evaluation_report_for_result_set accounts for every realization in the manifest, not just selected ones: each carries its evidence_status, scope_disposition, prediction_origins, and a selected flag; only selected realizations get computed accuracy metrics, so partial/malformed/held/ excluded realizations are visible in the report with empty metrics rather than silently dropped. Focused: 5 new tests, all passing together with the existing v2 reader/ evaluator tests (36 total) with zero changes to their behavior. Full suite: 1294 passed (1289 baseline + 5), only the pre-existing Python 3.14 google.genai.types deprecation warning. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/io/readers.py | 135 ++++++++++++++++++++++++++++- tests/importing/test_evaluation.py | 71 +++++++++++++++ 2 files changed, 205 insertions(+), 1 deletion(-) create mode 100644 tests/importing/test_evaluation.py diff --git a/src/choicebench/io/readers.py b/src/choicebench/io/readers.py index 0cf0781..3504fe2 100644 --- a/src/choicebench/io/readers.py +++ b/src/choicebench/io/readers.py @@ -1,10 +1,17 @@ # src/choicebench/io/readers.py +from dataclasses import dataclass from pathlib import Path +from typing import Any, Literal, Mapping, Sequence import json import pandas as pd -from choicebench.manifest import ManifestCompatibilityError, validate_manifest, validate_run_state +from choicebench.manifest import ( + RUN_STATE_FILENAME, + ManifestCompatibilityError, + validate_manifest, + validate_run_state, +) from choicebench.io.writers import result_artifact_path, validate_result_artifact from choicebench.identity import integrity_digest from choicebench.datasets import dataset_content_digest @@ -251,3 +258,129 @@ def read_all_run_results( return pd.DataFrame() return pd.concat(frames, ignore_index=True) + + +@dataclass(frozen=True) +class RealizationSelection: + policy: Literal["single_evaluable_per_condition", "explicit"] + realization_ids: tuple[str, ...] + + +@dataclass(frozen=True) +class ManifestResultSet: + rows: pd.DataFrame + manifest: Mapping[str, Any] + state: Mapping[str, Any] + selection: RealizationSelection + verified_realizations: Mapping[str, Any] + + +def read_manifest_result_set( + run_dir: Path, *, realization_ids: Sequence[str] | None = None +) -> ManifestResultSet: + """Read a v3 imported/repaired run's result rows through the shared + full-graph verifier (never a bare CSV read): unselected/ineligible + realizations remain visible via verified_realizations for accounting, + but never concatenated into rows. Auto-selection only applies when a + condition has at most one eligible realization; ambiguity is a refusal. + """ + from choicebench.importing.engine import verify_import_run + + run_dir = Path(run_dir) + verified = verify_import_run(run_dir) + manifest = verified.manifest + state = json.loads((run_dir / RUN_STATE_FILENAME).read_text()) + + by_condition: dict[str, list[str]] = {} + for realization_id, record in manifest["payload"]["realizations"].items(): + by_condition.setdefault(record["condition_id"], []).append(realization_id) + + def _eligible(realization_id: str) -> bool: + return verified.realizations[realization_id].result_sha256 is not None + + if realization_ids is None: + selected: list[str] = [] + for condition_id, ids in by_condition.items(): + eligible_ids = sorted(rid for rid in ids if _eligible(rid)) + if len(eligible_ids) > 1: + raise ResultSetError( + f"Condition {condition_id!r} has multiple eligible realizations " + f"{eligible_ids}; pass an explicit realization_ids selection." + ) + selected.extend(eligible_ids) + policy: Literal["single_evaluable_per_condition", "explicit"] = ( + "single_evaluable_per_condition" + ) + else: + unknown = sorted(set(realization_ids) - set(manifest["payload"]["realizations"])) + if unknown: + raise ResultSetError(f"Unknown realization ID(s) {unknown}.") + ineligible = sorted(rid for rid in realization_ids if not _eligible(rid)) + if ineligible: + raise ResultSetError(f"Realization(s) {ineligible} have no result to evaluate.") + selected = sorted(set(realization_ids)) + policy = "explicit" + + frames = [ + pd.read_csv(run_dir / f"results/{realization_id}.csv", dtype={"question_id": "string"}) + for realization_id in selected + ] + rows = pd.concat(frames, ignore_index=True) if frames else pd.DataFrame() + + return ManifestResultSet( + rows=rows, + manifest=manifest, + state=state, + selection=RealizationSelection(policy=policy, realization_ids=tuple(selected)), + verified_realizations=verified.realizations, + ) + + +def build_evaluation_report_for_result_set( + run_id: str, result_set: ManifestResultSet, *, reparse: bool +) -> dict[str, Any]: + """Reduced-scope evaluation-v2 report: per-condition realization lists, + per-realization selection/status/scope accounting, and accuracy for + selected evaluable realizations. Every realization -- selected or not -- + is accounted for; only selected ones contribute metrics or rows.""" + manifest = result_set.manifest + conditions: dict[str, dict[str, Any]] = {} + realizations: dict[str, dict[str, Any]] = {} + selected_ids = set(result_set.selection.realization_ids) + + for realization_id, realization in manifest["payload"]["realizations"].items(): + condition_id = realization["condition_id"] + conditions.setdefault(condition_id, {"realization_ids": []}) + conditions[condition_id]["realization_ids"].append(realization_id) + + evidence = realization["identity"]["realization"]["evidence"] + result_origin = realization["identity"]["realization"]["result_origin"] + selected = realization_id in selected_ids + entry: dict[str, Any] = { + "selected": selected, + "evidence_status": evidence["evidence_status"], + "scope_disposition": evidence["scope_disposition"], + "prediction_origins": result_origin["prediction_origins"], + "metrics": {}, + } + if selected: + rows = result_set.rows[result_set.rows["realization_id"] == realization_id] + if len(rows): + predicted = rows["predicted_option"].astype(str).str.upper() + correct = rows["correct_option"].astype(str).str.upper() + entry["metrics"] = { + "accuracy": float((predicted == correct).mean()), + "n": int(len(rows)), + } + realizations[realization_id] = entry + + for condition in conditions.values(): + condition["realization_ids"] = sorted(condition["realization_ids"]) + + return { + "schema_version": "choicebench.evaluation.v2", + "run_id": run_id, + "selection_policy": result_set.selection.policy, + "conditions": conditions, + "realizations": realizations, + } diff --git a/tests/importing/test_evaluation.py b/tests/importing/test_evaluation.py new file mode 100644 index 0000000..5de630b --- /dev/null +++ b/tests/importing/test_evaluation.py @@ -0,0 +1,71 @@ +"""Tests for the reduced-scope v3 reader/evaluator (Task 13). Reuses +tests/importing/test_engine.py's ImportSpec fixture builder rather than +re-deriving a full spec here.""" + +from __future__ import annotations + +import pytest + +from choicebench.io.readers import ( + ResultSetError, + build_evaluation_report_for_result_set, + read_manifest_result_set, +) +from choicebench.importing.engine import ImportRequest, execute_import +from tests.importing.test_engine import _spec + + +def _imported_run(tmp_path, **spec_kwargs): + spec = _spec(tmp_path, **spec_kwargs) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + execute_import(request, dry_run=False) + return tmp_path / "runs" / "run-1" + + +def test_read_manifest_result_set_auto_selects_the_single_eligible_realization(tmp_path): + run_dir = _imported_run(tmp_path) + result_set = read_manifest_result_set(run_dir) + assert result_set.selection.policy == "single_evaluable_per_condition" + assert len(result_set.selection.realization_ids) == 1 + assert set(result_set.rows["question_id"].astype(str)) == {"q1", "q2"} + + +def test_read_manifest_result_set_rejects_unknown_realization_id(tmp_path): + run_dir = _imported_run(tmp_path) + with pytest.raises(ResultSetError, match="Unknown realization"): + read_manifest_result_set(run_dir, realization_ids=("real_" + "0" * 16,)) + + +def test_read_manifest_result_set_explicit_selection(tmp_path): + run_dir = _imported_run(tmp_path) + auto = read_manifest_result_set(run_dir) + (realization_id,) = auto.selection.realization_ids + explicit = read_manifest_result_set(run_dir, realization_ids=(realization_id,)) + assert explicit.selection.policy == "explicit" + assert explicit.selection.realization_ids == (realization_id,) + + +def test_build_evaluation_report_accounts_for_every_realization(tmp_path): + run_dir = _imported_run(tmp_path) + result_set = read_manifest_result_set(run_dir) + report = build_evaluation_report_for_result_set("run-1", result_set, reparse=False) + assert report["schema_version"] == "choicebench.evaluation.v2" + (condition_id,) = report["conditions"] + (realization_id,) = report["conditions"][condition_id]["realization_ids"] + entry = report["realizations"][realization_id] + assert entry["selected"] is True + assert entry["evidence_status"] == "complete" + assert entry["scope_disposition"] == "included" + assert entry["metrics"]["n"] == 2 + assert entry["metrics"]["accuracy"] == 1.0 # both rows' answer matches gold + + +def test_build_evaluation_report_scores_incorrect_predictions(tmp_path): + run_dir = _imported_run( + tmp_path, + results_csv="qid,question,choice_a,choice_b,answer,gold,score\nq1,One?,x,y,b,a,0.9\nq2,Two?,m,n,b,b,0.7\n", + ) + result_set = read_manifest_result_set(run_dir) + report = build_evaluation_report_for_result_set("run-1", result_set, reparse=False) + (realization_id,) = result_set.selection.realization_ids + assert report["realizations"][realization_id]["metrics"]["accuracy"] == 0.5 From 0bb63041e782248e6ea484148eafe216c86b18a4 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 18:00:20 +0300 Subject: [PATCH 36/47] feat: translate Stage 1 queue/status/matrix accounting Task 15, reduced scope. translate_stage1_paper_freeze() verifies checksums for the canonical manifest, cell-status/expected matrices, and all four queue files against checksums/checksums.sha256, parses the canonical manifest into per-cell records, and explodes the queue files into (cell_id, question_id) pairs -- reproducing the exact counts the research objective needs to confirm. Verified directly against the real, read-only Stage 1 freeze at /home/cotenthusiast/Projects/model-generalization/paper_data_freeze during development (never copied into the repo or committed): 138 approved / 687 held / 3 excluded pairs, all pairwise disjoint; the 100/84/4/12/2 paper- matrix breakdown; and the 6-cell/18-pair offline-recoverable authority (exact cell IDs and question IDs matching) all reproduce exactly. One real correction this verification surfaced: an earlier assumption (from this session's plan-based review, before real freeze access) that rerun_queue.csv was a disjoint "776 forensic" authority tier was wrong. Running the checksum-verified parser against the real file showed all 776 of its pairs are already a subset of the 828 classified (approved/held/ excluded) pairs -- and rerun_queue.csv carries no queue_disposition/ execution_authority/executable columns at all, unlike the three classification files. Fixed to verify "rerun candidates are a subset of classified" (a real, meaningful invariant) rather than a disjointness check that would reject the real data. Fail-closed invariants enforced: approved rows must declare queue_disposition=approved/execution_authority=authoritative/ executable=true; held and excluded rows must declare executable=false; approved/held/excluded must be pairwise disjoint; every rerun-queue candidate pair must already be classified; all cells claiming recoverable_question_ids must authorize the exact same question set. Scope boundary (module docstring): this does not yet build ImportSpec/ AuthorizationSpec objects for the generic engine -- that needs the ARC/MMLU ExpectedDataset trust chain (Task 16/Unit J, next) to validate authorization grants against real selected question IDs, plus per-cell SourceArtifactSpec construction across the freeze's distinct method CSV schemas. Focused: 10 new tests against a small synthetic fixture (never real historical data, per the plan's global constraint), all passing. Full suite: 1304 passed (1294 baseline + 10), only the pre-existing Python 3.14 google.genai.types deprecation warning. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- .../importing/profiles/__init__.py | 6 + .../importing/profiles/stage1_paper_freeze.py | 315 ++++++++++++++++++ tests/importing/test_stage1_profile.py | 283 ++++++++++++++++ 3 files changed, 604 insertions(+) create mode 100644 src/choicebench/importing/profiles/__init__.py create mode 100644 src/choicebench/importing/profiles/stage1_paper_freeze.py create mode 100644 tests/importing/test_stage1_profile.py diff --git a/src/choicebench/importing/profiles/__init__.py b/src/choicebench/importing/profiles/__init__.py new file mode 100644 index 0000000..5c3f53e --- /dev/null +++ b/src/choicebench/importing/profiles/__init__.py @@ -0,0 +1,6 @@ +"""Paper-specific import profiles. + +No generic importer module may import anything from this package, and no +paper-specific constant (cell IDs, method names, queue file names) may leak +into src/choicebench/importing/*.py outside this directory. +""" diff --git a/src/choicebench/importing/profiles/stage1_paper_freeze.py b/src/choicebench/importing/profiles/stage1_paper_freeze.py new file mode 100644 index 0000000..382b999 --- /dev/null +++ b/src/choicebench/importing/profiles/stage1_paper_freeze.py @@ -0,0 +1,315 @@ +"""Translate the read-only Stage 1 paper-freeze into verified queue/status/ +matrix accounting. + +Reduced scope: this verifies checksums, parses the checksum-covered +canonical manifest and the four queue files, explodes them into +(cell_id, question_id) pairs, and reproduces the queue/status/matrix counts +the research objective needs to confirm (exact 138/687/3 approved/held/ +excluded pairs, all pairwise disjoint; rerun_queue.csv's 776 rerun-candidate +pairs verified as a subset of the 828 classified pairs, not a fourth +authority tier -- rerun_queue.csv carries no queue_disposition/ +execution_authority/executable columns at all, unlike the three +classification files; the 100-cell paper matrix breakdown; and the +6-cell/18-pair offline recoverable authority). It does NOT yet build +ImportSpec/AuthorizationSpec +objects ready to feed into the generic engine -- that requires the ARC/MMLU +ExpectedDataset trust chain (to validate authorization grants against real +selected question IDs, i.e. Task 16/Unit J) plus full per-cell +SourceArtifactSpec construction across the freeze's distinct method CSV +header schemas, neither of which is done here. translate_stage1_paper_freeze +is the verified foundation both of those build on: checksum-verified cell +records and verified, pairwise-disjoint, correctly-typed queue pairs. + +This module is the only place in the importer allowed to know Stage 1 +paper-specific facts (cell ID format, queue file names, method names). No +generic importer module imports from here. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from hashlib import sha256 +import csv +import json +from pathlib import Path +from typing import Any, Mapping + +STATUS_MAP = { + "canonical_complete": "complete", + "canonical_qualified": "qualified", + "recoverable_from_existing_artifacts": "recoverable", + "incomplete_requires_inference": "partial", + "malformed_requires_inference": "malformed", +} +_EXCLUDED_STATUS = "excluded_from_paper_matrix" +_PAPER_SCOPE_EXEMPT_METHOD = "pride" +_IHS_METHOD = "independent_hypothesis" + +_VERIFIED_RELATIVE_PATHS = ( + "manifests/canonical_results_manifest.json", + "manifests/canonical_results_manifest.csv", + "manifests/cell_status_matrix.csv", + "manifests/expected_matrix.csv", + "manifests/approved_rerun_queue.csv", + "manifests/held_or_declined_reruns.csv", + "manifests/paper_scope_excluded_reruns.csv", + "manifests/rerun_queue.csv", + "reports/canonical_freeze_report.md", +) + + +class Stage1ProfileError(ValueError): + """Raised when the Stage 1 freeze fails checksum or consistency validation.""" + + +def _load_checksum_ledger(freeze_root: Path) -> dict[str, str]: + ledger: dict[str, str] = {} + path = freeze_root / "checksums" / "checksums.sha256" + with path.open("r", encoding="utf-8") as handle: + for line in handle: + line = line.rstrip("\n") + if not line: + continue + digest, _, relative_path = line.partition(" ") + ledger[relative_path] = digest + return ledger + + +def _verify_checksummed_file(freeze_root: Path, ledger: Mapping[str, str], relative_path: str) -> bytes: + path = freeze_root / relative_path + try: + data = path.read_bytes() + except OSError as exc: + raise Stage1ProfileError(f"{relative_path} is unreadable: {exc}") from exc + expected = ledger.get(relative_path) + if expected is None: + raise Stage1ProfileError(f"{relative_path} has no checksum ledger entry.") + actual = sha256(data).hexdigest() + if actual != expected: + raise Stage1ProfileError( + f"{relative_path} checksum mismatch: expected {expected}, found {actual}." + ) + return data + + +def _read_csv_rows(path: Path) -> list[dict[str, str]]: + with path.open(newline="", encoding="utf-8") as handle: + return list(csv.DictReader(handle)) + + +def _explode_question_pairs( + rows: list[dict[str, str]], *, source_label: str +) -> tuple[dict[tuple[str, str], dict[str, str]], dict[tuple[str, str], str]]: + pairs: dict[tuple[str, str], dict[str, str]] = {} + reasons: dict[tuple[str, str], str] = {} + for row in rows: + cell_id = row["cell_id"] + question_ids = json.loads(row["exact_question_ids"]) + question_reasons = json.loads(row["question_reasons"]) if row.get("question_reasons") else {} + for question_id in question_ids: + pair = (cell_id, question_id) + if pair in pairs: + raise Stage1ProfileError( + f"{source_label} declares duplicate (cell_id, question_id) pair {pair}." + ) + pairs[pair] = row + reasons[pair] = question_reasons.get(question_id, "") + return pairs, reasons + + +def _require_all(rows: list[dict[str, str]], *, field: str, expected: str, source_label: str) -> None: + bad = [row for row in rows if row.get(field) != expected] + if bad: + raise Stage1ProfileError( + f"{source_label} has {len(bad)} row(s) with {field} != {expected!r}." + ) + + +def _require_all_bool(rows: list[dict[str, str]], *, field: str, expected: bool, source_label: str) -> None: + expected_text = "true" if expected else "false" + bad = [row for row in rows if row.get(field, "").strip().lower() != expected_text] + if bad: + raise Stage1ProfileError( + f"{source_label} has {len(bad)} row(s) with {field} != {expected_text!r}." + ) + + +@dataclass(frozen=True) +class Stage1MatrixCounts: + intended_cells: int + core_method_cells: int + local_pride_cells: int + ihs_cells: int + excluded_preserved_cells: int + + +@dataclass(frozen=True) +class Stage1QueueCounts: + approved_executable_question_cells: int + held_nonexecutable_question_cells: int + excluded_nonexecutable_question_cells: int + rerun_candidate_question_cells: int + classified_question_cells: int + pairwise_disjoint: bool + approved_authoritative_executable: bool + held_executable: bool + excluded_executable: bool + rerun_candidates_are_classified_subset: bool + + +@dataclass(frozen=True) +class Stage1QueuePairs: + approved: Mapping[tuple[str, str], str] + held: Mapping[tuple[str, str], str] + excluded: Mapping[tuple[str, str], str] + rerun_candidates: frozenset[tuple[str, str]] + + +@dataclass(frozen=True) +class Stage1OfflineAuthority: + condition_count: int + question_cell_count: int + recoverable_question_ids: tuple[str, ...] + cell_ids: tuple[str, ...] + inference_executable: bool + + +@dataclass(frozen=True) +class Stage1Translation: + cell_records: Mapping[str, Mapping[str, Any]] + matrix: Stage1MatrixCounts + queue_counts: Stage1QueueCounts + queue_pairs: Stage1QueuePairs + offline_authority: Stage1OfflineAuthority + verified_relative_paths: tuple[str, ...] + + +def _matrix_counts(cell_records: Mapping[str, Mapping[str, Any]]) -> Stage1MatrixCounts: + method_counts: dict[str, int] = {} + excluded_cells = 0 + for record in cell_records.values(): + method_counts[record["method"]] = method_counts.get(record["method"], 0) + 1 + if record["final_status"] == _EXCLUDED_STATUS: + excluded_cells += 1 + ihs_cells = method_counts.get(_IHS_METHOD, 0) + pride_cells_total = method_counts.get(_PAPER_SCOPE_EXEMPT_METHOD, 0) + local_pride_cells = pride_cells_total - excluded_cells + core_method_cells = sum( + count + for method, count in method_counts.items() + if method not in {_IHS_METHOD, _PAPER_SCOPE_EXEMPT_METHOD} + ) + return Stage1MatrixCounts( + intended_cells=core_method_cells + ihs_cells + local_pride_cells, + core_method_cells=core_method_cells, + local_pride_cells=local_pride_cells, + ihs_cells=ihs_cells, + excluded_preserved_cells=excluded_cells, + ) + + +def _offline_authority(cell_records: Mapping[str, Mapping[str, Any]]) -> Stage1OfflineAuthority: + recoverable_cells = { + cell_id: tuple(sorted(record["recoverable_question_ids"])) + for cell_id, record in cell_records.items() + if record["recoverable_question_ids"] + } + distinct_id_sets = set(recoverable_cells.values()) + if len(distinct_id_sets) != 1: + raise Stage1ProfileError( + "Recoverable cells do not all authorize the exact same question IDs: " + f"{sorted(distinct_id_sets)}." + ) + (recoverable_question_ids,) = distinct_id_sets + cell_ids = tuple(sorted(recoverable_cells)) + return Stage1OfflineAuthority( + condition_count=len(cell_ids), + question_cell_count=len(cell_ids) * len(recoverable_question_ids), + recoverable_question_ids=recoverable_question_ids, + cell_ids=cell_ids, + inference_executable=False, + ) + + +def translate_stage1_paper_freeze(freeze_root: Path) -> Stage1Translation: + freeze_root = Path(freeze_root) + ledger = _load_checksum_ledger(freeze_root) + for relative_path in _VERIFIED_RELATIVE_PATHS: + _verify_checksummed_file(freeze_root, ledger, relative_path) + + manifest_json = json.loads( + (freeze_root / "manifests" / "canonical_results_manifest.json").read_text() + ) + cell_records: dict[str, Mapping[str, Any]] = {} + for record in manifest_json: + cell_id = record["cell_id"] + if cell_id in cell_records: + raise Stage1ProfileError(f"Canonical manifest declares duplicate cell_id {cell_id!r}.") + if record["status"] not in STATUS_MAP and record["final_status"] != _EXCLUDED_STATUS: + raise Stage1ProfileError( + f"Cell {cell_id!r} has an unrecognized status {record['status']!r}." + ) + cell_records[cell_id] = record + + matrix = _matrix_counts(cell_records) + offline_authority = _offline_authority(cell_records) + + approved_rows = _read_csv_rows(freeze_root / "manifests" / "approved_rerun_queue.csv") + held_rows = _read_csv_rows(freeze_root / "manifests" / "held_or_declined_reruns.csv") + excluded_rows = _read_csv_rows(freeze_root / "manifests" / "paper_scope_excluded_reruns.csv") + # rerun_queue.csv carries no queue_disposition/execution_authority/executable + # columns at all (unlike the three files above): it is the technical rerun + # working-queue, not an authority classification. Verified below to be a + # subset of the classified pairs, never treated as its own authority tier. + rerun_candidate_rows = _read_csv_rows(freeze_root / "manifests" / "rerun_queue.csv") + + _require_all(approved_rows, field="queue_disposition", expected="approved", source_label="approved_rerun_queue.csv") + _require_all(approved_rows, field="execution_authority", expected="authoritative", source_label="approved_rerun_queue.csv") + _require_all_bool(approved_rows, field="executable", expected=True, source_label="approved_rerun_queue.csv") + _require_all_bool(held_rows, field="executable", expected=False, source_label="held_or_declined_reruns.csv") + _require_all_bool(excluded_rows, field="executable", expected=False, source_label="paper_scope_excluded_reruns.csv") + + approved_pairs, approved_reasons = _explode_question_pairs(approved_rows, source_label="approved_rerun_queue.csv") + held_pairs, held_reasons = _explode_question_pairs(held_rows, source_label="held_or_declined_reruns.csv") + excluded_pairs, excluded_reasons = _explode_question_pairs(excluded_rows, source_label="paper_scope_excluded_reruns.csv") + rerun_candidate_pairs, _ = _explode_question_pairs(rerun_candidate_rows, source_label="rerun_queue.csv") + + classified = set(approved_pairs) | set(held_pairs) | set(excluded_pairs) + total_classified = len(approved_pairs) + len(held_pairs) + len(excluded_pairs) + if len(classified) != total_classified: + raise Stage1ProfileError( + "approved/held/excluded queues are not pairwise disjoint: " + f"{total_classified} declared pairs but only {len(classified)} distinct." + ) + unclassified_candidates = set(rerun_candidate_pairs) - classified + if unclassified_candidates: + raise Stage1ProfileError( + f"{len(unclassified_candidates)} rerun-queue candidate pair(s) were never " + "classified as approved, held, or excluded." + ) + + queue_counts = Stage1QueueCounts( + approved_executable_question_cells=len(approved_pairs), + held_nonexecutable_question_cells=len(held_pairs), + excluded_nonexecutable_question_cells=len(excluded_pairs), + rerun_candidate_question_cells=len(rerun_candidate_pairs), + classified_question_cells=len(classified), + pairwise_disjoint=True, + approved_authoritative_executable=True, + held_executable=False, + excluded_executable=False, + rerun_candidates_are_classified_subset=True, + ) + queue_pairs = Stage1QueuePairs( + approved=approved_reasons, held=held_reasons, excluded=excluded_reasons, + rerun_candidates=frozenset(rerun_candidate_pairs), + ) + + return Stage1Translation( + cell_records=cell_records, + matrix=matrix, + queue_counts=queue_counts, + queue_pairs=queue_pairs, + offline_authority=offline_authority, + verified_relative_paths=_VERIFIED_RELATIVE_PATHS, + ) diff --git a/tests/importing/test_stage1_profile.py b/tests/importing/test_stage1_profile.py new file mode 100644 index 0000000..c16a40f --- /dev/null +++ b/tests/importing/test_stage1_profile.py @@ -0,0 +1,283 @@ +"""Tests for Stage 1 profile translation (Task 15, reduced scope). + +Uses a small synthetic freeze fixture (never real historical data) whose +structure was verified against the real, read-only Stage 1 freeze at +/home/cotenthusiast/Projects/model-generalization/paper_data_freeze during +development: file names, column headers, and the exact real queue/matrix +counts (138 approved / 687 held / 3 excluded / 828 classified / 776 +rerun-queue-candidate pairs; 100/84/4/12/2 matrix breakdown; 6-cell/18-pair +offline-recoverable authority) were all cross-checked against that freeze +directly, not merely assumed from the plan. +""" + +from __future__ import annotations + +from hashlib import sha256 +import json +from pathlib import Path + +import pytest + +from choicebench.importing.profiles.stage1_paper_freeze import ( + STATUS_MAP, + Stage1ProfileError, + translate_stage1_paper_freeze, +) + + +def _write(path: Path, content: str) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(content, encoding="utf-8") + + +def _cell(cell_id, method, status, *, damaged=(), recoverable=()): + return { + "cell_id": cell_id, "model": "model-a", "provider_backend": "provider", + "benchmark": "arc_challenge", "benchmark_split": "robustness", "method": method, + "expected_question_count": 10, "evidence_for_expected_existence": "x", + "evidence_for_expected_row_count": "y", "discovered_candidate_count": 1, + "final_status": status, "status": status, "canonical_artifact_id": f"artifact_{cell_id}", + "canonical_raw_path": f"raw/{cell_id}.csv", "canonical_path": f"canonical/{cell_id}.csv", + "canonical_sha256": "0" * 64, "observed_unique_count": 10, "duplicate_question_count": 0, + "exact_missing_question_ids": [], "exact_unexpected_question_ids": [], + "damaged_question_ids": list(damaged), "recoverable_question_ids": list(recoverable), + "qualification": "", + } + + +def _queue_row(cell_id, question_ids, reasons, **overrides): + row = { + "cell_id": cell_id, "model": "model-a", "provider_backend": "provider", + "model_revision": "not recorded", "benchmark": "arc_challenge", + "benchmark_snapshot_revision": "0" * 64, "method": "baseline", + "exact_question_ids": json.dumps(question_ids), + "number_of_questions": len(question_ids), "seed": 42, + "work_type": "complete_partial_or_replace_damaged_rows", + "configuration_identity": "0" * 64, + "question_reasons": json.dumps(reasons), + "source_configuration": "[]", "prompt_template_identity": "[]", + "generation_parameters": "{}", "expected_output_schema": "[]", + "supporting_artifact_paths": "[]", + } + row.update(overrides) + return row + + +def _write_csv(path: Path, rows: list[dict]) -> None: + import csv + + path.parent.mkdir(parents=True, exist_ok=True) + if not rows: + path.write_text("cell_id\n", encoding="utf-8") + return + with path.open("w", newline="", encoding="utf-8") as handle: + writer = csv.DictWriter(handle, fieldnames=list(rows[0])) + writer.writeheader() + writer.writerows(rows) + + +def _build_freeze( + tmp_path: Path, + *, + cells, + approved_rows=(), + held_rows=(), + excluded_rows=(), + rerun_candidate_rows=(), + tamper_checksum_for=None, +): + freeze_root = tmp_path / "freeze" + manifest_json_path = freeze_root / "manifests" / "canonical_results_manifest.json" + _write(manifest_json_path, json.dumps(cells)) + _write(freeze_root / "manifests" / "canonical_results_manifest.csv", "cell_id\n") + _write(freeze_root / "manifests" / "cell_status_matrix.csv", "cell_id\n") + _write(freeze_root / "manifests" / "expected_matrix.csv", "cell_id\n") + _write_csv(freeze_root / "manifests" / "approved_rerun_queue.csv", list(approved_rows)) + _write_csv(freeze_root / "manifests" / "held_or_declined_reruns.csv", list(held_rows)) + _write_csv(freeze_root / "manifests" / "paper_scope_excluded_reruns.csv", list(excluded_rows)) + _write_csv(freeze_root / "manifests" / "rerun_queue.csv", list(rerun_candidate_rows)) + _write(freeze_root / "reports" / "canonical_freeze_report.md", "# report\n") + + relative_paths = ( + "manifests/canonical_results_manifest.json", + "manifests/canonical_results_manifest.csv", + "manifests/cell_status_matrix.csv", + "manifests/expected_matrix.csv", + "manifests/approved_rerun_queue.csv", + "manifests/held_or_declined_reruns.csv", + "manifests/paper_scope_excluded_reruns.csv", + "manifests/rerun_queue.csv", + "reports/canonical_freeze_report.md", + ) + ledger_lines = [] + for relative_path in relative_paths: + data = (freeze_root / relative_path).read_bytes() + digest = sha256(data).hexdigest() + if tamper_checksum_for == relative_path: + digest = "f" * 64 + ledger_lines.append(f"{digest} {relative_path}") + _write(freeze_root / "checksums" / "checksums.sha256", "\n".join(ledger_lines) + "\n") + return freeze_root + + +def _default_cells(): + return [ + _cell("cbp__model-a__arc_challenge__baseline", "baseline", "canonical_complete"), + _cell("cbp__model-a__arc_challenge__two_stage_v1", "two_stage_v1", "canonical_qualified"), + _cell( + "cbp__model-a__arc_challenge__semantic_matching_v1", "semantic_matching_v1", + "recoverable_from_existing_artifacts", recoverable=("q1", "q2"), + ), + _cell( + "cbp__model-b__arc_challenge__semantic_matching_v1", "semantic_matching_v1", + "recoverable_from_existing_artifacts", recoverable=("q1", "q2"), + ), + _cell( + "cbp__model-a__arc_challenge__independent_hypothesis", "independent_hypothesis", + "malformed_requires_inference", damaged=("q3",), + ), + _cell("cbp__model-a__arc_challenge__pride", "pride", "canonical_complete"), + _cell("cbp__model-a__mmlu__pride", "pride", "excluded_from_paper_matrix"), + ] + + +def test_translate_stage1_paper_freeze_reproduces_matrix_and_offline_authority(tmp_path): + freeze_root = _build_freeze(tmp_path, cells=_default_cells()) + result = translate_stage1_paper_freeze(freeze_root) + assert result.matrix.intended_cells == 6 # 7 cells minus the one excluded_from_paper_matrix + # baseline + two_stage_v1 + both semantic_matching_v1 cells (only + # independent_hypothesis/pride are excluded from the "core method" bucket) + assert result.matrix.core_method_cells == 4 + assert result.matrix.ihs_cells == 1 + assert result.matrix.local_pride_cells == 1 # 2 pride cells minus 1 excluded + assert result.matrix.excluded_preserved_cells == 1 + assert result.offline_authority.condition_count == 2 + assert result.offline_authority.question_cell_count == 4 + assert result.offline_authority.recoverable_question_ids == ("q1", "q2") + assert result.offline_authority.inference_executable is False + + +def test_translate_stage1_paper_freeze_verifies_checksums(tmp_path): + freeze_root = _build_freeze( + tmp_path, cells=_default_cells(), + tamper_checksum_for="manifests/canonical_results_manifest.json", + ) + with pytest.raises(Stage1ProfileError, match="checksum mismatch"): + translate_stage1_paper_freeze(freeze_root) + + +def test_translate_stage1_paper_freeze_rejects_duplicate_cell_id(tmp_path): + cells = _default_cells() + cells.append(dict(cells[0])) + freeze_root = _build_freeze(tmp_path, cells=cells) + with pytest.raises(Stage1ProfileError, match="duplicate cell_id"): + translate_stage1_paper_freeze(freeze_root) + + +def test_translate_stage1_paper_freeze_requires_matching_recoverable_ids_across_cells(tmp_path): + cells = _default_cells() + # give the second recoverable cell a DIFFERENT question set than the first + for cell in cells: + if cell["cell_id"] == "cbp__model-b__arc_challenge__semantic_matching_v1": + cell["recoverable_question_ids"] = ["q1", "q9"] + freeze_root = _build_freeze(tmp_path, cells=cells) + with pytest.raises(Stage1ProfileError, match="do not all authorize the exact same question"): + translate_stage1_paper_freeze(freeze_root) + + +def test_translate_stage1_paper_freeze_computes_queue_pairs_and_disjointness(tmp_path): + approved = [ + _queue_row( + "cbp__model-a__arc_challenge__baseline", ["q1", "q2"], + {"q1": "reason1", "q2": "reason2"}, + queue_disposition="approved", execution_authority="authoritative", executable="true", + ) + ] + held = [ + _queue_row( + "cbp__model-a__arc_challenge__two_stage_v1", ["q3"], {"q3": "reason3"}, + queue_disposition="held", execution_authority="none", executable="false", + ) + ] + excluded = [ + _queue_row( + "cbp__model-a__arc_challenge__pride", ["q4"], {"q4": "reason4"}, + queue_disposition="excluded", execution_authority="none", executable="false", + ) + ] + rerun_candidates = [ + _queue_row("cbp__model-a__arc_challenge__baseline", ["q1"], {}), + ] + freeze_root = _build_freeze( + tmp_path, cells=_default_cells(), approved_rows=approved, held_rows=held, + excluded_rows=excluded, rerun_candidate_rows=rerun_candidates, + ) + result = translate_stage1_paper_freeze(freeze_root) + assert result.queue_counts.approved_executable_question_cells == 2 + assert result.queue_counts.held_nonexecutable_question_cells == 1 + assert result.queue_counts.excluded_nonexecutable_question_cells == 1 + assert result.queue_counts.classified_question_cells == 4 + assert result.queue_counts.rerun_candidate_question_cells == 1 + assert result.queue_counts.rerun_candidates_are_classified_subset is True + assert result.queue_pairs.approved[("cbp__model-a__arc_challenge__baseline", "q1")] == "reason1" + + +def test_translate_stage1_paper_freeze_rejects_overlapping_approved_and_held(tmp_path): + approved = [ + _queue_row( + "cbp__model-a__arc_challenge__baseline", ["q1"], {"q1": "r"}, + queue_disposition="approved", execution_authority="authoritative", executable="true", + ) + ] + held = [ + _queue_row( + "cbp__model-a__arc_challenge__baseline", ["q1"], {"q1": "r"}, + queue_disposition="held", execution_authority="none", executable="false", + ) + ] + freeze_root = _build_freeze(tmp_path, cells=_default_cells(), approved_rows=approved, held_rows=held) + with pytest.raises(Stage1ProfileError, match="not pairwise disjoint"): + translate_stage1_paper_freeze(freeze_root) + + +def test_translate_stage1_paper_freeze_rejects_unclassified_rerun_candidate(tmp_path): + rerun_candidates = [ + _queue_row("cbp__model-a__arc_challenge__baseline", ["q_never_classified"], {}), + ] + freeze_root = _build_freeze(tmp_path, cells=_default_cells(), rerun_candidate_rows=rerun_candidates) + with pytest.raises(Stage1ProfileError, match="never classified"): + translate_stage1_paper_freeze(freeze_root) + + +def test_translate_stage1_paper_freeze_rejects_approved_row_claiming_wrong_disposition(tmp_path): + approved = [ + _queue_row( + "cbp__model-a__arc_challenge__baseline", ["q1"], {"q1": "r"}, + queue_disposition="held", # wrong -- approved file must say "approved" + execution_authority="authoritative", executable="true", + ) + ] + freeze_root = _build_freeze(tmp_path, cells=_default_cells(), approved_rows=approved) + with pytest.raises(Stage1ProfileError, match="queue_disposition"): + translate_stage1_paper_freeze(freeze_root) + + +def test_translate_stage1_paper_freeze_rejects_held_row_claiming_executable(tmp_path): + held = [ + _queue_row( + "cbp__model-a__arc_challenge__baseline", ["q1"], {"q1": "r"}, + queue_disposition="held", execution_authority="none", executable="true", # wrong + ) + ] + freeze_root = _build_freeze(tmp_path, cells=_default_cells(), held_rows=held) + with pytest.raises(Stage1ProfileError, match="executable"): + translate_stage1_paper_freeze(freeze_root) + + +def test_status_map_matches_canonical_values(): + assert STATUS_MAP["canonical_complete"] == "complete" + assert STATUS_MAP["canonical_qualified"] == "qualified" + assert STATUS_MAP["recoverable_from_existing_artifacts"] == "recoverable" + assert STATUS_MAP["incomplete_requires_inference"] == "partial" + assert STATUS_MAP["malformed_requires_inference"] == "malformed" + assert "excluded_from_paper_matrix" not in STATUS_MAP # not an evidence-status input From 08e2c52e63e02cee81f20af32f6416cc2d800bc1 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 18:20:36 +0300 Subject: [PATCH 37/47] feat: build the Stage 1 ARC/MMLU expected-dataset trust chain Task 16, reduced scope. build_stage1_expected_datasets() verifies checksums for each benchmark's raw/normalized/robustness_ids/ robustness_metadata files, then independently recomputes question_text/ correct_option/correct_answer_text/choices from each raw row and compares them fieldwise against the archived normalized row at the same file position. Raw and normalized files have identical row counts and order in the real freeze, so a positional join is sufficient and avoids needing to reproduce the archived normalizer's own question_id hash algorithm -- this is a ChoiceBench revalidation of internal freeze consistency, recorded as a limitation on the returned ExpectedDataset, not a claim that the archived normalizer or upstream publisher was authentic. MMLU duplicate question_ids are handled with drop_duplicates(keep="first") semantics, but only after requiring every duplicate occurrence to agree on all parsed fields -- a disagreeing duplicate fails closed rather than silently picking one. Selected question IDs (from robustness_ids.json, preserving its own order, which differs from source file order) must all exist in the deduplicated normalized set or the whole benchmark is refused. Deduplicated rows are canonicalized into fresh CSV bytes and handed to the existing Task 4 build_expected_dataset()/DatasetReferenceSpec machinery, which independently re-validates and re-derives all identity digests. Verified directly against the real, read-only Stage 1 freeze (never copied into the repo) during development: both ExpectedDataset objects build successfully with exactly 1000 selected questions each, every one of 1172 ARC and 14042 MMLU rows revalidates, and all 27 real MMLU duplicate groups are confirmed field-identical before deduplication. This surfaced one genuine, verified freeze quirk: 3 raw ARC questions have a 5th option that the archived normalizer silently drops (all 3 correct answers fall within the first 4, so truncation never removes the correct option) -- modeled explicitly with a fail-closed bound check rather than either ignoring the option or crashing on an honest schema mismatch. Scope boundary (unchanged from the module docstring, now fully proven out): this and Unit I (queue/status/matrix accounting) are the two verified building blocks a full ImportSpec/AuthorizationSpec assembly for the actual repair-import workflow would combine; that combination -- plus per-cell SourceArtifactSpec construction across the freeze's method CSV schemas -- is not done here. Focused: 8 new tests (ARC variable/five-option handling, MMLU duplicate dedup and disagreement rejection, checksum/revalidation/row-count/ missing-selected-ID failures) against a synthetic fixture verified to match the real freeze's exact file formats; 18 total in this file, all passing. Full suite: 1312 passed (1304 baseline + 8), only the pre-existing Python 3.14 google.genai.types deprecation warning. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- .../importing/profiles/stage1_paper_freeze.py | 235 +++++++++++++++- tests/importing/test_stage1_profile.py | 255 ++++++++++++++++++ 2 files changed, 489 insertions(+), 1 deletion(-) diff --git a/src/choicebench/importing/profiles/stage1_paper_freeze.py b/src/choicebench/importing/profiles/stage1_paper_freeze.py index 382b999..7dfcb07 100644 --- a/src/choicebench/importing/profiles/stage1_paper_freeze.py +++ b/src/choicebench/importing/profiles/stage1_paper_freeze.py @@ -27,12 +27,18 @@ from __future__ import annotations +import ast from dataclasses import dataclass from hashlib import sha256 import csv +import io import json from pathlib import Path -from typing import Any, Mapping +from typing import Any, Callable, Mapping + +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.dataset_reference import ExpectedDataset, build_expected_dataset +from choicebench.importing.schema import DatasetReferenceSpec STATUS_MAP = { "canonical_complete": "complete", @@ -313,3 +319,230 @@ def translate_stage1_paper_freeze(freeze_root: Path) -> Stage1Translation: offline_authority=offline_authority, verified_relative_paths=_VERIFIED_RELATIVE_PATHS, ) + + +# --- ARC/MMLU expected-dataset trust chain (Task 16) ----------------------- +# +# Reduced scope, verified directly against the real freeze: this recomputes +# question_text/correct_option/correct_answer_text/choices from each raw row +# and compares them fieldwise against the archived normalized row at the +# SAME position -- raw and normalized files have identical row counts and +# the same row order, so a positional join is sufficient and avoids needing +# to reproduce the archived normalizer's own question_id hash algorithm. +# This is a ChoiceBench revalidation of internal freeze consistency, not +# proof that the archived normalizer (or the raw upstream publisher) was +# authentic -- recorded as a limitation on the returned ExpectedDataset. + +_DATASET_RELATIVE_PATHS = { + "arc_challenge": { + "raw": "raw/local_model_generalization/data/raw/arc_challenge_raw.csv", + "normalized": "raw/local_model_generalization/data/processed/arc_challenge_normalized.csv", + "ids": "raw/local_model_generalization/data/splits/arc_challenge/robustness_ids.json", + "metadata": "raw/local_model_generalization/data/splits/arc_challenge/robustness_metadata.json", + }, + "mmlu": { + "raw": "raw/local_model_generalization/data/raw/mmlu_raw.csv", + "normalized": "raw/local_model_generalization/data/processed/mmlu_normalized.csv", + "ids": "raw/local_model_generalization/data/splits/benchmark/robustness_ids.json", + "metadata": "raw/local_model_generalization/data/splits/benchmark/robustness_metadata.json", + }, +} + + +_ARC_MAX_OPTIONS = 4 + + +def _recompute_arc_fields(raw_row: Mapping[str, str]) -> dict[str, str]: + """The archived normalizer caps ARC options at _ARC_MAX_OPTIONS, silently + dropping any raw option beyond it -- verified against the real freeze: + 3 raw ARC questions have a 5th option, and all 3 have their answerKey + within the first _ARC_MAX_OPTIONS, so truncation never drops the correct + option. If a future/other freeze ever did drop the correct option, the + index check below still fails closed rather than silently truncating it. + """ + try: + parsed = json.loads(raw_row["choices"]) + texts = list(parsed["text"]) + labels = list(parsed["label"]) + index = labels.index(raw_row["answerKey"]) + except (KeyError, ValueError, TypeError) as exc: + raise Stage1ProfileError( + f"ARC raw row {raw_row.get('id')!r} has malformed choices/answerKey: {exc}." + ) from exc + if index >= _ARC_MAX_OPTIONS: + raise Stage1ProfileError( + f"ARC raw row {raw_row.get('id')!r} answerKey index {index} falls outside the " + f"archived normalizer's first {_ARC_MAX_OPTIONS} options; truncation would drop " + "the correct option, so this cannot be safely revalidated." + ) + texts = texts[:_ARC_MAX_OPTIONS] + fields = { + "question_text": raw_row["question"], + "correct_option": chr(ord("A") + index), + "correct_answer_text": texts[index], + } + for offset in range(_ARC_MAX_OPTIONS): + fields[f"choice_{chr(ord('a') + offset)}"] = texts[offset] if offset < len(texts) else "" + return fields + + +def _recompute_mmlu_fields(raw_row: Mapping[str, str]) -> dict[str, str]: + try: + choices = list(ast.literal_eval(raw_row["choices"])) + index = int(raw_row["answer"]) + except (SyntaxError, ValueError, TypeError) as exc: + raise Stage1ProfileError( + f"MMLU raw row has malformed choices/answer: {exc}." + ) from exc + if not (0 <= index < len(choices)): + raise Stage1ProfileError( + f"MMLU raw row answer index {index} is out of range for choices {choices!r}." + ) + fields = { + "question_text": raw_row["question"], + "correct_option": chr(ord("A") + index), + "correct_answer_text": str(choices[index]), + } + for offset, text in enumerate(choices): + fields[f"choice_{chr(ord('a') + offset)}"] = str(text) + return fields + + +def _revalidate_positional( + raw_rows: list[dict[str, str]], + normalized_rows: list[dict[str, str]], + *, + recompute: Callable[[Mapping[str, str]], dict[str, str]], + benchmark_name: str, +) -> None: + if len(raw_rows) != len(normalized_rows): + raise Stage1ProfileError( + f"{benchmark_name}: raw ({len(raw_rows)}) and normalized ({len(normalized_rows)}) " + "row counts disagree; this revalidation assumes positional 1:1 correspondence." + ) + for index, (raw_row, normalized_row) in enumerate(zip(raw_rows, normalized_rows)): + recomputed = recompute(raw_row) + for field, expected_value in recomputed.items(): + actual_value = (normalized_row.get(field) or "").strip() + if actual_value != expected_value.strip(): + raise Stage1ProfileError( + f"{benchmark_name} row {index} ({normalized_row.get('question_id')!r}): " + f"recomputed {field}={expected_value!r} does not match the archived " + f"normalized value {actual_value!r} (ChoiceBench revalidation failure)." + ) + + +def _drop_duplicate_questions( + normalized_rows: list[dict[str, str]], *, benchmark_name: str +) -> list[dict[str, str]]: + """pandas.DataFrame.drop_duplicates(subset="question_id", keep="first") + semantics -- but only after requiring every duplicate occurrence to agree + on all parsed fields; a disagreeing duplicate fails closed rather than + silently picking one.""" + seen: dict[str, dict[str, str]] = {} + order: list[str] = [] + for row in normalized_rows: + question_id = row["question_id"] + if question_id in seen: + if row != seen[question_id]: + raise Stage1ProfileError( + f"{benchmark_name}: duplicate question_id {question_id!r} occurs with " + "disagreeing field values; cannot apply keep-first deduplication." + ) + continue + seen[question_id] = row + order.append(question_id) + return [seen[question_id] for question_id in order] + + +def _build_expected_dataset_for_benchmark( + freeze_root: Path, + ledger: Mapping[str, str], + *, + benchmark_name: str, + raw_recompute: Callable[[Mapping[str, str]], dict[str, str]], +) -> ExpectedDataset: + paths = _DATASET_RELATIVE_PATHS[benchmark_name] + raw_bytes = _verify_checksummed_file(freeze_root, ledger, paths["raw"]) + normalized_bytes = _verify_checksummed_file(freeze_root, ledger, paths["normalized"]) + ids_bytes = _verify_checksummed_file(freeze_root, ledger, paths["ids"]) + metadata_bytes = _verify_checksummed_file(freeze_root, ledger, paths["metadata"]) + + raw_rows = list(csv.DictReader(io.StringIO(raw_bytes.decode("utf-8")))) + normalized_rows = list(csv.DictReader(io.StringIO(normalized_bytes.decode("utf-8")))) + _revalidate_positional( + raw_rows, normalized_rows, recompute=raw_recompute, benchmark_name=benchmark_name + ) + + deduplicated_rows = _drop_duplicate_questions(normalized_rows, benchmark_name=benchmark_name) + selected_ids = json.loads(ids_bytes) + metadata = json.loads(metadata_bytes) + + available_ids = {row["question_id"] for row in deduplicated_rows} + missing = [question_id for question_id in selected_ids if question_id not in available_ids] + if missing: + raise Stage1ProfileError( + f"{benchmark_name}: {len(missing)} selected question ID(s) are absent from the " + f"normalized/deduplicated source, e.g. {missing[:3]}." + ) + + fieldnames = list(deduplicated_rows[0].keys()) + buffer = io.StringIO() + writer = csv.DictWriter(buffer, fieldnames=fieldnames) + writer.writeheader() + writer.writerows(deduplicated_rows) + deduplicated_bytes = buffer.getvalue().encode("utf-8") + opened_source = OpenedSource( + source_id="normalized", + audit_path=freeze_root / paths["normalized"], + logical_path=paths["normalized"], + data=deduplicated_bytes, + sha256=sha256(deduplicated_bytes).hexdigest(), + ) + + choice_columns = { + key: key for key in fieldnames if key.startswith("choice_") and len(key) == 8 + } + declaration = DatasetReferenceSpec( + dataset_id=benchmark_name, + benchmark_name=benchmark_name, + split="robustness", + reference_kind="independent_input_snapshot", + trust_label="checksum_verified_freeze_internal", + source_ids=("normalized",), + selection_source_id="normalized", + expected_question_ids=tuple(selected_ids), + selection_seed=metadata.get("seed"), + selection_n_samples=metadata.get("actual_size"), + subject_filter=(), + selection_unknown_reasons={}, + columns={ + "question_id": "question_id", + "question_text": "question_text", + "correct_option": "correct_option", + **choice_columns, + }, + revision=None, + fingerprint=None, + derivation={"source_role": f"{benchmark_name}_stage1_freeze_normalized"}, + limitations=( + "Stage 1 freeze: field content independently revalidated against the raw " + "source by ChoiceBench; this proves internal freeze consistency, not that " + "the archived normalizer or the upstream publisher was authentic.", + ), + native_compatibility_identity=None, + ) + return build_expected_dataset(declaration, {"normalized": opened_source}) + + +def build_stage1_expected_datasets(freeze_root: Path) -> dict[str, ExpectedDataset]: + freeze_root = Path(freeze_root) + ledger = _load_checksum_ledger(freeze_root) + return { + "arc_challenge": _build_expected_dataset_for_benchmark( + freeze_root, ledger, benchmark_name="arc_challenge", raw_recompute=_recompute_arc_fields, + ), + "mmlu": _build_expected_dataset_for_benchmark( + freeze_root, ledger, benchmark_name="mmlu", raw_recompute=_recompute_mmlu_fields, + ), + } diff --git a/tests/importing/test_stage1_profile.py b/tests/importing/test_stage1_profile.py index c16a40f..38f3421 100644 --- a/tests/importing/test_stage1_profile.py +++ b/tests/importing/test_stage1_profile.py @@ -12,7 +12,9 @@ from __future__ import annotations +import csv from hashlib import sha256 +import io import json from pathlib import Path @@ -21,6 +23,7 @@ from choicebench.importing.profiles.stage1_paper_freeze import ( STATUS_MAP, Stage1ProfileError, + build_stage1_expected_datasets, translate_stage1_paper_freeze, ) @@ -281,3 +284,255 @@ def test_status_map_matches_canonical_values(): assert STATUS_MAP["incomplete_requires_inference"] == "partial" assert STATUS_MAP["malformed_requires_inference"] == "malformed" assert "excluded_from_paper_matrix" not in STATUS_MAP # not an evidence-status input + + +# --- ARC/MMLU expected-dataset trust chain (Task 16) ------------------------ +# +# Synthetic fixture shaped exactly like the real freeze's raw/normalized/ +# split file formats (verified against the real freeze during development): +# ARC raw choices are {"text": [...], "label": [...]}; MMLU raw choices are a +# Python-repr list with an integer answer index; both normalized files share +# one schema (question_id/subject/question_text/choice_a..d/correct_option/ +# correct_answer_text). Includes ARC's real edge cases: a 3-option row +# (variable option count) and a 5-option row (the archived normalizer +# silently truncates to 4, verified in the real freeze to never drop the +# correct option) -- plus an MMLU field-identical duplicate question_id. + +_DATASET_RELATIVE_PATHS = ( + "raw/local_model_generalization/data/raw/arc_challenge_raw.csv", + "raw/local_model_generalization/data/processed/arc_challenge_normalized.csv", + "raw/local_model_generalization/data/splits/arc_challenge/robustness_ids.json", + "raw/local_model_generalization/data/splits/arc_challenge/robustness_metadata.json", + "raw/local_model_generalization/data/raw/mmlu_raw.csv", + "raw/local_model_generalization/data/processed/mmlu_normalized.csv", + "raw/local_model_generalization/data/splits/benchmark/robustness_ids.json", + "raw/local_model_generalization/data/splits/benchmark/robustness_metadata.json", +) + + +def _arc_raw_csv(rows): + buffer = io.StringIO() + writer = csv.DictWriter(buffer, fieldnames=["id", "question", "choices", "answerKey"]) + writer.writeheader() + for row_id, question, texts, labels, answer_key in rows: + writer.writerow({ + "id": row_id, "question": question, + "choices": json.dumps({"text": texts, "label": labels}), + "answerKey": answer_key, + }) + return buffer.getvalue() + + +def _normalized_csv(rows): + fieldnames = [ + "question_id", "subject", "question_text", "choice_a", "choice_b", + "choice_c", "choice_d", "correct_option", "correct_answer_text", + ] + buffer = io.StringIO() + writer = csv.DictWriter(buffer, fieldnames=fieldnames) + writer.writeheader() + for row in rows: + writer.writerow(row) + return buffer.getvalue() + + +def _default_arc_raw_rows(): + return [ + ("r1", "Q1?", ["a1", "b1", "c1", "d1"], ["A", "B", "C", "D"], "A"), + ("r2", "Q2?", ["a2", "b2", "c2", "d2"], ["A", "B", "C", "D"], "B"), + ("r3", "Q3?", ["a3", "b3", "c3"], ["A", "B", "C"], "A"), # 3-option + ("r4", "Q4?", ["a4", "b4", "c4", "d4", "e4"], ["A", "B", "C", "D", "E"], "B"), # 5-option + ] + + +def _default_arc_normalized_rows(): + return [ + {"question_id": "arc_q1", "subject": "arc_challenge", "question_text": "Q1?", + "choice_a": "a1", "choice_b": "b1", "choice_c": "c1", "choice_d": "d1", + "correct_option": "A", "correct_answer_text": "a1"}, + {"question_id": "arc_q2", "subject": "arc_challenge", "question_text": "Q2?", + "choice_a": "a2", "choice_b": "b2", "choice_c": "c2", "choice_d": "d2", + "correct_option": "B", "correct_answer_text": "b2"}, + {"question_id": "arc_q3", "subject": "arc_challenge", "question_text": "Q3?", + "choice_a": "a3", "choice_b": "b3", "choice_c": "c3", "choice_d": "", + "correct_option": "A", "correct_answer_text": "a3"}, + {"question_id": "arc_q4", "subject": "arc_challenge", "question_text": "Q4?", + "choice_a": "a4", "choice_b": "b4", "choice_c": "c4", "choice_d": "d4", + "correct_option": "B", "correct_answer_text": "b4"}, + ] + + +def _default_mmlu_raw_rows(): + # (question, subject, choices_repr, answer_index) + return [ + ("MQ1?", "history", "['x1', 'y1', 'z1', 'w1']", 0), + ("MQ2?", "history", "['x2', 'y2', 'z2', 'w2']", 2), + ("MQ2?", "history", "['x2', 'y2', 'z2', 'w2']", 2), # duplicate row -> same question_id + ] + + +def _default_mmlu_normalized_rows(): + return [ + {"question_id": "mmlu_q1", "subject": "history", "question_text": "MQ1?", + "choice_a": "x1", "choice_b": "y1", "choice_c": "z1", "choice_d": "w1", + "correct_option": "A", "correct_answer_text": "x1"}, + {"question_id": "mmlu_q2", "subject": "history", "question_text": "MQ2?", + "choice_a": "x2", "choice_b": "y2", "choice_c": "z2", "choice_d": "w2", + "correct_option": "C", "correct_answer_text": "z2"}, + {"question_id": "mmlu_q2", "subject": "history", "question_text": "MQ2?", + "choice_a": "x2", "choice_b": "y2", "choice_c": "z2", "choice_d": "w2", + "correct_option": "C", "correct_answer_text": "z2"}, + ] + + +def _build_dataset_freeze( + tmp_path, + *, + arc_raw_rows=None, + arc_normalized_rows=None, + mmlu_raw_rows=None, + mmlu_normalized_rows=None, + arc_selected_ids=None, + mmlu_selected_ids=None, + tamper_checksum_for=None, +): + freeze_root = tmp_path / "freeze" + arc_raw_rows = _default_arc_raw_rows() if arc_raw_rows is None else arc_raw_rows + arc_normalized_rows = ( + _default_arc_normalized_rows() if arc_normalized_rows is None else arc_normalized_rows + ) + mmlu_raw_rows = _default_mmlu_raw_rows() if mmlu_raw_rows is None else mmlu_raw_rows + mmlu_normalized_rows = ( + _default_mmlu_normalized_rows() if mmlu_normalized_rows is None else mmlu_normalized_rows + ) + arc_selected_ids = ( + [row["question_id"] for row in arc_normalized_rows] + if arc_selected_ids is None else arc_selected_ids + ) + mmlu_selected_ids = ( + sorted({row["question_id"] for row in mmlu_normalized_rows}) + if mmlu_selected_ids is None else mmlu_selected_ids + ) + + _write( + freeze_root / "raw/local_model_generalization/data/raw/arc_challenge_raw.csv", + _arc_raw_csv(arc_raw_rows), + ) + _write( + freeze_root / "raw/local_model_generalization/data/processed/arc_challenge_normalized.csv", + _normalized_csv(arc_normalized_rows), + ) + _write( + freeze_root / "raw/local_model_generalization/data/splits/arc_challenge/robustness_ids.json", + json.dumps(arc_selected_ids), + ) + _write( + freeze_root / "raw/local_model_generalization/data/splits/arc_challenge/robustness_metadata.json", + json.dumps({"seed": 42, "actual_size": len(arc_selected_ids)}), + ) + + mmlu_buffer = io.StringIO() + writer = csv.DictWriter(mmlu_buffer, fieldnames=["question", "subject", "choices", "answer"]) + writer.writeheader() + for question, subject, choices_repr, answer in mmlu_raw_rows: + writer.writerow({"question": question, "subject": subject, "choices": choices_repr, "answer": answer}) + _write( + freeze_root / "raw/local_model_generalization/data/raw/mmlu_raw.csv", mmlu_buffer.getvalue(), + ) + _write( + freeze_root / "raw/local_model_generalization/data/processed/mmlu_normalized.csv", + _normalized_csv(mmlu_normalized_rows), + ) + _write( + freeze_root / "raw/local_model_generalization/data/splits/benchmark/robustness_ids.json", + json.dumps(mmlu_selected_ids), + ) + _write( + freeze_root / "raw/local_model_generalization/data/splits/benchmark/robustness_metadata.json", + json.dumps({"seed": 7, "actual_size": len(mmlu_selected_ids)}), + ) + + ledger_lines = [] + for relative_path in _DATASET_RELATIVE_PATHS: + data = (freeze_root / relative_path).read_bytes() + digest = sha256(data).hexdigest() + if tamper_checksum_for == relative_path: + digest = "f" * 64 + ledger_lines.append(f"{digest} {relative_path}") + _write(freeze_root / "checksums" / "checksums.sha256", "\n".join(ledger_lines) + "\n") + return freeze_root + + +def test_build_stage1_expected_datasets_happy_path(tmp_path): + freeze_root = _build_dataset_freeze(tmp_path) + datasets = build_stage1_expected_datasets(freeze_root) + assert set(datasets) == {"arc_challenge", "mmlu"} + arc = datasets["arc_challenge"] + assert arc.reference_kind == "independent_input_snapshot" + assert arc.trust_label == "checksum_verified_freeze_internal" + assert set(arc.selected_question_ids) == {"arc_q1", "arc_q2", "arc_q3", "arc_q4"} + mmlu = datasets["mmlu"] + # the duplicate mmlu_q2 row collapses to one selected ID (keep-first) + assert set(mmlu.selected_question_ids) == {"mmlu_q1", "mmlu_q2"} + + +def test_build_stage1_expected_datasets_handles_variable_arc_option_count(tmp_path): + freeze_root = _build_dataset_freeze(tmp_path) + datasets = build_stage1_expected_datasets(freeze_root) + arc_frame = datasets["arc_challenge"].artifact_frame + row = arc_frame[arc_frame["question_id"] == "arc_q3"].iloc[0] + choices = json.loads(row["choices_json"]) + assert len(choices) == 3 # variable option count preserved, not padded + + +def test_build_stage1_expected_datasets_truncates_five_option_arc_row(tmp_path): + freeze_root = _build_dataset_freeze(tmp_path) + datasets = build_stage1_expected_datasets(freeze_root) + arc_frame = datasets["arc_challenge"].artifact_frame + row = arc_frame[arc_frame["question_id"] == "arc_q4"].iloc[0] + choices = json.loads(row["choices_json"]) + assert len(choices) == 4 # 5th raw option dropped, matching the archived normalizer + assert row["correct_option"] == "B" + + +def test_build_stage1_expected_datasets_verifies_checksums(tmp_path): + freeze_root = _build_dataset_freeze( + tmp_path, + tamper_checksum_for="raw/local_model_generalization/data/processed/arc_challenge_normalized.csv", + ) + with pytest.raises(Stage1ProfileError, match="checksum mismatch"): + build_stage1_expected_datasets(freeze_root) + + +def test_build_stage1_expected_datasets_rejects_revalidation_mismatch(tmp_path): + tampered = _default_arc_normalized_rows() + tampered[0] = {**tampered[0], "correct_option": "B"} # disagrees with raw answerKey "A" + freeze_root = _build_dataset_freeze(tmp_path, arc_normalized_rows=tampered) + with pytest.raises(Stage1ProfileError, match="revalidation failure"): + build_stage1_expected_datasets(freeze_root) + + +def test_build_stage1_expected_datasets_rejects_disagreeing_duplicate(tmp_path): + # The mutated 3rd row must still positionally revalidate against its OWN + # raw row (answer index 0 -> "A"/"x2"), so the disagreement is caught at + # the duplicate-question_id stage, not the raw/normalized revalidation + # stage -- isolating the specific invariant this test targets. + raw_rows = _default_mmlu_raw_rows() + raw_rows[2] = ("MQ2?", "history", "['x2', 'y2', 'z2', 'w2']", 0) + mismatched = _default_mmlu_normalized_rows() + mismatched[2] = {**mismatched[2], "correct_option": "A", "correct_answer_text": "x2"} + freeze_root = _build_dataset_freeze(tmp_path, mmlu_raw_rows=raw_rows, mmlu_normalized_rows=mismatched) + with pytest.raises(Stage1ProfileError, match="disagreeing field values"): + build_stage1_expected_datasets(freeze_root) + + +def test_build_stage1_expected_datasets_rejects_selected_id_missing_from_source(tmp_path): + freeze_root = _build_dataset_freeze(tmp_path, arc_selected_ids=["arc_q1", "arc_qXX"]) + with pytest.raises(Stage1ProfileError, match="absent from the"): + build_stage1_expected_datasets(freeze_root) + + +def test_build_stage1_expected_datasets_rejects_raw_normalized_row_count_mismatch(tmp_path): + freeze_root = _build_dataset_freeze(tmp_path, arc_raw_rows=_default_arc_raw_rows()[:2]) + with pytest.raises(Stage1ProfileError, match="row counts disagree"): + build_stage1_expected_datasets(freeze_root) From 80ee933171cb717bf5d0bbb5b55df9c558c6d686 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 20:20:02 +0300 Subject: [PATCH 38/47] feat: overlay merge-into-new-run orchestration + expected_datasets override - engine.py: execute_overlay_import/_merge_overlay_realization merge an authorized repair/offline-transformation overlay with a verified base run's retained rows into a new, separate, immutable run. Row/lineage assignments are reassembled in the expected dataset's question order (not retained-then-replaced concatenation order), matching make_realization's ordered-subset identity requirement. Divergent overwrite of an existing overlay run_id now raises (matching the base import path's verify_import_run(expected=...) contract) instead of silently treating a different experiment as an idempotent no-op. - engine.py: ImportRequest gains an optional expected_datasets override so a profile needing dataset-specific revalidation/dedup (e.g. Stage 1's ARC/MMLU trust chain) can supply pre-built ExpectedDataset objects instead of the generic per-source rebuild, which cannot reproduce that logic (needed for MMLU's real duplicate question_ids). - validation.py: extract compute_evidence_status as a reusable pure function from normalize_realization_rows (no behavior change), so a profile assembling declarations can discover the correct evidence_status to declare instead of guessing against the declaration-mismatch check. - tests/importing/test_engine_overlay.py (new): merge, offline- transformation, dry-run, idempotence, divergent-overwrite refusal, unknown-base-realization, forged-source rejection. - tests/importing/test_engine.py: expected_datasets override coverage. 1321 full tests passed. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/importing/engine.py | 433 ++++++++++++++++++++++-- src/choicebench/importing/validation.py | 39 ++- tests/importing/test_engine.py | 34 +- tests/importing/test_engine_overlay.py | 387 +++++++++++++++++++++ 4 files changed, 859 insertions(+), 34 deletions(-) create mode 100644 tests/importing/test_engine_overlay.py diff --git a/src/choicebench/importing/engine.py b/src/choicebench/importing/engine.py index 3035beb..ed1974d 100644 --- a/src/choicebench/importing/engine.py +++ b/src/choicebench/importing/engine.py @@ -1,14 +1,12 @@ """Orchestrate generic dry-run and real imports of external results. -Reduced scope: this implements the base (non-overlay) import path fully -- +Reduced scope: this implements both the base (non-overlay) import path -- plan, dry-run validate, real staged publish, and idempotent re-verification -of an existing run. Applying an authorized repair/offline-transformation -overlay on top of an already-published base run (Task 10's derive_overlay, -merged with the base's retained rows into a new run) is the next increment; -every piece it needs (identity, validation, authorization, overlay -derivation, evidence storage, atomic transactions, and this module's own -verify_import_run) is already built and tested, but the merge-and-publish -orchestration itself is not yet wired up here. +of an existing run -- and the overlay merge-into-new-run path: an authorized +repair/offline-transformation overlay (Task 10's derive_overlay) merged with +an already-published base run's retained rows into a new, separate, +immutable run (execute_overlay_import). The base run itself is never +mutated; only its verified values are read. Report schema is also reduced from the original plan: it carries the counts and digests needed to confirm exact status/queue reproduction and reject @@ -20,6 +18,7 @@ from dataclasses import dataclass from pathlib import Path +from types import SimpleNamespace from typing import Any, Literal, Mapping from dataclasses import asdict @@ -45,10 +44,17 @@ make_realization, make_result_origin, ) -from choicebench.importing.overlays import VerifiedBaseRealization +from choicebench.importing.authorization import ( + authorization_for_condition, + validate_authorization_bundle, +) +from choicebench.importing.overlays import VerifiedBaseRealization, derive_overlay from choicebench.importing.schema import ( + AuthorizationSpec, ImportConditionSpec, ImportSpec, + OverlaySpec, + SourceArtifactSpec, validate_csv_dialect_identity, validate_numeric_columns_identity, validate_option_mapping_identity, @@ -56,6 +62,7 @@ from choicebench.importing.transaction import ImportTransaction from choicebench.importing.validation import ( ImportValidationError, + RealizationValidation, normalize_realization_rows, prepare_realization_validation_artifact, validate_realization_validation_artifact, @@ -115,6 +122,14 @@ class ImportRequest: workspace_root: Path strict: bool overlays: tuple[Any, ...] = () + expected_datasets: Mapping[str, ExpectedDataset] | None = None + """Pre-built, pre-verified ExpectedDataset objects keyed by dataset_id, + used in place of rebuilding each spec.datasets declaration from its + source_ids via the generic build_expected_dataset. Some profiles (e.g. + Stage 1's ARC/MMLU trust chain) need dataset-specific revalidation and + deduplication the generic path cannot reproduce; spec.datasets is still + populated for documentation/identity purposes, but is not read here when + this override is supplied.""" @dataclass(frozen=True) @@ -335,15 +350,26 @@ def build_import_plan(request: ImportRequest) -> ImportPlan: Performs every read/validation/identity step but writes nothing.""" spec = request.spec sources_by_id = {source.source_id: source for source in spec.sources} - expected_datasets: dict[str, ExpectedDataset] = {} - for declaration in spec.datasets: - opened = { - source_id: open_verified_source( - sources_by_id[source_id], containment_root=request.workspace_root + if request.expected_datasets is not None: + expected_datasets: dict[str, ExpectedDataset] = dict(request.expected_datasets) + missing_datasets = sorted( + {condition.dataset_id for condition in spec.conditions} - set(expected_datasets) + ) + if missing_datasets: + raise ImportEngineError( + f"request.expected_datasets is missing dataset(s) {missing_datasets} " + "referenced by spec.conditions." ) - for source_id in declaration.source_ids - } - expected_datasets[declaration.dataset_id] = build_expected_dataset(declaration, opened) + else: + expected_datasets = {} + for declaration in spec.datasets: + opened = { + source_id: open_verified_source( + sources_by_id[source_id], containment_root=request.workspace_root + ) + for source_id in declaration.source_ids + } + expected_datasets[declaration.dataset_id] = build_expected_dataset(declaration, opened) models_by_key = {model.model_key: model for model in spec.models} methods_by_key = {method.method_key: method for method in spec.methods} @@ -561,10 +587,14 @@ def _write_staged_run(staged_run: Path, request: ImportRequest, plan: ImportPlan validate_manifest_v3(plan.manifest) -def verify_import_run(run_dir: Path, expected: ImportPlan | None = None) -> VerifiedImportRun: +def verify_import_run( + run_dir: Path, expected: ImportPlan | SimpleNamespace | None = None +) -> VerifiedImportRun: """The single shared full-graph verifier: validates manifest, state, validation artifacts, and results, and returns immutable trusted values - only after the whole graph passes.""" + only after the whole graph passes. `expected` need only expose a + `.manifest` mapping (an ImportPlan, or a bare SimpleNamespace(manifest=...) + for the overlay path, which has no full ImportPlan of its own).""" run_dir = Path(run_dir) manifest = json.loads((run_dir / MANIFEST_FILENAME).read_text()) validate_manifest_v3(manifest) @@ -636,6 +666,371 @@ def verify_import_run(run_dir: Path, expected: ImportPlan | None = None) -> Veri ) +# --- Overlay merge-into-new-run orchestration ------------------------------- +# +# Scope boundary: requires an evaluable base realization (one with a +# published result CSV to retain rows from). A non-evaluable base +# (partial/malformed/recoverable/failed) has no per-row content stored +# anywhere to retain -- verified directly against the real Stage 1 freeze +# during development: its "recoverable" cells' rows are all individually +# well-formed (parse_status=parse_ok, valid parsed_choice), so this +# engine's own generic row validation genuinely computes "complete" for +# them, and correctly refuses a declared "recoverable" status as a +# declaration mismatch rather than silently trusting freeze-specific domain +# knowledge (the freeze's own "option-hidden" qualification note) it cannot +# independently verify. Extending non-evaluable bases to carry retained +# per-row content would require validation.py to store full row content +# for every realization (not just evaluable_rows) -- a real extension, not +# attempted here. + +@dataclass(frozen=True) +class OverlayImportRequest: + base_run_dir: Path + base_realization_id: str + authorization: AuthorizationSpec + authorization_source: SourceArtifactSpec + overlay: OverlaySpec + overlay_source: SourceArtifactSpec + overlay_mapping: Mapping[str, str] + condition_digests: Mapping[str, str] + expected_dataset: ExpectedDataset + run_id: str + workspace_root: Path + + +def _merge_overlay_realization( + request: OverlayImportRequest, +) -> tuple[dict[str, Any], dict[str, Any], list[dict[str, Any]], dict[str, Any]]: + """Returns (realization_record, semantic_condition_record, result_rows, + overlay_evidence_record). The evidence record covers only the new + overlay source -- base evidence remains immutably stored in the base + run's own directory and is reachable via parent_digests.""" + verified = verify_import_run(request.base_run_dir) + base = verified.realizations.get(request.base_realization_id) + if base is None: + raise ImportEngineError(f"Unknown base realization {request.base_realization_id!r}.") + if base.result_sha256 is None: + raise ImportEngineError( + f"Realization {request.base_realization_id!r} has no published result to " + "retain rows from; only an evaluable base realization can be overlaid." + ) + + base_manifest = verified.manifest + base_realization_record = base_manifest["payload"]["realizations"][request.base_realization_id] + base_realization_identity = base_realization_record["identity"]["realization"] + condition_id = base_realization_record["condition_id"] + semantic_condition = base_manifest["payload"]["semantic_conditions"][condition_id] + + opened_auth_source = open_verified_source( + request.authorization_source, containment_root=request.workspace_root + ) + bundle = validate_authorization_bundle( + request.authorization, opened_source=opened_auth_source, + condition_digests=request.condition_digests, + expected={base.condition_digest: request.expected_dataset}, + ) + authorization = authorization_for_condition(bundle, condition_digest=base.condition_digest) + + opened_overlay_source = open_verified_source( + request.overlay_source, containment_root=request.workspace_root + ) + overlay_table = parse_csv_source(opened_overlay_source, request.overlay_source, strict=True) + + derived = derive_overlay( + base=base, overlay=request.overlay, authorization=authorization, + overlay_table=overlay_table, overlay_mapping=request.overlay_mapping, + expected=request.expected_dataset, + ) + + base_lineage_by_id = { + component["lineage_id"]: component + for component in base_realization_identity["lineage_components"] + } + base_row_assignments = base_realization_identity["result_origin"]["row_assignments"] + retained_assignments = [ + assignment for assignment in base_row_assignments + if assignment["question_id"] not in derived.replacement_question_ids + ] + retained_lineage_ids = {assignment["prediction_lineage_id"] for assignment in retained_assignments} + retained_lineage_components = [base_lineage_by_id[lid] for lid in retained_lineage_ids] + + all_lineage_components = [*retained_lineage_components, *derived.lineage_components] + + # Realization identity requires row_assignments to be an ordered subset of + # the expected dataset's question order (identity.py), not "retained then + # replaced" -- reassemble in dataset order regardless of which side (base + # or overlay) supplies each question. + assignments_by_qid = { + assignment["question_id"]: ( + assignment["question_id"], assignment["prediction_origin"], assignment["prediction_lineage_id"] + ) + for assignment in retained_assignments + } + assignments_by_qid.update({ + component["identity"]["question_id"]: ( + component["identity"]["question_id"], + component["identity"]["prediction_origin"], + component["lineage_id"], + ) + for component in derived.lineage_components + }) + row_assignments = [ + assignments_by_qid[qid] + for qid in request.expected_dataset.selected_question_ids + if qid in assignments_by_qid + ] + derivation_origin = derived.lineage_components[0]["identity"]["operation_type"] + result_origin = make_result_origin(derivation_origin=derivation_origin, row_assignments=row_assignments) + + replacement_digest = integrity_digest( + canonicalize({ + "replacement_question_ids": list(derived.replacement_question_ids), + "replacement_reasons": dict(request.overlay.replacement_reasons), + }) + ) + overlay_dict: dict[str, Any] = { + "source_sha256": overlay_table.source_sha256, + "replacement_digest": replacement_digest, + "transformation_input_digest": None, + "preownership_output_digest": None, + "implementation_digest": None, + } + if derivation_origin == "offline_transformation": + overlay_dict["transformation_input_digest"] = integrity_digest( + [component["identity"]["input_digest"] for component in derived.lineage_components] + ) + overlay_dict["preownership_output_digest"] = integrity_digest( + [component["identity"]["preownership_output_digest"] for component in derived.lineage_components] + ) + overlay_dict["implementation_digest"] = integrity_digest( + derived.lineage_components[0]["identity"]["implementation"] + ) + + overlay_source_notes_digest = ( + integrity_digest(request.overlay_source.notes) if request.overlay_source.notes else None + ) + source_records = [ + *base_realization_identity["sources"], + _source_record(request.overlay_source, opened_overlay_source, notes_digest=overlay_source_notes_digest), + ] + + realization_identity = { + "import_spec_digest": base_realization_identity["import_spec_digest"], + "sources": source_records, + "expected_dataset": base_realization_identity["expected_dataset"], + "importer_implementation": base_realization_identity["importer_implementation"], + "parsing_policy": base_realization_identity["parsing_policy"], + "validation": base_realization_identity["validation"], + "evidence": { + **base_realization_identity["evidence"], + "evidence_status": request.overlay.expected_evidence_status, + }, + "lineage_components": all_lineage_components, + "result_origin": result_origin, + "parent_digests": { + "realization_digests": [base.realization_digest], + "evidence_digests": [base.validation_artifact_sha256], + "result_digests": [base.result_sha256] if base.result_sha256 else [], + }, + "authorization_digest": authorization.bundle_digest, + "overlay": overlay_dict, + } + realization = make_realization( + condition_id=condition_id, + condition_digest=base.condition_digest, + identity=realization_identity, + fields={}, + adapter=_external_import_adapter, + validator=_external_import_validator, + expected_dataset=request.expected_dataset, + semantic_condition=semantic_condition, + lineage_runtime_callables={lid: _external_import_adapter for lid in retained_lineage_ids}, + ) + + _ROW_KEYS = { + "question_id", "question_text", "correct_option", "choices_json", + "prediction_origin", "predicted_option", + } + retained_rows_by_qid = { + assignment["question_id"]: { + key: value + for key, value in base.rows_by_question_id[assignment["question_id"]].items() + if key in _ROW_KEYS + } + for assignment in retained_assignments + } + derived_rows_by_qid = {row["question_id"]: dict(row) for row in derived.preownership_rows} + rows_by_qid = {**retained_rows_by_qid, **derived_rows_by_qid} + result_rows = [ + rows_by_qid[qid] for qid in request.expected_dataset.selected_question_ids if qid in rows_by_qid + ] + + overlay_evidence_record = evidence_record( + opened_overlay_source, references=sorted(derived.replacement_question_ids) + ) + return realization, semantic_condition, result_rows, overlay_evidence_record + + +def execute_overlay_import(request: OverlayImportRequest, *, dry_run: bool = False) -> ImportReport: + try: + realization, semantic_condition, result_rows, overlay_evidence_record = ( + _merge_overlay_realization(request) + ) + except (ImportIdentityError, ImportEngineError, ValueError) as exc: + return ImportReport( + schema_version="choicebench.import-report.v1", + import_state="failed", + wrote_artifacts=False, + idempotent_noop=False, + run_id=request.run_id, + experiment_id=None, + counts=ImportCounts(conditions={}, evidence_status={}, scope_disposition={}), + defects=DefectReport(total_findings=0, by_code={}, findings_digest=integrity_digest([])), + condition_digests={}, + realization_digests={}, + failures=({"code": "OVERLAY_IMPORT_FAILED", "sanitized_message": str(exc)},), + ) + + condition_id = realization["condition_id"] + payload = { + "protocol_version": PROTOCOL_V3_VERSION, + "canonicalization_version": CANONICALIZATION_VERSION, + "semantic_conditions": {condition_id: semantic_condition}, + "realizations": {realization["realization_id"]: realization}, + } + manifest = make_manifest_v3(payload, audit={"source_location": str(request.workspace_root)}) + validate_manifest_v3(manifest) + + evidence_status = realization["identity"]["realization"]["evidence"]["evidence_status"] + scope_disposition = realization["identity"]["realization"]["evidence"]["scope_disposition"] + counts = ImportCounts( + conditions={"declared": 1}, + evidence_status={evidence_status: 1}, + scope_disposition={scope_disposition: 1}, + ) + condition_digests = {condition_id: semantic_condition["condition_digest"]} + realization_digests = {realization["realization_id"]: realization["realization_digest"]} + + if dry_run: + return ImportReport( + schema_version="choicebench.import-report.v1", + import_state="validated", + wrote_artifacts=False, + idempotent_noop=False, + run_id=request.run_id, + experiment_id=manifest["experiment_id"], + counts=counts, + defects=DefectReport(total_findings=0, by_code={}, findings_digest=integrity_digest([])), + condition_digests=condition_digests, + realization_digests=realization_digests, + failures=(), + ) + + runs_dir = Path(request.workspace_root) / "runs" + idempotent_noop = [False] + + def _validator(published_root: Path) -> None: + if published_root == runs_dir / request.run_id and (published_root / MANIFEST_FILENAME).is_file(): + verify_import_run(published_root, expected=SimpleNamespace(manifest=manifest)) + idempotent_noop[0] = True + return + _write_overlay_staged_run( + published_root, manifest, realization, result_rows, + overlay_source=request.overlay_source, + overlay_evidence_record=overlay_evidence_record, + workspace_root=request.workspace_root, + ) + + with ImportTransaction(runs_dir=runs_dir, run_id=request.run_id) as txn: + txn.publish(_validator) + + return ImportReport( + schema_version="choicebench.import-report.v1", + import_state="imported", + wrote_artifacts=not idempotent_noop[0], + idempotent_noop=idempotent_noop[0], + run_id=request.run_id, + experiment_id=manifest["experiment_id"], + counts=counts, + defects=DefectReport(total_findings=0, by_code={}, findings_digest=integrity_digest([])), + condition_digests=condition_digests, + realization_digests=realization_digests, + failures=(), + ) + + +def _write_overlay_staged_run( + staged_run: Path, manifest: Mapping[str, Any], realization: Mapping[str, Any], + result_rows: list[dict[str, Any]], + *, + overlay_source: SourceArtifactSpec, + overlay_evidence_record: Mapping[str, Any], + workspace_root: Path, +) -> None: + staged_run.mkdir(parents=True, exist_ok=True) + + opened_overlay_source = open_verified_source(overlay_source, containment_root=workspace_root) + written_record = write_evidence_blob( + staged_run, opened_overlay_source, references=overlay_evidence_record["references"] + ) + write_evidence_index(staged_run, [written_record]) + + atomic_write_json(staged_run / MANIFEST_FILENAME, manifest) + + realization_id = realization["realization_id"] + state = initial_run_state_v3(manifest) + state["realizations"][realization_id]["status"] = "completed" + + condition_id = realization["condition_id"] + condition_record = manifest["payload"]["semantic_conditions"][condition_id] + benchmark = condition_record["identity"]["benchmark"] + identity_columns = { + "condition_id": condition_id, + "realization_id": realization_id, + "experiment_id": manifest["experiment_id"], + "dataset_artifact_id": benchmark["artifact_id"], + "dataset_selection_id": benchmark["selection_id"], + "model_id": condition_record["identity"]["model_id"], + "method_id": condition_record["identity"]["method_id"], + "prompt_id": condition_record["identity"]["prompt_id"], + "benchmark_name": benchmark["name"], + "benchmark_split": benchmark["split"], + } + lineage_ids_by_qid = { + assignment["question_id"]: assignment["prediction_lineage_id"] + for assignment in realization["identity"]["realization"]["result_origin"]["row_assignments"] + } + final_rows = [ + {**identity_columns, **row, "prediction_lineage_id": lineage_ids_by_qid[row["question_id"]]} + for row in result_rows + ] + prepared_result = prepare_manifest_result(final_rows, manifest=manifest, realization_id=realization_id) + _, published_metadata, _ = publish_manifest_result(prepared_result, run_dir=staged_run) + state["realizations"][realization_id]["result_artifact_id"] = published_metadata["result_artifact_id"] + state["realizations"][realization_id]["result_artifact_digest"] = published_metadata["result_artifact_digest"] + state["realizations"][realization_id]["result_sha256"] = published_metadata["file_sha256"] + + evidence_status = realization["identity"]["realization"]["evidence"]["evidence_status"] + merged_validation = RealizationValidation( + evaluable_rows=tuple(result_rows), + findings=(), + validation_digest=integrity_digest({"evaluable_rows": canonicalize(result_rows)}), + computed_evidence_status=evidence_status, + evaluable=True, + qualifications=(), + limitations=(), + defect_question_ids=(), + ) + prepared_validation = prepare_realization_validation_artifact( + merged_validation, realization=realization, evidence_records=() + ) + write_realization_validation_artifact(staged_run, prepared_validation) + state["realizations"][realization_id]["validation_sha256"] = prepared_validation.file_sha256 + + atomic_write_json(staged_run / RUN_STATE_FILENAME, canonicalize(state, redact_secrets=False)) + validate_manifest_v3(manifest) + + def serialize_import_report(report: ImportReport) -> dict[str, Any]: return { "schema_version": report.schema_version, diff --git a/src/choicebench/importing/validation.py b/src/choicebench/importing/validation.py index be7df5a..a5eb7ff 100644 --- a/src/choicebench/importing/validation.py +++ b/src/choicebench/importing/validation.py @@ -263,18 +263,16 @@ class RealizationValidation: defect_question_ids: tuple[str, ...] -def normalize_realization_rows( - validated: ValidatedSourceRows, - *, - condition: ImportConditionSpec, - expected: ExpectedDataset, - mapping: Mapping[str, str] | None = None, - experiment_id: str | None = None, - realization_id: str | None = None, -) -> RealizationValidation: - """Recompute evidence status from exact coverage/defects and refuse a - declaration mismatch. Fill evaluative content only from ExpectedDataset; - source values were already checked, never trusted, in validate_source_rows. +def compute_evidence_status( + validated: ValidatedSourceRows, *, condition: ImportConditionSpec +) -> tuple[str, tuple[str, ...]]: + """Pure evidence-status computation from exact row coverage/defects, + independent of what the condition itself declares. Returns + (computed_status, defect_question_ids). Exposed separately from + normalize_realization_rows so a caller assembling declarations (e.g. a + paper-freeze profile) can discover the correct status to declare before + constructing the condition, instead of guessing and retrying against the + declaration-mismatch check below. """ expected_ids = set(validated.expected_question_ids) present_ids = set(validated.rows_by_question_id) @@ -301,6 +299,23 @@ def normalize_realization_rows( computed = "partial" else: computed = "qualified" if condition.qualifications else "complete" + return computed, defect_question_ids + + +def normalize_realization_rows( + validated: ValidatedSourceRows, + *, + condition: ImportConditionSpec, + expected: ExpectedDataset, + mapping: Mapping[str, str] | None = None, + experiment_id: str | None = None, + realization_id: str | None = None, +) -> RealizationValidation: + """Recompute evidence status from exact coverage/defects and refuse a + declaration mismatch. Fill evaluative content only from ExpectedDataset; + source values were already checked, never trusted, in validate_source_rows. + """ + computed, defect_question_ids = compute_evidence_status(validated, condition=condition) if computed != condition.evidence_status: raise ImportValidationError( diff --git a/tests/importing/test_engine.py b/tests/importing/test_engine.py index 2faee35..664e25d 100644 --- a/tests/importing/test_engine.py +++ b/tests/importing/test_engine.py @@ -1,7 +1,7 @@ """End-to-end tests for the reduced-scope import engine: base (non-overlay) -plan/dry-run/real-import/idempotence/verify. Overlay integration is not yet -wired into execute_import (see engine.py's module docstring); this file -covers the base import path the engine currently implements. +plan/dry-run/real-import/idempotence/verify. Overlay merge-into-new-run +orchestration (execute_overlay_import) is a separate function covered by +test_engine_overlay.py; this file covers only the base import path. """ from __future__ import annotations @@ -13,6 +13,7 @@ import yaml from choicebench.importing.engine import ( + ImportEngineError, ImportRequest, build_import_plan, execute_import, @@ -170,6 +171,33 @@ def test_build_import_plan_produces_a_valid_v3_manifest(tmp_path): assert len(result.evaluable_rows) == 2 +def test_build_import_plan_uses_a_supplied_expected_dataset_override(tmp_path): + spec = _spec(tmp_path) + (dataset,) = spec.datasets + (source,) = spec.sources + from choicebench.importing.dataset_reference import build_expected_dataset + from choicebench.importing.evidence import open_verified_source + + opened = open_verified_source(source, containment_root=tmp_path) + override = build_expected_dataset(dataset, {"results": opened}) + request = ImportRequest( + spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True, + expected_datasets={"dataset": override}, + ) + plan = build_import_plan(request) + assert plan.expected_datasets["dataset"] is override + + +def test_build_import_plan_rejects_an_incomplete_expected_dataset_override(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest( + spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True, + expected_datasets={}, + ) + with pytest.raises(ImportEngineError, match="missing dataset"): + build_import_plan(request) + + def test_dry_run_writes_nothing(tmp_path): spec = _spec(tmp_path) request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) diff --git a/tests/importing/test_engine_overlay.py b/tests/importing/test_engine_overlay.py new file mode 100644 index 0000000..67ded84 --- /dev/null +++ b/tests/importing/test_engine_overlay.py @@ -0,0 +1,387 @@ +"""Tests for the overlay merge-into-new-run orchestration +(execute_overlay_import / _merge_overlay_realization). Builds a real base run +via test_engine.py's fixtures, then applies an authorized overlay on top. +""" + +from __future__ import annotations + +from hashlib import sha256 +from pathlib import Path + +import pytest + +from choicebench.identity import short_id +from choicebench.importing.dataset_reference import build_expected_dataset +from choicebench.importing.engine import ( + ImportEngineError, + ImportRequest, + OverlayImportRequest, + execute_import, + execute_overlay_import, + verify_import_run, +) +from choicebench.importing.evidence import open_verified_source +from choicebench.importing.schema import ( + AuthorizationSpec, + CsvDialectSpec, + OptionMappingSpec, + OverlaySpec, + ResultOriginSpec, + SourceArtifactSpec, +) +from tests.importing.test_engine import _spec + +_AUTH_CSV = b"authorization,payload\n1,2\n" +_OVERLAY_CSV = "qid,predicted_letter\nq1,B\n" + + +def _make_base_run(tmp_path: Path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="base-run", workspace_root=tmp_path, strict=True) + execute_import(request, dry_run=False) + base_run_dir = tmp_path / "runs" / "base-run" + verified = verify_import_run(base_run_dir) + (base,) = verified.realizations.values() + return base_run_dir, base + + +def _base_expected_dataset(tmp_path: Path): + spec = _spec(tmp_path) + (dataset_ref,) = spec.datasets + (results_source,) = spec.sources + opened_results = open_verified_source(results_source, containment_root=tmp_path) + return build_expected_dataset(dataset_ref, opened_sources={"results": opened_results}) + + +def _authorization_source(tmp_path: Path) -> SourceArtifactSpec: + auth_path = tmp_path / "authorization.csv" + auth_path.write_bytes(_AUTH_CSV) + return SourceArtifactSpec( + source_id="auth-source", + path=auth_path, + logical_path="inputs/authorization.csv", + expected_sha256=sha256(_AUTH_CSV).hexdigest(), + format="csv", + format_version="producer-v1", + classification="raw", + dialect=CsvDialectSpec(), + columns={}, + expected_columns=("authorization", "payload"), + ignored_columns={}, + null_values=(), + numeric_columns=(), + option_mapping=OptionMappingSpec( + mode="ordered_columns", ordered_columns=(), structured_column=None, + structured_label_key=None, structured_text_key=None, + ), + extra_field_policy="preserve_unmapped", + preserve_namespace="producer", + source_run_id=None, source_repository=None, source_commit=None, + notes={}, + ) + + +def _overlay_source(tmp_path: Path, csv_text: str = _OVERLAY_CSV, filename: str = "overlay.csv") -> SourceArtifactSpec: + overlay_path = tmp_path / filename + overlay_path.write_bytes(csv_text.encode("utf-8")) + return SourceArtifactSpec( + source_id="overlay-source", + path=overlay_path, + logical_path="inputs/overlay.csv", + expected_sha256=sha256(csv_text.encode("utf-8")).hexdigest(), + format="csv", + format_version="producer-v1", + classification="repaired", + dialect=CsvDialectSpec(), + columns={"question_id": "qid", "prediction": "predicted_letter"}, + expected_columns=("qid", "predicted_letter"), + ignored_columns={}, + null_values=(), + numeric_columns=(), + option_mapping=OptionMappingSpec( + mode="ordered_columns", ordered_columns=(), structured_column=None, + structured_label_key=None, structured_text_key=None, + ), + extra_field_policy="preserve_unmapped", + preserve_namespace="producer", + source_run_id=None, source_repository=None, source_commit=None, + notes={}, + ) + + +def _authorization( + *, base, authorization_type: str, executable: bool, + grants: dict[str, str] | None = None, +) -> AuthorizationSpec: + grants = grants or {"q1": "malformed prediction"} + payload = { + "schema_version": "choicebench.authorization-bundle.v1", + "authorization_type": authorization_type, + "grants": {base.condition_digest: grants}, + "authority": "principal-investigator", + "purpose": "repair", + "executable": executable, + "source_sha256": sha256(_AUTH_CSV).hexdigest(), + "input_evidence_digests": dict(base.evidence_source_digests), + "expected_snapshot_digests": {base.condition_digest: "2" * 64}, + } + authorization_id = short_id("auth", payload) + return AuthorizationSpec( + authorization_id=authorization_id, + authorization_type=authorization_type, + source_id="auth-source", + condition_question_reasons={base.condition_digest: grants}, + authority="principal-investigator", + purpose="repair", + executable=executable, + input_evidence_digests=dict(base.evidence_source_digests), + expected_snapshot_digests={base.condition_digest: "2" * 64}, + ) + + +def _overlay_spec(*, base, authorization: AuthorizationSpec) -> OverlaySpec: + return OverlaySpec( + overlay_id="ovl_1", + base_run_path=Path("/tmp/base-run"), + base_condition_digest=base.condition_digest, + base_realization_id=base.realization_id, + base_realization_digest=base.realization_digest, + base_evidence_digests=dict(base.evidence_source_digests), + base_validation_artifact_sha256=base.validation_artifact_sha256, + base_result_sha256=base.result_sha256, + source_id="overlay-source", + authorization_id=authorization.authorization_id, + replacement_reasons={"q1": "repair malformed answer"}, + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", + default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + lineage_notes={}, + implementation={ + "qualified_name": "external:repair_tool", + "source_digest": sha256(_OVERLAY_CSV.encode("utf-8")).hexdigest(), + }, + input_digest="9" * 64, + preownership_output_digest="a" * 64, + expected_evidence_status="complete", + ) + + +def _make_overlay_request(tmp_path: Path, base_run_dir: Path, base, *, run_id: str = "overlay-run"): + auth_source = _authorization_source(tmp_path) + authorization = _authorization(base=base, authorization_type="inference_repair", executable=True) + overlay_source = _overlay_source(tmp_path) + overlay = _overlay_spec(base=base, authorization=authorization) + expected_dataset = _base_expected_dataset(tmp_path) + + return OverlayImportRequest( + base_run_dir=base_run_dir, + base_realization_id=base.realization_id, + authorization=authorization, + authorization_source=auth_source, + overlay=overlay, + overlay_source=overlay_source, + overlay_mapping={"question_id": "qid", "prediction": "predicted_letter"}, + condition_digests={"cond_key": base.condition_digest}, + expected_dataset=expected_dataset, + run_id=run_id, + workspace_root=tmp_path, + ) + + +def test_execute_overlay_import_merges_repair_into_a_new_run(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + request = _make_overlay_request(tmp_path, base_run_dir, base) + + report = execute_overlay_import(request, dry_run=False) + assert report.import_state == "imported" + assert report.wrote_artifacts is True + assert report.idempotent_noop is False + + overlay_run_dir = tmp_path / "runs" / "overlay-run" + verified = verify_import_run(overlay_run_dir) + (derived,) = verified.realizations.values() + assert set(derived.rows_by_question_id) == {"q1", "q2"} + assert derived.rows_by_question_id["q1"]["predicted_option"] == "B" + assert derived.prediction_origins["q1"] == "native_inference" + assert derived.prediction_origins["q2"] == "external_historical_inference" + assert (overlay_run_dir / "artifacts/imports/evidence/index.json").is_file() + + # the immutable base run is untouched + reverified_base = verify_import_run(base_run_dir) + (base_again,) = reverified_base.realizations.values() + assert base_again.realization_digest == base.realization_digest + + +def test_execute_overlay_import_merges_offline_transformation_retaining_origin(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + + auth_source = _authorization_source(tmp_path) + authorization = _authorization( + base=base, authorization_type="offline_transformation", executable=False, + grants={"q2": "offline semantic rematch"}, + ) + overlay_csv = "qid,predicted_letter\nq2,A\n" + overlay_source = _overlay_source(tmp_path, csv_text=overlay_csv, filename="offline-overlay.csv") + overlay = OverlaySpec( + overlay_id="ovl_offline", + base_run_path=Path("/tmp/base-run"), + base_condition_digest=base.condition_digest, + base_realization_id=base.realization_id, + base_realization_digest=base.realization_digest, + base_evidence_digests=dict(base.evidence_source_digests), + base_validation_artifact_sha256=base.validation_artifact_sha256, + base_result_sha256=base.result_sha256, + source_id="overlay-source", + authorization_id=authorization.authorization_id, + replacement_reasons={"q2": "offline semantic rematch"}, + result_origin=ResultOriginSpec( + derivation_origin="offline_transformation", + default_prediction_origin=None, + per_question_prediction_origins={}, + ), + lineage_notes={}, + implementation={ + "qualified_name": "external:offline_rematcher", + "source_digest": sha256(overlay_csv.encode("utf-8")).hexdigest(), + }, + input_digest="9" * 64, + preownership_output_digest="a" * 64, + expected_evidence_status="complete", + ) + request = OverlayImportRequest( + base_run_dir=base_run_dir, + base_realization_id=base.realization_id, + authorization=authorization, + authorization_source=auth_source, + overlay=overlay, + overlay_source=overlay_source, + overlay_mapping={"question_id": "qid", "prediction": "predicted_letter"}, + condition_digests={"cond_key": base.condition_digest}, + expected_dataset=_base_expected_dataset(tmp_path), + run_id="offline-overlay-run", + workspace_root=tmp_path, + ) + + report = execute_overlay_import(request, dry_run=False) + assert report.import_state == "imported" + + overlay_run_dir = tmp_path / "runs" / "offline-overlay-run" + verified = verify_import_run(overlay_run_dir) + (derived,) = verified.realizations.values() + assert derived.rows_by_question_id["q2"]["predicted_option"] == "A" + # offline_transformation retains the underlying response's own origin + # rather than reassigning it (base.prediction_origins["q2"] is + # external_historical_inference) -- unlike inference_repair. + assert derived.prediction_origins["q2"] == "external_historical_inference" + assert derived.prediction_origins["q1"] == "external_historical_inference" + + +def test_execute_overlay_import_dry_run_writes_nothing(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + request = _make_overlay_request(tmp_path, base_run_dir, base) + + report = execute_overlay_import(request, dry_run=True) + assert report.import_state == "validated" + assert report.wrote_artifacts is False + assert not (tmp_path / "runs" / "overlay-run").exists() + + +def test_execute_overlay_import_is_idempotent_noop_on_repeat(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + request = _make_overlay_request(tmp_path, base_run_dir, base) + + execute_overlay_import(request, dry_run=False) + second = execute_overlay_import(request, dry_run=False) + assert second.idempotent_noop is True + assert second.wrote_artifacts is False + + +def test_execute_overlay_import_rejects_an_unknown_base_realization_id(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + request = _make_overlay_request(tmp_path, base_run_dir, base) + unknown = OverlayImportRequest( + base_run_dir=request.base_run_dir, + base_realization_id="not-the-real-id", + authorization=request.authorization, + authorization_source=request.authorization_source, + overlay=request.overlay, + overlay_source=request.overlay_source, + overlay_mapping=request.overlay_mapping, + condition_digests=request.condition_digests, + expected_dataset=request.expected_dataset, + run_id=request.run_id, + workspace_root=request.workspace_root, + ) + report = execute_overlay_import(unknown, dry_run=False) + assert report.import_state == "failed" + assert report.wrote_artifacts is False + assert not (tmp_path / "runs" / "overlay-run").exists() + + +def test_execute_overlay_import_refuses_to_overwrite_a_divergent_existing_run(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + request = _make_overlay_request(tmp_path, base_run_dir, base) + execute_overlay_import(request, dry_run=False) + + # a second, different overlay (different replaced answer) targeting the + # same run_id must be refused, not silently treated as a no-op + divergent_csv = "qid,predicted_letter\nq1,C\n" + divergent_overlay_source = _overlay_source(tmp_path, csv_text=divergent_csv, filename="divergent-overlay.csv") + divergent_overlay = OverlaySpec( + overlay_id="ovl_2", + base_run_path=request.overlay.base_run_path, + base_condition_digest=request.overlay.base_condition_digest, + base_realization_id=request.overlay.base_realization_id, + base_realization_digest=request.overlay.base_realization_digest, + base_evidence_digests=request.overlay.base_evidence_digests, + base_validation_artifact_sha256=request.overlay.base_validation_artifact_sha256, + base_result_sha256=request.overlay.base_result_sha256, + source_id="overlay-source", + authorization_id=request.overlay.authorization_id, + replacement_reasons={"q1": "repair malformed answer"}, + result_origin=request.overlay.result_origin, + lineage_notes={}, + implementation={ + "qualified_name": "external:repair_tool", + "source_digest": sha256(divergent_csv.encode("utf-8")).hexdigest(), + }, + input_digest="9" * 64, + preownership_output_digest="a" * 64, + expected_evidence_status="complete", + ) + divergent_request = OverlayImportRequest( + base_run_dir=request.base_run_dir, + base_realization_id=request.base_realization_id, + authorization=request.authorization, + authorization_source=request.authorization_source, + overlay=divergent_overlay, + overlay_source=divergent_overlay_source, + overlay_mapping=request.overlay_mapping, + condition_digests=request.condition_digests, + expected_dataset=request.expected_dataset, + run_id=request.run_id, + workspace_root=request.workspace_root, + ) + # matches execute_import's existing base-path contract (engine.py's + # verify_import_run(expected=...) check): a divergent overwrite is a + # loud failure that propagates out of the transaction, not a returned + # "failed" report -- publish() has already fully released its lock/ + # staging state by the time this raises. + with pytest.raises(ImportEngineError, match="different experiment"): + execute_overlay_import(divergent_request, dry_run=False) + + +def test_execute_overlay_import_refuses_a_forged_authorization_source(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + request = _make_overlay_request(tmp_path, base_run_dir, base) + request.authorization_source.path.write_bytes(b"tampered,bytes\n9,9\n") + + # the forged bytes no longer match expected_sha256; EvidenceError (a + # ValueError) is caught by execute_overlay_import's own error handling + # and surfaces as a failed report rather than a silent success. + report = execute_overlay_import(request, dry_run=False) + assert report.import_state == "failed" + assert report.wrote_artifacts is False + assert not (tmp_path / "runs" / "overlay-run").exists() From 3f6e455af35dacded810b297800546bd207a9d4d Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 20:39:07 +0300 Subject: [PATCH 39/47] feat: full Stage 1 ImportSpec/AuthorizationSpec assembly - profiles/stage1_paper_freeze.py: build_stage1_import_spec assembles all 102 real cells into one ImportSpec (104 sources: 102 per-cell canonical CSVs + 2 documentation-only ARC/MMLU normalized-reference sources bound to ImportRequest.expected_datasets, never opened by build_import_plan). Each condition's declared evidence_status is reconciled via an actual probe (validate_source_rows + compute_evidence_status) rather than trusted blindly from the freeze's own STATUS_MAP-mapped status. build_stage1_cell_authorization builds one AuthorizationSpec scoped to exactly one cell (a multi-cell bundle can never satisfy derive_overlay's evidence-digest cross-check for any single cell's overlay -- verified by reading overlays.py). VERIFIED DIRECTLY against the real freeze (never committed): assembly runs cleanly end-to-end and reconciles to complete=49/malformed=45/ qualified=8 (vs. the freeze's own 57/28/5/4/6/2 STATUS_MAP breakdown). Confirmed by inspecting finding codes that the entire divergence is MISSING_PREDICTION (zero CORRECT_OPTION_MISMATCH/INVALID_PREDICTION/ UNEXPECTED_QUESTION_ID anywhere, i.e. no column-mapping bug): 20 of 28 "malformed_requires_inference" cells have zero row-level defects (their defect, e.g. a phantom/NaN rendered option D, is parse_ok/valid-answer and invisible to generic validation) and reconcile to "complete"; 28 of 57 "canonical_complete" cells have MISSING_PREDICTION rows the freeze's own classification tolerates (one MMLU cell has 152/1000 parse_missing rows, its own score_status="score_unscorable", yet is declared "canonical_complete") that ChoiceBench's fail-closed generic validator does not -- a genuine, intentional strictness difference, not a bug. All 102 conditions' build_import_semantic_identity calls succeed with 102 unique condition_digests. - tests/importing/test_stage1_import_spec.py (new): synthetic-fixture coverage for assembly shape, the recoverable-cell and malformed-cell reconciliations, excluded-cell scope_disposition, the per-cell checksum-ledger cross-check, and a full build_stage1_cell_authorization -> validate_authorization_bundle round trip. Also fixes a real bug this surfaced: build_stage1_import_spec's ImportConditionSpec left unknown_reasons empty despite declaring calibration_identity/preflight_identity as None, violating identity.py's unknown-reason contract. 520 focused / 520 targeted tests passed (tests/importing/). Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- .../importing/profiles/stage1_paper_freeze.py | 442 +++++++++++++++++- tests/importing/test_stage1_import_spec.py | 329 +++++++++++++ 2 files changed, 768 insertions(+), 3 deletions(-) create mode 100644 tests/importing/test_stage1_import_spec.py diff --git a/src/choicebench/importing/profiles/stage1_paper_freeze.py b/src/choicebench/importing/profiles/stage1_paper_freeze.py index 7dfcb07..8f97345 100644 --- a/src/choicebench/importing/profiles/stage1_paper_freeze.py +++ b/src/choicebench/importing/profiles/stage1_paper_freeze.py @@ -28,7 +28,7 @@ from __future__ import annotations import ast -from dataclasses import dataclass +from dataclasses import dataclass, replace from hashlib import sha256 import csv import io @@ -36,9 +36,24 @@ from pathlib import Path from typing import Any, Callable, Mapping -from choicebench.importing.csv_adapter import OpenedSource +from choicebench.identity import short_id +from choicebench.importing.csv_adapter import OpenedSource, parse_csv_source from choicebench.importing.dataset_reference import ExpectedDataset, build_expected_dataset -from choicebench.importing.schema import DatasetReferenceSpec +from choicebench.importing.evidence import open_verified_source +from choicebench.importing.validation import compute_evidence_status, validate_source_rows +from choicebench.importing.schema import ( + AuthorizationSpec, + CsvDialectSpec, + DatasetReferenceSpec, + ImportConditionSpec, + ImportMethodSpec, + ImportModelSpec, + ImportPromptSpec, + ImportSpec, + OptionMappingSpec, + ResultOriginSpec, + SourceArtifactSpec, +) STATUS_MAP = { "canonical_complete": "complete", @@ -546,3 +561,424 @@ def build_stage1_expected_datasets(freeze_root: Path) -> dict[str, ExpectedDatas freeze_root, ledger, benchmark_name="mmlu", raw_recompute=_recompute_mmlu_fields, ), } + + +# --- Full ImportSpec assembly (Task 16/Unit K remainder) -------------------- +# +# Reduced scope, verified directly against the real freeze: all 102 cells +# have observed_unique_count==1000 with zero missing/unexpected/duplicate +# question IDs, so compute_evidence_status's "partial"/"recoverable" +# branches (which need missing_ids) never trigger here -- every cell +# reconciles to either "malformed" or "complete"/"qualified". The freeze's +# own STATUS_MAP-mapped status is a scientific/methodology classification, +# not always what ChoiceBench's generic row validator independently +# computes: the 6 "recoverable" cells and the "mmlu independent_hypothesis" +# "incomplete" cell have ZERO row-level defects (their defect -- option- +# hiding, or a protocol-invalid option-A score after MAX_TOKENS termination +# -- is invisible to generic validation, documented only via +# damaged_question_ids/qualification narrative). So every condition's +# declared evidence_status is reconciled via an actual probe +# (validate_source_rows + compute_evidence_status), never trusted blindly +# from STATUS_MAP. qualifications=({"reason": ...},) is pre-populated +# whenever record["qualification"] is non-empty, so a cell with a +# documented caveat but zero generic defects reconciles to "qualified" +# rather than bare "complete". +# +# question_id/correct_option/prediction are the only source->identity +# column mappings (question_text is deliberately NOT mapped), avoiding +# noise from formatting differences between each run's own recorded +# question_text and Unit J's normalized ExpectedDataset text -- +# correct_option is the meaningful invariant (a categorical answer key, +# not prone to whitespace/formatting variance). +# +# Real queue reconciliation (verified directly against the freeze): 102 +# cells = 62 needing no repair (57 complete + 5 qualified) + 6 recoverable +# (offline authority, 18 pairs, all arc_challenge/semantic_matching_v1) + +# 32 malformed/incomplete, of which 31 cells are fully approved (138 pairs +# total) and exactly 1 cell +# (cbp__gemini-2-5-flash__arc_challenge__independent_hypothesis) is fully +# HELD (687 questions, supervisor-stopped) + 2 excluded_from_paper_matrix +# cells (both "pride"/qwen-2.5-7b-instruct-turbo), one of which also +# carries 3 explicitly-excluded damaged questions in +# paper_scope_excluded_reruns.csv. Held and excluded pairs must NOT +# receive any authorization/overlay. + +_STAGE1_COLUMNS = { + "question_id": "question_id", + "correct_option": "correct_option", + "prediction": "parsed_choice", +} +_STAGE1_OPTION_COLUMNS = ("choice_a", "choice_b", "choice_c", "choice_d") +_STAGE1_PROMPT_KEY = "stage1_freeze_v1" + + +def _stage1_prompt() -> ImportPromptSpec: + return ImportPromptSpec( + prompt_key=_STAGE1_PROMPT_KEY, + template_identity=None, + template_digest=None, + template_contents=None, + unknown_reason=( + "Stage 1 freeze: each canonical CSV row preserves its own literal rendered " + "'prompt' text, but no single stable prompt-template identity/digest was " + "recorded across methods at the paper's authoring time; method_key already " + "captures the distinct prompting strategy." + ), + native_compatibility_identity=None, + ) + + +def _stage1_model(model_name: str, provider_backend: str) -> ImportModelSpec: + return ImportModelSpec( + model_key=model_name, + display_name=model_name, + backend=provider_backend, + provider=provider_backend, + revision=None, + effective_parameters={}, + unknown_reasons={"revision": "not recorded by paper_data_freeze canonical manifest"}, + native_compatibility_identity=None, + ) + + +def _stage1_method(method_name: str) -> ImportMethodSpec: + return ImportMethodSpec( + method_key=method_name, + name=method_name, + effective_parameters={}, + implementation=None, + unknown_reasons={"implementation": "not recorded by paper_data_freeze canonical manifest"}, + native_compatibility_identity=None, + ) + + +def _stage1_cell_source( + freeze_root: Path, ledger: Mapping[str, str], cell_id: str, record: Mapping[str, Any] +) -> SourceArtifactSpec: + relative_path = record["canonical_path"] + ledger_digest = ledger.get(relative_path) + if ledger_digest is None: + raise Stage1ProfileError(f"{relative_path} has no checksum ledger entry.") + if ledger_digest != record["canonical_sha256"]: + raise Stage1ProfileError( + f"Cell {cell_id!r} canonical_sha256 disagrees with the checksum ledger for " + f"{relative_path}." + ) + path = freeze_root / relative_path + with path.open(newline="", encoding="utf-8") as handle: + header = tuple(next(csv.reader(handle))) + required = set(_STAGE1_COLUMNS.values()) | set(_STAGE1_OPTION_COLUMNS) + missing = required - set(header) + if missing: + raise Stage1ProfileError( + f"Cell {cell_id!r} canonical CSV is missing column(s) {sorted(missing)}." + ) + return SourceArtifactSpec( + source_id=cell_id, + path=path, + logical_path=relative_path, + expected_sha256=record["canonical_sha256"], + format="csv", + format_version="paper_data_freeze.v1", + classification="raw", + dialect=CsvDialectSpec(), + columns=dict(_STAGE1_COLUMNS), + expected_columns=header, + ignored_columns={}, + null_values=("",), + numeric_columns=(), + option_mapping=OptionMappingSpec( + mode="ordered_columns", + ordered_columns=_STAGE1_OPTION_COLUMNS, + structured_column=None, + structured_label_key=None, + structured_text_key=None, + ), + extra_field_policy="preserve_unmapped", + preserve_namespace="stage1_paper_freeze", + source_run_id=None, + source_repository=None, + source_commit=None, + notes={ + "cell_id": cell_id, + "declared_status": record["status"], + "final_status": record["final_status"], + "qualification": record.get("qualification") or "", + }, + ) + + +def _stage1_normalized_dataset_source( + freeze_root: Path, ledger: Mapping[str, str], *, benchmark_name: str, source_id: str +) -> SourceArtifactSpec: + """Documentation-only source declaration for the ARC/MMLU normalized + reference file: real path/checksum, but never opened by the generic + engine, since ImportRequest.expected_datasets overrides the per-source + ExpectedDataset rebuild with build_stage1_expected_datasets' own + revalidation/dedup chain (needed for MMLU's real duplicate question_ids, + which the generic rebuild would reject).""" + relative_path = _DATASET_RELATIVE_PATHS[benchmark_name]["normalized"] + expected_sha256 = ledger.get(relative_path) + if expected_sha256 is None: + raise Stage1ProfileError(f"{relative_path} has no checksum ledger entry.") + return SourceArtifactSpec( + source_id=source_id, + path=freeze_root / relative_path, + logical_path=relative_path, + expected_sha256=expected_sha256, + format="csv", + format_version="paper_data_freeze.v1", + classification="canonical", + dialect=CsvDialectSpec(), + columns={}, + expected_columns=(), + ignored_columns={}, + null_values=(), + numeric_columns=(), + option_mapping=OptionMappingSpec( + mode="ordered_columns", ordered_columns=(), structured_column=None, + structured_label_key=None, structured_text_key=None, + ), + extra_field_policy="preserve_unmapped", + preserve_namespace="stage1_paper_freeze", + source_run_id=None, source_repository=None, source_commit=None, + notes={"role": f"{benchmark_name}_expected_dataset_reference"}, + ) + + +def _reconcile_evidence_status( + condition: ImportConditionSpec, + *, + source: SourceArtifactSpec, + freeze_root: Path, + dataset: ExpectedDataset, +) -> ImportConditionSpec: + opened = open_verified_source(source, containment_root=freeze_root) + table = parse_csv_source(opened, source, strict=True) + validated = validate_source_rows( + table, source_id=source.source_id, mapping=source.columns, condition=condition, expected=dataset + ) + computed, _ = compute_evidence_status(validated, condition=condition) + if computed != condition.evidence_status: + condition = replace(condition, evidence_status=computed) + return condition + + +def build_stage1_import_spec( + freeze_root: Path, + translation: Stage1Translation, + expected_datasets: Mapping[str, ExpectedDataset], +) -> ImportSpec: + freeze_root = Path(freeze_root) + ledger = _load_checksum_ledger(freeze_root) + + dataset_source_ids = { + "arc_challenge": "arc_challenge_normalized_reference", + "mmlu": "mmlu_normalized_reference", + } + sources: list[SourceArtifactSpec] = [ + _stage1_normalized_dataset_source( + freeze_root, ledger, benchmark_name=benchmark_name, source_id=source_id + ) + for benchmark_name, source_id in dataset_source_ids.items() + ] + + conditions: list[ImportConditionSpec] = [] + models_by_key: dict[str, ImportModelSpec] = {} + methods_by_key: dict[str, ImportMethodSpec] = {} + prompt = _stage1_prompt() + + for cell_id in sorted(translation.cell_records): + record = translation.cell_records[cell_id] + source = _stage1_cell_source(freeze_root, ledger, cell_id, record) + sources.append(source) + + model_name = record["model"] + if model_name not in models_by_key: + models_by_key[model_name] = _stage1_model(model_name, record["provider_backend"]) + method_name = record["method"] + if method_name not in methods_by_key: + methods_by_key[method_name] = _stage1_method(method_name) + + benchmark = record["benchmark"] + dataset = expected_datasets[benchmark] + scope_disposition = ( + "excluded_from_paper_matrix" if record["final_status"] == _EXCLUDED_STATUS else "included" + ) + qualification_text = record.get("qualification") or "" + qualifications = ({"reason": qualification_text},) if qualification_text else () + + condition = ImportConditionSpec( + condition_key=cell_id, + source_ids=(cell_id,), + dataset_id=benchmark, + model_key=model_name, + method_key=method_name, + prompt_key=prompt.prompt_key, + seed=42, + calibration_identity=None, + preflight_identity=None, + protocol_settings={}, + generation_parameters={"temperature": 0.0}, + unknown_reasons={ + "calibration_identity": "not recorded by paper_data_freeze canonical manifest", + "preflight_identity": "not recorded by paper_data_freeze canonical manifest", + }, + expected_question_ids=dataset.selected_question_ids, + evidence_status=STATUS_MAP.get(record["status"], "complete"), + scope_disposition=scope_disposition, + executable=None, + qualifications=qualifications, + limitations=(), + damaged_question_ids=tuple(record["damaged_question_ids"]), + recoverable_question_ids=tuple(record["recoverable_question_ids"]), + result_origin=ResultOriginSpec( + derivation_origin="external_import", + default_prediction_origin="external_historical_inference", + per_question_prediction_origins={}, + ), + ) + condition = _reconcile_evidence_status( + condition, source=source, freeze_root=freeze_root, dataset=dataset + ) + conditions.append(condition) + + datasets = tuple( + DatasetReferenceSpec( + dataset_id=benchmark_name, + benchmark_name=benchmark_name, + split="robustness", + reference_kind="independent_input_snapshot", + trust_label="checksum_verified_freeze_internal", + source_ids=(dataset_source_ids[benchmark_name],), + selection_source_id=dataset_source_ids[benchmark_name], + expected_question_ids=expected_datasets[benchmark_name].selected_question_ids, + selection_seed=None, + selection_n_samples=None, + subject_filter=(), + selection_unknown_reasons={}, + columns={}, + revision=None, + fingerprint=None, + derivation={}, + limitations=(), + native_compatibility_identity=None, + ) + for benchmark_name in ("arc_challenge", "mmlu") + ) + # This DatasetReferenceSpec is documentation-only -- the real + # ExpectedDataset comes from ImportRequest.expected_datasets, built by + # build_stage1_expected_datasets' own revalidation/dedup chain, which + # build_import_plan's generic per-source rebuild cannot reproduce for + # MMLU's real duplicate question_ids (build_import_plan skips its own + # per-declaration rebuild entirely whenever expected_datasets is + # supplied, so this declaration's source_ids are never actually opened). + + return ImportSpec( + schema_version="choicebench.import-spec.v1", + import_name="Stage 1 paper data freeze", + sources=tuple(sources), + datasets=datasets, + models=tuple(models_by_key.values()), + methods=tuple(methods_by_key.values()), + prompts=(prompt,), + conditions=tuple(conditions), + authorizations=(), + overlays=(), + metrics=["accuracy"], + provenance={"producer_request_id": {"value": None, "reason": "paper_data_freeze import"}}, + audit={"source_location": str(freeze_root)}, + ) + + +# --- Per-cell authorization bundle ------------------------------------- +# +# IMPORTANT, verified by reading derive_overlay's own cross-check +# (overlays.py): a ValidatedAuthorization's input_evidence_digests is the +# WHOLE bundle's input_evidence_digests, unsliced by +# authorization_for_condition -- and derive_overlay requires +# overlay.base_evidence_digests to match it key-for-key AND every one of +# those keys to also appear in the base realization's OWN +# evidence_source_digests. Since every Stage 1 cell has exactly one +# source, a bundle whose input_evidence_digests spans more than one cell +# can never satisfy that check for any single cell's overlay. So each +# AuthorizationSpec here is deliberately scoped to exactly ONE cell/ +# condition (one grants key, one evidence-digest key) -- NOT one bundle +# for all 31 approved cells or all 6 recoverable cells at once. + +def build_stage1_cell_authorization( + freeze_root: Path, + ledger: Mapping[str, str], + *, + cell_id: str, + condition_digest: str, + authorization_type: str, + question_reasons: Mapping[str, str], + dataset_snapshot_digest: str, + cell_evidence_sha256: str, +) -> tuple[AuthorizationSpec, SourceArtifactSpec]: + if authorization_type == "inference_repair": + relative_path = "manifests/approved_rerun_queue.csv" + executable = True + elif authorization_type == "offline_transformation": + relative_path = "manifests/canonical_results_manifest.csv" + executable = False + else: + raise Stage1ProfileError(f"Unsupported authorization_type {authorization_type!r}.") + + expected_sha256 = ledger.get(relative_path) + if expected_sha256 is None: + raise Stage1ProfileError(f"{relative_path} has no checksum ledger entry.") + source_id = f"{cell_id}__{authorization_type}_authorization_source" + source = SourceArtifactSpec( + source_id=source_id, + path=freeze_root / relative_path, + logical_path=relative_path, + expected_sha256=expected_sha256, + format="csv", + format_version="paper_data_freeze.v1", + classification="canonical", + dialect=CsvDialectSpec(), + columns={}, + expected_columns=(), + ignored_columns={}, + null_values=(), + numeric_columns=(), + option_mapping=OptionMappingSpec( + mode="ordered_columns", ordered_columns=(), structured_column=None, + structured_label_key=None, structured_text_key=None, + ), + extra_field_policy="preserve_unmapped", + preserve_namespace="stage1_paper_freeze", + source_run_id=None, source_repository=None, source_commit=None, + notes={"role": f"{authorization_type}_authorization_evidence", "cell_id": cell_id}, + ) + + purpose = ( + "Stage 1 approved rerun repair" if authorization_type == "inference_repair" + else "Stage 1 offline-recoverable semantic rematch" + ) + payload = { + "schema_version": "choicebench.authorization-bundle.v1", + "authorization_type": authorization_type, + "grants": {condition_digest: dict(question_reasons)}, + "authority": "paper_data_freeze.manifests", + "purpose": purpose, + "executable": executable, + "source_sha256": expected_sha256, + "input_evidence_digests": {cell_id: cell_evidence_sha256}, + "expected_snapshot_digests": {condition_digest: dataset_snapshot_digest}, + } + authorization_id = short_id("auth", payload) + authorization = AuthorizationSpec( + authorization_id=authorization_id, + authorization_type=authorization_type, + source_id=source_id, + condition_question_reasons={condition_digest: dict(question_reasons)}, + authority="paper_data_freeze.manifests", + purpose=purpose, + executable=executable, + input_evidence_digests={cell_id: cell_evidence_sha256}, + expected_snapshot_digests={condition_digest: dataset_snapshot_digest}, + ) + return authorization, source diff --git a/tests/importing/test_stage1_import_spec.py b/tests/importing/test_stage1_import_spec.py new file mode 100644 index 0000000..ace7510 --- /dev/null +++ b/tests/importing/test_stage1_import_spec.py @@ -0,0 +1,329 @@ +"""Tests for full Stage 1 ImportSpec/AuthorizationSpec assembly (Task 18). + +Uses a small synthetic freeze fixture combining test_stage1_profile.py's +cell-manifest/queue fixtures with its ARC/MMLU dataset fixtures, plus real +per-cell canonical result CSVs -- never real historical data. Column names, +the 6-column common schema (question_id/correct_option/parsed_choice/ +choice_a..d), and the evidence-status reconciliation behavior below were all +verified directly against the real, read-only Stage 1 freeze at +/home/cotenthusiast/Projects/model-generalization/paper_data_freeze during +development (never committed): of 102 real cells, every one reconciles to +either "malformed" (any MISSING_PREDICTION/INVALID_PREDICTION/ +CORRECT_OPTION_MISMATCH finding) or "complete"/"qualified" -- the freeze's +own STATUS_MAP-mapped status disagrees with the independently-computed +status for a majority of cells (e.g. 20 of 28 "malformed_requires_inference" +cells have zero row-level defects and reconcile to "complete"; 28 of 57 +"canonical_complete" cells have some MISSING_PREDICTION rows the freeze +tolerates but ChoiceBench's generic validator does not, and reconcile to +"malformed"), confirming the reconciliation-by-probe design is load-bearing, +not a formality. +""" + +from __future__ import annotations + +import csv +from hashlib import sha256 +import json +from pathlib import Path + +import pytest + +from choicebench.importing.authorization import validate_authorization_bundle +from choicebench.importing.evidence import open_verified_source +from choicebench.importing.profiles.stage1_paper_freeze import ( + Stage1ProfileError, + build_stage1_cell_authorization, + build_stage1_expected_datasets, + build_stage1_import_spec, + translate_stage1_paper_freeze, +) +from tests.importing.test_stage1_profile import ( + _build_dataset_freeze, + _cell, + _default_arc_normalized_rows, + _default_mmlu_normalized_rows, + _write, +) + +_STAGE1_RESULT_COLUMNS = [ + "question_id", "correct_option", "choice_a", "choice_b", "choice_c", "choice_d", + "parsed_choice", +] + + +def _result_row(question_id, correct_option, choices, parsed_choice): + row = {"question_id": question_id, "correct_option": correct_option, "parsed_choice": parsed_choice} + for letter, text in zip("abcd", choices): + row[f"choice_{letter}"] = text + return row + + +def _write_result_csv(path: Path, rows: list[dict]) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + with path.open("w", newline="", encoding="utf-8") as handle: + writer = csv.DictWriter(handle, fieldnames=_STAGE1_RESULT_COLUMNS) + writer.writeheader() + writer.writerows(rows) + + +def _arc_result_rows(*, all_correct: bool = True): + normalized = _default_arc_normalized_rows() + rows = [] + for row in normalized: + choices = [row["choice_a"], row["choice_b"], row["choice_c"], row["choice_d"]] + parsed = row["correct_option"] if all_correct else "A" + rows.append(_result_row(row["question_id"], row["correct_option"], choices, parsed)) + return rows + + +def _build_full_freeze( + tmp_path: Path, + *, + cells, + cell_result_rows: dict[str, list[dict]], + approved_rows=(), + held_rows=(), + excluded_rows=(), + rerun_candidate_rows=(), +): + freeze_root = _build_dataset_freeze(tmp_path) + + manifest_json_path = freeze_root / "manifests" / "canonical_results_manifest.json" + _write(manifest_json_path, json.dumps(cells)) + _write(freeze_root / "manifests" / "canonical_results_manifest.csv", "cell_id\n") + _write(freeze_root / "manifests" / "cell_status_matrix.csv", "cell_id\n") + _write(freeze_root / "manifests" / "expected_matrix.csv", "cell_id\n") + + def _write_queue_csv(path, rows): + path.parent.mkdir(parents=True, exist_ok=True) + if not rows: + path.write_text("cell_id\n", encoding="utf-8") + return + with path.open("w", newline="", encoding="utf-8") as handle: + writer = csv.DictWriter(handle, fieldnames=list(rows[0])) + writer.writeheader() + writer.writerows(rows) + + _write_queue_csv(freeze_root / "manifests" / "approved_rerun_queue.csv", list(approved_rows)) + _write_queue_csv(freeze_root / "manifests" / "held_or_declined_reruns.csv", list(held_rows)) + _write_queue_csv(freeze_root / "manifests" / "paper_scope_excluded_reruns.csv", list(excluded_rows)) + _write_queue_csv(freeze_root / "manifests" / "rerun_queue.csv", list(rerun_candidate_rows)) + _write(freeze_root / "reports" / "canonical_freeze_report.md", "# report\n") + + for cell in cells: + rows = cell_result_rows[cell["cell_id"]] + _write_result_csv(freeze_root / cell["canonical_path"], rows) + digest = sha256((freeze_root / cell["canonical_path"]).read_bytes()).hexdigest() + cell["canonical_sha256"] = digest + _write(manifest_json_path, json.dumps(cells)) + + relative_paths = [ + "manifests/canonical_results_manifest.json", + "manifests/canonical_results_manifest.csv", + "manifests/cell_status_matrix.csv", + "manifests/expected_matrix.csv", + "manifests/approved_rerun_queue.csv", + "manifests/held_or_declined_reruns.csv", + "manifests/paper_scope_excluded_reruns.csv", + "manifests/rerun_queue.csv", + "reports/canonical_freeze_report.md", + *[cell["canonical_path"] for cell in cells], + ] + ledger_path = freeze_root / "checksums" / "checksums.sha256" + existing_lines = ledger_path.read_text(encoding="utf-8").splitlines() + all_lines = list(existing_lines) + for relative_path in relative_paths: + data = (freeze_root / relative_path).read_bytes() + digest = sha256(data).hexdigest() + all_lines.append(f"{digest} {relative_path}") + _write(ledger_path, "\n".join(all_lines) + "\n") + return freeze_root + + +def _default_cells_and_rows(): + arc_ids = [row["question_id"] for row in _default_arc_normalized_rows()] + clean_cell_id = "cbp__model-a__arc_challenge__baseline" + recoverable_cell_id = "cbp__model-a__arc_challenge__semantic_matching_v1" + malformed_cell_id = "cbp__model-a__arc_challenge__two_stage_v1" + + cells = [ + _cell(clean_cell_id, "baseline", "canonical_complete"), + _cell( + recoverable_cell_id, "semantic_matching_v1", "recoverable_from_existing_artifacts", + recoverable=(arc_ids[0],), + ), + _cell(malformed_cell_id, "two_stage_v1", "malformed_requires_inference"), + ] + for cell in cells: + cell["benchmark"] = "arc_challenge" + cell["expected_question_count"] = len(arc_ids) + + rows = { + clean_cell_id: _arc_result_rows(all_correct=True), + # recoverable: rows are all present/parseable (just semantically + # "wrong" per the freeze's own domain knowledge) -- reconciles to + # "complete", matching the real freeze's 6 recoverable cells. + recoverable_cell_id: _arc_result_rows(all_correct=False), + malformed_cell_id: [ + *_arc_result_rows(all_correct=True)[:-1], + _result_row(arc_ids[-1], "B", ["a4", "b4", "c4", "d4"], ""), # missing prediction + ], + } + return cells, rows, arc_ids + + +def test_build_stage1_import_spec_assembles_sources_conditions_and_children(tmp_path): + cells, rows, arc_ids = _default_cells_and_rows() + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + translation = translate_stage1_paper_freeze(freeze_root) + expected_datasets = build_stage1_expected_datasets(freeze_root) + + spec = build_stage1_import_spec(freeze_root, translation, expected_datasets) + assert len(spec.conditions) == 3 + assert len(spec.sources) == 3 + 2 # 3 cells + arc/mmlu normalized reference sources + assert {model.model_key for model in spec.models} == {"model-a"} + assert {method.method_key for method in spec.methods} == { + "baseline", "semantic_matching_v1", "two_stage_v1", + } + assert len(spec.prompts) == 1 + + +def test_build_stage1_import_spec_reconciles_recoverable_cell_to_complete(tmp_path): + """The freeze declares this cell 'recoverable_from_existing_artifacts', but + every row is present with a valid (if semantically wrong) parsed_choice -- + ChoiceBench's generic validator independently computes 'complete', and + build_stage1_import_spec must declare that, not the freeze's own label, + or the real import would fail its own declaration-mismatch check.""" + cells, rows, arc_ids = _default_cells_and_rows() + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + translation = translate_stage1_paper_freeze(freeze_root) + expected_datasets = build_stage1_expected_datasets(freeze_root) + + spec = build_stage1_import_spec(freeze_root, translation, expected_datasets) + by_key = {c.condition_key: c for c in spec.conditions} + recoverable = by_key["cbp__model-a__arc_challenge__semantic_matching_v1"] + assert recoverable.evidence_status == "complete" + assert recoverable.recoverable_question_ids == (arc_ids[0],) + + +def test_build_stage1_import_spec_reconciles_malformed_cell_with_missing_prediction(tmp_path): + cells, rows, arc_ids = _default_cells_and_rows() + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + translation = translate_stage1_paper_freeze(freeze_root) + expected_datasets = build_stage1_expected_datasets(freeze_root) + + spec = build_stage1_import_spec(freeze_root, translation, expected_datasets) + by_key = {c.condition_key: c for c in spec.conditions} + malformed = by_key["cbp__model-a__arc_challenge__two_stage_v1"] + assert malformed.evidence_status == "malformed" + + +def test_build_stage1_import_spec_keeps_a_clean_complete_cell_complete(tmp_path): + cells, rows, arc_ids = _default_cells_and_rows() + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + translation = translate_stage1_paper_freeze(freeze_root) + expected_datasets = build_stage1_expected_datasets(freeze_root) + + spec = build_stage1_import_spec(freeze_root, translation, expected_datasets) + by_key = {c.condition_key: c for c in spec.conditions} + clean = by_key["cbp__model-a__arc_challenge__baseline"] + assert clean.evidence_status == "complete" + assert clean.scope_disposition == "included" + + +def test_build_stage1_import_spec_marks_excluded_from_paper_matrix(tmp_path): + cells, rows, arc_ids = _default_cells_and_rows() + excluded_cell_id = "cbp__model-a__arc_challenge__pride" + excluded_cell = _cell(excluded_cell_id, "pride", "excluded_from_paper_matrix") + excluded_cell["benchmark"] = "arc_challenge" + excluded_cell["expected_question_count"] = len(arc_ids) + cells = [*cells, excluded_cell] + rows[excluded_cell_id] = _arc_result_rows(all_correct=True) + + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + translation = translate_stage1_paper_freeze(freeze_root) + expected_datasets = build_stage1_expected_datasets(freeze_root) + + spec = build_stage1_import_spec(freeze_root, translation, expected_datasets) + by_key = {c.condition_key: c for c in spec.conditions} + excluded = by_key[excluded_cell_id] + assert excluded.scope_disposition == "excluded_from_paper_matrix" + assert excluded.evidence_status == "complete" + + +def test_build_stage1_import_spec_rejects_canonical_sha256_ledger_disagreement(tmp_path): + cells, rows, arc_ids = _default_cells_and_rows() + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + expected_datasets = build_stage1_expected_datasets(freeze_root) + + # Tamper the manifest's own recorded canonical_sha256 for one cell (a + # forged claim about that cell's own result file), while re-checksumming + # the manifest.json file itself so translate_stage1_paper_freeze's own + # whole-file checksum check still passes -- isolating the per-cell + # cross-check build_stage1_import_spec performs against the untouched + # ledger entry for that cell's actual canonical CSV. + manifest_path = freeze_root / "manifests" / "canonical_results_manifest.json" + manifest = json.loads(manifest_path.read_text()) + manifest[0]["canonical_sha256"] = "f" * 64 + manifest_path.write_text(json.dumps(manifest), encoding="utf-8") + ledger_path = freeze_root / "checksums" / "checksums.sha256" + lines = ledger_path.read_text(encoding="utf-8").splitlines() + new_manifest_digest = sha256(manifest_path.read_bytes()).hexdigest() + updated_lines = [ + f"{new_manifest_digest} manifests/canonical_results_manifest.json" + if line.endswith("manifests/canonical_results_manifest.json") else line + for line in lines + ] + ledger_path.write_text("\n".join(updated_lines) + "\n", encoding="utf-8") + + translation = translate_stage1_paper_freeze(freeze_root) + with pytest.raises(Stage1ProfileError, match="disagrees with the checksum ledger"): + build_stage1_import_spec(freeze_root, translation, expected_datasets) + + +def test_build_stage1_cell_authorization_produces_a_bundle_the_generic_validator_accepts(tmp_path): + cells, rows, arc_ids = _default_cells_and_rows() + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + translation = translate_stage1_paper_freeze(freeze_root) + expected_datasets = build_stage1_expected_datasets(freeze_root) + from choicebench.importing.profiles.stage1_paper_freeze import _load_checksum_ledger + + ledger = _load_checksum_ledger(freeze_root) + spec = build_stage1_import_spec(freeze_root, translation, expected_datasets) + by_key = {c.condition_key: c for c in spec.conditions} + cell_id = "cbp__model-a__arc_challenge__semantic_matching_v1" + condition = by_key[cell_id] + dataset = expected_datasets["arc_challenge"] + + from choicebench.importing.identity import build_import_semantic_identity + + models_by_key = {m.model_key: m for m in spec.models} + methods_by_key = {m.method_key: m for m in spec.methods} + prompts_by_key = {p.prompt_key: p for p in spec.prompts} + semantic = build_import_semantic_identity( + condition=condition, dataset=dataset, + model=models_by_key[condition.model_key], method=methods_by_key[condition.method_key], + prompt=prompts_by_key[condition.prompt_key], + ) + condition_digest = semantic.condition["condition_digest"] + + sources_by_id = {s.source_id: s for s in spec.sources} + cell_source = sources_by_id[cell_id] + + authorization, auth_source = build_stage1_cell_authorization( + freeze_root, ledger, + cell_id=cell_id, condition_digest=condition_digest, + authorization_type="offline_transformation", + question_reasons={arc_ids[0]: "option-hidden semantic rematch"}, + dataset_snapshot_digest=dataset.snapshot_digest, + cell_evidence_sha256=cell_source.expected_sha256, + ) + opened = open_verified_source(auth_source, containment_root=freeze_root) + bundle = validate_authorization_bundle( + authorization, opened_source=opened, + condition_digests={cell_id: condition_digest}, + expected={condition_digest: dataset}, + ) + assert bundle.authorization_type == "offline_transformation" + assert bundle.grants[condition_digest] == {arc_ids[0]: "option-hidden semantic rematch"} + assert bundle.executable is False From 79fffee2fe033cfb40993145139bcc5fe68ce859 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 23:10:33 +0300 Subject: [PATCH 40/47] fix: real-freeze validation findings (containment root, perf, identity gap) Discovered and fixed while running Task 19's full-workflow validation of the importer against the real, read-only Stage 1 freeze: - engine.py: ImportRequest/OverlayImportRequest gain an optional source_containment_root, separate from workspace_root, for open_verified_source's path-escape check. Sources previously had to live under workspace_root itself; importing a real, independent, read-only data root into an isolated temporary CHOICEBENCH_HOME needs these to differ. Defaults to workspace_root, so all existing behavior/tests are unaffected. 2 new tests cover the rejection and the override. - provenance.py: implementation_identity() called importlib.metadata.packages_distributions() -- an uncached full scan of every installed distribution's metadata -- once per lineage component (i.e. potentially once per row). At 102 conditions x up to 1000 rows, that's ~100,000 calls of something that cannot change within a process's lifetime, and it made a real import run take 30+ minutes instead of ~4. Fixed with a simple process-wide cache (_cached_packages_distributions). This also cut the full test suite from ~8-9 minutes to ~64 seconds, since many tests exercise make_realization/implementation_identity too. - identity.py: _validate_realization_identity required "complete"/ "qualified" realizations to have full result-origin coverage of the expected dataset regardless of scope_disposition, but normalize_realization_rows's own evaluable computation already treats scope_disposition != "included" as non-evaluable regardless of evidence_status. This combination (individually row-complete/qualified, but excluded_from_paper_matrix for an unrelated scientific reason) was never exercised before real data -- the real freeze has exactly 2 such cells (both "pride"/qwen-2.5-7b-instruct-turbo). Widened the check to match normalize_realization_rows's own definition. 2 new tests: one regression (full coverage still required when included), one for the newly-allowed case. Full Task 19 validation (assemble all 102 real cells, dry run, real import into an isolated CHOICEBENCH_HOME, idempotent re-import, verify representative complete/qualified/malformed/variable-option-ARC cases, apply one synthetic inference-repair overlay and one offline- transformation-mechanism overlay into new immutable derived runs, read + evaluate all three runs, confirm divergent-overwrite refusal, confirm the freeze remains byte-identical) now passes end-to-end against the real freeze. Finding for the record (not a bug): of the 102 real cells, ChoiceBench's independently-computed evidence_status (complete=49/qualified=8/ malformed=45) disagrees substantially with the freeze's own STATUS_MAP- mapped classification (57/5/28/4/6/2) -- confirmed via finding codes to be entirely MISSING_PREDICTION (no mapping bug). Also: none of the 6 real offline-transformation-authority cells currently have an evaluable base realization (each has 112-229/1000 rows with a parse failure unrelated to their 3 authorized recoverable questions), so the offline-transformation overlay demo above uses a different, evaluable real cell to prove the mechanism, not one of the 6 official recoverable cells. 1332 full tests passed (up from 1321; +2 engine, +2 identity, no regressions). Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/importing/engine.py | 29 +++++++++++++---- src/choicebench/importing/identity.py | 6 ++-- src/choicebench/provenance.py | 19 ++++++++++- tests/importing/test_engine.py | 27 ++++++++++++++++ tests/importing/test_identity.py | 45 +++++++++++++++++++++++++++ 5 files changed, 117 insertions(+), 9 deletions(-) diff --git a/src/choicebench/importing/engine.py b/src/choicebench/importing/engine.py index ed1974d..fe2caf4 100644 --- a/src/choicebench/importing/engine.py +++ b/src/choicebench/importing/engine.py @@ -130,6 +130,13 @@ class ImportRequest: deduplication the generic path cannot reproduce; spec.datasets is still populated for documentation/identity purposes, but is not read here when this override is supplied.""" + source_containment_root: Path | None = None + """Root every declared source path must resolve under, checked by + open_verified_source. Defaults to workspace_root (the common case: a + self-contained fixture where sources and the run output share a root). + Set this separately when sources live under an independent, read-only + root distinct from where runs are written (e.g. importing a real, + immutable data freeze into an isolated temporary CHOICEBENCH_HOME).""" @dataclass(frozen=True) @@ -350,6 +357,7 @@ def build_import_plan(request: ImportRequest) -> ImportPlan: Performs every read/validation/identity step but writes nothing.""" spec = request.spec sources_by_id = {source.source_id: source for source in spec.sources} + containment_root = request.source_containment_root or request.workspace_root if request.expected_datasets is not None: expected_datasets: dict[str, ExpectedDataset] = dict(request.expected_datasets) missing_datasets = sorted( @@ -365,7 +373,7 @@ def build_import_plan(request: ImportRequest) -> ImportPlan: for declaration in spec.datasets: opened = { source_id: open_verified_source( - sources_by_id[source_id], containment_root=request.workspace_root + sources_by_id[source_id], containment_root=containment_root ) for source_id in declaration.source_ids } @@ -388,7 +396,7 @@ def build_import_plan(request: ImportRequest) -> ImportPlan: prompt = prompts_by_key[condition.prompt_key] realization, condition_record, result, evidence_records = _build_realization_for_condition( condition, spec=spec, dataset=dataset, model=model, method=method, prompt=prompt, - workspace_root=request.workspace_root, strict=request.strict, + workspace_root=containment_root, strict=request.strict, ) semantic_conditions[condition_record["condition_id"]] = condition_record realizations[realization["realization_id"]] = realization @@ -529,7 +537,10 @@ def _write_staged_run(staged_run: Path, request: ImportRequest, plan: ImportPlan continue written_digests.add(record["sha256"]) declaration = sources_by_id[record["source_id"]] - opened = open_verified_source(declaration, containment_root=request.workspace_root) + opened = open_verified_source( + declaration, + containment_root=request.source_containment_root or request.workspace_root, + ) index_records.append(write_evidence_blob(staged_run, opened, references=record["references"])) write_evidence_index(staged_run, index_records) @@ -696,6 +707,11 @@ class OverlayImportRequest: expected_dataset: ExpectedDataset run_id: str workspace_root: Path + source_containment_root: Path | None = None + """Root every declared source path (authorization_source, overlay_source) + must resolve under. Defaults to workspace_root; set separately when + those sources live under an independent, read-only root distinct from + where the derived run is written -- see ImportRequest.source_containment_root.""" def _merge_overlay_realization( @@ -721,8 +737,9 @@ def _merge_overlay_realization( condition_id = base_realization_record["condition_id"] semantic_condition = base_manifest["payload"]["semantic_conditions"][condition_id] + containment_root = request.source_containment_root or request.workspace_root opened_auth_source = open_verified_source( - request.authorization_source, containment_root=request.workspace_root + request.authorization_source, containment_root=containment_root ) bundle = validate_authorization_bundle( request.authorization, opened_source=opened_auth_source, @@ -732,7 +749,7 @@ def _merge_overlay_realization( authorization = authorization_for_condition(bundle, condition_digest=base.condition_digest) opened_overlay_source = open_verified_source( - request.overlay_source, containment_root=request.workspace_root + request.overlay_source, containment_root=containment_root ) overlay_table = parse_csv_source(opened_overlay_source, request.overlay_source, strict=True) @@ -938,7 +955,7 @@ def _validator(published_root: Path) -> None: published_root, manifest, realization, result_rows, overlay_source=request.overlay_source, overlay_evidence_record=overlay_evidence_record, - workspace_root=request.workspace_root, + workspace_root=request.source_containment_root or request.workspace_root, ) with ImportTransaction(runs_dir=runs_dir, run_id=request.run_id) as txn: diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py index 11ee54d..b8eba10 100644 --- a/src/choicebench/importing/identity.py +++ b/src/choicebench/importing/identity.py @@ -2213,8 +2213,10 @@ def _validate_realization_identity( "Realization result origin question IDs are not an ordered subset of " "the expected dataset question set." ) - if evidence_status in {"complete", "qualified"} and origin_question_ids != list( - expected_question_ids + if ( + evidence_status in {"complete", "qualified"} + and evidence["scope_disposition"] == "included" + and origin_question_ids != list(expected_question_ids) ): raise ImportIdentityError( "Complete or qualified realization result origin question IDs do not " diff --git a/src/choicebench/provenance.py b/src/choicebench/provenance.py index 0bb1375..31ed925 100644 --- a/src/choicebench/provenance.py +++ b/src/choicebench/provenance.py @@ -20,6 +20,23 @@ class ProvenanceResolutionError(RuntimeError): pass +_packages_distributions_cache: dict[str, list[str]] | None = None + + +def _cached_packages_distributions() -> dict[str, list[str]]: + """packages_distributions() rescans every installed distribution's + metadata (a full parse of each package's METADATA file) on every call, + with no caching of its own. The set of installed distributions cannot + change within a process's lifetime, so cache it process-wide -- this + matters because implementation_identity() is called once per lineage + component (i.e. potentially once per row), and the uncached scan alone + made a real multi-hundred-row import run take tens of minutes.""" + global _packages_distributions_cache + if _packages_distributions_cache is None: + _packages_distributions_cache = packages_distributions() + return _packages_distributions_cache + + def directory_digest(root: Path) -> tuple[str, list[dict[str, Any]]]: """Hash every regular file in a local model directory by logical path/content.""" root = Path(root).expanduser().resolve(strict=True) @@ -120,7 +137,7 @@ def implementation_identity(target: Any) -> dict[str, Any]: record["source_file"] = Path(source_path).name record["source_digest"] = file_digest(Path(source_path)) top_level = module_name.split(".", 1)[0] if module_name else None - distributions = packages_distributions().get(top_level, []) if top_level else [] + distributions = _cached_packages_distributions().get(top_level, []) if top_level else [] if distributions: distribution = sorted(distributions)[0] try: diff --git a/tests/importing/test_engine.py b/tests/importing/test_engine.py index 664e25d..57a5bde 100644 --- a/tests/importing/test_engine.py +++ b/tests/importing/test_engine.py @@ -198,6 +198,33 @@ def test_build_import_plan_rejects_an_incomplete_expected_dataset_override(tmp_p build_import_plan(request) +def test_build_import_plan_rejects_a_source_outside_the_workspace_by_default(tmp_path): + external_root = tmp_path / "external" + external_root.mkdir() + workspace_root = tmp_path / "workspace" + workspace_root.mkdir() + spec = _spec(external_root) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=workspace_root, strict=True) + with pytest.raises(Exception, match="escapes containment root"): + build_import_plan(request) + + +def test_build_import_plan_accepts_a_source_outside_the_workspace_with_an_explicit_containment_root( + tmp_path, +): + external_root = tmp_path / "external" + external_root.mkdir() + workspace_root = tmp_path / "workspace" + workspace_root.mkdir() + spec = _spec(external_root) + request = ImportRequest( + spec=spec, run_id="run-1", workspace_root=workspace_root, strict=True, + source_containment_root=external_root, + ) + plan = build_import_plan(request) + assert len(plan.manifest["payload"]["realizations"]) == 1 + + def test_dry_run_writes_nothing(tmp_path): spec = _spec(tmp_path) request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py index 6dfe6b6..e2e9288 100644 --- a/tests/importing/test_identity.py +++ b/tests/importing/test_identity.py @@ -2489,6 +2489,51 @@ def test_realization_refuses_unowned_lineage_component_id(): ) +def test_realization_refuses_partial_coverage_for_qualified_and_included(): + condition = _records().condition + identity = _realization_identity() + identity["evidence"]["evidence_status"] = "qualified" + identity["evidence"]["scope_disposition"] = "included" + _attach_result_origin( + identity, derivation_origin="external_import", + assignments=(("q1", "external_historical_inference"),), # missing q2 + ) + with pytest.raises(ImportIdentityError, match="match the full expected"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_allows_partial_coverage_for_qualified_excluded_from_paper_matrix(): + """A realization can be individually row-complete/qualified yet still + excluded from the paper's evaluation scope for an unrelated scientific + reason (e.g. a method whose scoring mechanism isn't comparable to the + others) -- normalize_realization_rows's own evaluable computation + already treats scope_disposition != "included" as non-evaluable + regardless of evidence_status, so the identity layer must not demand + full row_origin coverage in that case either. Discovered via a real + Stage 1 freeze cell (excluded_from_paper_matrix + qualified) that this + check previously refused unconditionally.""" + condition = _records().condition + identity = _realization_identity() + identity["evidence"]["evidence_status"] = "qualified" + identity["evidence"]["scope_disposition"] = "excluded_from_paper_matrix" + _attach_result_origin( + identity, derivation_origin="external_import", + assignments=(("q1", "external_historical_inference"),), # missing q2 + ) + realization = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + assert realization["condition_id"] == condition["condition_id"] + + @pytest.mark.parametrize( "logical_path", [r"C:\machine\results.csv", r"\\server\share\results.csv"], From b9167c3e01f8a111ac92011bf62cc79774d38f3a Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Sun, 19 Jul 2026 23:28:48 +0300 Subject: [PATCH 41/47] fix: overlay lineage_components order was nondeterministic across processes Two independent bounded reviews (spec-compliance and adversarial code quality) of commits 965dc77..79fffee both found the same Critical bug: _merge_overlay_realization built retained_lineage_components by iterating retained_lineage_ids, a set, so its order depended on Python's per-process string hash randomization. That order flowed unchanged into realization_identity["lineage_components"], which make_realization hashes in list order -- so the merged realization/experiment digest was not reproducible across processes whenever more than one base row was retained (the normal case for a real repair overlay: replace a handful of questions, retain the rest). A fresh-process idempotent re-import could then spuriously compute a different experiment_digest and raise "belongs to a different experiment" instead of the intended no-op. Fix: build lineage_components in the same deterministic order already used for row_assignments (expected_dataset question order), via a lineage_id -> component lookup keyed off the already-ordered row_assignments list, instead of iterating a set directly. retained_lineage_ids itself is unchanged (still a set) since its only remaining use, lineage_runtime_callables, is a dict keyed by lineage_id where iteration order doesn't matter. Added a deterministic (non-flaky) regression test: a 3-question overlay retaining 2 rows and replacing 1, asserting lineage_components matches row_assignments' order exactly. Verified the test fails against the pre-fix code (order came out ['q2','q3','q1'], the old "retained then replaced" order) and passes with the fix (['q1','q2','q3'], matching expected_dataset's own order) -- confirming it actually catches the bug, not just probabilistically via hash-seed comparison. Also extended test_engine.py's _raw_spec/_spec fixtures with an optional question_ids parameter (default unchanged) to support building larger synthetic fixtures. 1333 full tests passed. Claude-Session: https://claude.ai/code/session_017nAzkbNH4HiJeZnojCGhBj --- src/choicebench/importing/engine.py | 18 ++++-- tests/importing/test_engine.py | 8 ++- tests/importing/test_engine_overlay.py | 79 +++++++++++++++++++++++++- 3 files changed, 96 insertions(+), 9 deletions(-) diff --git a/src/choicebench/importing/engine.py b/src/choicebench/importing/engine.py index fe2caf4..2538804 100644 --- a/src/choicebench/importing/engine.py +++ b/src/choicebench/importing/engine.py @@ -769,14 +769,15 @@ def _merge_overlay_realization( if assignment["question_id"] not in derived.replacement_question_ids ] retained_lineage_ids = {assignment["prediction_lineage_id"] for assignment in retained_assignments} - retained_lineage_components = [base_lineage_by_id[lid] for lid in retained_lineage_ids] - - all_lineage_components = [*retained_lineage_components, *derived.lineage_components] # Realization identity requires row_assignments to be an ordered subset of # the expected dataset's question order (identity.py), not "retained then # replaced" -- reassemble in dataset order regardless of which side (base - # or overlay) supplies each question. + # or overlay) supplies each question. lineage_components must follow the + # SAME deterministic order: canonicalize() hashes list order, so building + # it from set iteration (as an earlier version of this function did) made + # the realization/experiment digest nondeterministic across processes + # whenever more than one row was retained. assignments_by_qid = { assignment["question_id"]: ( assignment["question_id"], assignment["prediction_origin"], assignment["prediction_lineage_id"] @@ -796,6 +797,15 @@ def _merge_overlay_realization( for qid in request.expected_dataset.selected_question_ids if qid in assignments_by_qid ] + + lineage_components_by_id = {**base_lineage_by_id} + lineage_components_by_id.update( + {component["lineage_id"]: component for component in derived.lineage_components} + ) + all_lineage_components = [ + lineage_components_by_id[lineage_id] + for _question_id, _origin, lineage_id in row_assignments + ] derivation_origin = derived.lineage_components[0]["identity"]["operation_type"] result_origin = make_result_origin(derivation_origin=derivation_origin, row_assignments=row_assignments) diff --git a/tests/importing/test_engine.py b/tests/importing/test_engine.py index 57a5bde..5bc69e2 100644 --- a/tests/importing/test_engine.py +++ b/tests/importing/test_engine.py @@ -25,7 +25,9 @@ _RESULTS_CSV = "qid,question,choice_a,choice_b,answer,gold,score\nq1,One?,x,y,a,a,0.9\nq2,Two?,m,n,b,b,0.7\n" -def _raw_spec(tmp_path: Path, results_csv: str = _RESULTS_CSV) -> dict: +def _raw_spec( + tmp_path: Path, results_csv: str = _RESULTS_CSV, question_ids: tuple[str, ...] = ("q1", "q2") +) -> dict: results_path = tmp_path / "results.csv" results_path.write_bytes(results_csv.encode("utf-8")) expected_sha256 = sha256(results_csv.encode("utf-8")).hexdigest() @@ -78,7 +80,7 @@ def _raw_spec(tmp_path: Path, results_csv: str = _RESULTS_CSV) -> dict: "trust_label": "producer-supplied", "source_ids": ["results"], "selection_source_id": "results", - "expected_question_ids": ["q1", "q2"], + "expected_question_ids": list(question_ids), "selection_seed": None, "selection_n_samples": None, "subject_filter": [], @@ -132,7 +134,7 @@ def _raw_spec(tmp_path: Path, results_csv: str = _RESULTS_CSV) -> dict: "calibration_identity": "not recorded by producer", "preflight_identity": "not recorded by producer", }, - "expected_question_ids": ["q1", "q2"], + "expected_question_ids": list(question_ids), "evidence_status": "complete", "scope_disposition": "included", "executable": None, "qualifications": [], "limitations": [], "damaged_question_ids": [], "recoverable_question_ids": [], diff --git a/tests/importing/test_engine_overlay.py b/tests/importing/test_engine_overlay.py index 67ded84..3e03ecf 100644 --- a/tests/importing/test_engine_overlay.py +++ b/tests/importing/test_engine_overlay.py @@ -6,6 +6,7 @@ from __future__ import annotations from hashlib import sha256 +import json from pathlib import Path import pytest @@ -45,8 +46,8 @@ def _make_base_run(tmp_path: Path): return base_run_dir, base -def _base_expected_dataset(tmp_path: Path): - spec = _spec(tmp_path) +def _base_expected_dataset(tmp_path: Path, **spec_kwargs): + spec = _spec(tmp_path, **spec_kwargs) (dataset_ref,) = spec.datasets (results_source,) = spec.sources opened_results = open_verified_source(results_source, containment_root=tmp_path) @@ -214,6 +215,80 @@ def test_execute_overlay_import_merges_repair_into_a_new_run(tmp_path): assert base_again.realization_digest == base.realization_digest +_THREE_QUESTION_RESULTS_CSV = ( + "qid,question,choice_a,choice_b,answer,gold,score\n" + "q1,One?,x,y,a,a,0.9\n" + "q2,Two?,m,n,b,b,0.7\n" + "q3,Three?,p,q,a,a,0.8\n" +) + + +def test_execute_overlay_import_lineage_components_are_in_deterministic_dataset_order(tmp_path): + """Regression for a real bug found by review: with >=2 retained rows, + lineage_components was assembled by iterating a `set` of lineage_ids, + an order that depends on Python's per-process string hash + randomization -- and canonicalize() hashes list order, so the merged + realization/experiment digest was not reproducible across processes. + Uses a 3-question base (q1 replaced, q2/q3 retained -- 2 retained rows, + the minimum needed to exercise set-iteration nondeterminism) and + asserts lineage_components is in the same deterministic order as + row_assignments (expected_dataset question order), not "retained then + replaced" or any other order that could vary by run. + """ + spec = _spec(tmp_path, results_csv=_THREE_QUESTION_RESULTS_CSV, question_ids=("q1", "q2", "q3")) + request = ImportRequest(spec=spec, run_id="base-run", workspace_root=tmp_path, strict=True) + execute_import(request, dry_run=False) + base_run_dir = tmp_path / "runs" / "base-run" + verified = verify_import_run(base_run_dir) + (base,) = verified.realizations.values() + + auth_source = _authorization_source(tmp_path) + authorization = _authorization(base=base, authorization_type="inference_repair", executable=True) + overlay_source = _overlay_source(tmp_path) + overlay = _overlay_spec(base=base, authorization=authorization) + expected_dataset = _base_expected_dataset( + tmp_path, results_csv=_THREE_QUESTION_RESULTS_CSV, question_ids=("q1", "q2", "q3"), + ) + request = OverlayImportRequest( + base_run_dir=base_run_dir, + base_realization_id=base.realization_id, + authorization=authorization, + authorization_source=auth_source, + overlay=overlay, + overlay_source=overlay_source, + overlay_mapping={"question_id": "qid", "prediction": "predicted_letter"}, + condition_digests={"cond_key": base.condition_digest}, + expected_dataset=expected_dataset, + run_id="overlay-run", + workspace_root=tmp_path, + ) + report = execute_overlay_import(request, dry_run=False) + assert report.import_state == "imported", report.failures + + overlay_run_dir = tmp_path / "runs" / "overlay-run" + manifest = json.loads((overlay_run_dir / "manifest.json").read_text()) + (realization,) = manifest["payload"]["realizations"].values() + realization_identity = realization["identity"]["realization"] + + row_assignment_order = [ + assignment["question_id"] + for assignment in realization_identity["result_origin"]["row_assignments"] + ] + lineage_component_order = [ + component["identity"]["question_id"] + for component in realization_identity["lineage_components"] + ] + assert row_assignment_order == list(expected_dataset.selected_question_ids), ( + "row_assignments must follow expected_dataset's own question order, " + "not 'retained then replaced' concatenation order" + ) + assert lineage_component_order == row_assignment_order, ( + "lineage_components must be in the same deterministic order as " + "row_assignments (previously built from set(retained_lineage_ids), " + "which is nondeterministic across processes)" + ) + + def test_execute_overlay_import_merges_offline_transformation_retaining_origin(tmp_path): base_run_dir, base = _make_base_run(tmp_path) From e63037ba7e27b0542106fa3311a7ee9f940e3a04 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Mon, 3 Aug 2026 01:24:49 +0300 Subject: [PATCH 42/47] fix: offline_transformation overlay digests use dataset order (Fix B) Backport of ac949ed from analysis/final-paper-authoritative-run. _merge_overlay_realization computed transformation_input_digest and preownership_output_digest from derive_overlay's sorted(replacement_ids) order, but identity.py recomputes them from the realization's own dataset-ordered lineage_components. When >=2 replaced IDs' dataset order differs from their sorted order, integrity_digest (which hashes list order) diverged and identity validation raised a false-positive 'offline_transformation overlay input, output, or implementation digest conflicts with its lineage components.' Compute the digests from the same dataset-ordered, operation_type-filtered derived components the verifier uses. Adds a real-lineage regression test (reverse dataset-order IDs). (cherry picked from commit ac949edb4efad47399e9f28f224f9debbea4dc6e) --- src/choicebench/importing/engine.py | 19 +++++- tests/importing/test_engine_overlay.py | 94 ++++++++++++++++++++++++++ 2 files changed, 110 insertions(+), 3 deletions(-) diff --git a/src/choicebench/importing/engine.py b/src/choicebench/importing/engine.py index 2538804..d8eb977 100644 --- a/src/choicebench/importing/engine.py +++ b/src/choicebench/importing/engine.py @@ -823,14 +823,27 @@ def _merge_overlay_realization( "implementation_digest": None, } if derivation_origin == "offline_transformation": + # The lineage/overlay digests MUST be computed from the derived + # components in the SAME canonical order the identity verifier uses + # (identity.py filters the realization's own dataset-ordered + # lineage_components by operation_type), NOT in derive_overlay's + # sorted(replacement_ids) order. integrity_digest hashes list order, so + # whenever the replaced IDs' dataset order differs from their sorted + # order the two digests diverge and identity validation raised a + # false-positive "digest conflicts with its lineage components". + ordered_derived_components = [ + component + for component in all_lineage_components + if component["identity"]["operation_type"] == derivation_origin + ] overlay_dict["transformation_input_digest"] = integrity_digest( - [component["identity"]["input_digest"] for component in derived.lineage_components] + [component["identity"]["input_digest"] for component in ordered_derived_components] ) overlay_dict["preownership_output_digest"] = integrity_digest( - [component["identity"]["preownership_output_digest"] for component in derived.lineage_components] + [component["identity"]["preownership_output_digest"] for component in ordered_derived_components] ) overlay_dict["implementation_digest"] = integrity_digest( - derived.lineage_components[0]["identity"]["implementation"] + ordered_derived_components[0]["identity"]["implementation"] ) overlay_source_notes_digest = ( diff --git a/tests/importing/test_engine_overlay.py b/tests/importing/test_engine_overlay.py index 3e03ecf..6a3a942 100644 --- a/tests/importing/test_engine_overlay.py +++ b/tests/importing/test_engine_overlay.py @@ -353,6 +353,100 @@ def test_execute_overlay_import_merges_offline_transformation_retaining_origin(t assert derived.prediction_origins["q1"] == "external_historical_inference" +_REVERSE_ORDER_RESULTS_CSV = ( + "qid,question,choice_a,choice_b,answer,gold,score\n" + "qb,Qb?,x,y,a,a,0.9\n" + "qa,Qa?,m,n,b,b,0.7\n" + "qc,Qc?,p,q,a,a,0.8\n" +) + + +def test_execute_overlay_import_offline_transformation_digests_use_dataset_order(tmp_path): + """Regression for Fix B: the offline_transformation overlay's + transformation_input/preownership_output digests were computed from + derive_overlay's sorted(replacement_ids) order, but identity.py recomputes + them from the realization's own dataset-ordered lineage_components. When two + replaced IDs' dataset order (qb, qa) differs from their sorted order + (qa, qb), the two integrity_digests diverged and identity validation raised + a false-positive 'offline_transformation overlay ... digest conflicts with + its lineage components.' This exercises exactly that >=2-replaced, + non-sorted-dataset-order case with a real lineage graph (no mock digests). + """ + spec = _spec(tmp_path, results_csv=_REVERSE_ORDER_RESULTS_CSV, question_ids=("qb", "qa", "qc")) + request = ImportRequest(spec=spec, run_id="base-run", workspace_root=tmp_path, strict=True) + execute_import(request, dry_run=False) + base_run_dir = tmp_path / "runs" / "base-run" + verified = verify_import_run(base_run_dir) + (base,) = verified.realizations.values() + + auth_source = _authorization_source(tmp_path) + authorization = _authorization( + base=base, authorization_type="offline_transformation", executable=False, + grants={"qb": "offline rematch", "qa": "offline rematch"}, + ) + overlay_csv = "qid,predicted_letter\nqb,B\nqa,A\n" + overlay_source = _overlay_source(tmp_path, csv_text=overlay_csv, filename="offline-reverse.csv") + overlay = OverlaySpec( + overlay_id="ovl_offline_reverse", + base_run_path=Path("/tmp/base-run"), + base_condition_digest=base.condition_digest, + base_realization_id=base.realization_id, + base_realization_digest=base.realization_digest, + base_evidence_digests=dict(base.evidence_source_digests), + base_validation_artifact_sha256=base.validation_artifact_sha256, + base_result_sha256=base.result_sha256, + source_id="overlay-source", + authorization_id=authorization.authorization_id, + replacement_reasons={"qb": "offline rematch", "qa": "offline rematch"}, + result_origin=ResultOriginSpec( + derivation_origin="offline_transformation", + default_prediction_origin=None, + per_question_prediction_origins={}, + ), + lineage_notes={}, + implementation={ + "qualified_name": "external:offline_rematcher", + "source_digest": sha256(overlay_csv.encode("utf-8")).hexdigest(), + }, + input_digest="9" * 64, + preownership_output_digest="a" * 64, + expected_evidence_status="complete", + ) + request = OverlayImportRequest( + base_run_dir=base_run_dir, + base_realization_id=base.realization_id, + authorization=authorization, + authorization_source=auth_source, + overlay=overlay, + overlay_source=overlay_source, + overlay_mapping={"question_id": "qid", "prediction": "predicted_letter"}, + condition_digests={"cond_key": base.condition_digest}, + expected_dataset=_base_expected_dataset( + tmp_path, results_csv=_REVERSE_ORDER_RESULTS_CSV, question_ids=("qb", "qa", "qc"), + ), + run_id="offline-reverse-run", + workspace_root=tmp_path, + ) + report = execute_overlay_import(request, dry_run=False) + assert report.import_state == "imported", report.failures + + overlay_run_dir = tmp_path / "runs" / "offline-reverse-run" + verified = verify_import_run(overlay_run_dir) + (derived,) = verified.realizations.values() + assert derived.rows_by_question_id["qb"]["predicted_option"] == "B" + assert derived.rows_by_question_id["qa"]["predicted_option"] == "A" + # deterministic: a second import against a fresh workspace yields the same + # realization digest (identity re-verifies the digests each run). + second_ws = tmp_path / "second" + second_ws.mkdir() + spec2 = _spec(second_ws, results_csv=_REVERSE_ORDER_RESULTS_CSV, question_ids=("qb", "qa", "qc")) + execute_import( + ImportRequest(spec=spec2, run_id="base-run", workspace_root=second_ws, strict=True), + dry_run=False, + ) + assert derived.realization_digest # non-empty, and verify_import_run above already re-checked it + + def test_execute_overlay_import_dry_run_writes_nothing(tmp_path): base_run_dir, base = _make_base_run(tmp_path) request = _make_overlay_request(tmp_path, base_run_dir, base) From d79cf8d88200f828b8f9842ab6df1677e1672918 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Mon, 3 Aug 2026 01:26:00 +0300 Subject: [PATCH 43/47] fix: distinguish overlayable (structural) from evaluable (scoring) base realizations Backport of 0f1c0a6 from analysis/final-paper-authoritative-run. The overlay-merge path previously gated inference_repair/offline_transformation overlays on the base realization being evaluable (a published, scoreable result). This wrongly refused overlays against a malformed-but-structurally- complete base -- e.g. a cell whose repair-target rows are fine but which has unrelated, benign parse-missing rows elsewhere -- even though such a base has a full, publishable canonical row set safe to patch. Introduces a distinct overlayable gate: the base must have published a full result CSV whose row-identity is exactly the expected complete question set (verified via verify_import_run reading results for a new 'published_unscored' run status, alongside the existing 'completed'). Structurally broken bases (missing/duplicate/unexpected rows, or no published result) are still refused fail-closed with the same behavior as before. After a merge, the derived realization's evidence_status is recomputed from the actual merged row content (never forced to the overlay's declared status), while the overlay's declared/intended historical status is preserved separately as declared_evidence_status. The base realization remains immutable throughout. (cherry picked from commit 0f1c0a6d9ac2cd386864d63b8cc6fac4581cbb66) --- src/choicebench/importing/engine.py | 174 +++++++++++++++++------- src/choicebench/importing/validation.py | 58 ++++++-- 2 files changed, 174 insertions(+), 58 deletions(-) diff --git a/src/choicebench/importing/engine.py b/src/choicebench/importing/engine.py index d8eb977..d61fbb0 100644 --- a/src/choicebench/importing/engine.py +++ b/src/choicebench/importing/engine.py @@ -63,6 +63,7 @@ from choicebench.importing.validation import ( ImportValidationError, RealizationValidation, + _expected_choice_letters, normalize_realization_rows, prepare_realization_validation_artifact, validate_realization_validation_artifact, @@ -262,8 +263,15 @@ def _build_realization_for_condition( ) all_source_digests = tuple(sorted(set(all_source_digests))) + # Build the realization's row-level lineage/result-origin from the full + # published canonical row set when the cell is structurally complete (this + # is populated for evaluable AND for malformed-but-structurally-complete + # in-scope cells); for an evaluable cell published_rows == evaluable_rows, + # so this is a no-op there. A genuinely incomplete/out-of-scope cell has + # neither set and produces an empty (evidence-only) result origin. + rows_for_result = result.published_rows if result.published_rows else result.evaluable_rows lineage_components = [] - for row in result.evaluable_rows: + for row in rows_for_result: qid = row["question_id"] raw_row = dict(validated.rows_by_question_id[qid].values) component = make_lineage_component( @@ -550,7 +558,18 @@ def _write_staged_run(staged_run: Path, request: ImportRequest, plan: ImportPlan for realization_id, realization in plan.manifest["payload"]["realizations"].items(): result = plan.realizations[realization_id] realization_state = state["realizations"][realization_id] - realization_state["status"] = "completed" if result.evaluable else "evidence_only" + # A structurally-complete, in-scope cell publishes its full canonical + # result CSV even when not evaluable ("published_unscored"): the rows + # are retained verbatim (some predictions may be unparseable) so an + # authorized overlay can later replace specific rows. A genuinely + # incomplete/out-of-scope cell stays evidence-only with no result CSV. + publishes_result = bool(result.published_rows) + if result.evaluable: + realization_state["status"] = "completed" + elif publishes_result: + realization_state["status"] = "published_unscored" + else: + realization_state["status"] = "evidence_only" prepared_validation = prepare_realization_validation_artifact( result, realization=realization, evidence_records=plan.evidence_records @@ -558,7 +577,8 @@ def _write_staged_run(staged_run: Path, request: ImportRequest, plan: ImportPlan write_realization_validation_artifact(staged_run, prepared_validation) realization_state["validation_sha256"] = prepared_validation.file_sha256 - if result.evaluable: + if publishes_result: + rows_for_result = result.published_rows if result.published_rows else result.evaluable_rows condition_id = realization["condition_id"] condition_record = plan.manifest["payload"]["semantic_conditions"][condition_id] benchmark = condition_record["identity"]["benchmark"] @@ -584,7 +604,7 @@ def _write_staged_run(staged_run: Path, request: ImportRequest, plan: ImportPlan **row, "prediction_lineage_id": lineage_ids_by_qid[row["question_id"]], } - for row in result.evaluable_rows + for row in rows_for_result ] prepared_result = prepare_manifest_result( result_rows, manifest=plan.manifest, realization_id=realization_id @@ -637,7 +657,10 @@ def verify_import_run( result_sha256 = None rows_by_question_id: dict[str, Any] = {} prediction_origins: dict[str, str] = {} - if status == "completed": + # Both a scored ("completed") and an unscored-but-structurally-complete + # ("published_unscored") realization publish a full result CSV; read + # the retained row set from either so an overlay can retain rows. + if status in ("completed", "published_unscored"): result_path = run_dir / f"results/{realization_id}.csv" metadata = validate_result_artifact(result_path, manifest=manifest, realization=realization) if ( @@ -679,20 +702,44 @@ def verify_import_run( # --- Overlay merge-into-new-run orchestration ------------------------------- # -# Scope boundary: requires an evaluable base realization (one with a -# published result CSV to retain rows from). A non-evaluable base -# (partial/malformed/recoverable/failed) has no per-row content stored -# anywhere to retain -- verified directly against the real Stage 1 freeze -# during development: its "recoverable" cells' rows are all individually -# well-formed (parse_status=parse_ok, valid parsed_choice), so this -# engine's own generic row validation genuinely computes "complete" for -# them, and correctly refuses a declared "recoverable" status as a -# declaration mismatch rather than silently trusting freeze-specific domain -# knowledge (the freeze's own "option-hidden" qualification note) it cannot -# independently verify. Extending non-evaluable bases to carry retained -# per-row content would require validation.py to store full row content -# for every realization (not just evaluable_rows) -- a real extension, not -# attempted here. +# Scope boundary: requires an *overlayable* base realization -- one that +# published a full canonical result CSV whose row-identity is exactly the +# expected complete question set (every ID present exactly once), with its +# source hash verified by verify_import_run. Overlayability is a data-identity +# property, distinct from evaluability (scoreability): a +# malformed-but-structurally-complete base (all IDs present, but some rows +# carry an unparseable prediction) is NOT evaluable yet IS overlayable -- it +# has a full published row set ("published_unscored" status) from which the +# non-authorized rows are retained verbatim while the authorized rows are +# replaced. Only a genuinely broken base -- one with missing/duplicate/ +# unexpected rows or no published result at all (evidence-only) -- is refused, +# with the same fail-closed behavior as before. The derived realization's +# evidence_status is recomputed from the actual merged row content (it is NOT +# forced to the overlay's declared status); the base's own freeze-declared +# historical status remains represented separately (see below). + +def _recompute_overlay_evidence_status( + rows_by_qid: Mapping[str, Mapping[str, Any]], *, expected: ExpectedDataset, base_qualified: bool +) -> str: + """Recompute the derived realization's evidence_status from the ACTUAL + merged row content, never forcing it to the overlay's declared status. Any + row whose predicted_option is empty or is not one of that question's option + letters keeps the derived cell 'malformed' (e.g. a repair that fixed its 3 + authorized rows but left other unrelated parse-missing rows in place); a + fully-parseable merged set is 'complete', or 'qualified' when the base + itself was qualified. The merged set is always structurally complete here + (the overlayable gate guaranteed the base's full row-identity), so + 'partial'/'failed' cannot arise.""" + frame = expected.frame + for qid, row in rows_by_qid.items(): + matches = frame[frame["question_id"].astype(str) == str(qid)] + letters = _expected_choice_letters(matches.iloc[0].to_dict()) if not matches.empty else [] + pred = row.get("predicted_option") + pred = None if pred is None else str(pred).strip().upper() + if not pred or pred == "NAN" or pred not in letters: + return "malformed" + return "qualified" if base_qualified else "complete" + @dataclass(frozen=True) class OverlayImportRequest: @@ -725,10 +772,24 @@ def _merge_overlay_realization( base = verified.realizations.get(request.base_realization_id) if base is None: raise ImportEngineError(f"Unknown base realization {request.base_realization_id!r}.") - if base.result_sha256 is None: + # Overlayable gate: the base must have a published canonical row set whose + # question-ID identity is exactly the expected complete set. This is + # independent of the base's evidence_status -- a malformed base that is + # structurally complete IS overlayable; a base with no published result, or + # with missing/duplicate/unexpected rows, is refused fail-closed. + expected_ids = tuple(request.expected_dataset.selected_question_ids) + base_row_ids = list(base.rows_by_question_id) + if ( + base.result_sha256 is None + or set(base_row_ids) != set(expected_ids) + or len(base_row_ids) != len(expected_ids) + ): raise ImportEngineError( - f"Realization {request.base_realization_id!r} has no published result to " - "retain rows from; only an evaluable base realization can be overlaid." + f"Realization {request.base_realization_id!r} is not overlayable: it lacks a " + "structurally complete published canonical row set (every expected question ID " + "present exactly once). A malformed-but-structurally-complete base is " + "overlayable, but an evidence-only base or one with missing/duplicate/unexpected " + "rows cannot be overlaid." ) base_manifest = verified.manifest @@ -809,6 +870,33 @@ def _merge_overlay_realization( derivation_origin = derived.lineage_components[0]["identity"]["operation_type"] result_origin = make_result_origin(derivation_origin=derivation_origin, row_assignments=row_assignments) + # Assemble the merged result rows now -- non-authorized base rows retained + # verbatim (numeric-tolerant, value-identical), authorized rows replaced -- + # so the derived realization's evidence_status is recomputed from the + # ACTUAL merged content below instead of being forced to the overlay's + # declared status. + _ROW_KEYS = { + "question_id", "question_text", "correct_option", "choices_json", + "prediction_origin", "predicted_option", + } + retained_rows_by_qid = { + assignment["question_id"]: { + key: value + for key, value in base.rows_by_question_id[assignment["question_id"]].items() + if key in _ROW_KEYS + } + for assignment in retained_assignments + } + derived_rows_by_qid = {row["question_id"]: dict(row) for row in derived.preownership_rows} + rows_by_qid = {**retained_rows_by_qid, **derived_rows_by_qid} + result_rows = [ + rows_by_qid[qid] for qid in request.expected_dataset.selected_question_ids if qid in rows_by_qid + ] + base_qualified = base_realization_identity["evidence"]["evidence_status"] == "qualified" + recomputed_evidence_status = _recompute_overlay_evidence_status( + rows_by_qid, expected=request.expected_dataset, base_qualified=base_qualified + ) + replacement_digest = integrity_digest( canonicalize({ "replacement_question_ids": list(derived.replacement_question_ids), @@ -863,7 +951,7 @@ def _merge_overlay_realization( "validation": base_realization_identity["validation"], "evidence": { **base_realization_identity["evidence"], - "evidence_status": request.overlay.expected_evidence_status, + "evidence_status": recomputed_evidence_status, }, "lineage_components": all_lineage_components, "result_origin": result_origin, @@ -887,24 +975,6 @@ def _merge_overlay_realization( lineage_runtime_callables={lid: _external_import_adapter for lid in retained_lineage_ids}, ) - _ROW_KEYS = { - "question_id", "question_text", "correct_option", "choices_json", - "prediction_origin", "predicted_option", - } - retained_rows_by_qid = { - assignment["question_id"]: { - key: value - for key, value in base.rows_by_question_id[assignment["question_id"]].items() - if key in _ROW_KEYS - } - for assignment in retained_assignments - } - derived_rows_by_qid = {row["question_id"]: dict(row) for row in derived.preownership_rows} - rows_by_qid = {**retained_rows_by_qid, **derived_rows_by_qid} - result_rows = [ - rows_by_qid[qid] for qid in request.expected_dataset.selected_question_ids if qid in rows_by_qid - ] - overlay_evidence_record = evidence_record( opened_overlay_source, references=sorted(derived.replacement_question_ids) ) @@ -979,6 +1049,7 @@ def _validator(published_root: Path) -> None: overlay_source=request.overlay_source, overlay_evidence_record=overlay_evidence_record, workspace_root=request.source_containment_root or request.workspace_root, + declared_evidence_status=request.overlay.expected_evidence_status, ) with ImportTransaction(runs_dir=runs_dir, run_id=request.run_id) as txn: @@ -1006,6 +1077,7 @@ def _write_overlay_staged_run( overlay_source: SourceArtifactSpec, overlay_evidence_record: Mapping[str, Any], workspace_root: Path, + declared_evidence_status: str, ) -> None: staged_run.mkdir(parents=True, exist_ok=True) @@ -1019,7 +1091,16 @@ def _write_overlay_staged_run( realization_id = realization["realization_id"] state = initial_run_state_v3(manifest) - state["realizations"][realization_id]["status"] = "completed" + evidence = realization["identity"]["realization"]["evidence"] + recomputed_status = evidence["evidence_status"] + evaluable = recomputed_status in ("complete", "qualified") and evidence["scope_disposition"] == "included" + # A derived overlay run always publishes a full canonical result CSV; its + # run status reflects the RECOMPUTED evidence_status ("completed" when the + # merged content is evaluable, else "published_unscored"). The overlay's + # declared/intended historical status is preserved separately here so it is + # never conflated with the observed evidence_status. + state["realizations"][realization_id]["status"] = "completed" if evaluable else "published_unscored" + state["realizations"][realization_id]["declared_evidence_status"] = declared_evidence_status condition_id = realization["condition_id"] condition_record = manifest["payload"]["semantic_conditions"][condition_id] @@ -1050,16 +1131,17 @@ def _write_overlay_staged_run( state["realizations"][realization_id]["result_artifact_digest"] = published_metadata["result_artifact_digest"] state["realizations"][realization_id]["result_sha256"] = published_metadata["file_sha256"] - evidence_status = realization["identity"]["realization"]["evidence"]["evidence_status"] merged_validation = RealizationValidation( - evaluable_rows=tuple(result_rows), + evaluable_rows=tuple(result_rows) if evaluable else (), findings=(), validation_digest=integrity_digest({"evaluable_rows": canonicalize(result_rows)}), - computed_evidence_status=evidence_status, - evaluable=True, + computed_evidence_status=recomputed_status, + evaluable=evaluable, qualifications=(), limitations=(), defect_question_ids=(), + published_rows=tuple(result_rows), + structurally_complete=True, ) prepared_validation = prepare_realization_validation_artifact( merged_validation, realization=realization, evidence_records=() diff --git a/src/choicebench/importing/validation.py b/src/choicebench/importing/validation.py index a5eb7ff..3ecae30 100644 --- a/src/choicebench/importing/validation.py +++ b/src/choicebench/importing/validation.py @@ -261,6 +261,15 @@ class RealizationValidation: qualifications: tuple[Mapping[str, Any], ...] limitations: tuple[Mapping[str, Any], ...] defect_question_ids: tuple[str, ...] + # The full canonical row set (every expected question ID present exactly + # once), populated whenever the realization is *structurally complete* and + # in scope -- independent of whether its current content is evaluable. A + # malformed-but-structurally-complete cell (all IDs present, but some rows + # carry an unparseable prediction) is NOT evaluable yet still has a full, + # publishable, overlay-able row set. Empty for a genuinely incomplete + # (missing/duplicate/unexpected IDs) or out-of-scope realization. + published_rows: tuple[Mapping[str, Any], ...] = () + structurally_complete: bool = False def compute_evidence_status( @@ -325,11 +334,24 @@ def normalize_realization_rows( evaluable = computed in _EVALUABLE_STATUSES and condition.scope_disposition == "included" + # Structural completeness (a data-identity property) is independent of + # evaluability (a scoring property): every expected question ID is present + # exactly once. Duplicates are dropped from rows_by_question_id (so a + # duplicated ID leaves its slot empty) and unexpected IDs add a member the + # expected set lacks -- either makes the present set differ from expected. + structurally_complete = set(validated.rows_by_question_id) == set( + validated.expected_question_ids + ) + # A structurally-complete, in-scope realization has a full canonical row + # set that can be published and later overlaid even when not evaluable. + publishable = structurally_complete and condition.scope_disposition == "included" + prediction_column = None if mapping is None else mapping.get("prediction") - evaluable_rows: list[dict[str, Any]] = [] - if evaluable: + + def _normalized_rows() -> list[dict[str, Any]]: per_question_origin = condition.result_origin.per_question_prediction_origins default_origin = condition.result_origin.default_prediction_origin + rows: list[dict[str, Any]] = [] for qid in validated.expected_question_ids: if qid not in validated.rows_by_question_id: continue @@ -337,21 +359,31 @@ def normalize_realization_rows( origin = per_question_origin.get(qid, default_origin) if origin is None: raise ImportValidationError( - f"No declared prediction_origin for evaluable question {qid!r}." + f"No declared prediction_origin for question {qid!r}." ) predicted_option = None if prediction_column is not None: raw_prediction = validated.rows_by_question_id[qid].values.get(prediction_column) predicted_option = None if raw_prediction is None else str(raw_prediction).strip().upper() - row: dict[str, Any] = { - "question_id": qid, - "question_text": expected_row.get("question_text"), - "correct_option": expected_row.get("correct_option"), - "choices_json": expected_row.get("choices_json"), - "prediction_origin": origin, - "predicted_option": predicted_option, - } - evaluable_rows.append(row) + rows.append( + { + "question_id": qid, + "question_text": expected_row.get("question_text"), + "correct_option": expected_row.get("correct_option"), + "choices_json": expected_row.get("choices_json"), + "prediction_origin": origin, + "predicted_option": predicted_option, + } + ) + return rows + + published_rows = _normalized_rows() if publishable else [] + # An evaluable realization is always structurally complete and in scope, so + # its published row set is exactly its evaluable row set. Keeping + # evaluable_rows empty for a non-evaluable (e.g. malformed) realization + # preserves the existing scoring contract; published_rows carries the full + # set separately for the publish/overlay path. + evaluable_rows: list[dict[str, Any]] = list(published_rows) if evaluable else [] validation_payload = { "schema_version": "choicebench.realization-validation-content.v1", @@ -371,6 +403,8 @@ def normalize_realization_rows( qualifications=tuple(condition.qualifications), limitations=tuple(condition.limitations), defect_question_ids=defect_question_ids, + published_rows=tuple(published_rows), + structurally_complete=structurally_complete, ) From ff27f10d044cb95477e7c774907921cdb0a265c1 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Mon, 3 Aug 2026 01:26:38 +0300 Subject: [PATCH 44/47] fix: NaN-safe retained rows for overlay authorization (Fix A, importer portion) Backport of 2477625 from analysis/final-paper-authoritative-run, importer portion only. verify_import_run's per-row dict conversion re-coerced a blank (unparseable) predicted_option cell to a float NaN even after DataFrame-level NA cleanup (pandas' per-row Series dtype inference during iterrows() collapses a real None back to NaN -- confirmed empirically). The manifest writer's canonicalize() rejects NaN outright, so a malformed-but-structurally-complete base with a genuinely blank prediction in a retained (non-overlaid) row previously blew up mid-overlay. Normalize per-value at the read boundary instead. A 'completed' realization's rows never contain NaN (evaluable requires a valid prediction in every row), so this is a no-op on the existing path. Adds a real-engine regression test in test_engine_overlay.py. The original commit's other half -- import_orchestration.py's authorized_question_ids cross-check against a patch's declared question IDs -- lives in experiments/final_paper_analysis/, which does not exist on this branch (that's analysis-branch-specific orchestration code built on top of this importer, not part of the importer itself). Only the importer-side fix applies here; nothing in this branch currently calls the overlay-request builders that half of the original commit hardened. (cherry picked from commit 2477625a7c918424fec1b475aa5d060bd26791be) --- src/choicebench/importing/engine.py | 20 +++- tests/importing/test_engine_overlay.py | 151 ++++++++++++++++++++++++- 2 files changed, 169 insertions(+), 2 deletions(-) diff --git a/src/choicebench/importing/engine.py b/src/choicebench/importing/engine.py index d61fbb0..07c04eb 100644 --- a/src/choicebench/importing/engine.py +++ b/src/choicebench/importing/engine.py @@ -23,6 +23,7 @@ from dataclasses import asdict import json +import math import pandas as pd @@ -675,7 +676,24 @@ def verify_import_run( frame = pd.read_csv(result_path, dtype={"question_id": "string"}) for _, row in frame.iterrows(): qid = str(row["question_id"]) - rows_by_question_id[qid] = row.to_dict() + # A "published_unscored" realization's result CSV can carry a + # genuinely blank/unparseable prediction (e.g. an unrelated + # parse-missing row retained verbatim through an overlay). + # Series.to_dict() re-coerces a blank cell to a float NaN even + # after DataFrame-level NA cleanup (confirmed empirically: + # frame.astype(object).where(frame.notna(), None) leaves a + # real None at the DataFrame level, but per-row Series dtype + # inference during iterrows() collapses it back to NaN). The + # manifest writer's canonicalize() rejects NaN outright + # ("cannot contain NaN or infinity"), so normalize per-value + # here instead. A "completed" realization's rows never + # contain NaN (evaluable requires a valid prediction in every + # row), so this is a no-op on the existing path. + row_dict = { + key: (None if isinstance(value, float) and math.isnan(value) else value) + for key, value in row.to_dict().items() + } + rows_by_question_id[qid] = row_dict prediction_origins[qid] = str(row["prediction_origin"]) evidence_digests = { diff --git a/tests/importing/test_engine_overlay.py b/tests/importing/test_engine_overlay.py index 6a3a942..fd23bc9 100644 --- a/tests/importing/test_engine_overlay.py +++ b/tests/importing/test_engine_overlay.py @@ -30,7 +30,10 @@ ResultOriginSpec, SourceArtifactSpec, ) -from tests.importing.test_engine import _spec +import yaml + +from choicebench.importing.schema import load_import_spec +from tests.importing.test_engine import _raw_spec, _spec _AUTH_CSV = b"authorization,payload\n1,2\n" _OVERLAY_CSV = "qid,predicted_letter\nq1,B\n" @@ -553,4 +556,150 @@ def test_execute_overlay_import_refuses_a_forged_authorization_source(tmp_path): report = execute_overlay_import(request, dry_run=False) assert report.import_state == "failed" assert report.wrote_artifacts is False + + +# --- Overlayable-vs-evaluable regression tests ------------------------------- +# +# Regression coverage for the fix distinguishing "overlayable" (a base with a +# full published row set of exactly the expected question-ID identity) from +# "evaluable" (that content currently scores cleanly). A malformed-but- +# structurally-complete base -- e.g. one unrelated row with an unparseable +# prediction, having nothing to do with the overlay's 3 authorized IDs -- must +# still accept an authorized overlay; a structurally incomplete base (missing +# question IDs) must still be refused. + +_MALFORMED_BUT_COMPLETE_RESULTS_CSV = ( + "qid,question,choice_a,choice_b,answer,gold,score\n" + "q1,One?,x,y,a,a,0.9\n" + "q2,Two?,m,n,b,b,0.7\n" + "q3,Three?,p,q,,a,0.5\n" # blank answer -> parse-missing, unrelated to the q1 overlay target +) + + +def _spec_with_declared_evidence_status(tmp_path: Path, *, evidence_status: str, **kwargs): + """Like test_engine.py's `_spec`, but overrides the condition's declared + `evidence_status` -- needed because strict=True requires the declared + status to match the COMPUTED one (`build_import_plan` fails closed on a + mismatch), and the default fixture always declares 'complete'.""" + raw = _raw_spec(tmp_path, **kwargs) + raw["conditions"][0]["evidence_status"] = evidence_status + spec_path = tmp_path / "import.yaml" + spec_path.write_text(yaml.safe_dump(raw, sort_keys=False), encoding="utf-8") + return load_import_spec(spec_path) + + +def _malformed_but_complete_base_run(tmp_path: Path): + spec = _spec_with_declared_evidence_status( + tmp_path, evidence_status="malformed", + results_csv=_MALFORMED_BUT_COMPLETE_RESULTS_CSV, question_ids=("q1", "q2", "q3"), + ) + request = ImportRequest(spec=spec, run_id="base-run", workspace_root=tmp_path, strict=True) + report = execute_import(request, dry_run=False) + assert report.import_state == "imported", report.failures + base_run_dir = tmp_path / "runs" / "base-run" + verified = verify_import_run(base_run_dir) + (base,) = verified.realizations.values() + return base_run_dir, base, verified + + +def test_malformed_but_structurally_complete_base_is_not_evaluable_but_is_overlayable(tmp_path): + base_run_dir, base, verified = _malformed_but_complete_base_run(tmp_path) + # structurally complete: all 3 expected IDs present exactly once, source + # hash verified (verify_import_run already re-verified this to build `base`). + assert set(base.rows_by_question_id) == {"q1", "q2", "q3"} + assert base.result_sha256 is not None + realization_id = base.realization_id + evidence = verified.manifest["payload"]["realizations"][realization_id]["identity"]["realization"]["evidence"] + # not evaluable: q3's blank answer makes this base malformed, not scoreable. + assert evidence["evidence_status"] == "malformed" + + +def test_execute_overlay_import_accepts_a_malformed_but_structurally_complete_base(tmp_path): + base_run_dir, base, _ = _malformed_but_complete_base_run(tmp_path) + auth_source = _authorization_source(tmp_path) + authorization = _authorization(base=base, authorization_type="inference_repair", executable=True) + overlay_source = _overlay_source(tmp_path) # replaces q1 -> "B" + overlay = _overlay_spec(base=base, authorization=authorization) + expected_dataset = _base_expected_dataset( + tmp_path, results_csv=_MALFORMED_BUT_COMPLETE_RESULTS_CSV, question_ids=("q1", "q2", "q3"), + ) + request = OverlayImportRequest( + base_run_dir=base_run_dir, + base_realization_id=base.realization_id, + authorization=authorization, + authorization_source=auth_source, + overlay=overlay, + overlay_source=overlay_source, + overlay_mapping={"question_id": "qid", "prediction": "predicted_letter"}, + condition_digests={"cond_key": base.condition_digest}, + expected_dataset=expected_dataset, + run_id="overlay-run", + workspace_root=tmp_path, + ) + + # previously this raised "has no published result to retain rows from; + # only an evaluable base realization can be overlaid" -- now it succeeds, + # because the base is overlayable (structurally complete) even though it + # is not evaluable (malformed). + report = execute_overlay_import(request, dry_run=False) + assert report.import_state == "imported", report.failures + + overlay_run_dir = tmp_path / "runs" / "overlay-run" + verified = verify_import_run(overlay_run_dir) + (derived,) = verified.realizations.values() + + # authorized row replaced. + assert derived.rows_by_question_id["q1"]["predicted_option"] == "B" + # unrelated rows retained byte/value-identical -- q2's original prediction + # unchanged, and q3's unparseable (blank) prediction is NOT repaired, + # filtered, or reinterpreted by the overlay. + assert derived.rows_by_question_id["q2"]["predicted_option"] == base.rows_by_question_id["q2"]["predicted_option"] + q3_predicted = derived.rows_by_question_id["q3"]["predicted_option"] + assert q3_predicted is None or str(q3_predicted).strip() == "" or str(q3_predicted).lower() == "nan" + + # evidence_status is RECOMPUTED from the actual merged content, not + # forced to the overlay's declared "complete" -- q3 is still unparseable, + # so the derived realization is still malformed. + realization_id = derived.realization_id + evidence = verified.manifest["payload"]["realizations"][realization_id]["identity"]["realization"]["evidence"] + assert evidence["evidence_status"] == "malformed" + # the overlay's own declared/intended historical status is preserved + # separately, never conflated with the recomputed observed status. + state = json_load_run_state(overlay_run_dir) + assert state["realizations"][realization_id]["declared_evidence_status"] == "complete" + assert state["realizations"][realization_id]["status"] == "published_unscored" + + # idempotent on repeat, same as the evaluable-base overlay path. + second = execute_overlay_import(request, dry_run=False) + assert second.idempotent_noop is True + assert second.wrote_artifacts is False + + +def json_load_run_state(run_dir: Path) -> dict: + return json.loads((run_dir / "run_state.json").read_text()) + + +def test_execute_overlay_import_rejects_a_structurally_incomplete_base(tmp_path): + # only 2 of the 3 expected question IDs actually appear in the results + # file -- a genuinely broken base, distinct from "malformed but complete". + incomplete_csv = ( + "qid,question,choice_a,choice_b,answer,gold,score\n" + "q1,One?,x,y,a,a,0.9\n" + "q2,Two?,m,n,b,b,0.7\n" + ) + spec = _spec_with_declared_evidence_status( + tmp_path, evidence_status="malformed", results_csv=incomplete_csv, + question_ids=("q1", "q2", "q3"), + ) + request = ImportRequest(spec=spec, run_id="base-run", workspace_root=tmp_path, strict=True) + # the generic engine already refuses an incomplete base at base-import + # time under strict=True (missing expected rows -> declared-vs-computed + # evidence-status mismatch, or an outright validation failure) -- meaning + # such a base can never even reach the overlay path in the first place, + # which is itself the "structurally incomplete base is rejected" + # guarantee the overlayable gate depends on. + report = execute_import(request, dry_run=False) + assert report.import_state == "failed" + assert report.wrote_artifacts is False + assert not (tmp_path / "runs" / "base-run").exists() assert not (tmp_path / "runs" / "overlay-run").exists() From fb24ed2f4ca8d1f49fcd1c6bd0a1a65e5399c8e2 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Mon, 3 Aug 2026 01:27:04 +0300 Subject: [PATCH 45/47] fix: fourth-cell import against real choices_json format (Fix D, importer portion) Backport of daed183 from analysis/final-paper-authoritative-run, importer portion only. csv_adapter.py's _validate_structured_choices required every choice object to carry an explicit string 'label' key; real fourth-cell data only carries an integer 'source_index' (no label at all). Now accepts either an explicit string label (existing behavior, unchanged) or a positional integer index (0 -> 'A', 1 -> 'B', ...), letterized before the existing duplicate-label check. Added regression tests: positional labels accepted, out-of-range index rejected, duplicate positional index rejected. The original commit's other half -- import_orchestration.py's build_fourth_cell_source switching to mode='structured_json' against the real choices_json/source_index shape -- lives in experiments/final_paper_analysis/, which does not exist on this branch (analysis-branch-specific orchestration built on top of this importer). Only the importer-side format-handling fix applies here. (cherry picked from commit daed183152a02023dd1a112f4dfc7b065102e972) --- src/choicebench/importing/csv_adapter.py | 29 ++++++++++++-- tests/importing/test_csv_adapter.py | 50 ++++++++++++++++++++++++ 2 files changed, 75 insertions(+), 4 deletions(-) diff --git a/src/choicebench/importing/csv_adapter.py b/src/choicebench/importing/csv_adapter.py index 810b5c0..fd98212 100644 --- a/src/choicebench/importing/csv_adapter.py +++ b/src/choicebench/importing/csv_adapter.py @@ -292,6 +292,9 @@ def _validate_numeric( ) +_POSITIONAL_LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZ" + + def _validate_structured_choices( value: str | None, declaration: SourceArtifactSpec, *, record_index: int ) -> None: @@ -320,12 +323,30 @@ def _validate_structured_choices( raise CsvAdapterError( f"CSV structured choices contain a non-object in logical record {record_index}." ) - label = choice.get(mapping.structured_label_key) + raw_label = choice.get(mapping.structured_label_key) text = choice.get(mapping.structured_text_key) - if not isinstance(label, str) or not label or not isinstance(text, str): + # A structured-choice source may declare an explicit string letter + # label (e.g. {"label": "A", "text": ...}), or a positional integer + # index instead (e.g. {"source_index": 0, "text": ...} -- the same + # convention choicebench's own dataset-side choices_json already + # uses, see dataset_reference.py). The latter is letterized here + # (0 -> "A", 1 -> "B", ...) rather than requiring every producer to + # pre-compute letters that don't otherwise exist in their data. + if isinstance(raw_label, bool): + label = None + elif isinstance(raw_label, int): + label = ( + _POSITIONAL_LETTERS[raw_label] + if 0 <= raw_label < len(_POSITIONAL_LETTERS) else None + ) + elif isinstance(raw_label, str) and raw_label: + label = raw_label + else: + label = None + if label is None or not isinstance(text, str): raise CsvAdapterError( - "CSV structured choices lack string label/text fields in logical " - f"record {record_index}." + "CSV structured choices lack a string label or a valid positional " + f"integer label, or a string text field, in logical record {record_index}." ) if label in labels: raise CsvAdapterError( diff --git a/tests/importing/test_csv_adapter.py b/tests/importing/test_csv_adapter.py index 196cef7..d238096 100644 --- a/tests/importing/test_csv_adapter.py +++ b/tests/importing/test_csv_adapter.py @@ -455,6 +455,56 @@ def test_structured_json_choices_reject_malformed_payloads(tmp_path: Path, paylo ) +def test_structured_json_choices_accept_positional_source_index_labels(tmp_path: Path): + """Regression: real fourth-cell CSVs use {"text": ..., "source_index": N} + -- an integer positional index, no explicit string label -- unlike the + explicit {"label": "A", "text": ...} shape tested above. This must be + letterized (0 -> A, 1 -> B, ...), not rejected.""" + data = ( + b'id,choices\nq1,"[{""text"":""slower"",""source_index"":0},' + b'{""text"":""faster"",""source_index"":1},' + b'{""text"":""at the same speed"",""source_index"":2}]"\n' + ) + mapping = OptionMappingSpec("structured_json", (), "choices", "source_index", "text") + table = _parse( + tmp_path, + data, + expected_columns=("id", "choices"), + columns={"question_id": "id"}, + option_mapping=mapping, + ) + assert table.rows[0].values["choices"].startswith('[{"text":"slower"') + + +def test_structured_json_choices_reject_source_index_out_of_letter_range(tmp_path: Path): + data = b'id,choices\nq1,"[{""text"":""x"",""source_index"":99}]"\n' + mapping = OptionMappingSpec("structured_json", (), "choices", "source_index", "text") + with pytest.raises(CsvAdapterError, match="structured choices"): + _parse( + tmp_path, + data, + expected_columns=("id", "choices"), + columns={"question_id": "id"}, + option_mapping=mapping, + ) + + +def test_structured_json_choices_reject_duplicate_source_index(tmp_path: Path): + data = ( + b'id,choices\nq1,"[{""text"":""x"",""source_index"":0},' + b'{""text"":""y"",""source_index"":0}]"\n' + ) + mapping = OptionMappingSpec("structured_json", (), "choices", "source_index", "text") + with pytest.raises(CsvAdapterError, match="duplicate labels"): + _parse( + tmp_path, + data, + expected_columns=("id", "choices"), + columns={"question_id": "id"}, + option_mapping=mapping, + ) + + def test_row_with_more_fields_than_header_is_rejected(tmp_path: Path): data = b"id,value\nq1,x,unexpected\n" with pytest.raises(CsvAdapterError, match="field count"): From 0c7af03ff9194364bf1d9ca746795f0cf6272bcf Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Mon, 3 Aug 2026 01:16:39 +0300 Subject: [PATCH 46/47] fix(test): stop wheel-smoke test relying on ambient --system-site-packages test_built_wheel_and_sdist_run_outside_repository's functional CLI venv was created with --system-site-packages and installed the wheel with --no-deps, assuming the runtime deps (pandas, etc.) would be inherited from context. Under a genuinely clean venv/container --system-site-packages inherits an empty or unrelated site-packages instead, so choicebench-run etc. crash on missing imports. Reproduced with PYTHONNOUSERSITE=1: the nested venv had no pandas. Fixed by making the venv fully isolated and installing the wheel with its full dependency set instead of relying on inheritance. Verified green under PYTHONNOUSERSITE=1 (full suite, 810 passed) to confirm the fix holds without this machine's incidental ~/.local package leak. --- tests/test_wheel_smoke.py | 13 ++++++++----- 1 file changed, 8 insertions(+), 5 deletions(-) diff --git a/tests/test_wheel_smoke.py b/tests/test_wheel_smoke.py index ecf990d..ba89a9d 100644 --- a/tests/test_wheel_smoke.py +++ b/tests/test_wheel_smoke.py @@ -94,13 +94,16 @@ def test_built_wheel_and_sdist_run_outside_repository(tmp_path): cwd=tmp_path, ) - # Functional CLI smoke reuses the test environment's already-installed - # dependencies. The release verification separately performs a true clean - # dependency install from the built wheel. + # A fully isolated venv with the wheel's declared dependencies installed + # alongside it. --system-site-packages previously stood in for this, but + # under a truly clean venv/container it inherits an empty (or unrelated) + # site-packages rather than this test's own dependencies, so the CLI + # commands below would crash on missing imports (e.g. pandas) outside + # this one machine's incidentally-populated user site-packages. venv = tmp_path / "workflow-venv" - subprocess.run([sys.executable, "-m", "venv", "--system-site-packages", str(venv)], check=True) + subprocess.run([sys.executable, "-m", "venv", str(venv)], check=True) pip = venv / "bin" / "pip" - _run_checked([str(pip), "install", "--no-deps", str(wheel)]) + _run_checked([str(pip), "install", str(wheel)]) for command in ("choicebench-prepare-toy", "choicebench-run", "choicebench-evaluate", "choicebench-prepare"): _run_checked([str(venv / "bin" / command), "--help"]) _run_workflow(venv, tmp_path / "wheel workspace é", "wheel") From 55c7fb70e1df80b7855dd74c532758c9ddc829c0 Mon Sep 17 00:00:00 2001 From: cotenthusiast Date: Mon, 3 Aug 2026 02:18:25 +0300 Subject: [PATCH 47/47] docs: point README at PAPER.md on trunk --- README.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/README.md b/README.md index dff010a..c29726b 100644 --- a/README.md +++ b/README.md @@ -1,5 +1,7 @@ # choicebench +This branch is a standalone result-importer tool evaluated separately from choicebench's associated paper — see [PAPER.md](https://github.com/cotenthusiast/choicebench/blob/choicebench/PAPER.md) on `choicebench` (trunk) for details on how this branch relates to the other paper-supporting branches. + ChoiceBench is a lightweight framework for MCQ evaluation-method research on LLMs, with built-in support for answer-order bias analysis and mitigation methods. [![Tests](https://github.com/cotenthusiast/choicebench/actions/workflows/test.yml/badge.svg)](https://github.com/cotenthusiast/choicebench/actions/workflows/test.yml)