diff --git a/README.md b/README.md index dff010a..c29726b 100644 --- a/README.md +++ b/README.md @@ -1,5 +1,7 @@ # choicebench +This branch is a standalone result-importer tool evaluated separately from choicebench's associated paper — see [PAPER.md](https://github.com/cotenthusiast/choicebench/blob/choicebench/PAPER.md) on `choicebench` (trunk) for details on how this branch relates to the other paper-supporting branches. + ChoiceBench is a lightweight framework for MCQ evaluation-method research on LLMs, with built-in support for answer-order bias analysis and mitigation methods. [![Tests](https://github.com/cotenthusiast/choicebench/actions/workflows/test.yml/badge.svg)](https://github.com/cotenthusiast/choicebench/actions/workflows/test.yml) diff --git a/docs/superpowers/plans/2026-07-18-external-results-importer.md b/docs/superpowers/plans/2026-07-18-external-results-importer.md new file mode 100644 index 0000000..1380219 --- /dev/null +++ b/docs/superpowers/plans/2026-07-18-external-results-importer.md @@ -0,0 +1,3260 @@ +# External Results Importer Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use +> `$subagent-driven-development` (recommended) or `$executing-plans` to +> implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for +> tracking. + +**Goal:** Add a reusable, provenance-preserving importer that validates +externally generated row-level results, publishes them as immutable +ChoiceBench-native-format imported result artifacts, and evaluates eligible +realizations without claiming ChoiceBench performed the original inference. + +**Architecture:** Manifest v3 separates stable scientific conditions from +source-sensitive realizations while manifest-v2 readers remain byte-compatible. +The paper-agnostic importer parses strict YAML, scans exact CSV logical records, +validates rows against an explicitly classified dataset reference, builds +identity and lineage, and atomically publishes a complete staged run. A thin +Stage 1 profile translates the immutable freeze into the same generic types. + +**Tech Stack:** Python 3.10+, dataclasses, `pathlib`, `csv`-compatible parsing, +pandas, PyYAML, SHA-256/canonical JSON through `choicebench.identity`, pytest, +setuptools, build, and Twine. The implementation adds no runtime dependency. + +## Global Constraints + +- Approved design: `docs/superpowers/specs/2026-07-18-external-results-importer-design.md`. +- Branch/worktree: `feat/external-results-importer` in + `/home/cotenthusiast/Projects/choicebench-external-results-importer`. +- Preserve the existing 808-test baseline; before every task or review-fix + commit run the targeted tests and `python -m pytest -q`, with zero failures. +- Use TDD for every behavior-changing task: add a failing behavioral test, run + it and observe the expected failure, implement the smallest slice, rerun + targeted tests, then the full suite, then commit. Test-only compatibility + characterization (Task 1) and documentation/installed-package + characterization (Task 18) intentionally begin with passing gates. +- Keep manifest-v2 validation, reading, evaluation identity, resume, and result + ownership behavior intact for legacy runs. Only a pre-existing v2 run may + continue writing v2-compatible native rows; every newly created run is v3 and + records structured result origin explicitly. +- Keep semantic `condition_id` aligned with ChoiceBench's existing benchmark + selection x model x method x prompt x seed/preflight/protocol identity. + Source bytes, parsing, mapping, status, scope, authorization, repair, and + lineage belong only to `realization_id` and downstream identities. +- Result-artifact identity is computed only after exact result bytes exist and + is stored in the v2 sidecar/run state, never in the immutable v3 manifest or + CSV. Pre-ownership transformation payloads and lineage-node IDs must not + depend on their child realization/result/experiment. +- The acyclic publication order is: semantic dataset/model/method/prompt and + condition records; source/evidence plus pre-ownership validation/lineage + digests; realization; manifest/experiment; exact evidence index and + realization-validation artifact; exact result CSV; result-artifact sidecar; + final realization-keyed run state. Only downstream records may reference an + earlier identity or exact file checksum. +- A manifest-v3 result path is + `results/.csv` with sidecar + `results/.artifact.json`. +- Published runs are immutable. A repair or offline transformation consumes a + verified base run as read-only input and publishes a different run ID, + experiment ID, realization ID, and result artifact. It never appends to, + replaces, or mutates the base run. +- Keep the generic importer free of paper names, the 100-cell matrix, ARC IDs, + Gemini/IHS/PriDe identities, and historical filename rules. All such knowledge + belongs in `importing/profiles/stage1_paper_freeze.py` and profile tests. +- The Stage 1 freeze is read-only external input. Never create, rewrite, rename, + chmod, or delete anything below + `/home/cotenthusiast/Projects/model-generalization/paper_data_freeze`. +- Tests use programmatically generated small synthetic fixtures. Do not copy + historical result rows, benchmark snapshots, cache trees, or reports into Git. +- Do not run model inference or any approved, held, excluded, or forensic queue. +- Explicit user-selected absolute source/output paths are valid after + canonicalization. Specifications cannot choose output roots. Reject traversal, + unsafe symlinks, recursive directory import, secrets, executable objects, and + divergent overwrite. +- Every high-risk gate named below gets an independent read-only review before + the next task. Resolve confirmed findings and rerun that gate's commands. +- Commit only the files named by the task. Never combine two task commits. + +## Target file map + +### New reusable importer modules + +- `src/choicebench/importing/__init__.py` — stable public importer exports. +- `src/choicebench/importing/schema.py` — strict dataclass schema and YAML loader. +- `src/choicebench/importing/csv_adapter.py` — deterministic UTF-8 logical-record + scanner and CSV adapter. +- `src/choicebench/importing/dataset_reference.py` — expected-question snapshot, + trust classification, and derivation verification. +- `src/choicebench/importing/identity.py` — semantic condition, realization, + lineage-node, origin assignment, and import-specification identities. +- `src/choicebench/importing/validation.py` — question-keyed row validation, + evidence findings, extension-field preservation, and evaluable row creation. +- `src/choicebench/importing/authorization.py` — typed authorization validation. +- `src/choicebench/importing/overlays.py` — immutable repair and offline + transformation derivation. +- `src/choicebench/importing/evidence.py` — run-local content-addressed evidence + and validation artifacts. +- `src/choicebench/importing/transaction.py` — same-filesystem staged publication + and verify-only existing-run behavior. +- `src/choicebench/importing/engine.py` — generic validate/import orchestration and + machine-readable report. +- `src/choicebench/importing/profiles/__init__.py` — profile registry. +- `src/choicebench/importing/profiles/stage1_paper_freeze.py` — Stage 1-only + translation, trust chain, statuses, queue ledgers, and schema maps. +- `src/choicebench/cli/import_results.py` — installed CLI with lazy workspace + bootstrap. + +### Existing modules changed deliberately + +- `src/choicebench/manifest.py` — version-dispatched v2/v3 validation, normalized + manifest view, v3 run state, and full child-digest/path verification. +- `src/choicebench/io/writers.py` — v2 writer compatibility plus non-circular v2 + result-artifact sidecars for v3 realizations. +- `src/choicebench/io/readers.py` — version-dispatched publication reader and + explicit realization selection. +- `src/choicebench/cli/evaluate_run.py` — legacy evaluation-v1 compatibility plus + realization-aware evaluation-v2. +- `src/choicebench/cli/run_experiment.py` — new native runs emit manifest v3 and + native realization/origin metadata; existing v2 resumes stay v2. +- `src/choicebench/infra/checkpoint.py` — realization-bound checkpoint-v2 with + legacy checkpoint-v1 resume compatibility. +- `src/choicebench/infra/artifacts.py` — no-replace directory publication and + directory fsync helper. +- `pyproject.toml` — `choicebench-import-results` entry point only; no dependency + or package-version change. +- `README.md`, `CHANGELOG.md`, and `docs/external-results-import.md` — user and + adapter documentation. + +### New and expanded tests + +- `tests/importing/conftest.py` — synthetic generic and miniature Stage 1 freeze + builders. +- `tests/importing/test_manifest_v2_compat.py` +- `tests/importing/test_manifest_v3.py` +- `tests/importing/test_schema.py` +- `tests/importing/test_csv_adapter.py` +- `tests/importing/test_dataset_reference.py` +- `tests/importing/test_identity.py` +- `tests/importing/test_result_artifact.py` +- `tests/importing/test_validation.py` +- `tests/importing/test_evidence_status.py` +- `tests/importing/test_authorization.py` +- `tests/importing/test_overlays.py` +- `tests/importing/test_evidence.py` +- `tests/importing/test_transaction.py` +- `tests/importing/test_engine.py` +- `tests/importing/test_evaluation.py` +- `tests/importing/test_native_v3.py` +- `tests/importing/test_stage1_profile.py` +- `tests/importing/test_cli.py` +- Expand `tests/test_publication_identity.py`, `tests/test_condition_grid.py`, + `tests/io/test_readers.py`, `tests/io/test_writers.py`, + `tests/scripts/test_evaluate_run.py`, `tests/scripts/test_reset_run.py`, + `tests/scripts/test_build_backend.py`, `tests/test_pride_reproduction_wiring.py`, + `tests/test_release_adversarial.py`, and `tests/test_wheel_smoke.py` only where + each task says. + +## Required review protocol + +For every task, the implementer commits only after targeted and full tests pass. +For Tasks 1, 3, 5, 6, 7, 10, 11, 13, 14, 16, 17, and 19, dispatch a fresh read-only +reviewer after the commit. Give the reviewer only the approved specification, +the task text, and the commit diff. Do not advance until all confirmed Critical +and Important findings are resolved in a focused follow-up commit and the same +validation is rerun. + +## Specification acceptance-test traceability + +The approved specification's 65 acceptance cases are assigned before any +production edit. The implementing worker must retain these numbered case IDs in +test docstrings or parametrization IDs so the final compliance reviewer can +prove that no requirement was lost: + +- Cases 1-6: Tasks 2, 7, 11, and 12 + (`test_schema.py`, `test_result_artifact.py`, `test_evidence.py`, + `test_transaction.py`, and `test_engine.py`) cover valid import, + deterministic identity, exact no-op, divergent refusal, checksum mismatch, + and source mutation. +- Cases 7-17: Tasks 3, 4, 8, and 9 (`test_csv_adapter.py`, + `test_dataset_reference.py`, `test_validation.py`, and + `test_evidence_status.py`) cover question membership, gold/options, + variable-option and three-option rows, null/NaN, predictions, correctness, + columns, and method extensions. +- Cases 18-22: Tasks 9 and 13 (`test_evidence_status.py` and + `test_evaluation.py`) cover complete/qualified/partial/malformed evidence, + imported evaluation, and non-metric accounting. +- Cases 23-27: Task 10 (`test_authorization.py` and `test_overlays.py`) covers + valid repair, unauthorized/conflicting replacements, repair identity/lineage, + and offline-transformation lineage. +- Cases 28-31: Tasks 15 and 16 (`test_stage1_profile.py`) cover profile + translation, hosted PriDe exclusion, held Gemini ARC IHS, and strict + separation of approved and forensic queues. +- Cases 32-36: Tasks 17-19 (`test_cli.py`, `test_wheel_smoke.py`, and + `test_release_adversarial.py`) cover dry-run, real import, outside-repository + operation, installed wheel/sdist behavior, and adversarial security. +- Cases 37-46: Tasks 2, 9-11, 13, and 17 cover orthogonal statuses, audit-only + identity exclusion, explicit new origin/legacy-only inference, evidence CAS, + absolute path safety, exact malformed evidence, evidence-only bases, + non-self-authorizing overlays, and derived-status recomputation. +- Cases 47-52: Tasks 5, 7, 10, and 13 cover the four identity layers, + condition-sharing alternatives, exact-byte result identity, constituent and + per-row origins, and offline derivation versus prediction origin. +- Cases 53-57: Tasks 4, 10, 15, and 16 cover typed inference/offline authority, + the exact six-condition/18-question-cell offline authority, distinct dataset + trust classes, and the frozen-input trust chain with its upstream limitation. +- Cases 58-62: Tasks 2, 3, and 16 cover every identity-bearing CSV setting, + strict UTF-8/BOM/dialect/newline behavior, exact embedded-newline spans, and + the three MMLU duplicate pairs with pre/post derivation digests. +- Cases 63-65: Tasks 5, 7, 11, and 12 cover acyclic publication identity, + crash-safe no-replace publication, staging invisibility, and existing-run + verify-only behavior. + +The mapping is a minimum, not permission to omit cross-layer integration tests +listed in the individual tasks. + +--- + +### Task 1: Freeze manifest-v2 compatibility before broad changes + +**Files:** + +- Create: `tests/importing/__init__.py` +- Create: `tests/importing/conftest.py` +- Create: `tests/importing/test_manifest_v2_compat.py` +- Inspect only: `src/choicebench/manifest.py`, `src/choicebench/io/writers.py`, + `src/choicebench/io/readers.py`, `src/choicebench/cli/evaluate_run.py` + +**Interfaces:** + +- Consumes unchanged v0.2 functions: `make_manifest`, `validate_manifest`, + `initial_run_state`, `write_run_results`, `read_manifest_results`, and + `build_evaluation_report`. +- Produces `synthetic_v2_run(tmp_path, monkeypatch) -> tuple[Path, dict, str]`, + a deterministic two-question native run with explicit source/environment + records and a hard-coded expected experiment/evaluation identity. + +- [ ] **Step 1: Confirm the untouched v0.2 baseline** + + ```bash + python -m pytest -q + ``` + + Expected: exactly 808 tests pass with zero failures. Record the duration and + commit SHA. If this does not hold, stop and diagnose the starting state before + writing any plan task files. + +- [ ] **Step 2: Add characterization tests before production changes** + + Create a literal two-question dataset/prompt/model/method/condition fixture. + Pass fixed `source` and `environment` to `build_manifest_payload`, write the + manifest, snapshots, result, sidecar, and state through the current v2 APIs, + and assert the exact current IDs. The test names and assertions must be: + + ```python + def test_v2_fixture_validates_without_rewrite(synthetic_v2_run): + run_dir, manifest, _ = synthetic_v2_run + before = {p.relative_to(run_dir): p.read_bytes() + for p in run_dir.rglob("*") if p.is_file()} + validate_manifest(manifest) + frame, loaded = read_manifest_results(run_dir) + after = {p.relative_to(run_dir): p.read_bytes() + for p in run_dir.rglob("*") if p.is_file()} + assert loaded["schema_version"] == "choicebench.manifest.v2" + assert frame["question_id"].astype(str).tolist() == ["q1", "q2"] + assert after == before + + + def test_v2_evaluation_identity_and_shape_are_frozen( + synthetic_v2_run, monkeypatch + ): + run_dir, manifest, expected_evaluation_id = synthetic_v2_run + monkeypatch.setattr(evaluate_run, "RUNS_DIR", run_dir.parent) + frame, _ = read_manifest_results(run_dir) + report = evaluate_run.build_evaluation_report( + run_dir.name, frame, manifest, reparse=False + ) + assert report["schema_version"] == "choicebench.evaluation.v1" + assert report["evaluation_id"] == expected_evaluation_id + assert set(report["conditions"]) == {"cond_legacyfixture"} + + + def test_v2_missing_result_origin_implies_native_only_in_compat_view( + synthetic_v2_run, + ): + _, manifest, _ = synthetic_v2_run + assert "result_origin" not in manifest["payload"]["conditions"][0] + ``` + + The fixture helper must assert its hard-coded experiment/evaluation IDs so an + implementer cannot update expected values casually after a regression. + +- [ ] **Step 3: Run the compatibility tests against unmodified v0.2 code** + + Run: + + ```bash + python -m pytest tests/importing/test_manifest_v2_compat.py -v + ``` + + Expected: all characterization tests PASS. If they do not, correct the + synthetic fixture; do not change production behavior in this task. + +- [ ] **Step 4: Reconfirm the baseline** + + Run: + + ```bash + python -m pytest -q + ``` + + Expected: the original 808 tests plus the new compatibility tests pass, with + zero failures. + +- [ ] **Step 5: Commit the compatibility boundary** + + ```bash + git add tests/importing/__init__.py tests/importing/conftest.py \ + tests/importing/test_manifest_v2_compat.py + git commit -m "test: freeze manifest v2 compatibility" + ``` + +**Intermediate gate:** A fresh reviewer compares the synthetic run and exact +assertions with current v2 construction, reading, and evaluation. No manifest-v3 +work begins until this test-only commit is approved. + +--- + +### Task 2: Add the strict, paper-agnostic import specification + +**Files:** + +- Create: `src/choicebench/importing/__init__.py` +- Create: `src/choicebench/importing/schema.py` +- Create: `tests/importing/test_schema.py` +- Inspect: `src/choicebench/config/schema.py`, `src/choicebench/identity.py` + +**Interfaces:** + +```python +ImportState = Literal["validated", "imported", "failed"] +EvidenceStatus = Literal[ + "complete", "qualified", "partial", "malformed", "recoverable", "failed" +] +ScopeDisposition = Literal[ + "included", "excluded_from_paper_matrix", "held", "superseded" +] +PredictionOrigin = Literal[ + "native_inference", + "external_historical_inference", + "external_repair_inference", +] +DerivationOrigin = Literal[ + "native_execution", "external_import", "repair_overlay", "offline_transformation" +] + + +@dataclass(frozen=True) +class CsvDialectSpec: + encoding: Literal["utf-8"] = "utf-8" + bom_policy: Literal["forbid", "strip_utf8_bom"] = "forbid" + decoding_errors: Literal["strict"] = "strict" + delimiter: str = "," + quote_character: str = '"' + escape_character: str | None = None + double_quote: bool = True + line_terminators: tuple[str, ...] = ("crlf", "lf", "cr") + mixed_line_terminators: Literal["allow", "forbid"] = "allow" + final_record_without_terminator: Literal["allow", "forbid"] = "allow" + blank_record_policy: Literal["reject"] = "reject" + skip_initial_space: bool = False + header: Literal["first_logical_record"] = "first_logical_record" + strict_syntax: bool = True + + +@dataclass(frozen=True) +class NumericColumnSpec: + source_column: str + value_type: Literal["integer", "float"] + null_allowed: bool + finite_only: bool = True + + +@dataclass(frozen=True) +class OptionMappingSpec: + mode: Literal["ordered_columns", "structured_json"] + ordered_columns: tuple[str, ...] + structured_column: str | None + structured_label_key: str | None + structured_text_key: str | None + + +@dataclass(frozen=True) +class SourceArtifactSpec: + source_id: str + path: Path # audit-only lookup location + logical_path: str # stable identity-bearing location + expected_sha256: str + format: Literal["csv"] + format_version: str + classification: Literal[ + "raw", "canonical", "derived", "repaired", "aggregate_only" + ] + dialect: CsvDialectSpec + columns: Mapping[str, str] + expected_columns: tuple[str, ...] + ignored_columns: Mapping[str, str] + null_values: tuple[str, ...] + numeric_columns: tuple[NumericColumnSpec, ...] + option_mapping: OptionMappingSpec + extra_field_policy: Literal["preserve_unmapped", "reject_unmapped"] + preserve_namespace: str + source_run_id: str | None + source_repository: str | None + source_commit: str | None + notes: Mapping[str, Any] + + +@dataclass(frozen=True) +class DatasetReferenceSpec: + dataset_id: str + benchmark_name: str + split: str + reference_kind: Literal[ + "independent_input_snapshot", "profile_derived_reference_snapshot" + ] + trust_label: str + source_ids: tuple[str, ...] + selection_source_id: str + expected_question_ids: tuple[str, ...] + selection_seed: int | None + selection_n_samples: int | None + subject_filter: tuple[str, ...] + selection_unknown_reasons: Mapping[str, str] + columns: Mapping[str, str] + revision: str | None + fingerprint: str | None + derivation: Mapping[str, Any] + limitations: tuple[str, ...] + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ImportModelSpec: + model_key: str + display_name: str + backend: str | None + provider: str | None + revision: str | None + effective_parameters: Mapping[str, Any] + unknown_reasons: Mapping[str, str] + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ImportMethodSpec: + method_key: str + name: str + effective_parameters: Mapping[str, Any] + implementation: Mapping[str, Any] | None + unknown_reasons: Mapping[str, str] + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ImportPromptSpec: + prompt_key: str + template_identity: str | None + template_digest: str | None + template_contents: Mapping[str, str] | None + unknown_reason: str | None + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ResultOriginSpec: + derivation_origin: DerivationOrigin + default_prediction_origin: PredictionOrigin | None + per_question_prediction_origins: Mapping[str, PredictionOrigin] + + +@dataclass(frozen=True) +class ImportConditionSpec: + condition_key: str + source_ids: tuple[str, ...] + dataset_id: str + model_key: str + method_key: str + prompt_key: str + seed: int | None + calibration_identity: Mapping[str, Any] | None + preflight_identity: Mapping[str, Any] | None + protocol_settings: Mapping[str, Any] + generation_parameters: Mapping[str, Any] + unknown_reasons: Mapping[str, str] + expected_question_ids: tuple[str, ...] + evidence_status: EvidenceStatus + scope_disposition: ScopeDisposition + executable: bool | None + qualifications: tuple[Mapping[str, Any], ...] + limitations: tuple[Mapping[str, Any], ...] + damaged_question_ids: tuple[str, ...] + recoverable_question_ids: tuple[str, ...] + result_origin: ResultOriginSpec + + +@dataclass(frozen=True) +class AuthorizationSpec: + authorization_id: str + authorization_type: Literal["inference_repair", "offline_transformation"] + source_id: str + condition_question_reasons: Mapping[str, Mapping[str, str]] + authority: str + purpose: str + executable: bool + input_evidence_digests: Mapping[str, str] + expected_snapshot_digests: Mapping[str, str] + + +@dataclass(frozen=True) +class OverlaySpec: + overlay_id: str + base_run_path: Path # audit-only lookup location + base_condition_digest: str + base_realization_id: str + base_realization_digest: str + base_evidence_digests: Mapping[str, str] + base_validation_artifact_sha256: str + base_result_sha256: str | None + source_id: str + authorization_id: str + replacement_reasons: Mapping[str, str] + result_origin: ResultOriginSpec + lineage_notes: Mapping[str, Any] + implementation: Mapping[str, Any] + input_digest: str + preownership_output_digest: str + expected_evidence_status: EvidenceStatus + + +@dataclass(frozen=True) +class ImportSpec: + schema_version: Literal["choicebench.import-spec.v1"] + import_name: str + sources: tuple[SourceArtifactSpec, ...] + datasets: tuple[DatasetReferenceSpec, ...] + models: tuple[ImportModelSpec, ...] + methods: tuple[ImportMethodSpec, ...] + prompts: tuple[ImportPromptSpec, ...] + conditions: tuple[ImportConditionSpec, ...] + authorizations: tuple[AuthorizationSpec, ...] + overlays: tuple[OverlaySpec, ...] + metrics: tuple[str, ...] # validated strictly against BUILTIN_METRICS + provenance: Mapping[str, Any] + audit: Mapping[str, Any] + + +def load_import_spec(path: Path) -> ImportSpec: ... +def stable_import_projection(spec: ImportSpec) -> dict[str, Any]: ... +def import_spec_digest(spec: ImportSpec) -> str: ... +``` + +The same file defines strict enums/dataclasses for dataset/model/method/prompt, +the orthogonal `import_state`, `evidence_status`, `scope_disposition`, structured +origins, expected question IDs, qualifications, and overlay references. The +schema has no output-root or output-path field. + +- [ ] **Step 1: Write failing strict-schema tests** + + Parameterize tests for: a valid minimal YAML; every unknown top-level/nested + key; non-mapping YAML; missing/duplicate references; strict booleans and finite + numerics; invalid SHA-256; invalid enum; arbitrary YAML object tag; recursive + credential-named keys; an attempted `output_root`; single-byte distinct ASCII + delimiter/quote/escape; fixed UTF-8/strict decoding; and explicit unknown + provenance as `{value: null, reason: "not recorded by producer"}`. + Reject every metric not present in the closed `BUILTIN_METRICS` registry, + especially `module:Class`, dotted import paths, entry points, and filesystem + paths. Import specifications can never select a dynamic import target. + Cover ordered-column and structured-JSON option declarations, typed finite + numeric policies, and preserve/reject extra-field policies; reject incomplete + or contradictory option declarations and duplicate numeric-column rules. + Validate optional native-compatibility payloads against exact current + dataset/model/method/prompt allowed keys and recomputed full/short identities; + reject arbitrary claimed IDs or partial/mismatched payloads. Stage 1 fixtures + leave these fields null rather than fabricating native equivalence. + + Add these identity assertions: + + ```python + def test_audit_locations_do_not_change_import_spec_digest( + minimal_spec, minimal_overlay_spec + ): + moved = replace( + minimal_spec, + audit={"source_path": "/different/host", "imported_at": "later"}, + ) + assert import_spec_digest(moved) == import_spec_digest(minimal_spec) + + moved_source = replace( + minimal_spec, + sources=(replace(minimal_spec.sources[0], path=Path("/other/freeze/source.csv")),), + ) + assert import_spec_digest(moved_source) == import_spec_digest(minimal_spec) + + moved_base = replace( + minimal_overlay_spec, + overlays=(replace(minimal_overlay_spec.overlays[0], + base_run_path=Path("/other/home/runs/base")),), + ) + assert import_spec_digest(moved_base) == import_spec_digest(minimal_overlay_spec) + + + @pytest.mark.parametrize( + "change", + ["mapping", "dialect", "numeric_policy", "option_policy", + "extra_field_policy", "status", "scope"], + ) + def test_stable_mapping_fields_change_import_spec_digest(minimal_spec, change): + changed = mutate_identity_field(minimal_spec, change) + assert import_spec_digest(changed) != import_spec_digest(minimal_spec) + ``` + +- [ ] **Step 2: Run tests and observe the missing module** + + Run: + + ```bash + python -m pytest tests/importing/test_schema.py -v + ``` + + Expected: FAIL during collection with + `ModuleNotFoundError: No module named 'choicebench.importing'`. + +- [ ] **Step 3: Implement strict dataclass parsing and stable projection** + + Use `yaml.safe_load`, explicit allowed-key sets, the existing credential-key + detection/canonicalization rules, and `integrity_digest`. Reject rather than + coerce booleans, integers, non-finite values, enum spellings, paths in stable + provenance, and unknown keys. Export only stable public types from + `importing/__init__.py`. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/config/test_schema.py tests/importing/test_schema.py -v + python -m pytest -q + ``` + + Expected: all schema tests pass; the original 808 tests remain green. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/importing/__init__.py \ + src/choicebench/importing/schema.py tests/importing/test_schema.py + git commit -m "feat: add strict external import specification" + ``` + +--- + +### Task 3: Implement exact UTF-8 CSV logical-record scanning + +**Files:** + +- Create: `src/choicebench/importing/csv_adapter.py` +- Create: `tests/importing/test_csv_adapter.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class OpenedSource: + source_id: str + audit_path: Path + logical_path: str + data: bytes + sha256: str + + +@dataclass(frozen=True) +class LogicalRecordSpan: + index: int + start: int + end: int + terminator: bytes + + +@dataclass(frozen=True) +class SourceRow: + values: Mapping[str, str | None] + span: LogicalRecordSpan + raw_sha256: str + + +@dataclass(frozen=True) +class AdaptedTable: + columns: tuple[str, ...] + rows: tuple[SourceRow, ...] + source_sha256: str + + +class SourceAdapter(Protocol): + def parse( + self, source: OpenedSource, declaration: SourceArtifactSpec, *, strict: bool + ) -> AdaptedTable: ... + + +def scan_csv_logical_records( + data: bytes, dialect: CsvDialectSpec +) -> tuple[LogicalRecordSpan, ...]: ... + + +def parse_csv_source( + source: OpenedSource, declaration: SourceArtifactSpec, *, strict: bool +) -> AdaptedTable: ... +``` + +- [ ] **Step 1: Write byte-literal failing tests** + + Use `Path.write_bytes` under `tmp_path`, never a committed CSV, for UTF-8 BOM, + invalid UTF-8, LF/CRLF/CR, mixed terminators, forbidden mixed terminators, + final record without terminator, blank records, doubled quotes, explicit + escape characters, quoted embedded CR/LF/CRLF, and malformed unclosed quotes. + Assert exact spans and row hashes: + + ```python + def test_embedded_newline_span_is_exact(): + data = b'id,text\r\nq1,"line one\r\nline two"\r\nq2,end\n' + spans = scan_csv_logical_records(data, CsvDialectSpec()) + assert data[spans[1].start:spans[1].end] == \ + b'q1,"line one\r\nline two"\r\n' + assert spans[1].terminator == b"\r\n" + + + def test_literal_nan_is_not_implicit_null(opened_csv, source_spec): + table = parse_csv_source(opened_csv(b"id,value\nq1,NaN\n"), source_spec, + strict=True) + assert table.rows[0].values["value"] == "NaN" + ``` + + Also cover header order, duplicate header refusal, extra columns in normal and + strict modes, explicit ignored-column reasons, empty string versus declared + null, each typed numeric policy and malformed numeric strings, ordered-column + and structured-JSON choices, three-option rows, and six-option rows. Assert + effective CLI strictness may tighten but never loosen the declared extra-field + policy and that the effective policy is realization-identity-bearing. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_csv_adapter.py -v + ``` + + Expected: FAIL because `choicebench.importing.csv_adapter` is absent. + +- [ ] **Step 3: Implement the binary scanner and adapter** + + Scan original bytes before decoding. Recognize ASCII quote/escape bytes and + CRLF longest-first; separators inside quoted fields remain field bytes. Decode + each exact logical span with strict UTF-8, then parse using the declared CSV + semantics. Never use pandas in this adapter and never normalize malformed + evidence. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_csv_adapter.py -v + python -m pytest -q + ``` + + Expected: every dialect/raw-span case passes; no baseline regression. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/importing/csv_adapter.py \ + tests/importing/test_csv_adapter.py + git commit -m "feat: add deterministic CSV source adapter" + ``` + +**Intermediate gate:** An adversarial read-only reviewer checks the state +machine against embedded newlines, escaped quotes, BOM offsets, malformed EOF, +and every identity-bearing dialect field. + +--- + +### Task 4: Build expected-dataset references and trust-qualified snapshots + +**Files:** + +- Create: `src/choicebench/importing/dataset_reference.py` +- Create: `tests/importing/test_dataset_reference.py` +- Inspect: `src/choicebench/datasets.py`, `src/choicebench/pipeline/options.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class ExpectedDataset: + dataset_id: str + benchmark_name: str + split: str + artifact_id: str + artifact_digest: str + selection_id: str + selection_digest: str + frame: pd.DataFrame + selected_question_ids: tuple[str, ...] + reference_kind: Literal[ + "independent_input_snapshot", "profile_derived_reference_snapshot" + ] + trust_label: str + question_set_digest: str + snapshot_digest: str + derivation: Mapping[str, Any] + derivation_digest: str + limitations: tuple[str, ...] + + +def build_expected_dataset( + declaration: DatasetReferenceSpec, + opened_sources: Mapping[str, OpenedSource], +) -> ExpectedDataset: ... + + +def write_expected_snapshot(staged_run: Path, dataset: ExpectedDataset) -> dict: ... +def validate_expected_snapshot(run_dir: Path, record: Mapping[str, Any]) -> None: ... +``` + +- [ ] **Step 1: Write failing trust and variable-option tests** + + Cover independently supplied input plus selection IDs; profile-derived + reference requiring two declared independent source groups; cross-source + question/gold/option disagreement refusal; full derivation digest; unknown + publisher revision limitation; duplicate selected ID refusal; missing selected + ID refusal; stable order; six options; and ARC-style three options without a + phantom empty fourth option. Recompute and verify full imported dataset- + artifact and selection digests plus their short IDs; the selection projection + follows ChoiceBench's existing artifact/content/sample-identity/seed/filter + semantics using explicit unknown/null values where historical sampling inputs + are unavailable. + + Assert two differently located/mapped/trust-classified references that produce + identical normalized semantic rows have the same artifact/selection IDs but + different reference/derivation digests. Changing gold, ordered option text, + membership, or selection order changes the applicable semantic IDs. + + ```python + def test_reference_kinds_are_not_equivalent(independent_decl, derived_decl, sources): + independent = build_expected_dataset(independent_decl, sources) + derived = build_expected_dataset(derived_decl, sources) + assert independent.reference_kind == "independent_input_snapshot" + assert derived.reference_kind == "profile_derived_reference_snapshot" + assert independent.derivation_digest != derived.derivation_digest + assert "not independently authenticated" in derived.limitations + ``` + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_dataset_reference.py -v + ``` + + Expected: FAIL because the dataset-reference module is absent. + +- [ ] **Step 3: Implement snapshot construction and self-validation** + + Use `build_option_map`, `correct_option_for_row`, + `dataset_content_digest`, and existing atomic JSON/CSV helpers. Populate the + run snapshot from the declared reference, never result-side fields. Preserve + reference kind/trust/limitations in the realization-facing reference record, + not in the semantic dataset artifact. Build full + dataset-artifact and selection records with the same semantic boundaries and + full-digest/short-ID pattern as current ChoiceBench dataset ownership; do not + put source paths, parsing/mapping policy, trust classification, or audit fields + into either semantic identity. The dataset artifact binds normalized semantic + question/gold/ordered-option content plus benchmark/split; selection binds the + exact selected semantic rows/question sample identities and known selection + semantics. Thus a source mapping/trust/provenance change that yields identical + semantic rows changes the realization but not artifact/selection/condition; + an actual semantic snapshot or selection change changes all applicable + semantic identities. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/pipeline/test_options.py \ + tests/importing/test_dataset_reference.py -v + python -m pytest -q + ``` + + Expected: trust distinctions, cross-source checks, and variable options pass; + no baseline regression. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/importing/dataset_reference.py \ + tests/importing/test_dataset_reference.py + git commit -m "feat: add trusted import dataset references" + ``` + +--- + +## Concrete compatibility resolution discovered during planning + +Current `build_execution_plan()` hashes the deterministic relative +`prompt_snapshot_path` (`artifacts/prompts/`) into `condition_id`. +The approved specification simultaneously requires v3 conditions to retain the +existing scientific ID and to exclude operational paths. Removing this field +would change every native condition ID. The implementation must therefore keep +this one prompt-ID-derived, machine-independent value as a frozen compatibility +discriminator in the v3 semantic condition payload. Tests prove it is exactly +derived from `prompt_id`; arbitrary/absolute paths remain forbidden. No other +path field is grandfathered. This is the only architecture clarification in the +plan and is directly required by inspected v0.2 code. + +--- + +### Task 5: Separate semantic condition, realization, lineage, and origin identity + +**Files:** + +- Create: `src/choicebench/importing/identity.py` +- Create: `tests/importing/test_identity.py` +- Inspect: `src/choicebench/identity.py`, + `src/choicebench/cli/run_experiment.py:980` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class ImportSemanticRecords: + dataset_artifact: Mapping[str, Any] + selection: Mapping[str, Any] + model: Mapping[str, Any] + method: Mapping[str, Any] + prompt: Mapping[str, Any] + condition: Mapping[str, Any] + + +def make_semantic_condition( + *, identity: Mapping[str, Any], fields: Mapping[str, Any] +) -> dict[str, Any]: ... + + +def build_import_semantic_identity( + *, + condition: ImportConditionSpec, + dataset: ExpectedDataset, + model: ImportModelSpec, + method: ImportMethodSpec, + prompt: ImportPromptSpec, +) -> ImportSemanticRecords: ... + + +def make_lineage_component( + *, + operation_type: str, + question_id: str, + parent_digests: Sequence[str], + source_digests: Sequence[str], + authorization_digest: str | None, + implementation: Mapping[str, Any], + parameters: Mapping[str, Any], + input_digest: str, + preownership_output_digest: str, + prediction_origin: str, +) -> dict[str, Any]: ... + + +def make_result_origin( + *, + derivation_origin: Literal[ + "native_execution", "external_import", "repair_overlay", + "offline_transformation" + ], + row_assignments: Sequence[tuple[str, str, str]], +) -> dict[str, Any]: ... + + +def importer_implementation_identity( + *, adapter: Callable[..., Any], validator: Callable[..., Any] +) -> dict[str, Any]: ... + + +def make_realization( + *, + condition_id: str, + condition_digest: str, + identity: Mapping[str, Any], + fields: Mapping[str, Any], +) -> dict[str, Any]: ... +``` + +`row_assignments` entries are `(question_id, prediction_origin, +prediction_lineage_id)`. Records contain full SHA-256 digests plus short IDs. + +- [ ] **Step 1: Write failing layered-identity tests** + + Use one base scientific identity and parameterize changes. Assert: + + ```python + @pytest.mark.parametrize( + "field", + ["source_sha256", "mapping", "dialect", "evidence_status", "scope", + "authorization_digest", "overlay_sha256", "prediction_origins"], + ) + def test_realization_changes_do_not_change_condition(base_records, field): + original_condition, original_realization = base_records + changed_condition, changed_realization = mutate_realization(base_records, field) + assert changed_condition["condition_id"] == original_condition["condition_id"] + assert changed_realization["realization_id"] != original_realization["realization_id"] + + + @pytest.mark.parametrize( + "field", ["selection_id", "model_id", "method_id", "prompt_id", "seed", + "protocol_settings"] + ) + def test_scientific_change_changes_condition(base_records, field): + changed = mutate_semantic_condition(base_records[0], field) + assert changed["condition_id"] != base_records[0]["condition_id"] + ``` + + Also test base/repair/transformation sharing one condition; rejection of bare + `mixed`; origin set/count and ordered mapping; offline rematching retaining + underlying prediction origin; lineage component exclusion of child IDs, + paths, timestamps, and final CSV hashes; audit changes having no effect; + importer package/version plus adapter/validator source identities present and + identity-bearing; the grandfathered prompt path preserving the exact current + `condition_id`; and arbitrary path injection refusal. + + Add a generic-spec integration fixture that builds dataset-artifact, + selection, external model, historical method, prompt, and condition records + from the Task 2 dataclasses. Assert the final condition payload has the exact + current keys (`benchmark` with artifact/selection IDs, `preflight`, + `model_id`, `method_id`, `prompt_id`, frozen prompt discriminator, `seed`, and + applicable `protocol_settings`). Scientific input changes alter the proper + child/full/condition digests; source bytes, dialect, mapping, status, scope, + and audit-only changes leave that payload fixed and alter only realization + and downstream identities. Unknown historical implementation/revision/prompt + values remain explicit null+reason records and are never replaced by native + ChoiceBench identities. + + Freeze these canonical child payload contracts in literal tests: + + ```python + imported_dataset_payload = { + "schema_version": "choicebench.semantic-dataset.v1", + "benchmark": benchmark_name, + "split": split, + "content_digest": normalized_semantic_content_digest, + } + imported_selection_payload = { + "artifact_id": artifact_id, + "content_digest": selected_semantic_content_digest, + "sample_identities": ordered_sample_identities, + "seed": selection_seed_or_explicit_unknown, + "n_samples": selection_n_samples_or_explicit_unknown, + "subject_filter": sorted_subject_filter, + } + imported_model_payload = { + "schema_version": "choicebench.semantic-model.v1", + "backend": backend_or_explicit_unknown, + "provider": provider_or_explicit_unknown, + "model": display_name, + "revision": revision_or_explicit_unknown, + "effective_parameters": canonical_parameters, + } + imported_method_payload = { + "schema_version": "choicebench.semantic-method.v1", + "name": exact_historical_name, + "effective_params": canonical_parameters, + "preflight": preflight_identity, + "implementation": implementation_or_explicit_unknown, + } + imported_prompt_payload = { + "schema_version": "choicebench.semantic-prompt.v1", + "template_identity": identity_or_explicit_unknown, + "template_digest": digest_or_explicit_unknown, + "template_contents": contents_or_explicit_unknown, + } + ``` + + Full digests cover these exact payloads; short prefixes remain `ds`, `sel`, + `model`, `method`, and `prompt`. These imported payloads are the documented + fallback when no exact current-native identity is recoverable. When a + synthetic external specification supplies a fully validated + `native_compatibility_identity` for every child, use the exact current native + canonical payload instead, compare its child and condition identities with + `build_execution_plan()`, and require equality. + When historical fields are unknown or the imported semantic-dataset schema is + necessarily content-derived rather than current `DatasetSpec`-derived, require + explicit documented divergence while preserving the same condition payload + structure and never pretending equivalence. Mapping, trust, and source + provenance remain realization-only in either case. + + Parameterize every `CsvDialectSpec` field individually through final + realization construction and require each parsing-affecting change to alter + `realization_id` while preserving `condition_id`. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_identity.py -v + ``` + + Expected: FAIL because the importer identity module is absent. + +- [ ] **Step 3: Implement canonical full/short identity records** + + Reuse `canonicalize`, `stable_digest`, `integrity_digest`, and `short_id`. + Reuse `choicebench.provenance.implementation_identity` for registered importer, + adapter, validator, and transformation callables; never accept their active + code identity from untrusted YAML. + Explicitly allow only the declared identity keys at every layer. Compute in + this order: semantic condition; parent-derived lineage components; + pre-ownership output; row-origin mapping; realization. Never accept a current + experiment ID, result-artifact ID, final CSV SHA, output path, or audit field + in realization identity. + + Build the imported dataset/model/method/prompt records with the same semantic + payload boundaries and full-digest/short-ID pattern as native records. + Historical method identity is declared provenance, not a claim that the + current ChoiceBench method implementation produced the rows. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/test_condition_grid.py tests/importing/test_identity.py -v + python -m pytest -q + ``` + + Expected: all layered-identity cases pass and every existing native short ID + test remains unchanged. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/importing/identity.py tests/importing/test_identity.py + git commit -m "feat: separate condition and realization identity" + ``` + +**Intermediate gate:** Independent identity/provenance review. The reviewer must +construct the identity dependency graph and confirm there is no child/self edge +and no source/provenance field in the semantic condition. + +--- + +### Task 6: Introduce manifest v3 while preserving manifest v2 exactly + +**Files:** + +- Modify: `src/choicebench/manifest.py` +- Create: `tests/importing/test_manifest_v3.py` +- Modify: `tests/importing/test_manifest_v2_compat.py` +- Modify: `tests/test_publication_identity.py` + +**Interfaces:** + +```python +PROTOCOL_VERSION = "choicebench.protocol.v2" # frozen public legacy default +PROTOCOL_V3_VERSION = "choicebench.protocol.v3" +MANIFEST_SCHEMA_VERSION = "choicebench.manifest.v2" # frozen public legacy default +MANIFEST_V3_SCHEMA_VERSION = "choicebench.manifest.v3" +RUN_STATE_SCHEMA_VERSION = "choicebench.run-state.v2" # frozen public legacy default +RUN_STATE_V3_SCHEMA_VERSION = "choicebench.run-state.v3" + + +@dataclass(frozen=True) +class ManifestView: + manifest: Mapping[str, Any] + schema_version: str + semantic_conditions: Mapping[str, Mapping[str, Any]] + realizations: Mapping[str, Mapping[str, Any]] + realization_ids_by_condition: Mapping[str, tuple[str, ...]] + legacy_v2: bool + + +def make_manifest_v3( + payload: Mapping[str, Any], *, audit: Mapping[str, Any] | None = None +) -> dict[str, Any]: ... + + +def validate_manifest_v2(manifest: Mapping[str, Any]) -> None: ... +def validate_manifest_v3(manifest: Mapping[str, Any]) -> None: ... +def validate_manifest(manifest: Mapping[str, Any]) -> None: ... +def normalize_manifest(manifest: Mapping[str, Any]) -> ManifestView: ... +``` + +Keep current `PROTOCOL_VERSION`, `MANIFEST_SCHEMA_VERSION`, +`RUN_STATE_SCHEMA_VERSION`, `build_manifest_payload()`, and `make_manifest()` as +v2-compatible public entry points even after Task 14. Importer and new native +creation call explicit v3 functions/constants. `normalize_manifest()` synthesizes exactly one in-memory +native realization for v2 without rewriting bytes or changing stored IDs. + +- [ ] **Step 1: Write failing v3 and expanded v2 compatibility tests** + + Test `semantic_conditions` and `realizations` as separate unique tables; + condition full-digest recomputation; realization full-digest recomputation; + stable `semantic_grid_digest`; realization-to-condition full-digest binding; + fixed realization result/sidecar/validation paths; semantic validation digest; + explicit structured origin; rejection + of result-artifact ID/final CSV SHA in a manifest; audit exclusion from + experiment identity; unsafe paths; duplicate IDs; and v3 run state keyed by + realization. + + Extend the Task 1 tests to ensure v2 normalization preserves stored + experiment/condition/result paths, infers native origin only in memory, and + leaves all run bytes unchanged. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_manifest_v3.py \ + tests/importing/test_manifest_v2_compat.py tests/test_publication_identity.py -v + ``` + + Expected: new v3 tests fail because the symbols/schema are absent; every Task + 1 v2 test still passes. + +- [ ] **Step 3: Implement exact schema dispatch** + + Dispatch only on the exact `schema_version`, never on missing fields. Preserve + current v2 validation code as `validate_manifest_v2`. Add child-ID/full-digest + recomputation for v3, safe fixed paths, separate audit integrity, and + realization-keyed v3 `initial_run_state`/`validate_run_state`. Add normalized + access helpers instead of rewriting callers prematurely. + +- [ ] **Step 4: Run compatibility, publication, and full tests** + + ```bash + python -m pytest tests/importing/test_manifest_v3.py \ + tests/importing/test_manifest_v2_compat.py tests/test_publication_identity.py \ + tests/test_release_adversarial.py -v + python -m pytest -q + ``` + + Expected: v2 byte/identity tests and v3 structural tests pass; no baseline + failures. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/manifest.py tests/importing/test_manifest_v3.py \ + tests/importing/test_manifest_v2_compat.py tests/test_publication_identity.py + git commit -m "feat: add manifest v3 compatibility layer" + ``` + +**Intermediate gate:** Independent manifest review covering v2 dispatch, +full-child-digest verification, path safety, audit exclusion, and the absence of +result-artifact identity from the immutable manifest. + +--- + +### Task 7: Publish non-circular realization result artifacts + +**Files:** + +- Modify: `src/choicebench/io/writers.py` +- Create: `tests/importing/test_result_artifact.py` +- Modify: `tests/io/test_writers.py` +- Modify: `tests/test_publication_identity.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class PreparedResultArtifact: + result_path: str + metadata_path: str + csv_bytes: bytes + metadata: Mapping[str, Any] + + +def prepare_manifest_result( + results: Sequence[Mapping[str, Any]], + *, + manifest: Mapping[str, Any], + realization_id: str, +) -> PreparedResultArtifact: ... + + +def publish_manifest_result( + prepared: PreparedResultArtifact, *, run_dir: Path +) -> tuple[Path, Mapping[str, Any], bool]: ... + + +def validate_result_artifact( + result_path: Path, + *, + manifest: Mapping[str, Any] | None = None, + realization: Mapping[str, Any] | None = None, +) -> dict[str, Any]: ... +``` + +The returned boolean means an identical artifact already existed. Preserve the +legacy/programmatic `write_run_results()` contract and result-artifact-v1 +validation for v2. + +Result-artifact-v2 metadata contains exact experiment/condition/realization full +digests; result file SHA-256/size/row and ordered-column/question digests; +dataset artifact/selection/expected-snapshot digests; model/method/prompt full +digests; import-spec and source-byte digests; evidence-index digest; +realization-validation semantic digest and exact file SHA-256; authorization and +lineage digests; structured derivation origin; constituent prediction-origin +set/count and ordered per-question origin/lineage digest; result-artifact full +digest/short ID; and a sidecar self-digest. Audit locations/timestamps are not +included. Imported realizations require the import/source/evidence bindings; +native-v3 realizations encode those fields as explicit not-applicable records +with reasons and still require dataset, validation, origin, lineage, and result +bindings. They are never silently omitted. + +- [ ] **Step 1: Write failing result-artifact-v2 tests** + + Test exact CSV rendering in memory; mandatory uniform condition, realization, + experiment, dataset/model/method/prompt, split, `prediction_origin`, and + `prediction_lineage_id`; result ID/digest bound to the already fixed + experiment/realization and exact file bytes; ordered columns/question/origin + digests; no result ID inside CSV/manifest; correct realization paths; full + sidecar self-validation; identical no-op; one-byte divergence refusal before + writing; missing half-pair refusal; and v1 allowed only for legacy v2. Add a + tampering/mismatch case for every v2 metadata binding above and compare it + against the manifest realization, validation artifact, evidence index, exact + CSV, and run state rather than merely trusting the sidecar's own digest. + + ```python + def test_result_identity_is_computed_after_manifest(v3_manifest, native_rows): + prepared = prepare_manifest_result( + native_rows, manifest=v3_manifest, + realization_id=native_rows[0]["realization_id"], + ) + assert "result_artifact_id" not in v3_manifest + assert prepared.metadata["experiment_id"] == v3_manifest["experiment_id"] + assert prepared.metadata["file_sha256"] == hashlib.sha256( + prepared.csv_bytes + ).hexdigest() + assert prepared.metadata["result_artifact_id"].startswith("result_") + ``` + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_result_artifact.py \ + tests/io/test_writers.py tests/test_publication_identity.py -v + ``` + + Expected: new API tests fail; legacy writer tests pass. + +- [ ] **Step 3: Implement preparation, v2 sidecar, and collision refusal** + + Render once to UTF-8 bytes, validate rows, compute the sidecar payload and + `result_artifact_digest`, then add its short ID and sidecar integrity digest. + `publish_manifest_result` may write only inside a staging tree or to an absent + pair. If either target exists, validate and compare the entire proposed graph; + return no-op only for exact equality. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_result_artifact.py \ + tests/io/test_writers.py tests/test_publication_identity.py -v + python -m pytest -q + ``` + + Expected: v1/v2 sidecar dispatch and non-circular identity cases pass; no + baseline regression. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/io/writers.py tests/importing/test_result_artifact.py \ + tests/io/test_writers.py tests/test_publication_identity.py + git commit -m "feat: publish realization result artifacts" + ``` + +**Intermediate gate:** Independent provenance/collision review. Require a +reviewer-produced dependency graph proving +condition -> realization -> experiment -> CSV/result artifact, never the +reverse. + +--- + +### Task 8: Validate source rows by question identity + +**Files:** + +- Create: `src/choicebench/importing/validation.py` +- Create: `tests/importing/test_validation.py` +- Inspect: `src/choicebench/pipeline/options.py`, + `src/choicebench/scoring/scorer.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class ValidationFinding: + code: str + source_id: str + field: str | None + question_id: str | None + record_span: LogicalRecordSpan | None + raw_row_sha256: str | None + message: str + + +@dataclass(frozen=True) +class ValidatedSourceRows: + source_id: str + rows_by_question_id: Mapping[str, SourceRow] + expected_question_ids: tuple[str, ...] + findings: tuple[ValidationFinding, ...] + findings_digest: str + + +def validate_source_rows( + table: AdaptedTable, + *, + condition: ImportConditionSpec, + expected: ExpectedDataset, +) -> ValidatedSourceRows: ... +``` + +- [ ] **Step 1: Write failing membership and scientific-content tests** + + Add one focused test per required failure: missing required field; null ID; + duplicate ID with both raw spans; missing ID; unexpected ID; gold mismatch; + option count mismatch; ordered option text mismatch; correct-option mismatch; + invalid parsed prediction for the row's actual option set; recomputable + correctness mismatch; malformed numeric/NaN/infinity; model/benchmark/split/ + method ownership mismatch; status inconsistent with observed coverage; and + source-schema/extra-column violation. + + Parameterize option coverage with three-, four-, and six-option rows. Assert no + input row is dropped, padded, deduplicated, reordered, repaired, or converted + from malformed to valid. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_validation.py -v + ``` + + Expected: FAIL because `validation.py` is absent. + +- [ ] **Step 3: Implement the fail-closed validator** + + Join only by exact string `question_id`. Compare all source-provided + question/gold/option fields with the expected snapshot, but never let source + values override it. Accumulate all deterministic findings in source-record + order and raise a single actionable validation error for undeclared defects. + Sanitize source text in messages while preserving field, question, span, and + digest evidence. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_validation.py \ + tests/pipeline/test_options.py tests/scoring/test_scorer.py -v + python -m pytest -q + ``` + + Expected: all question-keyed failures are detected and the baseline remains + green. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/importing/validation.py \ + tests/importing/test_validation.py + git commit -m "feat: validate imported rows by question identity" + ``` + +--- + +### Task 9: Preserve evidence status, malformed bytes, and extension fields + +**Files:** + +- Modify: `src/choicebench/importing/validation.py` +- Create: `tests/importing/test_evidence_status.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class RealizationValidation: + evaluable_rows: tuple[Mapping[str, Any], ...] + findings: tuple[ValidationFinding, ...] + validation_digest: str + computed_evidence_status: Literal[ + "complete", "qualified", "partial", "malformed", "recoverable", "failed" + ] + evaluable: bool + qualifications: tuple[Mapping[str, Any], ...] + limitations: tuple[Mapping[str, Any], ...] + defect_question_ids: tuple[str, ...] + + +@dataclass(frozen=True) +class PreparedValidationArtifact: + relative_path: str + json_bytes: bytes + file_sha256: str + validation_digest: str + + +def normalize_realization_rows( + validated: ValidatedSourceRows, + *, + condition: ImportConditionSpec, + expected: ExpectedDataset, + experiment_id: str | None = None, + realization_id: str | None = None, +) -> RealizationValidation: ... + + +def prepare_realization_validation_artifact( + validation: RealizationValidation, + *, + realization: Mapping[str, Any], + evidence_records: Sequence[Mapping[str, Any]], +) -> PreparedValidationArtifact: ... + + +def write_realization_validation_artifact( + staged_run: Path, prepared: PreparedValidationArtifact +) -> tuple[Path, bool]: ... + + +def validate_realization_validation_artifact( + path: Path, + *, + manifest: Mapping[str, Any], + realization: Mapping[str, Any], + expected_sha256: str, +) -> Mapping[str, Any]: ... +``` + +The first call before manifest construction produces a canonical pre-ownership +payload/digest. The second call after experiment and realization IDs are fixed +injects ownership without changing predictions or evidence decisions. + +- [ ] **Step 1: Write failing status/extension tests** + + Cover complete, qualified, partial, malformed, recoverable, and failed + evidence; complete+excluded and malformed+excluded; qualifications preserved + beside metrics; exact declared defect IDs and reasons; undeclared defect + refusal; no evaluable rows for partial/malformed/recoverable/failed/excluded/ + held/superseded or `aggregate_only` sources; unknown/extra fields preserved in namespaced + `external_fields_json`; explicit ignored fields and reasons; credential-shaped + field refusal; literal null/NaN preservation; and method-specific nested JSON + round-trip. + + For malformed evidence assert: + + ```python + assert finding.record_span == LogicalRecordSpan(2, 19, 47, b"\r\n") + assert finding.raw_row_sha256 == sha256(source_bytes[19:47]).hexdigest() + assert validation.evaluable_rows == () + ``` + + Assert deterministic JSON bytes at + `artifacts/imports/validation/.json`, self-digest validation, + exact realization/evidence/finding bindings, checksum change on any finding or + raw-row digest change, atomic temp-file publication inside staging, identical + no-op/divergent refusal, and refusal for traversal, unsafe path, + realization/evidence mismatch, tampering, or a raw-span digest that does not + reproduce from the archived evidence bytes. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_evidence_status.py -v + ``` + + Expected: FAIL because normalization/status behavior is not implemented. + +- [ ] **Step 3: Implement normalization without evidence repair** + + Fill evaluative question text/options/gold only from `ExpectedDataset`; retain + original mapped and extension fields separately. Recompute evidence status + from exact coverage/defects and refuse a declaration mismatch. Produce + evaluable rows only for imported+included complete/qualified realizations. + Never fabricate request IDs, timestamps, usage, prompts, responses, provider + fields, commits, revisions, or seeds. + + Prepare one exact validation artifact for every realization, including + evidence-only realizations. Its semantic `validation_digest` is available to + realization identity before ownership; after the realization is fixed, its + exact serialized file SHA-256 is recorded downstream in run state and any + result sidecar, never back-propagated into the realization/manifest identity. + +- [ ] **Step 4: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_validation.py \ + tests/importing/test_evidence_status.py -v + python -m pytest -q + ``` + + Expected: status eligibility and exact malformed evidence pass; no baseline + regression. + +- [ ] **Step 5: Commit** + + ```bash + git add src/choicebench/importing/validation.py \ + tests/importing/test_evidence_status.py + git commit -m "feat: preserve imported evidence and status" + ``` + +--- + +### Task 10: Validate typed authorization and derive immutable overlays + +**Files:** + +- Create: `src/choicebench/importing/authorization.py` +- Create: `src/choicebench/importing/overlays.py` +- Create: `tests/importing/test_authorization.py` +- Create: `tests/importing/test_overlays.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class ValidatedAuthorizationBundle: + authorization_id: str + authorization_digest: str + authorization_type: Literal["inference_repair", "offline_transformation"] + grants: Mapping[str, Mapping[str, str]] # condition_digest -> qid -> reason + authority: str + purpose: str + executable: bool + source_sha256: str + input_evidence_digests: Mapping[str, str] + expected_snapshot_digests: Mapping[str, str] + + +@dataclass(frozen=True) +class ValidatedAuthorization: + bundle_id: str + bundle_digest: str + authorization_type: Literal["inference_repair", "offline_transformation"] + condition_digest: str + question_reasons: Mapping[str, str] + authority: str + purpose: str + executable: bool + source_sha256: str + input_evidence_digests: Mapping[str, str] + expected_snapshot_digest: str + + +@dataclass(frozen=True) +class VerifiedBaseRealization: + condition_id: str + condition_digest: str + realization_id: str + realization_digest: str + evidence_index_digest: str + evidence_source_digests: Mapping[str, str] + validation_artifact_sha256: str + result_sha256: str | None + rows_by_question_id: Mapping[str, Mapping[str, Any]] + prediction_origins: Mapping[str, str] + + +@dataclass(frozen=True) +class DerivedRealizationPayload: + condition_id: str + condition_digest: str + preownership_rows: tuple[Mapping[str, Any], ...] + lineage_components: tuple[Mapping[str, Any], ...] + result_origin: Mapping[str, Any] + evidence_status: str + replacement_question_ids: tuple[str, ...] + + +def validate_authorization_bundle( + declaration: AuthorizationSpec, + *, + opened_source: OpenedSource, + condition_digests: Mapping[str, str], + expected: Mapping[str, ExpectedDataset], +) -> ValidatedAuthorizationBundle: ... + + +def authorization_for_condition( + bundle: ValidatedAuthorizationBundle, *, condition_digest: str +) -> ValidatedAuthorization: ... + + +def derive_overlay( + *, + base: VerifiedBaseRealization, + overlay: OverlaySpec, + authorization: ValidatedAuthorization, + overlay_table: AdaptedTable, + expected: ExpectedDataset, +) -> DerivedRealizationPayload: ... +``` + +- [ ] **Step 1: Write failing typed-authorization tests** + + `inference_repair` must require exact condition/question membership, + `queue_disposition=approved`, `execution_authority=authoritative`, and + `executable=true`. Assert refusal for held, excluded, forensic, absent, + false-executable, and offline records. `offline_transformation` must require a + separately opened checksum-verified source with exact IDs/reasons/authority/ + purpose, input-evidence digests, expected-snapshot digest, and + `inference_executable=false`; an overlay/spec cannot authorize its own IDs. + Assert the types cannot substitute for one another or bind a different base + condition/snapshot/evidence graph. + + `derive_overlay` requires exact equality between + `overlay.base_evidence_digests`, the authorization slice's applicable + `input_evidence_digests`, and the realization-scoped source/blob digest mapping + returned by `verify_import_run`; no evidence-index digest alone is accepted as + a substitute. Refuse missing, extra, or cross-realization evidence digests. + + Test both a one-condition bundle and a single content-addressed six-condition/ + 18-pair bundle. The aggregate digest binds every condition digest, question, + reason, evidence/snapshot digest, purpose, authority, and executable flag; + each condition-scoped slice retains the same immutable bundle digest and + cannot add or borrow grants from another condition. + +- [ ] **Step 2: Write failing overlay/lineage tests** + + Cover a valid evidence-only base with no result checksum; a valid evaluable + base; exact replacement subset; unauthorized, duplicate, unexpected, missing, + and conflicting replacement IDs; multiple overlays targeting one ID; + gold/options/ownership mismatch; status recomputation; base evidence/result + checksum preservation; base -> authorization -> overlay/transformation -> + resulting lineage; mixed historical/native repair origins with exact per-row + mapping; external repair origin; offline rematching preserving underlying + prediction origin; base condition/realization/validation-artifact digest + verification; and transformation pre-ownership I/O/code digests. Require a + declared origin assignment for every replacement ID: native repair rows use + `native_inference`, externally generated repair rows use + `external_repair_inference`, and offline rematching retains the base response's + prediction origin while recording `offline_transformation` as derivation. + Reject missing/extra assignments and any bare `mixed` value; identity-bind the + complete per-row mapping and stable lineage notes. + + `derive_overlay` accepts only a `VerifiedBaseRealization`, never a filesystem + path or unverified caller dictionary. Unit tests construct this frozen trusted + value and prove the pure function cannot mutate it. Task 12 integration + snapshots every base-run file before/after full-graph verification and + derivation and requires byte equality. The derived payload shares the semantic + condition only when model/dataset/method/prompt/protocol are unchanged. + +- [ ] **Step 3: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_authorization.py \ + tests/importing/test_overlays.py -v + ``` + + Expected: FAIL because authorization/overlay modules are absent. + +- [ ] **Step 4: Implement typed verification and pure derivation** + + Bundle verification consumes an independently opened authorization artifact. + The overlay function is pure with respect to the filesystem and accepts only + the trusted base value produced later by the engine's full graph verifier. It + returns a derived pre-ownership payload and never opens/writes a base run or + performs inference. Use the Task 5 lineage and origin builders, excluding + child IDs from lineage identity. + +- [ ] **Step 5: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_authorization.py \ + tests/importing/test_overlays.py tests/importing/test_identity.py -v + python -m pytest -q + ``` + + Expected: typed authority, overlay conflicts, mixed origins, and base + immutability all pass. + +- [ ] **Step 6: Commit** + + ```bash + git add src/choicebench/importing/authorization.py \ + src/choicebench/importing/overlays.py \ + tests/importing/test_authorization.py tests/importing/test_overlays.py + git commit -m "feat: add authorized immutable import overlays" + ``` + +**Intermediate gate:** Independent authorization/overlay review. The reviewer +must attempt self-authorization, cross-type authorization, duplicate/conflicting +replacement, base mutation, and a bare `mixed` origin. + +--- + +### Task 11: Add verified evidence storage and atomic staged publication + +**Files:** + +- Modify: `src/choicebench/infra/artifacts.py` +- Create: `src/choicebench/importing/evidence.py` +- Create: `src/choicebench/importing/transaction.py` +- Create: `tests/importing/test_evidence.py` +- Create: `tests/importing/test_transaction.py` +- Modify: `tests/test_release_adversarial.py` + +**Interfaces:** + +```python +def open_verified_source( + declaration: SourceArtifactSpec, + *, + containment_root: Path | None = None, +) -> OpenedSource: ... + + +def evidence_blob_path(staged_run: Path, sha256: str, format_name: str) -> Path: ... +def write_evidence_blob( + staged_run: Path, source: OpenedSource, references: Sequence[str] +) -> dict[str, Any]: ... +def validate_evidence_blob(run_dir: Path, record: Mapping[str, Any]) -> None: ... +def write_evidence_index( + staged_run: Path, records: Sequence[Mapping[str, Any]] +) -> Mapping[str, Any]: ... +def validate_evidence_index( + run_dir: Path, expected_digest: str +) -> Mapping[str, Any]: ... + + +def fsync_tree(root: Path) -> None: ... +def atomic_publish_directory_no_replace(source: Path, destination: Path) -> None: ... + + +class ImportTransaction: + def __init__(self, *, runs_dir: Path, run_id: str): ... + def __enter__(self) -> "ImportTransaction": ... + @property + def staged_run(self) -> Path: ... + def publish(self, validator: Callable[[Path], None]) -> Path: ... + def __exit__(self, exc_type, exc, traceback) -> None: ... +``` + +- [ ] **Step 1: Write failing verified-source/evidence tests** + + Cover independently computed checksum mismatch; source mutation/replacement; + same opened bytes used for hash and parse; absolute user source; profile-root + containment; traversal, directory, device, and unsafe symlink refusal; suffix + derived from validated format; within-run dedup; cross-run independent copy; + tampered blob/sidecar/index; divergent digest-path refusal; exact malformed row + evidence; no recursive import; and secret-shaped metadata refusal. + +- [ ] **Step 2: Write failing transaction tests** + + Assert the established `.locks/.manifest.lock`; same-filesystem sibling + owner-marked staging; full validator invoked before publication; destination + absent; real atomic no-replace behavior; fsync calls; failure leaves final path + absent; stale staging is never recognized as a run; safe cleanup touches only + its owner-marked directory; existing final run is verify-only and creates no + staging; identical final graph no-ops; divergent graph refuses unchanged. + + Linux implementation tests should exercise libc `renameat2(..., + RENAME_NOREPLACE)` through a small private wrapper. Mock absence of that symbol + and assert fail-closed behavior; do not fall back to `os.replace` or an + existence-check race. + +- [ ] **Step 3: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_evidence.py \ + tests/importing/test_transaction.py tests/test_release_adversarial.py -v + ``` + + Expected: FAIL because evidence/transaction APIs do not exist. + +- [ ] **Step 4: Implement same-bytes source opening and run-local CAS** + + Open only declared regular files; read once; hash `data`; pass `data` forward. + Store only referenced row-level bytes at + `artifacts/imports/evidence/sha256//.` + inside staging. Sidecars bind byte size/format/references; identical bytes in + one run reuse the blob. Write each blob, sidecar, and the self-digested + evidence index through a temporary file, flush/fsync, revalidate exact bytes, + and atomically rename within staging. Never accept a divergent existing blob, + sidecar, or index. + +- [ ] **Step 5: Implement staged no-replace publication** + + Acquire the manifest lock before examining final state. For a new run, build + under a sibling staging directory, validate/fsync the complete tree, and call + the no-replace primitive once. For an existing run, invoke the supplied full + graph verifier only; exact equality no-ops and divergence raises. A report + outside the run cannot affect commit state. + +- [ ] **Step 6: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_evidence.py \ + tests/importing/test_transaction.py tests/test_release_adversarial.py -v + python -m pytest -q + ``` + + Expected: source, CAS, crash, symlink, no-replace, and idempotence tests pass; + no baseline regression. + +- [ ] **Step 7: Commit** + + ```bash + git add src/choicebench/infra/artifacts.py \ + src/choicebench/importing/evidence.py \ + src/choicebench/importing/transaction.py \ + tests/importing/test_evidence.py tests/importing/test_transaction.py \ + tests/test_release_adversarial.py + git commit -m "feat: stage external imports atomically" + ``` + +**Intermediate gate:** Independent filesystem/security review on source opening, +staging containment, owner markers, symlink behavior, no-replace semantics, +cleanup scope, and verify-only final runs. + +--- + +### Task 12: Orchestrate generic dry-run and real imports + +**Files:** + +- Create: `src/choicebench/importing/engine.py` +- Modify: `src/choicebench/importing/__init__.py` +- Create: `tests/importing/test_engine.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class ImportRequest: + spec: ImportSpec + run_id: str + workspace_root: Path + strict: bool + overlays: tuple[OverlaySpec, ...] = () + + +@dataclass(frozen=True) +class ImportPlan: + manifest: Mapping[str, Any] + expected_datasets: Mapping[str, ExpectedDataset] + realizations: Mapping[str, RealizationValidation] + evidence_records: tuple[Mapping[str, Any], ...] + report_identity: Mapping[str, Any] + + +@dataclass(frozen=True) +class VerifiedImportRun: + manifest: Mapping[str, Any] + manifest_digest: str + realizations: Mapping[str, VerifiedBaseRealization] + + +@dataclass(frozen=True) +class ImportCounts: + conditions: Mapping[str, int] # exact keys listed below + realizations: Mapping[str, int] + rows: Mapping[str, int] + evidence_status: Mapping[str, int] + scope_disposition: Mapping[str, int] + prediction_origin: Mapping[str, int] + derivation_origin: Mapping[str, int] + source_classification: Mapping[str, int] + column_disposition: Mapping[str, int] + authorization: Mapping[str, int] + + +@dataclass(frozen=True) +class DefectReport: + total_findings: int + by_code: Mapping[str, int] + question_ids_by_code: Mapping[str, tuple[str, ...]] + findings_digest: str + + +@dataclass(frozen=True) +class ChecksumRecord: + logical_id: str + expected_sha256: str | None + actual_sha256: str + matches: bool | None + + +@dataclass(frozen=True) +class ArtifactChecksumRecord: + logical_path: str + file_sha256: str + size: int + semantic_digest: str | None + + +@dataclass(frozen=True) +class ImportChecksumReport: + all_referenced_sources_match: bool + sources: Mapping[str, ChecksumRecord] + expected_snapshots: Mapping[str, ArtifactChecksumRecord] + evidence_index: ArtifactChecksumRecord | None + validation_artifacts: Mapping[str, ArtifactChecksumRecord] + results: Mapping[str, ArtifactChecksumRecord] + + +@dataclass(frozen=True) +class ResultIdentityReport: + result_artifact_id: str + result_artifact_digest: str + file_sha256: str + + +@dataclass(frozen=True) +class ImportIdentityProjection: + import_spec_digest: str + semantic_grid_digest: str + experiment_id: str | None + experiment_digest: str | None + condition_digests: Mapping[str, str] + realization_digests: Mapping[str, str] + result_artifacts: Mapping[str, ResultIdentityReport] + authorization_bundle_digests: Mapping[str, str] + lineage_digests: Mapping[str, str] + + +@dataclass(frozen=True) +class ImportFailureReport: + code: str + source_id: str | None + field: str | None + question_id: str | None + record_span: Mapping[str, int | str] | None + raw_row_sha256: str | None + sanitized_message: str + + +@dataclass(frozen=True) +class ImportAuditReport: + source_location: str | None + output_root: str | None + imported_at: str + + +@dataclass(frozen=True) +class ImportReport: + schema_version: Literal["choicebench.import-report.v1"] + import_state: Literal["validated", "imported", "failed"] + wrote_artifacts: bool + idempotent_noop: bool + run_id: str + experiment_id: str | None + counts: ImportCounts + defects: DefectReport + checksums: ImportChecksumReport + identity_projection: ImportIdentityProjection + profile: Mapping[str, Any] + failures: tuple[ImportFailureReport, ...] + audit: ImportAuditReport + + +def build_import_plan(request: ImportRequest) -> ImportPlan: ... +def verify_import_run( + run_dir: Path, expected: ImportPlan | None = None +) -> VerifiedImportRun: ... +def execute_import(request: ImportRequest, *, dry_run: bool = False) -> ImportReport: ... +def serialize_import_report(report: ImportReport) -> dict[str, Any]: ... +``` + +The serialized JSON uses these fields exactly at top level. `profile` is `{}` +for generic imports; the Stage 1 profile uses stable `matrix`, `queues`, +`offline_authority`, `datasets`, and `source_schema_count` keys. `failures` +contains deterministic code/source/field/question/span/digest plus sanitized +message entries and is empty on success. Tracebacks, credentials, raw untrusted +text, absolute temporary paths, and timestamps appear in neither failures nor +identity projection; safe machine-local locations/timestamps may appear only +under `audit`. + +Nested report-v1 keys are also closed and exact: + +```text +counts.conditions: declared, intended, preserved +counts.realizations: planned, evaluable, evidence_only, selected +counts.rows: source, expected, validated, evaluable, defective, replaced +counts.evidence_status: complete, qualified, partial, malformed, recoverable, failed +counts.scope_disposition: included, excluded_from_paper_matrix, held, superseded +counts.prediction_origin: native_inference, external_historical_inference, + external_repair_inference +counts.derivation_origin: native_execution, external_import, repair_overlay, + offline_transformation +counts.source_classification: raw, canonical, derived, repaired, aggregate_only +counts.column_disposition: mapped, namespaced_preserved, ignored_with_reason +counts.authorization: inference_repair_bundles, offline_transformation_bundles, + granted_pairs, executable_pairs, nonexecutable_pairs +defects: total_findings, by_code, question_ids_by_code, findings_digest +checksums: all_referenced_sources_match, sources, expected_snapshots, + evidence_index, validation_artifacts, results +identity_projection: import_spec_digest, semantic_grid_digest, + experiment_id, experiment_digest, condition_digests, + realization_digests, result_artifacts, + authorization_bundle_digests, lineage_digests +identity_projection.result_artifacts[realization_id]: result_artifact_id, + result_artifact_digest, file_sha256 +profile: matrix, queues, offline_authority, datasets, source_schema_count +failures[]: code, source_id, field, question_id, record_span, + raw_row_sha256, sanitized_message +audit: source_location, output_root, imported_at +``` + +Every mapping has unknown-key rejection in report self-validation. Empty +categories remain explicit zero/empty mappings. Generic success, validation +failure, imported success, exact no-op, repair overlay, offline transformation, +and Stage 1 profile tests assert the entire nested key set and cross-total +invariants; report identity excludes `audit` and `import_state` but includes all +stable semantic/provenance/lineage projections. +`record_span` has exactly `record_index`, `start`, `end`, and +`terminator_hex`. Tests delete, add, and tamper every typed leaf in turn and +require report self-validation failure. + +- [ ] **Step 1: Write failing end-to-end synthetic engine tests** + + Cover valid generic CSV import; deterministic IDs; dry run writing nothing; + real complete/qualified artifacts; partial/malformed/recoverable evidence-only + realizations; source checksum mismatch and mutation; exact repeated import + no-op with a byte-identical tree; mapping/source/status/overlay divergence + refusal; result artifact/run-state linkage; atomic failure; import report + counts; audit paths excluded from identities; and import/evaluation outside the + repository using an absolute workspace. + + Assert run state uses `status=completed` for evaluable realizations and + `status=evidence_only` for preserved ineligible realizations. Imported + realization records explicitly use `import_state=imported`; a failed + transaction publishes no realization/run. Dry-run reports use + `import_state=validated` without changing planned realization IDs. + + Add a `load_import_spec()` -> `build_import_plan()` integration assertion for + the exact dataset/model/method/prompt/semantic-condition bridge from Task 5, + including the scientific-versus-realization-only mutation matrix. Assert a + failed validation serializes `import_state=failed`, both operation booleans + false, no experiment ID, and structured sanitized failures without publishing + a run. + + Add base-plus-overlay integration only here, after the full verifier exists: + verify the immutable base into `VerifiedImportRun`, pass the selected frozen + realization to Task 10's pure derivation, and publish the derived result under + a different run/experiment. Snapshot every base byte before/after; cover valid, + evidence-only, unauthorized, conflict, mixed-origin, offline, and divergent + cases. No Task 10 function opens the base filesystem directly. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_engine.py -v + ``` + + Expected: FAIL because engine orchestration is absent. + +- [ ] **Step 3: Implement planning and dry-run** + + Resolve sources/datasets/references, adapt rows, validate, build conditions and + realizations, resolve every importer metric only through the closed + `BUILTIN_METRICS` registry and bind its registered implementation identity, + then create manifest v3. Dry run executes every read/validation/ + identity step and returns the report without creating workspace directories, + evidence, staging, results, state, or manifest. It still renders eligible + final CSV bytes and prepares validation/result sidecars in memory after the + planned manifest is fixed, so report-v1 includes the same per-realization + result artifact IDs/digests/file SHA-256 as a real identical import. + +- [ ] **Step 4: Implement staged real publication and graph verification** + + Use `ImportTransaction`; write snapshots/evidence/validation artifacts; + normalize eligible rows with final ownership; prepare/publish result artifacts + in staging; write final realization-keyed state with exact evidence-index and + validation-artifact SHA-256 for every realization (plus result identity/digest/ + SHA-256 only when evaluable); validate the entire staged + graph; publish once. Existing final runs call `verify_import_run` and either + exact no-op or refuse. + + `verify_import_run` is the single shared full-graph verifier used for existing- + run idempotence, overlay base input, and Task 13 publication reading. It + validates manifest/state/snapshots/evidence/index/validation artifacts/results/ + sidecars, derives each realization's exact logical-source/blob SHA-256 mapping, + and returns immutable trusted values only after the whole graph passes. + +- [ ] **Step 5: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_engine.py \ + tests/importing/test_transaction.py tests/importing/test_result_artifact.py -v + python -m pytest -q + ``` + + Expected: all generic dry-run/import/idempotence/status cases pass; no baseline + regression. + +- [ ] **Step 6: Commit** + + ```bash + git add src/choicebench/importing/engine.py \ + src/choicebench/importing/__init__.py tests/importing/test_engine.py + git commit -m "feat: import external results into immutable runs" + ``` + +--- + +### Task 13: Read and evaluate explicit realizations + +**Files:** + +- Modify: `src/choicebench/io/readers.py` +- Modify: `src/choicebench/cli/evaluate_run.py` +- Create: `tests/importing/test_evaluation.py` +- Modify: `tests/io/test_readers.py` +- Modify: `tests/scripts/test_evaluate_run.py` +- Modify: `tests/test_publication_identity.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class RealizationSelection: + policy: Literal["single_evaluable_per_condition", "explicit"] + realization_ids: tuple[str, ...] + + +@dataclass(frozen=True) +class ManifestResultSet: + rows: pd.DataFrame + manifest: Mapping[str, Any] + view: ManifestView + state: Mapping[str, Any] + selection: RealizationSelection + artifacts: Mapping[str, Mapping[str, Any]] + accounting: Mapping[str, Mapping[str, Any]] + + +def read_manifest_result_set( + run_dir: Path, *, realization_ids: Sequence[str] | None = None +) -> ManifestResultSet: ... + + +def read_manifest_results(run_dir: Path) -> tuple[pd.DataFrame, dict]: ... + + +def build_evaluation_report_for_result_set( + run_id: str, + result_set: ManifestResultSet, + *, + reparse: bool, +) -> dict[str, Any]: ... +``` + +`read_manifest_results` remains the v2-compatible wrapper. For v3 it may +auto-select only when each semantic condition has at most one eligible +realization; ambiguity is a refusal. Add repeatable CLI +`choicebench-evaluate --realization-id `. + +- [ ] **Step 1: Write failing realization-reader tests** + + Cover manifest/state/evidence/snapshot/realization-validation-artifact/sidecar full-graph validation; undeclared + result refusal; result for evidence-only realization refusal; exact row + ownership; question membership and uniqueness; and fieldwise join of + `question_text`, serialized choices, `correct_option`, and correct answer text + against the expected snapshot. Include one variable-option ARC-style result. + + Add two eligible realizations for one condition and assert default refusal, + explicit selection of exactly one, unknown/ineligible ID refusal, and no + concatenation. All unselected/ineligible realizations remain in accounting. + +- [ ] **Step 2: Write failing evaluation-v2 tests** + + Assert complete/qualified selected metrics; qualifications/limitations beside + qualified metrics; partial/malformed/recoverable/failed/excluded/held/ + superseded accounting with null result/metrics; complete+excluded and + malformed+excluded orthogonality; condition grouping separate from realization + metrics; selection policy and every accounted digest in evaluation identity; + audit path/timestamp changes excluded; result byte change included; and + selection change produces a new evaluation ID. Add an importer-manifest metric + changed to `module:Class` and prove refusal occurs before `importlib` is called; + separately retain a legacy native custom-metric compatibility test. + + Report shape must include: + + ```python + assert report["conditions"][condition_id]["realization_ids"] == [real_a, real_b] + assert report["realizations"][real_a]["selected"] is True + assert report["realizations"][real_a]["benchmark_name"] == "arc_challenge" + assert report["realizations"][real_a]["evidence_status"] == "complete" + assert report["realizations"][real_a]["scope_disposition"] == "included" + assert report["realizations"][real_a]["prediction_origins"] == [ + "external_historical_inference" + ] + assert report["realizations"][real_a]["option_count_distribution"] == { + "3": 1, + "4": 1, + } + assert report["realizations"][real_b]["selected"] is False + assert report["realizations"][malformed]["metrics"] == {} + ``` + + Re-run the Task 1 v2 fixture and require the exact evaluation-v1 identity and + shape to remain unchanged. + +- [ ] **Step 3: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_evaluation.py tests/io/test_readers.py \ + tests/scripts/test_evaluate_run.py \ + tests/importing/test_manifest_v2_compat.py -v + ``` + + Expected: new v3 selection/evaluation tests fail; legacy tests pass. + +- [ ] **Step 4: Implement version-dispatched reading** + + Preserve the current v2 path in a private helper. For v3, use + `verify_import_run` rather than duplicating graph verification, then enforce + row ownership, status eligibility, and explicit selection over its trusted + realization records. + Never call `read_all_run_results`. + +- [ ] **Step 5: Implement evaluation-v2 and keep evaluation-v1** + + Branch on manifest schema. Keep the existing v1 identity/report function + byte-for-byte for legacy input. Build v2 units sorted by full realization + identity and bind selection, statuses, qualifications, evidence, + authorization, lineage, result IDs/digests/checksums, metric implementations, + and optional reparsing implementations. + + For importer-created manifests, instantiate only the manifest-bound built-in + metric name after rechecking it against `BUILTIN_METRICS` and its recorded + implementation identity; never pass an importer metric string to the current + dynamic `_load_metric` path. Preserve the existing trusted native v2/v3 custom- + metric behavior without making it reachable from an import specification. + +- [ ] **Step 6: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_evaluation.py tests/io/test_readers.py \ + tests/scripts/test_evaluate_run.py tests/test_publication_identity.py \ + tests/importing/test_manifest_v2_compat.py -v + python -m pytest -q + ``` + + Expected: v2 identity compatibility and v3 selection/accounting both pass; + the full baseline is green. + +- [ ] **Step 7: Commit** + + ```bash + git add src/choicebench/io/readers.py \ + src/choicebench/cli/evaluate_run.py tests/importing/test_evaluation.py \ + tests/io/test_readers.py tests/scripts/test_evaluate_run.py \ + tests/test_publication_identity.py + git commit -m "feat: evaluate explicit result realizations" + ``` + +**Intermediate gate:** Independent reader/evaluation review. Require tests or a +manual adversarial construction demonstrating that two alternatives cannot be +silently combined and that v2 evaluation identity remains exact. + +--- + +### Task 14: Migrate newly created native runs to manifest v3 + +**Files:** + +- Modify: `src/choicebench/cli/run_experiment.py` +- Modify: `src/choicebench/infra/checkpoint.py` +- Modify: `tests/test_condition_grid.py` +- Modify: `tests/infra/test_checkpoint.py` +- Modify: `tests/scripts/test_reset_run.py` +- Modify: `tests/scripts/test_build_backend.py` +- Modify: `tests/test_pride_reproduction_wiring.py` +- Modify: `tests/test_wheel_smoke.py` +- Create: `tests/importing/test_native_v3.py` + +**Interfaces:** + +```python +class ExecutionPlan: + selections: Sequence[BenchmarkSelection] + preflights: Mapping[tuple[int, int], Any] + semantic_conditions: Mapping[str, Mapping[str, Any]] + realizations: Mapping[str, Mapping[str, Any]] + jobs: Mapping[tuple[int, int, int], Mapping[str, Any]] + manifest: Mapping[str, Any] + + +def build_execution_plan( + config: ExperimentConfig, + *, + manifest_schema: Literal["v3", "legacy_v2_resume"] = "v3", +) -> ExecutionPlan: ... +``` + +New checkpoint-v2 records use `checkpoints/.json` and bind +experiment, condition, realization, and selection. Existing checkpoint-v1 and +programmatic behavior remain readable only in legacy-v2 resume mode. +`build_execution_plan(..., manifest_schema="v3")` calls the explicit v3 +manifest/state builders; it does not change the legacy public v2 constants or +the default behavior of `make_manifest()` for external programmatic callers. + +- [ ] **Step 1: Write failing native-v3 tests before changing the runner** + + Assert a new dry execution plan has manifest v3; unchanged current + `condition_id`; one `native_execution` realization per condition; explicit + `native_inference` row assignment/lineage; realization-keyed run state; + realization-addressed checkpoint/result/validation paths; deterministic native + validation artifact with no external-evidence claim; result-artifact-v2; and native + evaluation-v2. Assert no missing result origin in a new manifest. + + Add a pre-existing v2 run/resume test: detect the existing manifest before + candidate creation, rebuild a v2-compatible plan, verify exact experiment/ + condition/path ownership, resume/skip without rewriting completed bytes, and + never upgrade the manifest/state/sidecar in place. A new absent run may never + request `legacy_v2_resume`. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_native_v3.py \ + tests/test_condition_grid.py tests/infra/test_checkpoint.py \ + tests/scripts/test_reset_run.py -v + ``` + + Expected: new-run v3 assertions fail; all existing runner tests pass. + +- [ ] **Step 3: Split runtime jobs from semantic/realization manifest records** + + Preserve the current condition identity payload including the single frozen + prompt-path compatibility discriminator. Build one explicit native realization + per job. Runtime jobs may merge descriptive condition/realization fields for + execution, but the manifest tables stay separate and semantic records contain + no realization back-reference. + +- [ ] **Step 4: Inject explicit native row origin and use v3 result publication** + + Prefer the existing `condition_metadata` injection after each runner batch so + method implementations remain untouched. Inject condition/realization and + ownership IDs, `prediction_origin=native_inference`, and a non-null stable + `prediction_lineage_id`. Validate the completed native rows against the bound + dataset snapshot, write the same self-digested realization-validation artifact + with an empty external-evidence reference set, then publish through + `prepare_manifest_result`/`publish_manifest_result`; update realization-keyed + run state with validation SHA-256 plus result ID/digest/file checksum. Native + failure/gate accounting retains its existing evidence artifact and emits no + fake imported source/evidence record. + +- [ ] **Step 5: Add versioned checkpoint/resume behavior** + + New checkpoints bind `realization_id`; legacy-v2 resume continues using the + old condition checkpoint/result contract and implicit native origin. Do not + combine v2 checkpoint rows with v3 rows. Preserve current reset safety. + +- [ ] **Step 6: Run focused native compatibility tests** + + ```bash + python -m pytest tests/importing/test_native_v3.py \ + tests/test_condition_grid.py tests/infra/test_checkpoint.py \ + tests/scripts/test_reset_run.py tests/scripts/test_build_backend.py \ + tests/test_pride_reproduction_wiring.py tests/test_wheel_smoke.py -v + ``` + + Expected: new native runs are v3, legacy v2 resumes remain byte-compatible, + and native dummy-backend workflows pass. + +- [ ] **Step 7: Run the full suite** + + ```bash + python -m pytest -q + ``` + + Expected: all original 808 tests plus importer tests pass with zero failures. + +- [ ] **Step 8: Commit** + + ```bash + git add src/choicebench/cli/run_experiment.py \ + src/choicebench/infra/checkpoint.py \ + tests/importing/test_native_v3.py tests/test_condition_grid.py \ + tests/infra/test_checkpoint.py tests/scripts/test_reset_run.py \ + tests/scripts/test_build_backend.py tests/test_pride_reproduction_wiring.py \ + tests/test_wheel_smoke.py + git commit -m "feat: emit native manifest v3 realizations" + ``` + +**Intermediate gate:** Independent compatibility review must exercise one new +v3 native run and one pre-existing v2 resume. Stop if either changes scientific +condition IDs or rewrites legacy bytes. + +--- + +### Task 15: Translate Stage 1 manifests and typed authorities + +**Files:** + +- Create: `src/choicebench/importing/profiles/__init__.py` +- Create: `src/choicebench/importing/profiles/stage1_paper_freeze.py` +- Extend: `tests/importing/conftest.py` +- Create: `tests/importing/test_stage1_profile.py` + +**Interfaces:** + +```python +@dataclass(frozen=True) +class ProfileTranslation: + spec: ImportSpec + report: Mapping[str, Any] + + +def translate_stage1_paper_freeze(freeze_root: Path) -> ProfileTranslation: ... +``` + +The profile registry exposes the name `stage1-paper-freeze`. No generic module +may import this profile or contain any Stage 1 constant. + +- [ ] **Step 1: Build a programmatic miniature freeze fixture** + + In `tests/importing/conftest.py`, write only synthetic metadata/one-row source + CSVs under `tmp_path` with the same headers and joins as: + `expected_matrix.csv`, canonical manifest JSON/CSV, status matrix, + `artifact_inventory.csv`, `frozen_artifact_index.csv`, approved/held/excluded/ + forensic queues, report, and checksum ledger. Exercise the authoritative join + chain explicitly: + + ```text + canonical_results_manifest.canonical_artifact_id + -> artifact_inventory.artifact_id + -> frozen raw path and checksum + -> frozen_artifact_index canonical path and cell_id + ``` + + A missing, duplicate, or cross-cell link at any hop fails closed. Generate + checksum values only for the miniature equivalents of files covered by the + real ledger, including `reports/canonical_freeze_report.md`; intentionally + leave only artifact inventory/index unledgered and assert those two remain + independently hashed corroboration. Assert a one-byte canonical-report change + fails its ledger check. Include one included + complete, qualified, recoverable, partial, and malformed cell; both hosted + PriDe exclusion combinations; one approved inference repair; one held record; + one excluded record; and one forensic-only record. + +- [ ] **Step 2: Write failing translation/authority tests** + + Require joins by explicit `cell_id`/`canonical_artifact_id`, never filename. + Test exact status mapping: + + ```python + STATUS_MAP = { + "canonical_complete": "complete", + "canonical_qualified": "qualified", + "recoverable_from_existing_artifacts": "recoverable", + "incomplete_requires_inference": "partial", + "malformed_requires_inference": "malformed", + } + ``` + + `excluded_from_paper_matrix` is not an evidence-status input to this map. For + the two hosted PriDe rows, assign only + `scope_disposition=excluded_from_paper_matrix`, then run ordinary row/defect + validation to compute MMLU `complete` and ARC `malformed`. Assert these two + dimensions remain independent and the records are preserved but not intended. + + Assert approved/authoritative/executable inference authority is separate from + 687-style held false, excluded false, and forensic-only records. For every + queue row, expand exact `(cell_id, question_id)` pairs; require membership in + the canonical manifest's missing/damaged IDs as applicable; require the three + queue ledgers to be pairwise disjoint; and validate `queue_disposition`, + `execution_authority`, `executable`, and `supersedes_for_execution`. Assert all + 828 classified queue pairs equal 138 approved + 687 held + 3 excluded and the + 776 forensic pairs never become authority. + + Assert method identities and all eight exact canonical source-header schemas + map explicitly, with every column assigned to mapped, namespaced-preserved, or + explicitly ignored-with-reason. Cover `semantic_matching_v1`, + `text_extraction`, `two_stage_v1/v2/v3`, `independent_hypothesis`, + `cyclic_generation_majority`, `PriDe`, and historical `baseline` without + renaming/merging. + +- [ ] **Step 3: Write the exact offline-authority test** + + The fixture must include these six historical semantic-matching cell IDs and + authorize the same three question IDs for each cell: + + ```python + RECOVERABLE_CELLS = ( + "cbp__gemini-2-5-flash__arc_challenge__semantic_matching_v1", + "cbp__gpt-4-1-mini__arc_challenge__semantic_matching_v1", + "cbp__llama-3-1-8b-instant__arc_challenge__semantic_matching_v1", + "cbp__meta-llama-llama-3-1-8b-instruct__arc_challenge__semantic_matching_v1", + "cbp__qwen-qwen2-5-7b-instruct-turbo__arc_challenge__semantic_matching_v1", + "cbp__qwen-qwen2-5-7b-instruct__arc_challenge__semantic_matching_v1", + ) + ``` + + ```python + RECOVERABLE_QUESTION_IDS = ( + "79e8c959bbeb74a0", + "ad6b5d46ae54842c", + "c30e75b011696a95", + ) + assert report["offline_authority"]["condition_count"] == 6 + assert report["offline_authority"]["question_cell_count"] == 18 + assert report["offline_authority"]["inference_executable"] is False + ``` + + Assert each cell has exactly `RECOVERABLE_QUESTION_IDS` in + `recoverable_question_ids`; preserve every question reason; and bind the + canonical source SHA-256/base-evidence digest, full semantic condition digest, + expected-snapshot digest, allowed semantic-rematching purpose, and + `inference_executable=false`. Verify the authority is a deterministic + content-addressed projection of checksum-covered canonical-manifest JSON + entries cross-checked fieldwise with the CSV, and that + `arc_question_audit.csv` cannot be its trust anchor. + + Emit one `AuthorizationSpec` bundle containing the six condition-key grants, + not six unrelated self-authorizing records. Assert its aggregate bundle digest + is reproduced by all six condition-scoped validated slices. + +- [ ] **Step 4: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_stage1_profile.py -v + ``` + + Expected: FAIL because the profile modules are absent. + +- [ ] **Step 5: Implement manifest/checksum/queue translation** + + Independently hash every opened authoritative file against + `checksums.sha256` when it has a ledger entry. Root identity/status/source + authority in the checksum-covered canonical manifest and independently + computed canonical/source-artifact bytes. Treat `artifact_inventory.csv`, + `frozen_artifact_index.csv`, and reports as corroborating cross-reference + evidence when the ledger does not cover them; record their independently + computed audit digests but never promote them to a stronger trust anchor. Use + exact method/header-signature mapping tables in the + profile; do not inspect filenames for identity. Preserve original absolute + paths only as audit metadata and freeze-relative paths as stable provenance. + +- [ ] **Step 6: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_stage1_profile.py \ + tests/importing/test_authorization.py tests/importing/test_schema.py -v + python -m pytest -q + ``` + + Expected: profile translation, exclusions, queue separation, offline authority, + and paper-agnostic core checks pass. + +- [ ] **Step 7: Commit** + + ```bash + git add src/choicebench/importing/profiles/__init__.py \ + src/choicebench/importing/profiles/stage1_paper_freeze.py \ + tests/importing/conftest.py tests/importing/test_stage1_profile.py + git commit -m "feat: translate Stage 1 import manifests" + ``` + +--- + +### Task 16: Implement the Stage 1 expected-dataset trust chain + +**Files:** + +- Modify: `src/choicebench/importing/profiles/stage1_paper_freeze.py` +- Extend: `tests/importing/test_stage1_profile.py` + +**Interfaces:** + +```python +def build_stage1_expected_datasets( + freeze_root: Path, + checksum_ledger: Mapping[str, str], +) -> tuple[ExpectedDataset, ExpectedDataset, Mapping[str, Any]]: ... +``` + +- [ ] **Step 1: Extend the synthetic freeze with input/split artifacts** + + Create small synthetic ARC raw/normalized plus selection/metadata and MMLU + raw/normalized plus selection/metadata files under these exact Stage 1 + relative paths: + + ```text + raw/local_model_generalization/data/raw/arc_challenge_raw.csv + raw/local_model_generalization/data/processed/arc_challenge_normalized.csv + raw/local_model_generalization/data/splits/arc_challenge/robustness_ids.json + raw/local_model_generalization/data/splits/arc_challenge/robustness_metadata.json + raw/local_model_generalization/data/raw/mmlu_raw.csv + raw/local_model_generalization/data/processed/mmlu_normalized.csv + raw/local_model_generalization/data/splits/benchmark/robustness_ids.json + raw/local_model_generalization/data/splits/benchmark/robustness_metadata.json + ``` + + ARC includes one three-option row. MMLU includes the exact three duplicate IDs + from the approved spec, each occurring twice with identical parsed fields. + Make selection order deliberately differ from normalized source order so the + source-order keep-first rule and final selection-order projection are tested + separately. Add every byte checksum to the synthetic ledger. + +- [ ] **Step 2: Write failing ARC/MMLU trust-chain tests** + + Assert `reference_kind=independent_input_snapshot` and + `trust_label=checksum_verified_freeze_internal`; raw/normalized/selection/ + metadata checksum verification; result independence; selection membership; + stable snapshot order; variable ARC options; unknown upstream Hugging Face + revision/authenticity limitation; and byte-identical copies/result agreement + as corroboration only. + + For MMLU assert source-order, field-identical duplicate pairs, exact duplicate + row digests, keep-first semantics equivalent to + `drop_duplicates(subset="question_id", keep="first")`, and identity-bound + pre/post ordered digests. Mutate the second occurrence of each pair in turn and + require fail-closed conflict. A different duplicate policy must change the + derivation/realization identity. + +- [ ] **Step 3: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_stage1_profile.py \ + -k "dataset or duplicate or option or trust" -v + ``` + + Expected: new trust-chain assertions fail. + +- [ ] **Step 4: Implement the frozen-input derivation** + + Verify the ledger against the exact opened bytes. Use the generic CSV adapter + and dataset-reference builder. In registered profile code, deterministically + derive the comparable semantic fields and question IDs from the raw rows and + compare them fieldwise with the checksum-bound normalized rows; report this as + a ChoiceBench revalidation, not proof that the archived normalizer was the + producer. Reproduce the archived MMLU first-occurrence rule, but first require + every duplicate occurrence to agree on all parsed fields. Record the archived + script path/checksum as provenance evidence and the ChoiceBench implementation + identity as the active verification/transformation identity. Never execute + the archived script or any other freeze content. Preserve the limitation that + this proves internal freeze consistency, not upstream publisher authenticity. + +- [ ] **Step 5: Run focused, profile, and full tests** + + ```bash + python -m pytest tests/importing/test_stage1_profile.py \ + tests/importing/test_dataset_reference.py \ + tests/importing/test_csv_adapter.py -v + python -m pytest -q + ``` + + Expected: trust labels, three-option ARC, exact duplicate derivation, and + profile translation pass; no baseline regression. + +- [ ] **Step 6: Commit** + + ```bash + git add src/choicebench/importing/profiles/stage1_paper_freeze.py \ + tests/importing/test_stage1_profile.py + git commit -m "feat: verify Stage 1 dataset trust chain" + ``` + +**Intermediate gate:** Independent Stage 1 review against the immutable freeze. +The reviewer independently checks ledger coverage, selected-ID counts, all three +MMLU duplicate pairs, variable ARC options, 100/2 scope split, status counts, +138/687/3 queue counts, and six-cell/18-row offline authority. Review is dry-run +and read-only. + +--- + +### Task 17: Add the installed CLI, reports, and security regression suite + +**Files:** + +- Create: `src/choicebench/cli/import_results.py` +- Modify: `pyproject.toml` +- Create: `tests/importing/test_cli.py` +- Modify: `tests/test_release_adversarial.py` + +**Interfaces:** + +```python +def parse_args(argv: Sequence[str] | None = None) -> argparse.Namespace: ... +def write_import_report(path: Path, report: Mapping[str, Any]) -> None: ... +def main(argv: Sequence[str] | None = None) -> int: ... +``` + +At module scope import only stdlib modules and `choicebench` package metadata; +do not import `choicebench.config.paths`, the engine, profiles, readers, or any +provider. `main()` parses and validates `--output-root`, sets +`CHOICEBENCH_HOME`, and only then imports workspace-dependent modules. + +- [ ] **Step 1: Write failing subprocess CLI tests** + + Cover generic spec import; `--dry-run` and `--validate-only` equivalence; + `--strict`; absolute `--output-root`; `CHOICEBENCH_HOME` fallback; current + directory fallback; required safe run ID; profile dispatch; repeated + `--overlay`; human summary; machine report; exact no-op; divergent refusal; + sanitized errors; and help text. Run output-root tests in fresh subprocesses + with `PYTHONPATH` controlled so module caching cannot hide import order. + + Machine report assertions use stable fields: + + ```python + assert report["counts"]["scope_disposition"]["included"] >= 1 + assert report["counts"]["evidence_status"]["complete"] >= 1 + assert report["wrote_artifacts"] is expected_write + assert report["idempotent_noop"] is expected_noop + assert report["failures"] == [] + assert report["audit"]["source_location"] == str(source.resolve()) + assert "audit" not in report["identity_projection"] + ``` + +- [ ] **Step 2: Extend adversarial/security tests before implementation** + + Add YAML executable-tag refusal; credential metadata; source traversal; + specification-controlled output; unsafe source/output symlinks; recursive + directory; source change after validation; malicious CSV error text; + malformed byte sequences; result/evidence/manifest/state tampering; divergent + report overwrite; overlay self-authorization; and response-cache forest + refusal. Explicitly accept user-selected absolute source and output roots. + +- [ ] **Step 3: Run and observe failure** + + ```bash + python -m pytest tests/importing/test_cli.py \ + tests/test_release_adversarial.py -v + ``` + + Expected: CLI tests fail because the module/entry point is absent. + +- [ ] **Step 4: Implement bootstrap parsing and command dispatch** + + Positional input is a generic spec path unless `--profile` is present, in + which case it is the profile root. Required options/aliases are: + + ```text + choicebench-import-results INPUT --run-id ID + [--profile stage1-paper-freeze] + [--dry-run | --validate-only] + [--strict] + [--output-root PATH] + [--overlay PATH ...] + [--report PATH] + ``` + + Output-root selection comes only from CLI/environment. Canonicalize it, refuse + unsafe symlinks, set the environment, then lazily import and execute. Reports + are atomic; an identical existing report is a no-op and divergent content is + refused. Audit locations/timestamps remain outside report identity. + +- [ ] **Step 5: Register the command without changing dependencies/version** + + Add exactly: + + ```toml + choicebench-import-results = "choicebench.cli.import_results:main" + ``` + + Do not change `version = "0.2.0"` or package data. + +- [ ] **Step 6: Run targeted and full tests** + + ```bash + python -m pytest tests/importing/test_cli.py \ + tests/test_release_adversarial.py -v + python -m pytest -q + ``` + + Expected: CLI/environment/security tests pass and the full baseline is green. + +- [ ] **Step 7: Commit** + + ```bash + git add src/choicebench/cli/import_results.py pyproject.toml \ + tests/importing/test_cli.py tests/test_release_adversarial.py + git commit -m "feat: add external results import command" + ``` + +**Intermediate gate:** Independent adversarial review of CLI import order, path +trust boundaries, YAML/CSV handling, report collision behavior, secret rejection, +and absence of inference execution. + +--- + +### Task 18: Document and package the importer + +**Files:** + +- Create: `docs/external-results-import.md` +- Modify: `README.md` +- Modify: `CHANGELOG.md` +- Modify: `tests/test_wheel_smoke.py` + +**Interfaces:** No new runtime interface. Documentation examples use only the +installed `choicebench-import-results` and `choicebench-evaluate` commands. + +- [ ] **Step 1: Extend the installed wheel/sdist smoke test first** + + In each installed environment, assert importer modules and console entry point + exist; run `--help`; create a two-question generic CSV/spec outside the + repository; dry-run; real import; exact repeated no-op; and evaluation. Assert + the installed wheel/sdist contains no run, report, historical, cache, or test + fixture data. + +- [ ] **Step 2: Run and observe failure** + + ```bash + python -m pytest tests/test_wheel_smoke.py -v + ``` + + Expected: PASS because Tasks 2-17 already supplied and registered the runtime + functionality. Treat this as a packaging characterization gate before the + documentation-only changes; if it fails, diagnose and fix the responsible + earlier runtime/packaging task in a focused commit rather than hiding the + problem in documentation. + +- [ ] **Step 3: Write complete user/adapter documentation** + + `docs/external-results-import.md` must cover: external versus native inference; + supported UTF-8 CSV/dialect fields; every import-spec section; generic and + dry-run examples; `CHOICEBENCH_HOME`/absolute output; checksums and four + identity layers; orthogonal statuses; expected-dataset trust; method-specific + extension data; mixed prediction origins; immutable new-run overlays; typed + inference/offline authorization; evaluation selection/accounting; security + boundary; limitations; future adapter protocol; and a complete two-question + example. + + Link the guide from README. Add an `Unreleased` changelog entry consistent + with the existing newest-first format. State explicitly that no historical + data or model inference is included. Do not change the package version. + +- [ ] **Step 4: Run documentation-facing and full tests** + + ```bash + python -m pytest tests/test_wheel_smoke.py tests/importing/test_cli.py -v + python -m pytest -q + ``` + + Expected: wheel and sdist workflows work outside the repository; full suite + passes. + +- [ ] **Step 5: Commit** + + ```bash + git add docs/external-results-import.md README.md CHANGELOG.md \ + tests/test_wheel_smoke.py + git commit -m "docs: document and package external result imports" + ``` + +--- + +### Task 19: Full validation, independent reviews, push, and unmerged PR + +**Files:** + +- Modify only files required to resolve confirmed review findings. +- Create outside Git: Stage 1 JSON reports, temporary workspaces/builds, and PR + body. + +- [ ] **Step 1: Review the complete diff and repository cleanliness** + + ```bash + git status --short + git diff --check origin/choicebench...HEAD + git diff --stat origin/choicebench...HEAD + git diff --name-only origin/choicebench...HEAD + ``` + + Expected: only source, tests, documentation, and small synthetic test support; + no imported runs, reports, historical data, caches, credentials, build output, + or temporary files. + +- [ ] **Step 2: Run all focused importer/compatibility/security tests** + + ```bash + python -m pytest tests/importing tests/io/test_readers.py \ + tests/io/test_writers.py tests/infra/test_checkpoint.py \ + tests/test_condition_grid.py tests/test_publication_identity.py \ + tests/test_release_adversarial.py tests/scripts/test_evaluate_run.py \ + tests/scripts/test_reset_run.py -v + ``` + + Expected: all focused tests pass with zero failures. + +- [ ] **Step 3: Run the full ChoiceBench suite** + + ```bash + python -m pytest tests/ -v + ``` + + Expected: the existing 808 tests plus all new tests pass; record the exact + final count and duration for the PR. + +- [ ] **Step 4: Confirm the repository has no configured lint/type gate** + + ```bash + rg -n "ruff|mypy|pyright|black|flake8|pylint" pyproject.toml \ + .github/workflows tests + ``` + + Expected: no configured command. Do not invent a new formatter/type-check gate + in this PR; report that pytest is the repository's configured CI check. + +- [ ] **Step 5: Build and check wheel/sdist outside the repository** + + ```bash + IMPORT_BUILD_DIST="$(mktemp -d)" + python -m build --outdir "$IMPORT_BUILD_DIST" + python -m twine check "$IMPORT_BUILD_DIST"/* + python -m pytest tests/test_wheel_smoke.py -v + ``` + + Expected: build exits 0; Twine reports both distributions `PASSED`; installed + wheel and sdist generic import/evaluation smoke passes outside the repository. + +- [ ] **Step 6: Record a read-only freeze inventory before Stage 1 validation** + + ```bash + FREEZE_ROOT=/home/cotenthusiast/Projects/model-generalization/paper_data_freeze + FREEZE_METADATA_BEFORE="$(mktemp)" + FREEZE_CONTENT_BEFORE="$(mktemp)" + find "$FREEZE_ROOT" -printf '%y\t%P\t%m\t%U\t%G\t%s\t%T@\t%l\n' \ + | sort > "$FREEZE_METADATA_BEFORE" + find "$FREEZE_ROOT" -type f -print0 | sort -z | xargs -0 sha256sum \ + > "$FREEZE_CONTENT_BEFORE" + ``` + + Expected: metadata, symlink targets, and content digests for the entire freeze + are written outside both repositories. Do not execute any freeze script and + do not change file permissions or metadata. + +- [ ] **Step 7: Run the full Stage 1 strict dry run** + + ```bash + STAGE1_DRY_REPORT_DIR="$(mktemp -d)" + STAGE1_DRY_REPORT="$STAGE1_DRY_REPORT_DIR/stage1-dry-run.json" + choicebench-import-results "$FREEZE_ROOT" \ + --profile stage1-paper-freeze \ + --run-id paper-stage1-validation \ + --dry-run --strict \ + --report "$STAGE1_DRY_REPORT" + ``` + + Validate the report with: + + ```bash + jq -e ' + .profile.matrix.intended_cells == 100 and + .profile.matrix.core_method_cells == 84 and + .profile.matrix.local_pride_cells == 4 and + .profile.matrix.ihs_cells == 12 and + .profile.matrix.excluded_preserved_cells == 2 and + .counts.evidence_status.complete == 58 and + .counts.evidence_status.qualified == 5 and + .counts.evidence_status.recoverable == 6 and + .counts.evidence_status.partial == 4 and + .counts.evidence_status.malformed == 29 and + .profile.matrix.intended_evidence.complete == 57 and + .profile.matrix.intended_evidence.qualified == 5 and + .profile.matrix.intended_evidence.recoverable == 6 and + .profile.matrix.intended_evidence.partial == 4 and + .profile.matrix.intended_evidence.malformed == 28 and + .profile.queues.approved_executable_question_cells == 138 and + .profile.queues.held_nonexecutable_question_cells == 687 and + .profile.queues.excluded_nonexecutable_question_cells == 3 and + .profile.queues.forensic_question_cells == 776 and + .profile.queues.classified_question_cells == 828 and + .profile.queues.pairwise_disjoint == true and + .profile.queues.approved_authoritative_executable == true and + .profile.queues.held_executable == false and + .profile.queues.excluded_executable == false and + .profile.source_schema_count == 8 and + .profile.offline_authority.condition_count == 6 and + .profile.offline_authority.question_cell_count == 18 and + .profile.offline_authority.inference_executable == false and + .profile.datasets.arc.reference_kind == "independent_input_snapshot" and + .profile.datasets.arc.trust_label == "checksum_verified_freeze_internal" and + .profile.datasets.arc.selected_question_count == 1000 and + .profile.datasets.arc.option_count_distribution["3"] == 3 and + .profile.datasets.arc.option_count_distribution["4"] == 997 and + .profile.datasets.mmlu.reference_kind == "independent_input_snapshot" and + .profile.datasets.mmlu.trust_label == "checksum_verified_freeze_internal" and + .profile.datasets.mmlu.selected_question_count == 1000 and + .profile.datasets.mmlu.pre_dedup_row_count == 1003 and + .profile.datasets.mmlu.post_dedup_row_count == 1000 and + .profile.datasets.mmlu.duplicate_question_ids == + ["79686d32dfe155ea", "2f7aa3c7ebb98cfe", "74f7227e190200ac"] and + (.profile.datasets.mmlu.pre_dedup_digest | length) == 64 and + (.profile.datasets.mmlu.post_dedup_digest | length) == 64 and + .profile.datasets.arc.publisher_authenticated == false and + .profile.datasets.mmlu.publisher_authenticated == false and + .checksums.all_referenced_sources_match == true + ' "$STAGE1_DRY_REPORT" + ``` + + Expected: importer reports `validated`, writes no run, and every assertion is + true. Held/excluded records remain non-executable. The profile performs the + same sealed invariants through registered ChoiceBench code; it does not run + `scripts/validate_stage1_freeze.py` or any other external executable content. + +- [ ] **Step 8: Perform an isolated real Stage 1 import and evaluation** + + ```bash + STAGE1_IMPORT_ROOT="$(mktemp -d)" + STAGE1_IMPORT_REPORT_DIR="$(mktemp -d)" + STAGE1_IMPORT_REPORT="$STAGE1_IMPORT_REPORT_DIR/stage1-import.json" + choicebench-import-results "$FREEZE_ROOT" \ + --profile stage1-paper-freeze \ + --run-id paper-stage1-import \ + --strict \ + --output-root "$STAGE1_IMPORT_ROOT" \ + --report "$STAGE1_IMPORT_REPORT" + CHOICEBENCH_HOME="$STAGE1_IMPORT_ROOT" \ + choicebench-evaluate --run-id paper-stage1-import + STAGE1_EVAL_REPORT="$(find "$STAGE1_IMPORT_ROOT/reports" -maxdepth 1 \ + -type f -name 'paper-stage1-import_*_metrics.json' -print -quit)" + jq -e ' + any(.realizations[]; + .selected == true and .evidence_status == "complete" and + .scope_disposition == "included" and (.metrics | length) > 0 and + (.prediction_origins | all(. == "external_historical_inference"))) and + any(.realizations[]; + .selected == true and .evidence_status == "qualified" and + .scope_disposition == "included" and (.qualifications | length) > 0 and + (.metrics | length) > 0) and + any(.realizations[]; + (.evidence_status == "partial" or .evidence_status == "malformed") and + (.metrics | length) == 0) and + any(.realizations[]; + .benchmark_name == "arc_challenge" and .selected == true and + .option_count_distribution["3"] == 3 and (.metrics | length) > 0) and + all(.realizations[]; + if (.scope_disposition == "excluded_from_paper_matrix" or + .scope_disposition == "held" or + .scope_disposition == "superseded" or + .evidence_status == "failed") + then (.metrics | length) == 0 else true end) + ' "$STAGE1_EVAL_REPORT" + ``` + + Expected: a complete immutable run exists only below the temporary root; + included complete and qualified realizations have metrics; incomplete/ + malformed/recoverable/excluded records are accounted without metrics; at + least one evaluated ARC realization contains all three variable-option rows; + origins remain external; no inference is invoked. Record experiment, + realization, result, and evaluation IDs and the evaluation report path. + +- [ ] **Step 9: Verify the freeze stayed unchanged** + + ```bash + FREEZE_METADATA_AFTER="$(mktemp)" + FREEZE_CONTENT_AFTER="$(mktemp)" + find "$FREEZE_ROOT" -printf '%y\t%P\t%m\t%U\t%G\t%s\t%T@\t%l\n' \ + | sort > "$FREEZE_METADATA_AFTER" + find "$FREEZE_ROOT" -type f -print0 | sort -z | xargs -0 sha256sum \ + > "$FREEZE_CONTENT_AFTER" + cmp "$FREEZE_METADATA_BEFORE" "$FREEZE_METADATA_AFTER" + cmp "$FREEZE_CONTENT_BEFORE" "$FREEZE_CONTENT_AFTER" + ``` + + Expected: both `cmp` commands exit 0. The import root and reports remain + outside Git. + +- [ ] **Step 10: Run independent specification-compliance review** + + Give a fresh read-only reviewer the approved specification, implementation + plan, full diff, and validation evidence. Require a line-by-line verdict on + generic core, v2/v3 compatibility, provenance/identity, statuses, overlays, + evaluation, Stage 1 profile, CLI, tests, docs, and boundaries. Resolve all + confirmed Critical/Important findings in focused commits; rerun affected tests; + return current HEAD/evidence to that reviewer until it gives an explicit + approval on the current commit. + +- [ ] **Step 11: Run independent code-quality review** + + A different fresh reviewer inspects cohesion, API/type consistency, error + clarity, duplication, maintenance cost, performance/memory on Stage 1, and + paper leakage. Resolve confirmed blockers, rerun focused/full tests, and obtain + that reviewer's explicit approval of the resulting current commit. + +- [ ] **Step 12: Run independent adversarial/security review** + + A different reviewer attacks YAML/CSV parsing, byte spans, checksum races, + paths/symlinks, staging/no-replace, idempotence, secret handling, manifest/ + sidecar/state tampering, authorization escalation, base-run mutation, and + malformed evidence normalization. Resolve blockers and rerun security plus + full tests, then obtain that reviewer's explicit approval of current HEAD. + +- [ ] **Step 13: Run fresh-context final review** + + Give a final reviewer only the user requirements, approved spec/plan, final + diff, test/build/Twine/Stage 1 evidence, and earlier resolved-finding commits. + Require a clear PR-ready/not-ready verdict. Do not proceed on Critical or + Important findings. If it finds any, resolve them in focused commits, repeat + affected/full verification, and send the new current HEAD back to the same + fresh-context reviewer for a renewed verdict. Repeat until it explicitly says + PR-ready for the exact commit that will be pushed. + +- [ ] **Step 14: Repeat final verification after all review fixes** + + Confirm the specification, quality, security, and final reviewers all approved + the current HEAD, then repeat Steps 1–9. If this verification causes any code, + test, or documentation change, invalidate all four approvals and repeat Steps + 10–13 before proceeding. Expected: approvals and validation evidence refer to + the exact same commit. + +- [ ] **Step 15: Reconfirm remote target and branch ancestry** + + ```bash + git fetch origin + git remote show origin + git merge-base --is-ancestor origin/choicebench HEAD + git status --short --branch + ``` + + Expected: worktree clean; feature contains the synchronized integration base. + Confirm the repository's actual integration/default target from remote state + rather than assuming its name. If it differs from the branch used to start + this work or ancestry is false, stop and report instead of rebasing/retargeting + silently. + +- [ ] **Step 16: Push the feature branch** + + ```bash + git push -u origin feat/external-results-importer + ``` + + Expected: GitHub confirms the remote feature branch. Do not force-push. + +- [ ] **Step 17: Create the unmerged pull request** + + Create `/tmp/choicebench-external-results-pr.md` with `apply_patch`, containing + motivation; architecture; imported-versus-native semantics; generic importer; + Stage 1 profile; overlay/authorization design; security; tests and exact final + count; Stage 1 dry-run and real-import evidence; build/Twine; limitations; + explicit no historical data/inference statement; base/feature branch; and + exact commit hashes. Then run: + + ```bash + gh pr create \ + --base \ + --head feat/external-results-importer \ + --title "feat: add provenance-preserving external results importer" \ + --body-file /tmp/choicebench-external-results-pr.md + ``` + + Expected: GitHub returns a PR number and URL. Verify with `gh pr view`. Leave + it open; do not merge, tag, or release. If authentication prevents creation, + preserve the pushed branch and report the exact GitHub CLI error without + claiming a PR exists. + +--- + +## Plan completion condition + +Implementation is complete only when every task/checkbox is satisfied, all +independent reviews pass, the full validation evidence is current, the feature +branch is pushed, and GitHub confirms an open unmerged PR. Generated runs, +reports, historical data, caches, temporary build/workspaces, and the external +PR-body file remain intentionally excluded from Git. diff --git a/docs/superpowers/specs/2026-07-18-external-results-importer-design.md b/docs/superpowers/specs/2026-07-18-external-results-importer-design.md new file mode 100644 index 0000000..3c08651 --- /dev/null +++ b/docs/superpowers/specs/2026-07-18-external-results-importer-design.md @@ -0,0 +1,1344 @@ +# External Results Importer Design + +**Date:** 2026-07-18 +**Status:** Independently reviewed; awaiting implementation-planning approval +**Target:** ChoiceBench after v0.2.0 +**Feature branch:** `feat/external-results-importer` +**Base:** `origin/choicebench` at `a01d85b1297bd7b557e0d2a16ad51c8c1c38e866` + +## Purpose + +ChoiceBench needs a first-class way to import row-level experiment results that +were generated outside ChoiceBench. The imported data must participate in +ChoiceBench's deterministic identity, provenance, manifest, result ownership, +integrity, status accounting, reading, and evaluation systems without implying +that ChoiceBench performed the original model inference. + +The first real consumer is the immutable Stage 1 paper-data freeze at +`/home/cotenthusiast/Projects/model-generalization/paper_data_freeze`. The core +design is intentionally paper-agnostic. A thin profile translates the Stage 1 +manifests into the same generic import specifications accepted for any other +external source. + +This stage imports and validates existing row-level evidence. It does not run a +model, execute repair queues, alter historical predictions, regenerate the +freeze, or copy the 3.4 GiB raw/cache evidence forest. + +## Existing architecture and design constraints + +ChoiceBench v0.2 is manifest-first: + +- canonical JSON plus SHA-256 supplies deterministic identities; +- dataset artifacts and selected rows have verified content identities; +- models, methods, prompts, conditions, and experiments have distinct + identities; +- in manifest v2, a condition exclusively owns + `results/.csv` and its sidecar; +- immutable manifests bind the condition grid and artifact paths; +- run state accounts for operational completion, gating, and failure; +- publication-grade readers revalidate manifests, snapshots, result sidecars, + ownership fields, and question coverage; +- evaluation verifies metric implementation identity and binds consumed result + bytes and condition accounting into the evaluation identity; +- `CHOICEBENCH_HOME`, or the current directory when it is unset, is the + writable workspace; +- installed commands are strict `argparse` entry points named + `choicebench-*`. + +The importer extends these systems. It must not use the legacy filename-based +result aggregator, fabricate an inference configuration, or route rows through +a fake model backend. + +Two existing boundaries require importer-specific hardening: + +1. The generic result writer can overwrite a same-condition file after only + verifying its sidecar. Import transactions must instead perform a full + identical-content no-op or refuse divergent content. +2. The existing reader validates result question IDs but does not compare + result-side question text, options, and gold labels with the archived input + snapshot. Imported evaluable rows must be joined to a declared expected + snapshot with an explicit trust classification and checked field by field + before they can become evaluable imported result artifacts. + +## Chosen architecture + +The importer creates ordinary ChoiceBench run directories and manifests. It +adds explicit import records, evidence artifacts, provenance dimensions, and +lineage while reusing existing identity, locking, atomic-write, manifest, +sidecar, reader, and evaluator facilities. + +The main components are: + +1. **Strict import schema** — dataclass-based YAML loading with the same + unknown-field rejection and primitive validation style as experiment + configuration. +2. **Source-adapter protocol** — an internal boundary that returns a stable + tabular representation plus exact source-column information. CSV is the + first implementation. +3. **Expected-dataset loader** — resolves the declared question set and + trust-qualified question/gold/option data, creates a ChoiceBench run snapshot, + and binds its digest to the realization identity. +4. **Identity/realization builder** — preserves the existing scientific + `condition_id` while assigning every imported, repaired, transformed, or + native execution record a separate immutable `realization_id`. +5. **Row validator and normalizer** — joins by `question_id`, validates source + evidence, and creates ChoiceBench-native-format rows only when the evidence + is eligible for evaluation. +6. **Evidence store** — atomically archives only referenced row-level source + bytes in a content-addressed store and reuses identical blobs. +7. **Lineage/overlay engine** — validates repair and offline-transformation + overlays without modifying the base import. +8. **Import transaction** — locks a run, constructs and verifies the manifest, + performs collision-safe writes, and produces a machine-readable report. +9. **Manifest-aware evaluation extensions** — account for imported evidence + and compute metrics only for eligible conditions. +10. **Stage 1 profile** — translates sealed paper manifests and queue ledgers + into generic specifications without adding paper knowledge to the core. +11. **Installed CLI** — `choicebench-import-results` with a profile switch, + dry-run mode, strict validation, output-root selection, and reports. + +Likely module boundaries are: + +```text +src/choicebench/importing/ + schema.py + adapters.py + validation.py + evidence.py + lineage.py + engine.py + profiles/stage1_paper_freeze.py +src/choicebench/cli/import_results.py +``` + +Exact filenames may change during planning if existing module cohesion calls +for a smaller layout. The adapter/profile boundary and the absence of +paper-specific branches in the reusable engine are mandatory. + +## Provenance dimensions + +Import outcome, evidence quality, paper/execution scope, and inference origin +are orthogonal. No single status field may collapse them. + +### Structured result origin and derivation + +A single scalar origin is insufficient because a repaired realization can +contain predictions from several producers. Every newly written manifest uses +a structured `result_origin` record with separate derivation and prediction +origin information: + +```yaml +result_origin: + derivation_origin: repair_overlay + prediction_origins: + - external_historical_inference + - native_inference + ordered_row_origin_digest: + lineage_component_origin_digest: +``` + +`derivation_origin` describes how the realization was created: + +- `native_execution` +- `external_import` +- `repair_overlay` +- `offline_transformation` + +`prediction_origin` describes who or what produced each prediction: + +- `native_inference` +- `external_historical_inference` +- `external_repair_inference` + +Every evaluable row has a non-null `prediction_origin` and +`prediction_lineage_id`. The realization record and result sidecar contain the +sorted set/counts of constituent origins plus a digest of the ordered per-row +assignment. When more than one origin occurs, the per-row mapping is mandatory; +the importer never substitutes an uninformative `mixed` label. + +An offline transformation is a derivation, not new model inference. A rematched +row retains the prediction origin of its underlying model response while its +lineage points to an `offline_transformation` component and its implementation +identity. Every lineage node (base evidence, repair evidence, authorization, +and transformation) records its own origin. + +Missing `result_origin` may imply a homogeneous native realization only while +reading legacy manifest-v2 records. All newly written native and imported +manifest-v3 realizations record the structured field explicitly. A +ChoiceBench-native-format imported result artifact retains its external +prediction origins; the format and evaluator do not change who performed the +inference. + +### Import state + +`import_state` describes the import operation only: + +- `validated` — validation succeeded in a dry-run/validate-only report; +- `imported` — the verified artifact transaction was committed; +- `failed` — validation or the transaction failed. + +A failed transaction does not leave a partially published final run. It may +leave only a clearly marked unpublished staging directory after process or host +failure; readers never resolve staging paths. Failure details are recorded in +the import report. An already committed run may account for a source-level +failed condition through `evidence_status`; that is distinct from an importer +failure. + +### Evidence status + +`evidence_status` describes scientific usability: + +- `complete` +- `qualified` +- `partial` +- `malformed` +- `recoverable` +- `failed` + +Qualifications, limitations, damaged IDs, recoverable IDs, and failure reasons +are explicit structured evidence. The existence or row count of a CSV never +upgrades evidence status. + +### Scope disposition + +`scope_disposition` describes inclusion and preservation policy: + +- `included` +- `excluded_from_paper_matrix` +- `held` +- `superseded` + +Excluded, held, failed, and superseded entries may reference genuine preserved +source evidence and lineage. They receive no new evaluable imported result +artifact or metric, but their evidence is not erased. + +### Executability + +`executable` is an explicit boolean wherever a repair/work authorization is +represented. It is not evaluation eligibility and the importer never interprets +it as permission to run model inference. In the Stage 1 profile, historical +condition import records are non-executing, while queue records reproduce the +authoritative true/false execution classification for later stages. + +The evaluator's eligibility predicate is based on successful import, +`evidence_status in {complete, qualified}`, and an included scope disposition. +It does not use `executable`. + +## Identity layers and audit provenance + +The importer separates semantic condition identity, realization identity, +result artifact identity, evaluation identity, and audit-only provenance. These +layers are related but never interchangeable. + +### Semantic condition identity + +`condition_id` remains the scientific grid-cell identity already constructed by +ChoiceBench v0.2. Its canonical payload contains only scientifically meaningful +condition metadata: + +- dataset artifact and exact selection identities, benchmark, and split; +- model identity, including the backend/provider and known consumed generation + settings or revision; +- method identity, including the historical method name, effective parameters, + preflight identity, and method implementation identity where recoverable; +- prompt/template identity; +- seed, calibration/preflight identity, and applicable protocol settings. + +This matches the current model, method, prompt, dataset-selection, and condition +construction rather than defining an importer-only scientific identity. The +canonical payload has a full `condition_digest`; `condition_id` is its existing +filesystem-safe short form. + +Source CSV bytes, CSV dialect/mapping, importer code, evidence status, scope, +authorization, overlay bytes, repair IDs, and lineage are not semantic +condition fields. Changing any of them alone leaves `condition_id` unchanged. +Changing the dataset selection, model, method, prompt, generation semantics, +seed, or protocol settings changes the semantic condition. + +### Import/result realization identity + +A `realization_id` identifies one immutable evidence/provenance record that +asserts or derives results for a semantic condition. Multiple historical, +repaired, transformed, or native realizations may share one `condition_id`. +The canonical realization payload includes: + +- the full semantic `condition_digest`; +- stable import-specification digest; +- logical source records, source classifications, formats, exact source-byte + SHA-256 values, and stable source provenance; +- expected-input snapshot and question-set digests; +- adapter/importer implementation identity, complete CSV dialect, source + mapping, null/numeric/option/extra-field policies; +- validation-findings digest, evidence status, qualification/limitation/defect + digests, and scope disposition; +- structured result origin and ordered prediction-origin assignment digest; +- parent realization/evidence/result digests; +- typed authorization digest; +- overlay source bytes, replacement IDs/reasons, and transformation + input/pre-ownership-output/code digests. + +The output digest in a realization is never the final ChoiceBench CSV digest. +It is either the checksum of a separately supplied precomputed overlay/output +source, or a canonical transformation payload produced before ChoiceBench row +ownership is injected. That canonical payload excludes the child +`realization_id`, result-artifact and experiment identities, output paths, +timestamps, and machine locations. The final CSV checksum appears only in the +post-manifest result-artifact identity, sidecar, and run state. + +Lineage-component IDs included in the realization are likewise computed only +from their operation type, parent/source/evidence/authorization digests, +implementation identity, stable parameters, question ID, and pre-ownership +input/output digests. They never include the child realization, child result +artifact, or current experiment identity. The realization can therefore bind +the ordered row-to-lineage-component mapping without a self-reference. + +`import_state` is not identity-bearing: a dry-run realization that moves from +`validated` to atomically committed `imported` keeps the same planned identity. +A failed transaction publishes no realization. + +Changing source bytes, source checksum, CSV dialect, source mapping, stable +provenance, importer/adapter implementation, evidence classification, scope, +authorization, overlay, repaired IDs, prediction-origin assignment, or +transformation always changes `realization_id`. It changes `condition_id` only +when the scientific condition payload also changes. + +A repaired or transformed output that implements the same intended dataset x +model x method x prompt protocol is a derived realization of the same semantic +condition. If the repair changes the model, prompt, method algorithm, generation +semantics, dataset selection, or protocol, it belongs to a new semantic +condition. Deterministic rematching that reconstructs the declared +`semantic_matching_v1` protocol from preserved Stage-1 responses shares the +base semantic condition; a different matcher definition would change the +method and condition identities. + +### Manifest and experiment identity + +Manifest schema v3 separates `semantic_conditions` from `realizations`. +Semantic-condition records contain no realization back-reference in their +condition identity; grouping is derived from the realization table. Realizations +reference a semantic `condition_id` and own fixed result/evidence +paths plus an expected result ownership/schema contract. The immutable manifest +does not contain the result-artifact ID or digest, which cannot be known until +the exact result bytes exist. New native runs normally have one realization per +condition; imported runs may retain multiple alternative or derived +realizations. Run state is keyed by `realization_id`, not by semantic condition. + +The existing `experiment_id` remains the immutable manifest/run identity. Its +identity payload includes the semantic grid and the selected realization +records, so changing source bytes or lineage changes the experiment identity +without changing the shared scientific condition. A separate +`semantic_grid_digest` over datasets/models/methods/prompts/semantic conditions +supports cross-realization comparison. Legacy manifest-v2 records retain their +existing experiment and condition IDs and are normalized in memory as one +native realization per condition. + +### Result artifact identity + +An evaluable CSV has a separate `result_artifact_id` and full +`result_artifact_digest`. The digest binds: + +- semantic condition and realization full digests; +- manifest experiment ID; +- exact CSV SHA-256, row count, ordered columns, and columns digest; +- row-ownership summary and ordered question-identity digest; +- structured prediction-origin counts and ordered assignment digest; +- evidence and lineage digests. + +The result ID is calculated after the exact CSV bytes are produced. It is stored +in the integrity sidecar and run state, not inside the CSV or immutable manifest. +This is an explicit two-phase boundary: first the manifest/experiment fixes the +semantic conditions, realizations, output contracts, and safe paths; then the +atomic result publication computes and records the result-artifact identity. +The result-artifact digest may therefore bind the already fixed experiment ID +without an identity cycle. Manifest-v3 result paths are fixed by realization, +`results/.csv` and +`results/.artifact.json`; a second realization therefore cannot +collide with the shared semantic condition. Legacy v2 paths remain +`results/.csv`. + +### Evaluation identity + +Evaluation units are realizations grouped under semantic conditions. Alternative +realizations are never silently concatenated. The evaluation identity binds: + +- experiment digest and explicit realization-selection policy; +- every accounted semantic condition and realization full digest; +- evidence/scope/qualification/limitation and stable lineage digests; +- applicable result artifact ID/digest and exact consumed CSV SHA-256; +- metric, parser/scorer/postprocessing, and evaluator implementation identities. + +Ineligible evidence-only realizations remain identity-bound accounting units +with null result fields. Reports expose both +`conditions[condition_id].realization_ids` and realization-level metrics/status. +Changing source/mapping/provenance changes the realization and applicable +result/evaluation identities even when the semantic condition is unchanged. + +Imported sources, datasets, models, methods, prompts, semantic conditions, +realizations, and results carry full SHA-256 digests in addition to short IDs. +Readers verify the full digest before resolving any imported artifact path. + +### Audit-only provenance + +The following are recorded for traceability but excluded from semantic +condition, realization, result-artifact, experiment, and evaluation identities: + +- import timestamp; +- the machine-local absolute source location used during this invocation; +- an absolute `CHOICEBENCH_HOME` or CLI output root; +- temporary paths; +- hostname and other machine-local execution details; +- machine-local import-report location. + +Stage 1's historical original path is preserved as audit provenance. Its +freeze-relative canonical/raw path is stable logical provenance. The absolute +path used to locate the freeze on this machine is audit-only. + +An import-specification digest is computed from a documented stable projection +that excludes operational locations and audit fields. Moving byte-identical +sources or the specification between machines does not change identity. + +Unknown provenance stays explicit. An unknown field contains a null value and a +reason such as `not recorded by producer`; it is never replaced with an +inferred request ID, time, token count, prompt, response, provider field, +commit, seed, revision, or configuration value. + +## Generic import specification + +The generic schema follows existing strict ChoiceBench configuration patterns. +It contains these logical sections: + +- schema/protocol version; +- stable import name and optional source-run declaration; +- one or more source artifacts; +- expected dataset declarations; +- model declarations; +- method declarations; +- prompt/template declarations; +- condition declarations; +- realization and structured-origin declarations; +- typed authorization declarations; +- optional repair or transformation overlays; +- metric selection; +- notes and provenance evidence; +- audit-only location bindings. + +A source artifact declaration can express: + +- user-selected source location; +- expected SHA-256; +- stable logical path/name; +- format and format version; +- the complete identity-bearing decoding and CSV-dialect declaration; +- source run, repository, and commit declarations; +- raw/canonical/derived/repaired/aggregate-only classification; +- exact or allowed source schema; +- source column mapping; +- null and numeric policies; +- extra-field policy; +- arbitrary safe notes and evidence. + +An expected-dataset declaration also states its reference kind and trust basis, +as defined below, rather than representing every reference snapshot as equally +trusted. + +A condition declaration can express: + +- source and expected-dataset references; +- model/backend/provider identity; +- benchmark and split identity; +- snapshot revision/fingerprint; +- historical method identity; +- prompt/template identity; +- known generation parameters; +- the four provenance/status dimensions and `executable` where relevant; +- expected question set; +- qualifications, limitations, and exact damaged/recoverable IDs; +- optional overlay references. + +Credential-named keys are refused by the same identity canonicalization rules +used elsewhere in ChoiceBench. Arbitrary executable Python objects, pickle, and +source-controlled dynamic import targets are not supported. + +## Path and source safety + +Absolute paths are not inherently unsafe. These explicitly user-selected paths +are valid after canonicalization and containment/symlink checks: + +- an absolute `CHOICEBENCH_HOME`; +- an absolute CLI-selected output root; +- an explicitly declared read-only external source such as the Stage 1 freeze. + +The trust boundary is who selects the output, not whether the path begins at the +filesystem root. + +An untrusted import specification cannot redirect writes. Output root selection +comes only from `CHOICEBENCH_HOME` or an explicit CLI option. Specifications +contain source locations and logical artifact relationships, but never an +output destination. Every derived output path is a fixed relative path under +the selected root. + +The importer: + +- resolves and canonicalizes the selected roots before writing; +- rejects `..` traversal or any derived output escaping the output root; +- rejects an output root or run directory that resolves through unsafe + symlinks; +- refuses source symlinks when their target/containment cannot be safely + established; +- never follows a source path into a directory tree for recursive import; +- opens only explicitly referenced row-level artifacts; +- computes SHA-256 independently from the opened bytes; +- parses the same bytes that were hashed, avoiding a hash/parse time-of-check + race; +- rejects a source whose bytes differ from the specification checksum; +- never executes artifact content; +- sanitizes untrusted exception text before logging or reporting it. + +## Evidence store and deduplication + +Every referenced row-level source is archived atomically in the run's import +evidence store. Storage is addressed by the actual source-byte SHA-256, for +example: + +```text +artifacts/imports/evidence/sha256/ce/ced2c5....csv +``` + +The suffix comes from the validated adapter format, not an untrusted filename. +An integrity sidecar records the digest, size, format, and stable source +references. Identical bytes referenced by multiple source declarations or +conditions within the same run reuse one evidence blob. A pre-existing blob is +accepted only after full byte digest and sidecar validation; divergent overwrite +is refused. Separate immutable runs archive their own content-addressed copy even +when bytes match; the initial design deliberately avoids a mutable global store +or cross-run hardlink trust boundary. + +Evidence snapshots are written to a temporary file under the destination, +flushed, verified, and atomically renamed. The evidence index is also atomic and +self-digested. No raw repository, response-cache directory, checkpoint forest, +archive collection, or other unreferenced Stage 1 file is copied. + +For malformed evidence, the byte-identical source snapshot is authoritative. +The CSV adapter retains each logical record's exact raw byte span, including +quoting, delimiters, line endings, and embedded newlines. The +realization-validation artifact records each affected question ID, the byte +offsets, the SHA-256 of those exact source-row bytes, and the validation failure. A +canonical source-row digest may be recorded as additional search/index data, +but never substitutes for the raw-row-byte digest. The importer does not coerce +a malformed prediction into an apparently valid normalized prediction. + +## CSV adapter and extension fields + +The first adapter supports CSV/tabular row sources. It reads source bytes as +data, never as code, preserves header order, and produces source cells without +allowing pandas' default NaN coercion to erase the distinction between an empty +cell and a literal `NaN` string. The import specification or profile declares +the accepted null representation and numeric parsing rules. + +### Deterministic decoding and dialect + +Parsing behavior is fully declared and identity-bearing in the realization +payload because it determines logical records and cell values. The initial CSV +adapter has these explicit defaults: + +```yaml +csv: + encoding: utf-8 + bom_policy: forbid + decoding_errors: strict + delimiter: "," + quote_character: '"' + escape_character: null + double_quote: true + line_terminators: [crlf, lf, cr] + mixed_line_terminators: allow + final_record_without_terminator: allow + blank_record_policy: reject + skip_initial_space: false + header: first_logical_record + strict_syntax: true +``` + +The initial adapter accepts only `encoding: utf-8`. `bom_policy` may instead be +explicitly set to `strip_utf8_bom`; stripping is limited to one leading UTF-8 +BOM and the byte offsets still refer to the original source. Decoding always +fails closed on invalid byte sequences; a replacement-character or ignore +policy is not supported. Delimiter, quote, and non-null escape characters are +restricted in the first adapter to distinct single-byte ASCII characters. +`double_quote` controls whether two consecutive quote characters inside a +quoted field represent one literal quote. + +The listed line terminators are the only recognized record separators and are +recognized only outside a quoted field. Their order is canonical and longest +first, so CRLF is one terminator rather than CR followed by LF. The mixed-line +policy, acceptance of an unterminated final logical record, and blank-record +policy are explicit. Raw spans include the record terminator when one is +present. A binary logical-record scanner uses the declared quote, escape, and +doubled-quote rules before text decoding, retains start/end byte offsets in the +original source, and treats newlines inside quoted fields as field bytes. It +then decodes and parses those exact spans under the same declaration. +Consequently the malformed-row raw-byte-span guarantee remains well-defined for +quoted records containing embedded CR, LF, or CRLF. The fixed UTF-8 encoding +declaration and every declared BOM, delimiter, quoting, escaping, +doubled-quote, terminator, mixed-terminator, final/blank-record, whitespace, +header, decoding-error, and strictness field are identity-bearing. A future +adapter version that supports another encoding necessarily produces a distinct +realization identity, though not a different semantic condition. + +Variable-option questions are supported through either: + +- an explicitly mapped structured choices column; or +- an ordered set/pattern of option columns whose present values determine the + row's option count. + +The adapter never assumes four choices. A three-option ARC row remains a +three-option row; an empty trailing option is not converted into a phantom +choice. + +Every source column has one disposition: + +1. mapped to a ChoiceBench common field; +2. preserved in `external_fields_json` under a source namespace; or +3. explicitly ignored with a documented reason. + +Normal mode preserves safe unmapped fields and reports them. Strict mode +requires every column to be declared and rejects extras. Neither mode silently +drops a source column. Credential-named or secret-bearing metadata is rejected +in both modes. + +`external_fields_json` preserves the original source field names, string/null +values, and method-specific diagnostics without forcing unrelated methods into +a lossy shared schema. This covers Stage-1/Stage-2 responses, matching output, +option scores and parse flags, permutations, votes, provider finish reasons, +failure counts, historical parse data, IHS per-option fields, and other safe +diagnostics. + +Historical method identity is taken from the specification/profile, not renamed +from the source filename or simplified to a current built-in method. The Stage +1 profile preserves at least: + +- `semantic_matching_v1` +- `text_extraction` +- `two_stage_v1` +- `two_stage_v2` +- `two_stage_v3` +- `independent_hypothesis` +- `cyclic_generation_majority` +- `PriDe` + +Source-row aliases such as `cyclic`, `twostage_semantic_match`, and +`two_prompt` are validation evidence interpreted only through explicit profile +mapping. + +## Expected-dataset reference trust + +The reusable core distinguishes two reference kinds: + +- `independent_input_snapshot` — benchmark input artifacts and selection records + that existed independently of the result rows being imported. The declaration + binds their exact checksums, selection identity, transformation chain, and any + independently known publisher revision/fingerprint. Validation may claim + question, option, and gold consistency against this snapshot, while separately + stating the authenticity limit of its provenance chain. +- `profile_derived_reference_snapshot` — a reference reconstructed only from + result-side evidence because no independent input artifact is available. It + requires at least two explicitly declared independent source groups when + possible, exact cross-source agreement for question/gold/options, a complete + derivation record, conflict refusal, and a stable derivation digest. It may + support internal consistency and membership validation, but reports must not + claim independent benchmark truth or upstream dataset authenticity. + +The reference kind, source/checksum chain, declared trust level, cross-source +policy, and derivation digest are realization- and evaluation-identity-bearing. +They do not change semantic dataset identity unless the selected questions or +their semantic content changes. Profiles may not silently upgrade a +result-derived reference to an independent snapshot merely because several +result files agree. + +## Row-level validation + +Validation is by question identity, never by row count alone. All errors name +the condition, realization, source, field, and question ID where possible. No +imported result row is silently dropped, padded, deduplicated, reinterpreted, or +repaired. Any declared expected-dataset derivation is a separate, identity-bound +input transformation and cannot be used to excuse duplicate result rows. + +The validator checks: + +- required mapped fields and source schema; +- unique, non-null question IDs; +- exact expected membership for complete/qualified evidence; +- declared subset membership for partial/malformed/recoverable evidence; +- exact missing and unexpected IDs; +- duplicate IDs, including all duplicate locations; +- declared reference-snapshot gold answer equality, qualified by its trust kind; +- option count and ordered option identity/text where available; +- correct-option mapping into the actual variable-size option set; +- parsed prediction validity or an exact declared damaged/failure record; +- correctness consistency when recomputable; +- method-specific numeric fields for malformed values, NaN, and infinity; +- row ownership by condition; +- source-row model, benchmark, method, and split consistency when those fields + are present; +- evidence status against observed coverage and declared damaged rows; +- all source/checksum/specification digests; +- unknown/extra column policy; +- path and symlink safety. + +The declared expected snapshot supplies the evaluative question text, choices, +correct option, and gold answer. Source-provided versions are compared with it +but never override it. The report qualifies the resulting validation claim by +the snapshot's declared reference kind and trust chain. + +For `complete` and `qualified` evidence, every expected question must have one +valid evaluable prediction. A declared provider failure may occupy a source row, +but that realization is not complete until a valid authorized repair produces a +derived realization whose recomputed coverage supports that status. + +For `partial`, `malformed`, or `recoverable` evidence, known defects are legal +only when the exact IDs and reasons are declared. Any additional defect fails +validation. These sources can be archived and accounted, but they do not +produce an evaluable imported result artifact. + +## ChoiceBench-native-format imported result artifacts + +Only realizations with: + +- `import_state=imported`; +- `evidence_status=complete` or `qualified`; and +- `scope_disposition=included` + +produce an evaluable imported result artifact under +`results/.csv`. + +Its row ownership fields use the ordinary ChoiceBench experiment, condition, +dataset, model, method, prompt, benchmark, and split identities and add +`realization_id`, `prediction_origin`, and `prediction_lineage_id`. The result +sidecar additionally binds the full realization, result-artifact, source, +authorization, lineage, and ordered row-origin digests and explicitly records +the structured `result_origin`. It is a ChoiceBench-native-format imported +result artifact in format and validation only; it never claims ChoiceBench +executed the model. + +Non-evaluable realizations retain their semantic-condition reference, evidence +references, exact validation artifacts, and lineage. They do not receive +placeholder result CSVs or empty metrics. + +## Idempotence and collision safety + +An import transaction holds the normal per-run process lock for validation of +existing state through final atomic publication. + +For a new run, it reuses ChoiceBench's parent-level +`.locks/.manifest.lock`, creates a uniquely named sibling staging +directory on the same filesystem as `runs/`, and writes the complete +manifest, content-addressed evidence, validation artifacts, evaluable results, +sidecars, and final run state there. Every file and cross-reference is validated +from the staged tree; files and directories are flushed/fsynced where the +platform supports it. Publication is one atomic, no-replace directory rename +from the staged tree to the previously absent final run path while the lock is +held. The implementation never uses a replacement rename for a run directory. +If no safe atomic no-replace publication is available on the platform, it fails +closed rather than falling back to incremental final-path writes. + +Validation failure removes only the importer-owned staging tree when safe. A +crash may leave an owner-marked staging tree, but it is not a run, is never +enumerated by readers/evaluators, and can be verified and garbage-collected by a +separate safe maintenance operation. Because the content-addressed evidence +store is inside that staging run, no externally published blob points at an +uncommitted run. Machine-readable reports written outside the run are published +separately and cannot make a failed run appear committed. + +For an already existing final run, the importer performs verify-only behavior: +it validates the entire manifest, state, evidence, result, and sidecar graph. An +exact match is an idempotent no-op; any mismatch is a refusal. It never stages +over or replaces an existing run. + +Repeating an import with the same stable specification projection, source +bytes, expected snapshot, mapping, status dimensions, origins, authorization, +and lineage yields the same semantic condition, realization, result-artifact, +experiment, and evaluation identities. If all existing manifest, evidence, +result, and sidecar bytes validate and match, the command reports an idempotent +no-op. + +If a run ID, realization ID, result-artifact path, source-digest path, or report +identity already exists with different verified content, the importer refuses +the overwrite and names the differing identity-bearing sections. A shared +semantic `condition_id` is not itself a collision: it may legitimately group +several separately addressed realizations. The importer does not offer a silent +reset or destructive replacement path. + +## Repair overlays and offline transformations + +An overlay declaration contains: + +- immutable base realization and semantic-condition full digests; +- mandatory base evidence and realization-validation-artifact checksums; +- a base evaluable-result checksum only when the base realization has one; +- overlay source and independently computed checksum; +- exact replacement question IDs; +- an authorization type and reference to a separately declared, immutable, + checksum-verified authorization record; +- one replacement reason per question; +- output source classification; +- structured result origin and per-row prediction origins; +- transformation or repair implementation identity; +- stable lineage notes. + +An overlay cannot declare or expand its own authorization set. Before overlay +validation, the importer independently validates the referenced authorization +artifact, its checksum, its scope to the base semantic condition and realization, +its exact authorized IDs and reasons, its authority, and the declared operation +type. The transformation/overlay specification may only reference authorization; +it cannot serve as the authority for its own IDs. + +Two authorization types are distinct and non-interchangeable: + +- `inference_repair` authorizes new model inference. For the Stage 1 profile it + must resolve every condition/question pair to the checksum-verified + `approved_rerun_queue.csv` record with `queue_disposition=approved`, + `execution_authority=authoritative`, and `executable=true`. Held, excluded, + forensic, absent, or false-executable records cannot authorize inference. +- `offline_transformation` authorizes deterministic processing of preserved + evidence without model execution. Its immutable authorization source binds + the exact condition/question IDs, reasons, authority, allowed transformation + purpose, input-evidence digests, and expected-dataset digest. It explicitly + records `inference_executable=false`. It neither requires nor fabricates an + approved-rerun-queue entry. + +An authorization record can permit only its named operation type. An +`offline_transformation` record cannot authorize inference, and an +`inference_repair` record does not implicitly authorize unrelated rematching or +postprocessing. + +The overlay engine requires every replacement ID to be in that independently +validated authorization record, the expected dataset, and the base expected +set. It rejects duplicate IDs, unauthorized IDs, unexpected IDs, duplicate +overlay rows, multiple overlays that replace the same ID, inconsistent +gold/options/ownership, and conflicting overlay declarations. It validates +replacement rows with the same rules as base rows. + +Applying an overlay never changes the base evidence snapshot, semantic +condition, base realization, or optional base result. It creates a derived +realization and result artifact with new identities and complete base -> +authorization -> overlay/transformation -> resulting-artifact lineage. It keeps +the shared semantic condition when the scientific protocol is unchanged. The +specification declares the expected derived evidence status; the importer +recomputes it from validated coverage and remaining defects and refuses a +mismatch. A declaration of complete or qualified succeeds only if all remaining +defects are resolved. + +For a repair that combines historical base rows with newly generated rows, the +derived realization records `derivation_origin=repair_overlay`, retains +`external_historical_inference` on unchanged rows, assigns `native_inference` or +`external_repair_inference` to each replacement row as applicable, and records +the full constituent set/counts and ordered mapping. It never erases component +origins or reduces them to `mixed`. + +Offline transformations use the same derived-artifact mechanism. They record +input and pre-ownership output checksums plus transformation code identity under +the non-circular rules above. The final ChoiceBench CSV checksum is added only +after the realization and experiment are fixed. A specification cannot ask +ChoiceBench to import and execute arbitrary code. A transformation is either: + +- a registered ChoiceBench implementation whose code identity is computed by + ChoiceBench; or +- a precomputed external output whose producing code identity/digest is + declared and whose output is independently validated. + +Deterministic rematching of surviving Stage-1 responses is represented with +`derivation_origin=offline_transformation`, while the rematched row retains the +origin of the inference response being rematched. It is not new model inference. + +### Stage 1 offline-transformation authority + +The six recoverable ARC `semantic_matching_v1` cells are authorized separately +from the rerun queue. The Stage 1 profile derives and content-addresses an +immutable `offline_transformation` authorization record from the +checksum-covered recoverable entries in +`manifests/canonical_results_manifest.json` (cross-checked against its CSV form). +The record identifies that manifest as the user-designated Stage 1 authority and +names these exact cells: + +- `cbp__gemini-2-5-flash__arc_challenge__semantic_matching_v1` +- `cbp__gpt-4-1-mini__arc_challenge__semantic_matching_v1` +- `cbp__llama-3-1-8b-instant__arc_challenge__semantic_matching_v1` +- `cbp__meta-llama-llama-3-1-8b-instruct__arc_challenge__semantic_matching_v1` +- `cbp__qwen-qwen2-5-7b-instruct-turbo__arc_challenge__semantic_matching_v1` +- `cbp__qwen-qwen2-5-7b-instruct__arc_challenge__semantic_matching_v1` + +The profile's manifest translation binds each historical `cell_id` to its full +ChoiceBench semantic `condition_digest`; authorization validation checks both +identities. For each it authorizes exactly these three question IDs: + +- `79e8c959bbeb74a0` +- `ad6b5d46ae54842c` +- `c30e75b011696a95` + +Thus it binds exactly six conditions and 18 condition/question pairs, their +manifest reasons, base-source checksums, expected-snapshot digest, the allowed +semantic-rematching purpose, and `inference_executable=false`. The generated +generic transformation specification references this independent authorization +record and digest; it cannot add IDs. `arc_question_audit.csv` may corroborate +row-level evidence but, because it is not itself covered by the freeze checksum +ledger, it is not the authorization trust anchor. + +## Manifest, reader, and evaluation behavior + +Manifest v3 retains the existing models, methods, prompts, datasets, semantic +conditions, and experiment semantics, and adds an identity-bearing realization +table. Each realization references one semantic condition and carries the +source, validation, orthogonal provenance dimensions, structured origin, +authorization, evidence, lineage, and planned result contract/path. After an +evaluable artifact is atomically written, its sidecar and run-state entry—not +the immutable manifest—reference its separate result-artifact identity. +Audit-only provenance is stored in a separate non-identity section whose +exclusion is explicit and validated. + +Manifest validation recomputes imported child full digests and verifies their +short IDs and fixed artifact paths. Result-side identity summaries and column +digests are compared with the CSV and manifest rather than merely stored. + +The publication-grade reader: + +- validates content-addressed evidence snapshots and realization validation + artifacts; +- validates evaluable imported result artifacts through the normal sidecar and + row-ownership path; +- refuses result artifacts for ineligible realizations; +- selects evaluation realizations explicitly and never concatenates alternative + realizations that share a semantic condition; +- returns evaluable rows plus manifest accounting for all imported semantic + conditions and realizations; +- preserves compatibility with legacy native v0.2 manifests. + +Evaluation: + +- computes configured metrics for explicitly selected included + complete/qualified imported realizations; +- includes qualifications and limitations beside qualified metrics; +- accounts for partial, malformed, recoverable, failed, excluded, held, and + superseded evidence without metrics; +- reports semantic-condition and realization counts by each orthogonal dimension + rather than one lossy status tally; +- never estimates metrics for missing, invalid, excluded, or held rows; +- continues to validate dataset snapshots, metric implementation identity, row + ownership, and result checksums. + +Evaluation identity includes only stable semantic provenance and lineage: + +- experiment and semantic-condition full digests; +- explicit realization-selection policy and every accounted realization digest; +- evidence status and scope disposition; +- stable qualification/limitation digest; +- result-artifact, evidence, authorization, structured-origin, and lineage + digests as applicable; +- configured metric and postprocessing implementation identities; +- exact consumed evaluable result checksums. + +It excludes timestamps, absolute paths, output roots, report locations, +hostnames, and temporary paths. + +## CLI and reports + +The installed command follows current project naming conventions: + +```bash +choicebench-import-results SPEC.yaml --run-id external-study + +choicebench-import-results SPEC.yaml \ + --run-id external-study \ + --dry-run \ + --strict \ + --report /tmp/external-study-validation.json + +choicebench-import-results /path/to/paper_data_freeze \ + --profile stage1-paper-freeze \ + --run-id paper-stage1 \ + --dry-run +``` + +Required options/behavior include: + +- real import; +- `--dry-run`/`--validate-only` aliases with no result/evidence writes; +- `--output-root` as an explicit trusted alternative to `CHOICEBENCH_HOME`; +- a required safe run ID; +- `--strict` source-column/schema validation; +- repeated explicit overlay arguments or overlays in the specification; +- human-readable summaries; +- optional machine-readable JSON report; +- mandatory independent checksum verification; +- sanitized, actionable errors. + +`--output-root` is parsed before importing modules that resolve +`CHOICEBENCH_HOME`. An absolute user-selected output root is valid. The import +specification cannot set or override it. + +The machine report includes: + +- stable import-specification, semantic-grid, condition, realization, + result-artifact (when evaluable), and experiment IDs; +- audit-only input/output locations; +- counts by import state, evidence status, scope disposition, and origin; +- evaluable versus evidence-only realization and semantic-condition counts; +- row coverage and exact defect IDs; +- source/checksum/schema results; +- extra-column dispositions; +- overlay authorizations and lineage; +- Stage 1 queue classifications when the profile is used; +- whether the operation wrote artifacts or was an idempotent no-op. + +## Stage 1 paper profile + +The profile is selected through the same command: + +```bash +choicebench-import-results FREEZE_ROOT \ + --profile stage1-paper-freeze \ + --run-id stage1-paper-import +``` + +It reads the current sealed manifests and validator-enforced invariants. It does +not regenerate them from `build_stage1_audit.py`, which predates the final +local-only PriDe amendment. + +Authoritative inputs include, as relevant: + +- `manifests/expected_matrix.csv` +- `manifests/canonical_results_manifest.json` and CSV cross-check +- `manifests/cell_status_matrix.csv` +- `manifests/frozen_artifact_index.csv` +- `manifests/approved_rerun_queue.csv` +- `manifests/held_or_declined_reruns.csv` +- `manifests/paper_scope_excluded_reruns.csv` +- `reports/canonical_freeze_report.md` +- `checksums/checksums.sha256` + +The profile joins by explicit `cell_id` and `canonical_artifact_id`; it does not +infer scientific identity from filenames. Freeze-relative paths are resolved +beneath the explicitly selected absolute freeze root and checked against +traversal and symlinks. + +The profile produces 102 preserved condition/evidence records: + +- exactly 100 have `scope_disposition=included` and constitute the intended + paper matrix; +- exactly two hosted-Qwen PriDe records have + `scope_disposition=excluded_from_paper_matrix` and are never counted as + intended paper cells. + +This permits the hosted Qwen PriDe MMLU record to remain complete and excluded, +while the hosted Qwen PriDe ARC record remains malformed and excluded. + +The intended-matrix evidence counts must be exactly: + +- 57 complete; +- 5 qualified; +- 6 recoverable; +- 4 partial, mapped explicitly from Stage 1 + `incomplete_requires_inference`; +- 28 malformed. + +Across all 102 preserved records, the two exclusions add one complete hosted +PriDe MMLU record and one malformed hosted PriDe ARC record. Thus the preserved +evidence totals are 58 complete, 5 qualified, 6 recoverable, 4 partial, and 29 +malformed, while the scope totals remain 100 included and two excluded. + +The profile verifies all source checksums used by its generated specifications. +It preserves the eight distinct historical CSV schemas and their method-specific +fields. + +Inference-repair queue authority is explicit: + +- `approved_rerun_queue.csv`: 138 question-cells, executable true; +- `held_or_declined_reruns.csv`: 687 Gemini ARC IHS question-cells, + executable false and held; +- `paper_scope_excluded_reruns.csv`: three hosted PriDe question-cells, + executable false and excluded; +- `rerun_queue.csv`: historical forensic evidence only and never execution + authority. + +These queue classifications authorize or decline new model inference only; they +do not authorize offline transformations. The profile imports no repair output +and executes none of these queues. It records the typed authorization boundaries +needed for future overlays and the separate offline-transformation authority +defined above. + +### Stage 1 expected-dataset trust anchor + +Stage 1 does not need to reconstruct expected questions from result CSVs. Its +trust anchor is the class of pre-inference benchmark-input and split artifacts +under `raw/local_model_generalization/data/`, all bound by +`checksums/checksums.sha256`: + +- ARC raw/normalized inputs: + `data/raw/arc_challenge_raw.csv` and + `data/processed/arc_challenge_normalized.csv`; +- ARC selection and metadata: + `data/splits/arc_challenge/robustness_ids.json` and + `robustness_metadata.json`; +- MMLU raw/normalized inputs: + `data/raw/mmlu_raw.csv` and `data/processed/mmlu_normalized.csv`; +- MMLU selection and metadata: + `data/splits/benchmark/robustness_ids.json` and + `robustness_metadata.json`. + +The profile classifies these as `independent_input_snapshot` because they are +benchmark inputs independent of the result rows, with the more precise trust +label `checksum_verified_freeze_internal`. It verifies the freeze checksum +ledger against the opened bytes, parses the raw/normalized data under an +explicit adapter declaration, verifies the split-selection and metadata +digests, applies the declared selection derivation, constructs the ChoiceBench +snapshot, and then cross-checks result-side question, gold, and option evidence. + +The ARC normalized input resolves its 1,000 unique selected IDs directly and +retains 997 four-option and three three-option rows. The MMLU normalized input +contains 1,003 rows for its 1,000 unique selected IDs because each of +`79686d32dfe155ea`, `2f7aa3c7ebb98cfe`, and `74f7227e190200ac` occurs twice. +This is part of the frozen input evidence, not silently invalidated or ignored. +For MMLU the profile reproduces the archived split-construction rule from +`raw/local_model_generalization/scripts/prepare_data.py`: preserve source order +and keep the first occurrence under +`drop_duplicates(subset="question_id", keep="first")`. Before applying it, the +profile requires every duplicate occurrence to agree exactly on all parsed +fields; any conflict fails closed. The registered profile implementation +identity, ordered pre-dedup row digest, exact duplicate-ID/row digests, +first-occurrence policy, and ordered post-dedup digest are identity-bearing. +After that declared derivation, every selected MMLU ID resolves exactly once and +all 1,000 selected rows have four options. + +This chain establishes internal freeze consistency and independence from result +CSVs. It does not establish upstream publisher authenticity: the freeze does not +record a verifiable Hugging Face revision/commit/fingerprint or publisher-signed +checksum for these bytes. The profile therefore preserves unknown upstream +revision fields, reports that limitation, and never labels the snapshot as +publisher-authenticated. Byte-identical copies elsewhere in the freeze and +cross-result agreement are corroboration only, not the trust basis. If these +input artifacts were absent, the profile would have to use the weaker +`profile_derived_reference_snapshot` rules and correspondingly limited claims. + +The generated generic specifications contain the exact reference kind, trust +label, source and selection checksums, transformation/selection derivation, +expected question sets, and snapshot digests. The reusable core receives no +paper-specific inference rule. + +## Security boundaries + +The importer maintains ChoiceBench's fail-closed posture: + +- no pickle or executable-object deserialization; +- no execution of artifact content; +- no arbitrary dynamic transformation import from untrusted YAML; +- no source-provided checksum trust without independent hashing; +- no untrusted specification-controlled output path; +- no traversal or unsafe symlink resolution; +- no credential-named or secret-bearing persisted metadata; +- no divergent overwrite; +- no weakening of canonicalization, manifest, dataset, or evaluation checks; +- no recursive import of repositories, caches, archives, or directory forests; +- sanitized untrusted error text; +- process locking and atomic writes for every published artifact/index. + +Checksums prove internal byte consistency, not producer authenticity. Declared +source repository/model/provider facts remain declared unless independently +resolved from supplied immutable evidence. + +## Testing strategy + +Implementation follows test-driven development with small synthetic fixtures. +No historical dataset or large artifact is copied into the repository. + +Focused tests cover: + +1. Valid generic CSV import. +2. Deterministic import identity. +3. Idempotent repeated import. +4. Divergent overwrite refusal. +5. Source checksum mismatch. +6. Source mutation after specification creation. +7. Duplicate question IDs. +8. Missing question IDs. +9. Unexpected question IDs. +10. Gold-answer mismatch. +11. Variable option counts. +12. Three-option ARC-style rows. +13. Null and NaN behavior. +14. Invalid parsed predictions. +15. Correctness mismatch. +16. Unknown and extra fields in normal/strict modes. +17. Method-specific extra-field preservation. +18. Complete imported evidence. +19. Qualified imported evidence. +20. Partial and malformed imported evidence. +21. Imported-condition evaluation. +22. Incomplete-condition accounting. +23. Base plus valid repair overlay. +24. Unauthorized repair-ID rejection. +25. Conflicting repair-overlay rejection. +26. Repair lineage and identity. +27. Offline-transformation lineage. +28. Paper-profile manifest translation. +29. Hosted PriDe exclusions are not intended paper cells. +30. Held Gemini ARC IHS remains non-executable. +31. Approved queue is not confused with the forensic queue. +32. CLI dry-run. +33. CLI real import. +34. Import and evaluation outside the repository. +35. Wheel/sdist installation retains importer functionality. +36. Security/adversarial cases matching existing release tests. +37. Orthogonal status combinations, including complete+excluded and + malformed+excluded. +38. Audit paths/timestamps do not change import or evaluation identity. +39. New manifests always write explicit structured `result_origin`. +40. Legacy manifests alone may infer native origin. +41. Content-addressed evidence deduplication and divergent-blob refusal. +42. Explicit absolute source/output roots are accepted while spec-controlled + output and traversal are rejected. +43. Malformed row evidence remains byte/digest exact and non-evaluable. +44. Repair overlays on evidence-only bases do not require a nonexistent base + result checksum. +45. Overlay-declared authorization cannot self-authorize replacement IDs. +46. Declared derived evidence status must equal the recomputed status. +47. Source bytes/mapping/provenance change realization and result/evaluation + identities without changing an otherwise identical semantic condition. +48. Scientifically meaningful model/dataset/method/prompt changes do change the + semantic condition. +49. Base, repaired, and transformed realizations share a semantic condition + when protocol semantics are unchanged, and alternatives are never + concatenated for evaluation. +50. Result-artifact identity changes when exact result bytes change. +51. A repaired artifact records constituent origins and the exact per-row origin + mapping; no bare `mixed` origin is accepted. +52. Offline transformation preserves underlying prediction origin while + recording a separate derivation origin and component lineage. +53. `inference_repair` requires an exact approved executable queue binding and + cannot use held, excluded, forensic, or offline authority. +54. `offline_transformation` requires its separate checksum-verified authority, + cannot self-authorize, and does not require executable inference authority. +55. The Stage 1 offline authority contains exactly six semantic-matching cells, + 18 condition/question pairs, and the three declared ARC IDs per cell. +56. Independent-input and profile-derived dataset references retain distinct + trust labels, derivations, and validation claims. +57. The Stage 1 profile anchors expected data to the checksum-verified frozen + benchmark inputs/splits and reports the unknown upstream revision. +58. Every decoding/dialect field participates in realization identity. +59. UTF-8 BOM forbid/strip behavior and strict decoding-error refusal. +60. Delimiter, quote, escape, doubled-quote, whitespace, header, and strict CSV + behavior, including invalid-declaration rejection. +61. CRLF, LF, CR, and mixed-line policies, with exact raw byte spans for quoted + records containing embedded newlines. +62. The Stage 1 MMLU snapshot verifies the three exact duplicate pairs, rejects + conflicting duplicates, applies stable first-occurrence deduplication, and + identity-binds both pre- and post-dedup ordered row digests. +63. Manifest, realization, lineage-component, transformation-output, and + result-artifact identities can be computed in order with no self-reference. +64. Failure/crash before the no-replace staging rename leaves no published final + run, and readers ignore owner-marked staging directories. +65. Existing final runs are verify-only: exact graphs no-op and any divergent + graph is refused without staging over the run. + +Existing native-run tests remain unchanged or receive compatibility coverage. +The full suite must continue to pass. + +## Stage 1 end-to-end validation + +After synthetic tests pass: + +1. Run the Stage 1 profile in dry-run mode across the entire intended matrix. +2. Confirm 100 intended and two excluded preserved cells. +3. Confirm the exact evidence-status and queue counts listed above. +4. Confirm held and excluded work stays non-executable. +5. Confirm every referenced canonical source checksum. +6. Verify both frozen benchmark-input/split chains, their trust labels, all + 2,000 selected IDs, and the three variable-option ARC rows. +7. Verify the separate offline-transformation authorization has exactly six + conditions and 18 authorized condition/question pairs and grants no inference + execution. +8. Perform a real import into a temporary absolute `CHOICEBENCH_HOME` outside + both repositories. +9. Verify the freeze has no filesystem changes. +10. Use native ChoiceBench reading/evaluation to demonstrate: + - one complete imported condition; + - one qualified imported condition; + - one incomplete or malformed evidence-only condition; + - one variable-option ARC condition. +11. Write the end-to-end report outside the repository. + +Generated imports, reports, build artifacts, temporary workspaces, and +historical evidence remain uncommitted. + +## Documentation and packaging + +User documentation will explain: + +- what external import means and why it is not native inference; +- supported CSV format and extension-field preservation; +- import schema fields and examples; +- generic, dry-run, strict, output-root, and Stage 1 profile usage; +- deterministic identity versus audit provenance; +- semantic condition, realization, result-artifact, and evaluation identity; +- checksum and idempotence behavior; +- orthogonal status dimensions and evaluation eligibility; +- mixed prediction origins, repair lineage, and separately authorized offline + transformations; +- expected-dataset trust kinds and deterministic CSV decoding/dialect behavior; +- security boundaries and limitations; +- how to implement a future adapter/profile. + +A concise changelog entry will be added because the repository records notable +features there. The package version will not change. + +Before the pull request, validation includes targeted tests, the full suite, +repository lint/type/static checks that actually exist, security/adversarial +tests, `python -m build`, `python -m twine check dist/*`, the existing isolated +wheel/sdist smoke test, the Stage 1 dry run, the isolated real import, and +representative native evaluation. + +## Review workflow + +The work follows these independent gates: + +1. architecture reconstruction and design review; +2. independent written-specification review; +3. test-driven implementation by a fresh implementer where tasks are safely + separable; +4. independent specification-compliance review; +5. independent code-quality review; +6. independent adversarial/security review; +7. fresh-context final review of the complete diff and validation evidence. + +Reviewers are read-only and do not inherit implementer conclusions. Confirmed +blockers are fixed and relevant validation is rerun before advancing. + +## Out of scope + +This stage does not: + +- modify Stage 1 source artifacts; +- run model inference; +- execute any repair, held, or excluded queue; +- start the final 2x2 experiment; +- rewrite the paper; +- alter historical predictions; +- implement unrelated evaluation modules; +- create a general data-platform abstraction; +- import response caches, checkpoints, or raw repository forests; +- commit imported outputs or historical data; +- change the package version; +- merge the eventual pull request or create a release/tag. + +## Success criteria + +The design is successful when a strict, reusable specification can import +external CSV rows into collision-safe ChoiceBench manifests and +ChoiceBench-native-format imported result artifacts; semantic condition, +realization, result-artifact, and evaluation identities have the specified +separation; every artifact retains structured derivation and exact constituent +prediction origins; non-evaluable evidence remains preserved and honestly +accounted; deterministic identities exclude machine-local audit data; repair +and offline-transformation lineage uses the correct immutable typed authority; +dataset trust and CSV parsing behavior are explicit; the Stage 1 profile +reproduces all sealed counts without modifying the freeze; native +reading/evaluation and installed-package workflows work outside the repository; +and the feature is submitted on its isolated branch as an unmerged pull request. diff --git a/src/choicebench/identity.py b/src/choicebench/identity.py index 2ee381e..ec3a492 100644 --- a/src/choicebench/identity.py +++ b/src/choicebench/identity.py @@ -167,6 +167,8 @@ def canonicalize(value: Any, *, redact_secrets: bool = True) -> Any: return sorted(items, key=lambda item: json.dumps(item, sort_keys=True, separators=(",", ":"), ensure_ascii=True)) if isinstance(value, Path): return value.as_posix() + if isinstance(value, (bytes, bytearray)): + return bytes(value).hex() if isinstance(value, str): text = unicodedata.normalize("NFC", value) return redact_text(text) if redact_secrets else text diff --git a/src/choicebench/importing/__init__.py b/src/choicebench/importing/__init__.py new file mode 100644 index 0000000..378d594 --- /dev/null +++ b/src/choicebench/importing/__init__.py @@ -0,0 +1,51 @@ +"""Public types for strict external result import declarations.""" + +from choicebench.importing.schema import ( + AuthorizationSpec, + CsvDialectSpec, + DatasetReferenceSpec, + DerivationOrigin, + EvidenceStatus, + ImportConditionSpec, + ImportMethodSpec, + ImportModelSpec, + ImportPromptSpec, + ImportSpec, + ImportSpecError, + ImportState, + NumericColumnSpec, + OptionMappingSpec, + OverlaySpec, + PredictionOrigin, + ResultOriginSpec, + ScopeDisposition, + SourceArtifactSpec, + import_spec_digest, + load_import_spec, + stable_import_projection, +) + +__all__ = [ + "AuthorizationSpec", + "CsvDialectSpec", + "DatasetReferenceSpec", + "DerivationOrigin", + "EvidenceStatus", + "ImportConditionSpec", + "ImportMethodSpec", + "ImportModelSpec", + "ImportPromptSpec", + "ImportSpec", + "ImportSpecError", + "ImportState", + "NumericColumnSpec", + "OptionMappingSpec", + "OverlaySpec", + "PredictionOrigin", + "ResultOriginSpec", + "ScopeDisposition", + "SourceArtifactSpec", + "import_spec_digest", + "load_import_spec", + "stable_import_projection", +] diff --git a/src/choicebench/importing/authorization.py b/src/choicebench/importing/authorization.py new file mode 100644 index 0000000..00c968e --- /dev/null +++ b/src/choicebench/importing/authorization.py @@ -0,0 +1,191 @@ +"""Validate typed repair/offline-transformation authorization bundles. + +An authorization bundle grants specific (condition_digest, question_id) pairs +permission to be replaced, typed as either `inference_repair` (executable -- +a real model may be run) or `offline_transformation` (non-executing semantic +rematching only). This module never runs inference or opens a filesystem +path itself; it validates an already-opened source and already-known +condition digests/expected datasets. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from typing import Mapping + +from choicebench.identity import integrity_digest, short_id +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.dataset_reference import ExpectedDataset +from choicebench.importing.schema import AuthorizationSpec + + +class AuthorizationError(ValueError): + """Raised when an authorization bundle or its use is invalid or unsafe.""" + + +@dataclass(frozen=True) +class ValidatedAuthorizationBundle: + authorization_id: str + authorization_digest: str + authorization_type: str + grants: Mapping[str, Mapping[str, str]] # condition_digest -> question_id -> reason + authority: str + purpose: str + executable: bool + source_sha256: str + input_evidence_digests: Mapping[str, str] + expected_snapshot_digests: Mapping[str, str] + + +@dataclass(frozen=True) +class ValidatedAuthorization: + bundle_id: str + bundle_digest: str + authorization_type: str + condition_digest: str + question_reasons: Mapping[str, str] + authority: str + purpose: str + executable: bool + source_sha256: str + input_evidence_digests: Mapping[str, str] + expected_snapshot_digest: str + + +def validate_authorization_bundle( + declaration: AuthorizationSpec, + *, + opened_source: OpenedSource, + condition_digests: Mapping[str, str], + expected: Mapping[str, ExpectedDataset], +) -> ValidatedAuthorizationBundle: + """Validate one authorization artifact against an independently opened + source and the caller's own known condition digests/expected datasets. + + Fails closed on: source tampering, an authorization_type/executable + mismatch, a grant referencing an unknown condition or an out-of-selection + question, an empty reason, or a self-assigned authorization_id that does + not match the bundle's own recomputed digest. + """ + if opened_source.source_id != declaration.source_id: + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} declares source " + f"{declaration.source_id!r} but was validated against a differently " + f"declared opened source {opened_source.source_id!r}." + ) + if declaration.authorization_type == "inference_repair" and not declaration.executable: + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} declares inference_repair " + "but executable=false." + ) + if declaration.authorization_type == "offline_transformation" and declaration.executable: + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} declares " + "offline_transformation but executable=true." + ) + + grants: dict[str, dict[str, str]] = {} + for condition_digest, question_reasons in declaration.condition_question_reasons.items(): + if condition_digest not in condition_digests.values(): + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} grants a condition " + f"{condition_digest!r} that is not among the caller's known conditions." + ) + dataset = expected.get(condition_digest) + if not isinstance(question_reasons, Mapping) or not question_reasons: + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} has an empty grant for " + f"condition {condition_digest!r}." + ) + selected = set(dataset.selected_question_ids) if dataset is not None else None + normalized_reasons: dict[str, str] = {} + for question_id, reason in question_reasons.items(): + if not isinstance(reason, str) or not reason.strip(): + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} grant for " + f"{condition_digest!r}/{question_id!r} has no explicit reason." + ) + if selected is not None and question_id not in selected: + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} grants question " + f"{question_id!r} that is not in condition {condition_digest!r}'s " + "expected selection." + ) + normalized_reasons[question_id] = reason + grants[condition_digest] = normalized_reasons + if not grants: + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} grants no conditions." + ) + + bundle_payload = { + "schema_version": "choicebench.authorization-bundle.v1", + "authorization_type": declaration.authorization_type, + "grants": grants, + "authority": declaration.authority, + "purpose": declaration.purpose, + "executable": declaration.executable, + "source_sha256": opened_source.sha256, + "input_evidence_digests": dict(declaration.input_evidence_digests), + "expected_snapshot_digests": dict(declaration.expected_snapshot_digests), + } + authorization_digest = integrity_digest(bundle_payload) + expected_id = short_id("auth", bundle_payload) + if declaration.authorization_id != expected_id: + raise AuthorizationError( + f"Authorization ID {declaration.authorization_id!r} does not match its own " + f"recomputed digest (expected {expected_id!r}); it cannot self-assign an ID." + ) + if set(declaration.expected_snapshot_digests) != set(grants): + raise AuthorizationError( + f"Authorization {declaration.authorization_id!r} expected_snapshot_digests " + "do not exactly cover its granted conditions." + ) + + return ValidatedAuthorizationBundle( + authorization_id=expected_id, + authorization_digest=authorization_digest, + authorization_type=declaration.authorization_type, + grants=grants, + authority=declaration.authority, + purpose=declaration.purpose, + executable=declaration.executable, + source_sha256=opened_source.sha256, + input_evidence_digests=dict(declaration.input_evidence_digests), + expected_snapshot_digests=dict(declaration.expected_snapshot_digests), + ) + + +def authorization_for_condition( + bundle: ValidatedAuthorizationBundle, *, condition_digest: str +) -> ValidatedAuthorization: + """Slice one condition's grant out of a bundle without weakening its + aggregate identity: bundle_id/bundle_digest are the whole bundle's, so a + slice can be traced back to (and cannot silently diverge from) the + exact bundle it came from, and it can never borrow another condition's + grants.""" + question_reasons = bundle.grants.get(condition_digest) + if question_reasons is None: + raise AuthorizationError( + f"Authorization {bundle.authorization_id!r} grants no questions for " + f"condition {condition_digest!r}; this condition cannot self-authorize." + ) + expected_snapshot_digest = bundle.expected_snapshot_digests.get(condition_digest) + if expected_snapshot_digest is None: + raise AuthorizationError( + f"Authorization {bundle.authorization_id!r} has no expected snapshot digest " + f"for condition {condition_digest!r}." + ) + return ValidatedAuthorization( + bundle_id=bundle.authorization_id, + bundle_digest=bundle.authorization_digest, + authorization_type=bundle.authorization_type, + condition_digest=condition_digest, + question_reasons=dict(question_reasons), + authority=bundle.authority, + purpose=bundle.purpose, + executable=bundle.executable, + source_sha256=bundle.source_sha256, + input_evidence_digests=dict(bundle.input_evidence_digests), + expected_snapshot_digest=expected_snapshot_digest, + ) diff --git a/src/choicebench/importing/csv_adapter.py b/src/choicebench/importing/csv_adapter.py new file mode 100644 index 0000000..fd98212 --- /dev/null +++ b/src/choicebench/importing/csv_adapter.py @@ -0,0 +1,438 @@ +"""Deterministic byte-preserving adapter for external CSV result sources.""" + +from __future__ import annotations + +import csv +from dataclasses import asdict, dataclass +from hashlib import sha256 +import io +import json +import math +from pathlib import Path +import re +from typing import Mapping, Protocol + +from choicebench.identity import canonicalize, is_credential_key +from choicebench.importing.schema import CsvDialectSpec, SourceArtifactSpec + + +_UTF8_BOM = b"\xef\xbb\xbf" +_TERMINATORS = {"crlf": b"\r\n", "lf": b"\n", "cr": b"\r"} +_ACTUAL_TERMINATORS = (b"\r\n", b"\n", b"\r") +_INTEGER_RE = re.compile(r"[+-]?\d+\Z") +_FLOAT_RE = re.compile( + r"[+-]?(?:(?:\d+(?:\.\d*)?)|(?:\.\d+))(?:[eE][+-]?\d+)?\Z" +) +_NONFINITE_FLOAT_RE = re.compile( + r"[+-]?(?:nan|inf(?:inity)?)\Z", flags=re.IGNORECASE +) + + +class CsvAdapterError(ValueError): + """Raised when source CSV bytes do not satisfy their declaration.""" + + +@dataclass(frozen=True) +class OpenedSource: + source_id: str + audit_path: Path + logical_path: str + data: bytes + sha256: str + + +@dataclass(frozen=True) +class LogicalRecordSpan: + index: int + start: int + end: int + terminator: bytes + + +@dataclass(frozen=True) +class SourceRow: + values: Mapping[str, str | None] + span: LogicalRecordSpan + raw_sha256: str + + +@dataclass(frozen=True) +class AdaptedTable: + columns: tuple[str, ...] + rows: tuple[SourceRow, ...] + source_sha256: str + + +class SourceAdapter(Protocol): + def parse( + self, source: OpenedSource, declaration: SourceArtifactSpec, *, strict: bool + ) -> AdaptedTable: ... + + +def _dialect_bytes(dialect: CsvDialectSpec) -> tuple[int, int, int | None]: + values = (dialect.delimiter, dialect.quote_character, dialect.escape_character) + encoded: list[int | None] = [] + for name, value in zip(("delimiter", "quote", "escape"), values, strict=True): + if value is None: + encoded.append(None) + continue + raw = value.encode("utf-8") + if len(raw) != 1 or raw[0] >= 128 or raw in {b"\x00", b"\r", b"\n"}: + raise CsvAdapterError(f"CSV {name} must be one usable ASCII byte.") + encoded.append(raw[0]) + non_null = [value for value in encoded if value is not None] + if len(non_null) != len(set(non_null)): + raise CsvAdapterError("CSV delimiter, quote, and escape bytes must be distinct.") + return encoded[0], encoded[1], encoded[2] # type: ignore[return-value] + + +def _declared_terminators(dialect: CsvDialectSpec) -> tuple[bytes, ...]: + try: + declared = tuple(_TERMINATORS[name] for name in dialect.line_terminators) + except KeyError as exc: + raise CsvAdapterError(f"Unsupported CSV line terminator {exc.args[0]!r}.") from exc + if not declared or len(declared) != len(set(declared)): + raise CsvAdapterError("CSV line terminators must be a non-empty unique list.") + return tuple(sorted(declared, key=len, reverse=True)) + + +def scan_csv_logical_records( + data: bytes, dialect: CsvDialectSpec +) -> tuple[LogicalRecordSpan, ...]: + """Return exact source-byte spans for logical CSV records.""" + if not isinstance(data, bytes): + raise TypeError(f"data must be bytes; got {type(data).__name__}.") + delimiter, quote, escape = _dialect_bytes(dialect) + declared_terminators = set(_declared_terminators(dialect)) + spans: list[LogicalRecordSpan] = [] + observed_terminators: set[bytes] = set() + start = 0 + if data.startswith(_UTF8_BOM): + if dialect.bom_policy == "forbid": + raise CsvAdapterError("Source contains a forbidden UTF-8 BOM.") + index = len(_UTF8_BOM) + else: + index = 0 + content_start = index + in_quotes = False + at_field_start = True + + while index < len(data): + byte = data[index] + if in_quotes: + if escape is not None and byte == escape: + if index + 1 >= len(data): + raise CsvAdapterError( + f"CSV escape byte at byte {index} has no following byte." + ) + index += 2 + continue + if byte == quote: + if ( + dialect.double_quote + and index + 1 < len(data) + and data[index + 1] == quote + ): + index += 2 + continue + in_quotes = False + index += 1 + continue + + if escape is not None and byte == escape: + if index + 1 >= len(data): + raise CsvAdapterError( + f"CSV escape byte at byte {index} has no following byte." + ) + at_field_start = False + index += 2 + continue + terminator = next( + ( + candidate + for candidate in _ACTUAL_TERMINATORS + if data.startswith(candidate, index) + ), + None, + ) + if terminator is not None: + if terminator not in declared_terminators: + raise CsvAdapterError( + f"CSV contains undeclared line terminator {terminator!r} " + f"at byte {index}." + ) + end = index + len(terminator) + if index == content_start: + raise CsvAdapterError(f"CSV blank logical record at byte {start} is forbidden.") + observed_terminators.add(terminator) + if ( + dialect.mixed_line_terminators == "forbid" + and len(observed_terminators) > 1 + ): + raise CsvAdapterError("CSV contains forbidden mixed line terminators.") + spans.append(LogicalRecordSpan(len(spans), start, end, terminator)) + start = end + index = end + content_start = end + at_field_start = True + continue + if byte == delimiter: + at_field_start = True + elif byte == quote and at_field_start: + in_quotes = True + at_field_start = False + elif not (dialect.skip_initial_space and at_field_start and byte == 0x20): + at_field_start = False + index += 1 + + if in_quotes: + raise CsvAdapterError( + f"CSV has an unclosed quote in logical record starting at byte {start}." + ) + if start < len(data): + if dialect.final_record_without_terminator == "forbid": + raise CsvAdapterError("CSV final logical record has no declared terminator.") + spans.append(LogicalRecordSpan(len(spans), start, len(data), b"")) + return tuple(spans) + + +def effective_csv_adapter_projection( + declaration: SourceArtifactSpec, *, strict: bool +) -> dict[str, object]: + """Return parsing controls that must participate in realization identity.""" + effective_extra_policy = ( + "reject_unmapped" + if strict or declaration.extra_field_policy == "reject_unmapped" + else "preserve_unmapped" + ) + return canonicalize( + { + "adapter": "choicebench.csv.v1", + "dialect": asdict(declaration.dialect), + "declared_extra_field_policy": declaration.extra_field_policy, + "cli_strict": strict, + "effective_extra_field_policy": effective_extra_policy, + } + ) + + +def _decode_record( + data: bytes, + span: LogicalRecordSpan, + dialect: CsvDialectSpec, + *, + strip_bom: bool, +) -> list[str]: + record = data[span.start : span.end - len(span.terminator) if span.terminator else span.end] + stripped_prefix_size = 0 + if strip_bom: + stripped_prefix_size = len(_UTF8_BOM) + record = record[len(_UTF8_BOM) :] + try: + text = record.decode("utf-8", errors="strict") + except UnicodeDecodeError as exc: + raise CsvAdapterError( + f"Invalid UTF-8 in CSV logical record {span.index} at source byte " + f"{span.start + stripped_prefix_size + exc.start}." + ) from exc + try: + parsed = list( + csv.reader( + io.StringIO(text, newline=""), + delimiter=dialect.delimiter, + quotechar=dialect.quote_character, + escapechar=dialect.escape_character, + doublequote=dialect.double_quote, + skipinitialspace=dialect.skip_initial_space, + strict=dialect.strict_syntax, + ) + ) + except csv.Error as exc: + raise CsvAdapterError( + f"Malformed CSV syntax in logical record {span.index}: {type(exc).__name__}." + ) from exc + if len(parsed) != 1: + raise CsvAdapterError( + f"CSV logical record {span.index} parsed into {len(parsed)} physical rows." + ) + return parsed[0] + + +def _validate_numeric( + value: str | None, *, source_column: str, value_type: str, + null_allowed: bool, finite_only: bool, record_index: int +) -> None: + if value is None: + if not null_allowed: + raise CsvAdapterError( + f"CSV null is forbidden for numeric column {source_column!r} " + f"in logical record {record_index}." + ) + return + if value_type == "integer": + if _INTEGER_RE.fullmatch(value) is None: + raise CsvAdapterError( + f"CSV integer column {source_column!r} has a malformed value in " + f"logical record {record_index}." + ) + return + if _FLOAT_RE.fullmatch(value) is not None: + number = float(value) + elif _NONFINITE_FLOAT_RE.fullmatch(value) is not None: + number = float(value) + else: + raise CsvAdapterError( + f"CSV finite float column {source_column!r} has a malformed value in " + f"logical record {record_index}." + ) + if finite_only and not math.isfinite(number): + raise CsvAdapterError( + f"CSV finite float column {source_column!r} has a non-finite value in " + f"logical record {record_index}." + ) + + +_POSITIONAL_LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZ" + + +def _validate_structured_choices( + value: str | None, declaration: SourceArtifactSpec, *, record_index: int +) -> None: + mapping = declaration.option_mapping + if mapping.mode != "structured_json": + return + if value is None: + raise CsvAdapterError( + f"CSV structured choices are null in logical record {record_index}." + ) + try: + choices = json.loads(value) + except (json.JSONDecodeError, UnicodeError) as exc: + raise CsvAdapterError( + f"CSV structured choices are malformed in logical record {record_index}." + ) from exc + if not isinstance(choices, list) or not choices: + raise CsvAdapterError( + f"CSV structured choices must be a non-empty list in logical record {record_index}." + ) + assert mapping.structured_label_key is not None + assert mapping.structured_text_key is not None + labels: set[str] = set() + for choice in choices: + if not isinstance(choice, Mapping): + raise CsvAdapterError( + f"CSV structured choices contain a non-object in logical record {record_index}." + ) + raw_label = choice.get(mapping.structured_label_key) + text = choice.get(mapping.structured_text_key) + # A structured-choice source may declare an explicit string letter + # label (e.g. {"label": "A", "text": ...}), or a positional integer + # index instead (e.g. {"source_index": 0, "text": ...} -- the same + # convention choicebench's own dataset-side choices_json already + # uses, see dataset_reference.py). The latter is letterized here + # (0 -> "A", 1 -> "B", ...) rather than requiring every producer to + # pre-compute letters that don't otherwise exist in their data. + if isinstance(raw_label, bool): + label = None + elif isinstance(raw_label, int): + label = ( + _POSITIONAL_LETTERS[raw_label] + if 0 <= raw_label < len(_POSITIONAL_LETTERS) else None + ) + elif isinstance(raw_label, str) and raw_label: + label = raw_label + else: + label = None + if label is None or not isinstance(text, str): + raise CsvAdapterError( + "CSV structured choices lack a string label or a valid positional " + f"integer label, or a string text field, in logical record {record_index}." + ) + if label in labels: + raise CsvAdapterError( + f"CSV structured choices contain duplicate labels in logical record {record_index}." + ) + labels.add(label) + + +def parse_csv_source( + source: OpenedSource, declaration: SourceArtifactSpec, *, strict: bool +) -> AdaptedTable: + """Validate and parse one already-opened, checksum-bound CSV source.""" + if source.source_id != declaration.source_id: + raise CsvAdapterError("Opened source_id does not match its source declaration.") + if source.logical_path != declaration.logical_path: + raise CsvAdapterError("Opened source logical_path does not match its declaration.") + actual_sha256 = sha256(source.data).hexdigest() + if source.sha256 != actual_sha256: + raise CsvAdapterError("Opened source checksum does not match its bytes.") + if declaration.expected_sha256 != actual_sha256: + raise CsvAdapterError("Source checksum does not match the import declaration.") + if declaration.format != "csv": + raise CsvAdapterError(f"CSV adapter cannot parse format {declaration.format!r}.") + if source.data.startswith(_UTF8_BOM) and declaration.dialect.bom_policy == "forbid": + raise CsvAdapterError("Source contains a forbidden UTF-8 BOM.") + + spans = scan_csv_logical_records(source.data, declaration.dialect) + if not spans: + raise CsvAdapterError("CSV source has no header logical record.") + header = tuple( + _decode_record( + source.data, + spans[0], + declaration.dialect, + strip_bom=( + declaration.dialect.bom_policy == "strip_utf8_bom" + and source.data.startswith(_UTF8_BOM) + ), + ) + ) + duplicates = sorted({column for column in header if header.count(column) > 1}) + if duplicates: + raise CsvAdapterError(f"CSV duplicate header columns are forbidden: {duplicates}.") + credential_columns = sorted(column for column in header if is_credential_key(column)) + if credential_columns: + raise CsvAdapterError( + f"CSV credential-named columns are forbidden: {credential_columns}." + ) + missing = sorted(set(declaration.expected_columns) - set(header)) + if missing: + raise CsvAdapterError(f"CSV is missing declared source columns: {missing}.") + extras = sorted(set(header) - set(declaration.expected_columns)) + effective = effective_csv_adapter_projection(declaration, strict=strict) + if extras and effective["effective_extra_field_policy"] == "reject_unmapped": + raise CsvAdapterError(f"CSV has unexpected source columns: {extras}.") + + rows: list[SourceRow] = [] + for span in spans[1:]: + cells = _decode_record( + source.data, span, declaration.dialect, strip_bom=False + ) + if len(cells) != len(header): + raise CsvAdapterError( + f"CSV logical record {span.index} field count {len(cells)} does not " + f"match header field count {len(header)}." + ) + values: dict[str, str | None] = { + column: None if cell in declaration.null_values else cell + for column, cell in zip(header, cells, strict=True) + } + for numeric in declaration.numeric_columns: + _validate_numeric( + values[numeric.source_column], + source_column=numeric.source_column, + value_type=numeric.value_type, + null_allowed=numeric.null_allowed, + finite_only=numeric.finite_only, + record_index=span.index, + ) + if declaration.option_mapping.mode == "structured_json": + assert declaration.option_mapping.structured_column is not None + _validate_structured_choices( + values[declaration.option_mapping.structured_column], + declaration, + record_index=span.index, + ) + raw = source.data[span.start : span.end] + rows.append(SourceRow(values, span, sha256(raw).hexdigest())) + return AdaptedTable(header, tuple(rows), actual_sha256) diff --git a/src/choicebench/importing/dataset_reference.py b/src/choicebench/importing/dataset_reference.py new file mode 100644 index 0000000..6f48d5a --- /dev/null +++ b/src/choicebench/importing/dataset_reference.py @@ -0,0 +1,912 @@ +"""Trust-qualified expected-dataset snapshots for external result imports.""" + +from __future__ import annotations + +from dataclasses import dataclass +from hashlib import sha256 +from io import BytesIO +import json +from pathlib import Path, PurePosixPath +from typing import Any, Literal, Mapping + +import pandas as pd + +from choicebench.datasets import ( + dataset_content_digest, + dataset_sample_identities, + validate_normalized_dataset, +) +from choicebench.identity import canonicalize, integrity_digest, short_id +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.schema import DatasetReferenceSpec +from choicebench.infra.artifacts import atomic_write_json, atomic_write_text +from choicebench.pipeline.options import build_option_map, correct_option_for_row + + +class DatasetReferenceError(ValueError): + """Raised when an expected-dataset reference is incomplete or inconsistent.""" + + +@dataclass(frozen=True) +class ExpectedDataset: + dataset_id: str + benchmark_name: str + split: str + artifact_id: str + artifact_digest: str + artifact_payload: Mapping[str, Any] + selection_id: str + selection_digest: str + selection_payload: Mapping[str, Any] + selection_semantics: Mapping[str, Any] + selection_unknown_reasons: Mapping[str, str] + identity_mode: Literal["imported_semantic_fallback", "native_compatibility"] + artifact_frame: pd.DataFrame + frame: pd.DataFrame + selected_question_ids: tuple[str, ...] + reference_kind: Literal[ + "independent_input_snapshot", "profile_derived_reference_snapshot" + ] + trust_label: str + question_set_digest: str + snapshot_digest: str + derivation: Mapping[str, Any] + derivation_digest: str + limitations: tuple[str, ...] + + +_SNAPSHOT_SCHEMA = "choicebench.expected-dataset.v1" +_REFERENCE_SCHEMA = "choicebench.dataset-reference.v1" +_REQUIRED_FIELDS = ("question_id", "question_text", "correct_option") +_NATIVE_DATASET_SPEC_KEYS = { + "benchmark", + "split", + "hf_path", + "hf_subset", + "source_revision", + "normalization_version", + "transforms", + "output_name", +} +_SNAPSHOT_RECORD_KEYS = { + "schema_version", + "dataset_id", + "benchmark_name", + "split", + "artifact_id", + "artifact_digest", + "artifact_payload", + "selection_id", + "selection_digest", + "selection_payload", + "selection_semantics", + "selection_unknown_reasons", + "identity_mode", + "selected_question_ids", + "question_set_digest", + "snapshot_digest", + "reference_kind", + "trust_label", + "derivation", + "derivation_digest", + "limitations", + "run_snapshot_path", + "reference_metadata_path", + "snapshot_sha256", + "row_count", + "record_digest", +} + + +def _source_frame(source_id: str, source: OpenedSource) -> pd.DataFrame: + if source.source_id != source_id: + raise DatasetReferenceError( + f"Expected dataset source key {source_id!r} does not match opened source_id " + f"{source.source_id!r}." + ) + actual = sha256(source.data).hexdigest() + if source.sha256 != actual: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} checksum does not match its bytes." + ) + try: + return pd.read_csv( + BytesIO(source.data), dtype=str, keep_default_na=False, na_filter=False + ) + except (UnicodeError, pd.errors.ParserError, pd.errors.EmptyDataError) as exc: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} is not a valid UTF-8 CSV." + ) from exc + + +def _choice_keys(columns: Mapping[str, str]) -> tuple[str, ...]: + keys = tuple( + sorted( + (key for key in columns if key.startswith("choice_") and len(key) == 8), + key=lambda key: key.removeprefix("choice_"), + ) + ) + return keys + + +def _normalize_source( + declaration: DatasetReferenceSpec, source_id: str, source: OpenedSource +) -> pd.DataFrame: + raw = _source_frame(source_id, source) + mapping = dict(declaration.columns) + missing_common = sorted(set(_REQUIRED_FIELDS) - set(mapping)) + if missing_common: + raise DatasetReferenceError( + f"Expected dataset mapping is missing required fields {missing_common}." + ) + missing_source = sorted(set(mapping.values()) - set(raw.columns)) + if missing_source: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} is missing mapped columns {missing_source}." + ) + if len(mapping.values()) != len(set(mapping.values())): + raise DatasetReferenceError("Expected dataset mapping reuses a source column.") + + records: list[dict[str, Any]] = [] + choice_keys = _choice_keys(mapping) + has_structured_choices = "choices_json" in mapping + if not choice_keys and not has_structured_choices: + raise DatasetReferenceError("Expected dataset mapping declares no ordered options.") + if choice_keys: + expected_keys = tuple( + f"choice_{chr(ord('a') + index)}" for index in range(len(choice_keys)) + ) + if choice_keys != expected_keys: + raise DatasetReferenceError( + "Expected dataset ordered choice keys must be contiguous from choice_a." + ) + + ordinary_keys = sorted( + key for key in mapping if key not in choice_keys and key != "choices_json" + ) + for row_number, raw_row in raw.iterrows(): + record = {key: raw_row[mapping[key]] for key in ordinary_keys} + question_id = str(record["question_id"]).strip() + if not question_id: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has an empty question_id at row " + f"{row_number + 2}." + ) + record["question_id"] = question_id + record["correct_option"] = str(record["correct_option"]).strip().upper() + for integer_field in ("correct_index", "n_choices"): + if integer_field not in record: + continue + raw_integer = str(record[integer_field]).strip() + try: + parsed_integer = int(raw_integer) + except ValueError as exc: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has malformed {integer_field} " + f"for question_id {question_id!r}." + ) from exc + if str(parsed_integer) != raw_integer or parsed_integer < 0: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has malformed {integer_field} " + f"for question_id {question_id!r}." + ) + record[integer_field] = parsed_integer + + if has_structured_choices: + raw_choices = raw_row[mapping["choices_json"]] + try: + choices = json.loads(raw_choices) + except (json.JSONDecodeError, TypeError) as exc: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has malformed ordered options " + f"for question_id {question_id!r}." + ) from exc + if not isinstance(choices, list): + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has malformed ordered options " + f"for question_id {question_id!r}." + ) + normalized_choices = [] + source_indices: set[int] = set() + for index, item in enumerate(choices): + if not isinstance(item, Mapping) or not isinstance(item.get("text"), str): + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has malformed ordered options " + f"for question_id {question_id!r}." + ) + raw_source_index = item.get("source_index", index) + try: + if isinstance(raw_source_index, bool): + raise ValueError + source_index = int(raw_source_index) + except (TypeError, ValueError) as exc: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has an invalid source_index " + f"for question_id {question_id!r}." + ) from exc + if source_index < 0 or source_index in source_indices: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has an invalid source_index " + f"for question_id {question_id!r}." + ) + if not item["text"].strip(): + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has an empty structured " + f"option for question_id {question_id!r}." + ) + source_indices.add(source_index) + normalized_choices.append( + { + "text": item["text"], + "source_index": source_index, + } + ) + else: + choice_values = [str(raw_row[mapping[key]]) for key in choice_keys] + nonempty_indices = [ + index for index, value in enumerate(choice_values) if value.strip() + ] + if nonempty_indices: + last_nonempty = nonempty_indices[-1] + if any(not value.strip() for value in choice_values[:last_nonempty]): + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has an empty option followed " + f"by a populated option for question_id {question_id!r}." + ) + choice_values = choice_values[: last_nonempty + 1] + normalized_choices = [ + {"text": value, "source_index": index} + for index, value in enumerate(choice_values) + if value.strip() + ] + record["choices_json"] = json.dumps( + normalized_choices, ensure_ascii=True, separators=(",", ":") + ) + try: + options = build_option_map(record) + record["correct_option"] = correct_option_for_row(record, options) + except (TypeError, ValueError, json.JSONDecodeError) as exc: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has invalid ordered options or " + f"correct_option for question_id {question_id!r}: {exc}" + ) from exc + records.append(record) + + frame = pd.DataFrame(records) + if frame.empty: + raise DatasetReferenceError(f"Expected dataset source {source_id!r} has no rows.") + duplicated = frame["question_id"].astype(str).duplicated(keep=False) + if duplicated.any(): + duplicate_ids = sorted(frame.loc[duplicated, "question_id"].astype(str).unique()) + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} has duplicate question_id values " + f"{duplicate_ids}." + ) + return frame + + +def _selected_frame( + frame: pd.DataFrame, expected_question_ids: tuple[str, ...] +) -> pd.DataFrame: + if len(expected_question_ids) != len(set(expected_question_ids)): + duplicates = sorted( + question_id + for question_id in set(expected_question_ids) + if expected_question_ids.count(question_id) > 1 + ) + raise DatasetReferenceError( + f"Expected dataset declaration has duplicate selected question ID(s) {duplicates}." + ) + indexed = frame.set_index(frame["question_id"].astype(str), drop=False) + missing = [question_id for question_id in expected_question_ids if question_id not in indexed.index] + if missing: + raise DatasetReferenceError( + f"Expected dataset is missing selected question ID(s) {missing}." + ) + selected = indexed.loc[list(expected_question_ids)].reset_index(drop=True) + return selected + + +def _semantic_disagreement(left: pd.DataFrame, right: pd.DataFrame) -> str | None: + for left_row, right_row in zip( + left.to_dict("records"), right.to_dict("records"), strict=True + ): + question_id = str(left_row["question_id"]) + if left_row["question_text"] != right_row["question_text"]: + return f"question_text disagreement for question_id {question_id!r}" + if left_row["correct_option"] != right_row["correct_option"]: + return f"correct_option disagreement for question_id {question_id!r}" + if build_option_map(left_row) != build_option_map(right_row): + return f"ordered options disagreement for question_id {question_id!r}" + return None + + +def _profile_groups(declaration: DatasetReferenceSpec) -> tuple[tuple[str, ...], ...]: + raw = declaration.derivation.get("independent_source_groups") + if not isinstance(raw, (list, tuple)): + raise DatasetReferenceError( + "Profile-derived reference requires at least two independent source groups." + ) + groups: list[tuple[str, ...]] = [] + for item in raw: + if not isinstance(item, (list, tuple)) or not item: + raise DatasetReferenceError( + "Profile-derived reference requires at least two independent source groups." + ) + groups.append(tuple(str(source_id) for source_id in item)) + if len(groups) < 2: + raise DatasetReferenceError( + "Profile-derived reference requires at least two independent source groups." + ) + flattened = [source_id for group in groups for source_id in group] + if len(flattened) != len(set(flattened)) or set(flattened) != set(declaration.source_ids): + raise DatasetReferenceError( + "Profile-derived reference independent source groups must be disjoint and " + "cover every declared source." + ) + return tuple(groups) + + +def _fallback_artifact_payload( + frame: pd.DataFrame, declaration: DatasetReferenceSpec +) -> dict[str, Any]: + return { + "schema_version": "choicebench.semantic-dataset.v1", + "benchmark": declaration.benchmark_name, + "split": declaration.split, + "content_digest": dataset_content_digest(frame), + } + + +def _selection_payload( + frame: pd.DataFrame, declaration: DatasetReferenceSpec, artifact_id: str +) -> dict[str, Any]: + return { + "artifact_id": artifact_id, + "content_digest": dataset_content_digest(frame), + "sample_identities": dataset_sample_identities(frame), + "seed": declaration.selection_seed, + "n_samples": declaration.selection_n_samples, + "subject_filter": sorted(declaration.subject_filter), + } + + +def _validated_identity_records( + artifact_frame: pd.DataFrame, + selected_frame: pd.DataFrame, + declaration: DatasetReferenceSpec, +) -> tuple[ + Mapping[str, Any], + str, + str, + Mapping[str, Any], + str, + str, + Literal["imported_semantic_fallback", "native_compatibility"], +]: + native = declaration.native_compatibility_identity + if native is None: + artifact_payload = canonicalize( + _fallback_artifact_payload(artifact_frame, declaration) + ) + artifact_digest = integrity_digest(artifact_payload) + artifact_id = short_id("ds", artifact_payload) + selection_payload = canonicalize( + _selection_payload(selected_frame, declaration, artifact_id) + ) + return ( + artifact_payload, + artifact_digest, + artifact_id, + selection_payload, + integrity_digest(selection_payload), + short_id("sel", selection_payload), + "imported_semantic_fallback", + ) + + required = { + "artifact_payload", + "artifact_digest", + "artifact_id", + "selection_payload", + "selection_digest", + "selection_id", + } + if not isinstance(native, Mapping) or set(native) != required: + raise DatasetReferenceError( + "Expected dataset native compatibility identity has invalid fields." + ) + artifact_payload = canonicalize(native["artifact_payload"]) + selection_payload = canonicalize(native["selection_payload"]) + artifact_digest = integrity_digest(artifact_payload) + artifact_id = short_id("ds", artifact_payload) + selection_digest = integrity_digest(selection_payload) + selection_id = short_id("sel", selection_payload) + if ( + native["artifact_digest"] != artifact_digest + or native["artifact_id"] != artifact_id + or native["selection_digest"] != selection_digest + or native["selection_id"] != selection_id + ): + raise DatasetReferenceError( + "Expected dataset native compatibility identity claims are invalid." + ) + if not isinstance(artifact_payload, Mapping) or set(artifact_payload) != { + "spec", + "content_digest", + "source", + }: + raise DatasetReferenceError( + "Expected dataset native compatibility artifact payload is invalid." + ) + spec = artifact_payload.get("spec") + if ( + not isinstance(spec, Mapping) + or set(spec) != _NATIVE_DATASET_SPEC_KEYS + or not isinstance(artifact_payload.get("source"), Mapping) + or spec.get("benchmark") != declaration.benchmark_name + or spec.get("split") != declaration.split + or artifact_payload.get("content_digest") + != dataset_content_digest(artifact_frame) + ): + raise DatasetReferenceError( + "Expected dataset native compatibility content does not own the selected rows." + ) + expected_selection = _selection_payload(selected_frame, declaration, artifact_id) + if selection_payload != canonicalize(expected_selection): + raise DatasetReferenceError( + "Expected dataset native compatibility selection does not own the selected rows." + ) + return ( + artifact_payload, + artifact_digest, + artifact_id, + selection_payload, + selection_digest, + selection_id, + "native_compatibility", + ) + + +def build_expected_dataset( + declaration: DatasetReferenceSpec, + opened_sources: Mapping[str, OpenedSource], +) -> ExpectedDataset: + """Build one semantic snapshot plus its trust-qualified reference record.""" + missing_sources = sorted(set(declaration.source_ids) - set(opened_sources)) + if missing_sources: + raise DatasetReferenceError( + f"Expected dataset declaration is missing opened source(s) {missing_sources}." + ) + if declaration.selection_source_id not in declaration.source_ids: + raise DatasetReferenceError("selection_source_id is not a declared dataset source.") + + normalized = { + source_id: _normalize_source(declaration, source_id, opened_sources[source_id]) + for source_id in declaration.source_ids + } + for source_id, source_frame in normalized.items(): + try: + validate_normalized_dataset( + source_frame, source=f"expected dataset source {source_id!r}" + ) + except Exception as exc: + raise DatasetReferenceError( + f"Expected dataset source {source_id!r} is invalid: {exc}" + ) from exc + selected = { + source_id: _selected_frame(frame, declaration.expected_question_ids) + for source_id, frame in normalized.items() + } + if declaration.reference_kind == "profile_derived_reference_snapshot": + _profile_groups(declaration) + baseline = selected[declaration.selection_source_id] + for source_id in declaration.source_ids: + disagreement = _semantic_disagreement(baseline, selected[source_id]) + if disagreement is not None: + raise DatasetReferenceError( + f"Profile-derived reference source {source_id!r} has {disagreement}." + ) + + artifact_frame = normalized[declaration.selection_source_id].copy() + frame = selected[declaration.selection_source_id].copy() + try: + validate_normalized_dataset(frame, source="expected dataset reference") + except Exception as exc: + raise DatasetReferenceError(f"Expected dataset reference is invalid: {exc}") from exc + + unknown_reasons = canonicalize(dict(declaration.selection_unknown_reasons)) + expected_unknown_keys = { + field + for field, value in ( + ("selection_seed", declaration.selection_seed), + ("selection_n_samples", declaration.selection_n_samples), + ) + if value is None + } + if set(unknown_reasons) != expected_unknown_keys or not all( + isinstance(reason, str) and reason.strip() + for reason in unknown_reasons.values() + ): + raise DatasetReferenceError( + "Expected dataset selection unknown reasons do not match null semantics." + ) + if ( + declaration.selection_n_samples is not None + and declaration.selection_n_samples != len(frame) + ): + raise DatasetReferenceError( + "Expected dataset selection n_samples does not match selected question " + "coverage." + ) + ( + artifact_payload, + artifact_digest, + artifact_id, + selection_payload, + selection_digest, + selection_id, + identity_mode, + ) = _validated_identity_records(artifact_frame, frame, declaration) + selection_semantics = canonicalize( + { + "seed": declaration.selection_seed, + "n_samples": declaration.selection_n_samples, + "subject_filter": sorted(declaration.subject_filter), + } + ) + + limitations = list(declaration.limitations) + if declaration.revision is None and "publisher revision was not recorded" not in limitations: + limitations.append("publisher revision was not recorded") + if ( + declaration.reference_kind == "profile_derived_reference_snapshot" + and "not independently authenticated" not in limitations + ): + limitations.append("not independently authenticated") + + source_chain = [ + { + "source_id": source_id, + "logical_path": opened_sources[source_id].logical_path, + "sha256": opened_sources[source_id].sha256, + } + for source_id in declaration.source_ids + ] + derivation = canonicalize( + { + "schema_version": _REFERENCE_SCHEMA, + "dataset_id": declaration.dataset_id, + "reference_kind": declaration.reference_kind, + "trust_label": declaration.trust_label, + "source_chain": source_chain, + "selection_source_id": declaration.selection_source_id, + "columns": dict(declaration.columns), + "revision": declaration.revision, + "fingerprint": declaration.fingerprint, + "declared_derivation": dict(declaration.derivation), + "selected_question_ids": list(declaration.expected_question_ids), + "selection_semantics": selection_semantics, + "selection_unknown_reasons": unknown_reasons, + "identity_mode": identity_mode, + "limitations": limitations, + } + ) + derivation_digest = integrity_digest(derivation) + question_set_digest = integrity_digest(list(declaration.expected_question_ids)) + snapshot_digest = integrity_digest( + { + "artifact_digest": artifact_digest, + "selection_digest": selection_digest, + "question_set_digest": question_set_digest, + "derivation_digest": derivation_digest, + "reference_kind": declaration.reference_kind, + "trust_label": declaration.trust_label, + } + ) + return ExpectedDataset( + dataset_id=declaration.dataset_id, + benchmark_name=declaration.benchmark_name, + split=declaration.split, + artifact_id=artifact_id, + artifact_digest=artifact_digest, + artifact_payload=artifact_payload, + selection_id=selection_id, + selection_digest=selection_digest, + selection_payload=selection_payload, + selection_semantics=selection_semantics, + selection_unknown_reasons=unknown_reasons, + identity_mode=identity_mode, + artifact_frame=artifact_frame, + frame=frame, + selected_question_ids=tuple(declaration.expected_question_ids), + reference_kind=declaration.reference_kind, + trust_label=declaration.trust_label, + question_set_digest=question_set_digest, + snapshot_digest=snapshot_digest, + derivation=derivation, + derivation_digest=derivation_digest, + limitations=tuple(limitations), + ) + + +def _safe_relative_path(value: Any, field: str) -> Path: + if not isinstance(value, str): + raise DatasetReferenceError(f"Expected dataset {field} must be a relative path.") + pure = PurePosixPath(value) + if pure.is_absolute() or ".." in pure.parts: + raise DatasetReferenceError(f"Expected dataset {field} is unsafe.") + return Path(*pure.parts) + + +def _contained_path(root: Path, value: Any, field: str) -> Path: + relative = _safe_relative_path(value, field) + try: + canonical_root = root.resolve(strict=True) + except OSError as exc: + raise DatasetReferenceError("Expected dataset output root is unreadable.") from exc + candidate = canonical_root / relative + current = canonical_root + for part in relative.parts: + current = current / part + if current.is_symlink(): + raise DatasetReferenceError( + f"Expected dataset {field} contains an unsafe symlink." + ) + try: + candidate.resolve(strict=False).relative_to(canonical_root) + except (OSError, ValueError) as exc: + raise DatasetReferenceError(f"Expected dataset {field} is unsafe.") from exc + return candidate + + +def _snapshot_record(dataset: ExpectedDataset) -> dict[str, Any]: + snapshot_path = f"artifacts/datasets/{dataset.selection_id}.csv" + metadata_path = f"artifacts/datasets/{dataset.selection_id}.reference.json" + csv_bytes = dataset.frame.to_csv(index=False, lineterminator="\n").encode("utf-8") + record: dict[str, Any] = { + "schema_version": _SNAPSHOT_SCHEMA, + "dataset_id": dataset.dataset_id, + "benchmark_name": dataset.benchmark_name, + "split": dataset.split, + "artifact_id": dataset.artifact_id, + "artifact_digest": dataset.artifact_digest, + "artifact_payload": dataset.artifact_payload, + "selection_id": dataset.selection_id, + "selection_digest": dataset.selection_digest, + "selection_payload": dataset.selection_payload, + "selection_semantics": dataset.selection_semantics, + "selection_unknown_reasons": dataset.selection_unknown_reasons, + "identity_mode": dataset.identity_mode, + "selected_question_ids": list(dataset.selected_question_ids), + "question_set_digest": dataset.question_set_digest, + "snapshot_digest": dataset.snapshot_digest, + "reference_kind": dataset.reference_kind, + "trust_label": dataset.trust_label, + "derivation": dataset.derivation, + "derivation_digest": dataset.derivation_digest, + "limitations": list(dataset.limitations), + "run_snapshot_path": snapshot_path, + "reference_metadata_path": metadata_path, + "snapshot_sha256": sha256(csv_bytes).hexdigest(), + "row_count": len(dataset.frame), + } + record["record_digest"] = integrity_digest(record) + return record + + +def write_expected_snapshot(staged_run: Path, dataset: ExpectedDataset) -> dict[str, Any]: + """Atomically write the selected semantic CSV and its self-digesting record.""" + root = Path(staged_run) + record = _snapshot_record(dataset) + snapshot_path = _contained_path(root, record["run_snapshot_path"], "run_snapshot_path") + metadata_path = _contained_path( + root, + record["reference_metadata_path"], "reference_metadata_path" + ) + atomic_write_text(snapshot_path, dataset.frame.to_csv(index=False, lineterminator="\n")) + atomic_write_json(metadata_path, record) + return record + + +def _validate_snapshot_record_shape(record: Mapping[str, Any]) -> None: + if set(record) != _SNAPSHOT_RECORD_KEYS: + raise DatasetReferenceError("Expected dataset reference record fields are invalid.") + if record["schema_version"] != _SNAPSHOT_SCHEMA: + raise DatasetReferenceError("Expected dataset reference schema is unsupported.") + nonempty_strings = { + "dataset_id", + "benchmark_name", + "split", + "artifact_id", + "artifact_digest", + "selection_id", + "selection_digest", + "question_set_digest", + "snapshot_digest", + "trust_label", + "derivation_digest", + "run_snapshot_path", + "reference_metadata_path", + "snapshot_sha256", + "record_digest", + } + if any( + not isinstance(record[field], str) or not record[field] + for field in nonempty_strings + ): + raise DatasetReferenceError("Expected dataset reference record fields are invalid.") + if record["reference_kind"] not in { + "independent_input_snapshot", + "profile_derived_reference_snapshot", + } or record["identity_mode"] not in { + "imported_semantic_fallback", + "native_compatibility", + }: + raise DatasetReferenceError("Expected dataset reference record fields are invalid.") + if not all( + isinstance(record[field], Mapping) + for field in ( + "artifact_payload", + "selection_payload", + "selection_semantics", + "selection_unknown_reasons", + "derivation", + ) + ): + raise DatasetReferenceError("Expected dataset reference record fields are invalid.") + selected = record["selected_question_ids"] + limitations = record["limitations"] + if ( + not isinstance(selected, list) + or not selected + or not all(isinstance(value, str) and value for value in selected) + or len(selected) != len(set(selected)) + or not isinstance(limitations, list) + or not all(isinstance(value, str) for value in limitations) + or isinstance(record["row_count"], bool) + or not isinstance(record["row_count"], int) + or record["row_count"] < 1 + ): + raise DatasetReferenceError("Expected dataset reference record fields are invalid.") + + +def validate_expected_snapshot(run_dir: Path, record: Mapping[str, Any]) -> None: + """Recompute every snapshot, semantic identity, and reference digest.""" + _validate_snapshot_record_shape(record) + root = Path(run_dir) + raw_record = dict(record) + claimed_digest = raw_record.pop("record_digest", None) + if claimed_digest != integrity_digest(raw_record): + raise DatasetReferenceError("Expected dataset reference record integrity failed.") + snapshot_path = _contained_path(root, record["run_snapshot_path"], "run_snapshot_path") + metadata_path = _contained_path( + root, record["reference_metadata_path"], "reference_metadata_path" + ) + try: + persisted = json.loads(metadata_path.read_text(encoding="utf-8")) + except (OSError, json.JSONDecodeError, UnicodeError) as exc: + raise DatasetReferenceError("Expected dataset reference metadata is unreadable.") from exc + if canonicalize(persisted) != canonicalize(dict(record)): + raise DatasetReferenceError("Expected dataset reference metadata integrity failed.") + try: + snapshot_bytes = snapshot_path.read_bytes() + except OSError as exc: + raise DatasetReferenceError("Expected dataset snapshot is unreadable.") from exc + if sha256(snapshot_bytes).hexdigest() != record.get("snapshot_sha256"): + raise DatasetReferenceError("Expected dataset snapshot integrity failed.") + try: + frame = pd.read_csv(BytesIO(snapshot_bytes), dtype=str, keep_default_na=False, na_filter=False) + validate_normalized_dataset(frame, source="expected dataset snapshot") + except Exception as exc: + raise DatasetReferenceError("Expected dataset snapshot integrity failed.") from exc + selected_ids = tuple(frame["question_id"].astype(str)) + if selected_ids != tuple(record.get("selected_question_ids", ())): + raise DatasetReferenceError("Expected dataset snapshot question ownership is invalid.") + if len(frame) != record.get("row_count"): + raise DatasetReferenceError("Expected dataset snapshot row count is invalid.") + if integrity_digest(list(selected_ids)) != record.get("question_set_digest"): + raise DatasetReferenceError("Expected dataset snapshot question-set identity is invalid.") + derivation = record.get("derivation") + if integrity_digest(derivation) != record.get("derivation_digest"): + raise DatasetReferenceError("Expected dataset derivation integrity failed.") + + artifact_payload = canonicalize(record["artifact_payload"]) + if record["identity_mode"] == "imported_semantic_fallback": + if ( + not isinstance(artifact_payload, Mapping) + or set(artifact_payload) + != {"schema_version", "benchmark", "split", "content_digest"} + or artifact_payload["schema_version"] + != "choicebench.semantic-dataset.v1" + or artifact_payload["benchmark"] != record["benchmark_name"] + or artifact_payload["split"] != record["split"] + ): + raise DatasetReferenceError("Expected dataset artifact identity is invalid.") + else: + if not isinstance(artifact_payload, Mapping) or set(artifact_payload) != { + "spec", + "content_digest", + "source", + }: + raise DatasetReferenceError("Expected dataset artifact identity is invalid.") + native_spec = artifact_payload["spec"] + if ( + not isinstance(native_spec, Mapping) + or set(native_spec) != _NATIVE_DATASET_SPEC_KEYS + or not isinstance(artifact_payload["source"], Mapping) + or native_spec.get("benchmark") != record["benchmark_name"] + or native_spec.get("split") != record["split"] + ): + raise DatasetReferenceError("Expected dataset artifact identity is invalid.") + content_digest = artifact_payload.get("content_digest") + if ( + not isinstance(content_digest, str) + or len(content_digest) != 64 + or any(character not in "0123456789abcdef" for character in content_digest) + ): + raise DatasetReferenceError("Expected dataset artifact identity is invalid.") + artifact_digest = integrity_digest(artifact_payload) + artifact_id = short_id("ds", artifact_payload) + if artifact_digest != record.get("artifact_digest") or artifact_id != record.get("artifact_id"): + raise DatasetReferenceError("Expected dataset artifact identity is invalid.") + selection_semantics = record["selection_semantics"] + if ( + not isinstance(selection_semantics, Mapping) + or set(selection_semantics) != {"seed", "n_samples", "subject_filter"} + or ( + selection_semantics["seed"] is not None + and not isinstance(selection_semantics["seed"], int) + ) + or ( + selection_semantics["n_samples"] is not None + and not isinstance(selection_semantics["n_samples"], int) + ) + or not isinstance(selection_semantics["subject_filter"], list) + or not all( + isinstance(subject, str) + for subject in selection_semantics["subject_filter"] + ) + ): + raise DatasetReferenceError("Expected dataset selection semantics are invalid.") + selection_payload = { + "artifact_id": artifact_id, + "content_digest": dataset_content_digest(frame), + "sample_identities": dataset_sample_identities(frame), + "seed": selection_semantics["seed"], + "n_samples": selection_semantics["n_samples"], + "subject_filter": selection_semantics["subject_filter"], + } + if canonicalize(record["selection_payload"]) != selection_payload: + raise DatasetReferenceError("Expected dataset selection identity is invalid.") + unknown_reasons = record["selection_unknown_reasons"] + expected_unknown_keys = { + field + for field, value in ( + ("selection_seed", selection_semantics["seed"]), + ("selection_n_samples", selection_semantics["n_samples"]), + ) + if value is None + } + if set(unknown_reasons) != expected_unknown_keys or not all( + isinstance(reason, str) and reason.strip() + for reason in unknown_reasons.values() + ): + raise DatasetReferenceError("Expected dataset selection semantics are invalid.") + if ( + integrity_digest(selection_payload) != record.get("selection_digest") + or short_id("sel", selection_payload) != record.get("selection_id") + ): + raise DatasetReferenceError("Expected dataset selection identity is invalid.") + expected_snapshot_digest = integrity_digest( + { + "artifact_digest": record.get("artifact_digest"), + "selection_digest": record.get("selection_digest"), + "question_set_digest": record.get("question_set_digest"), + "derivation_digest": record.get("derivation_digest"), + "reference_kind": record.get("reference_kind"), + "trust_label": record.get("trust_label"), + } + ) + if expected_snapshot_digest != record.get("snapshot_digest"): + raise DatasetReferenceError("Expected dataset reference snapshot identity is invalid.") diff --git a/src/choicebench/importing/engine.py b/src/choicebench/importing/engine.py new file mode 100644 index 0000000..07c04eb --- /dev/null +++ b/src/choicebench/importing/engine.py @@ -0,0 +1,1195 @@ +"""Orchestrate generic dry-run and real imports of external results. + +Reduced scope: this implements both the base (non-overlay) import path -- +plan, dry-run validate, real staged publish, and idempotent re-verification +of an existing run -- and the overlay merge-into-new-run path: an authorized +repair/offline-transformation overlay (Task 10's derive_overlay) merged with +an already-published base run's retained rows into a new, separate, +immutable run (execute_overlay_import). The base run itself is never +mutated; only its verified values are read. + +Report schema is also reduced from the original plan: it carries the counts +and digests needed to confirm exact status/queue reproduction and reject +held/excluded work, not the full historical checksum-report/column- +disposition/source-classification breakdown. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from pathlib import Path +from types import SimpleNamespace +from typing import Any, Literal, Mapping + +from dataclasses import asdict +import json +import math + +import pandas as pd + +from choicebench.identity import CANONICALIZATION_VERSION, canonicalize, integrity_digest +from choicebench.importing.csv_adapter import parse_csv_source +from choicebench.importing.dataset_reference import ExpectedDataset, build_expected_dataset +from choicebench.importing.evidence import ( + evidence_record, + open_verified_source, + validate_evidence_index, + write_evidence_blob, + write_evidence_index, +) +from choicebench.importing.identity import ( + ImportIdentityError, + build_import_semantic_identity, + importer_implementation_identity, + make_lineage_component, + make_realization, + make_result_origin, +) +from choicebench.importing.authorization import ( + authorization_for_condition, + validate_authorization_bundle, +) +from choicebench.importing.overlays import VerifiedBaseRealization, derive_overlay +from choicebench.importing.schema import ( + AuthorizationSpec, + ImportConditionSpec, + ImportSpec, + OverlaySpec, + SourceArtifactSpec, + validate_csv_dialect_identity, + validate_numeric_columns_identity, + validate_option_mapping_identity, +) +from choicebench.importing.transaction import ImportTransaction +from choicebench.importing.validation import ( + ImportValidationError, + RealizationValidation, + _expected_choice_letters, + normalize_realization_rows, + prepare_realization_validation_artifact, + validate_realization_validation_artifact, + validate_source_rows, + write_realization_validation_artifact, +) +from choicebench.io.writers import ( + prepare_manifest_result, + publish_manifest_result, + validate_result_artifact, +) +from choicebench.infra.artifacts import atomic_write_json +from choicebench.manifest import ( + MANIFEST_FILENAME, + PROTOCOL_V3_VERSION, + RUN_STATE_FILENAME, + initial_run_state_v3, + make_manifest_v3, + validate_manifest_v3, + validate_run_state_v3, +) + +_ADAPTER_STRICT_KEY = "strict" + + +class ImportEngineError(ValueError): + """Raised when an import request or an existing run fails validation.""" + + +def _external_import_adapter(*, strict: bool) -> None: + """Identity marker for this engine's CSV-adaptation stage. + + Deliberately a trivial, side-effect-free function rather than binding + directly to parse_csv_source/validate_source_rows: those reference many + stdlib globals (hashlib.sha256, re, etc.) that are neither plain data nor + inspectable user-level code, and Task 5's runtime-callable identity fails + closed on exactly that shape rather than silently ignoring it. Binding + here instead still captures "did this engine's own orchestration change" + (implementation_identity hashes this whole file); the installed + choicebench package version already in importer_implementation_identity's + "package" field is the primary signal for "did the installed release + (including csv_adapter.py/validation.py) change." + """ + return None + + +def _external_import_validator() -> None: + """Identity marker for this engine's row-validation stage. See + _external_import_adapter for why this is a marker, not the real function.""" + return None + + +@dataclass(frozen=True) +class ImportRequest: + spec: ImportSpec + run_id: str + workspace_root: Path + strict: bool + overlays: tuple[Any, ...] = () + expected_datasets: Mapping[str, ExpectedDataset] | None = None + """Pre-built, pre-verified ExpectedDataset objects keyed by dataset_id, + used in place of rebuilding each spec.datasets declaration from its + source_ids via the generic build_expected_dataset. Some profiles (e.g. + Stage 1's ARC/MMLU trust chain) need dataset-specific revalidation and + deduplication the generic path cannot reproduce; spec.datasets is still + populated for documentation/identity purposes, but is not read here when + this override is supplied.""" + source_containment_root: Path | None = None + """Root every declared source path must resolve under, checked by + open_verified_source. Defaults to workspace_root (the common case: a + self-contained fixture where sources and the run output share a root). + Set this separately when sources live under an independent, read-only + root distinct from where runs are written (e.g. importing a real, + immutable data freeze into an isolated temporary CHOICEBENCH_HOME).""" + + +@dataclass(frozen=True) +class ImportPlan: + manifest: Mapping[str, Any] + expected_datasets: Mapping[str, ExpectedDataset] + realizations: Mapping[str, Any] + evidence_records: tuple[Mapping[str, Any], ...] + findings: tuple[Any, ...] + + +@dataclass(frozen=True) +class VerifiedImportRun: + manifest: Mapping[str, Any] + manifest_digest: str + realizations: Mapping[str, VerifiedBaseRealization] + + +@dataclass(frozen=True) +class ImportCounts: + conditions: Mapping[str, int] + evidence_status: Mapping[str, int] + scope_disposition: Mapping[str, int] + + +@dataclass(frozen=True) +class DefectReport: + total_findings: int + by_code: Mapping[str, int] + findings_digest: str + + +@dataclass(frozen=True) +class ImportReport: + schema_version: Literal["choicebench.import-report.v1"] + import_state: Literal["validated", "imported", "failed"] + wrote_artifacts: bool + idempotent_noop: bool + run_id: str + experiment_id: str | None + counts: ImportCounts + defects: DefectReport + condition_digests: Mapping[str, str] + realization_digests: Mapping[str, str] + failures: tuple[Mapping[str, Any], ...] + + +def _source_record(declaration, opened, *, notes_digest: str | None) -> dict[str, Any]: + unknown_reasons = { + field: "not declared in generic import specification" + for field, value in ( + ("source_run_id", declaration.source_run_id), + ("source_repository", declaration.source_repository), + ("source_commit", declaration.source_commit), + ) + if value is None + } + return { + "source_id": declaration.source_id, + "logical_path": declaration.logical_path, + "classification": declaration.classification, + "format": declaration.format, + "format_version": declaration.format_version, + "sha256": opened.sha256, + "provenance": { + "source_run_id": declaration.source_run_id, + "source_repository": declaration.source_repository, + "source_commit": declaration.source_commit, + "notes_digest": notes_digest, + "evidence_digest": opened.sha256, + "unknown_reasons": unknown_reasons, + }, + } + + +def _build_realization_for_condition( + condition: ImportConditionSpec, + *, + spec: ImportSpec, + dataset: ExpectedDataset, + model, + method, + prompt, + workspace_root: Path, + strict: bool, +) -> tuple[dict[str, Any], Any, tuple[dict[str, Any], ...]]: + """Returns (realization_record, RealizationValidation, evidence_records).""" + sources_by_id = {source.source_id: source for source in spec.sources} + semantic = build_import_semantic_identity( + condition=condition, dataset=dataset, model=model, method=method, prompt=prompt + ) + + opened_by_source: dict[str, Any] = {} + tables_by_source: dict[str, Any] = {} + for source_id in condition.source_ids: + declaration = sources_by_id[source_id] + opened = open_verified_source(declaration, containment_root=workspace_root) + opened_by_source[source_id] = opened + tables_by_source[source_id] = parse_csv_source(opened, declaration, strict=strict) + + primary_source_id = condition.source_ids[0] + primary_declaration = sources_by_id[primary_source_id] + primary_table = tables_by_source[primary_source_id] + validated = validate_source_rows( + primary_table, source_id=primary_source_id, mapping=primary_declaration.columns, + condition=condition, expected=dataset, + ) + result = normalize_realization_rows( + validated, condition=condition, expected=dataset, mapping=primary_declaration.columns, + ) + + source_records = [] + evidence_records: list[dict[str, Any]] = [] + all_source_digests = [] + for source_id in condition.source_ids: + declaration = sources_by_id[source_id] + opened = opened_by_source[source_id] + notes_digest = integrity_digest(declaration.notes) if declaration.notes else None + source_records.append(_source_record(declaration, opened, notes_digest=notes_digest)) + all_source_digests.append(opened.sha256) + evidence_records.append( + evidence_record(opened, references=sorted(validated.rows_by_question_id)) + ) + all_source_digests = tuple(sorted(set(all_source_digests))) + + # Build the realization's row-level lineage/result-origin from the full + # published canonical row set when the cell is structurally complete (this + # is populated for evaluable AND for malformed-but-structurally-complete + # in-scope cells); for an evaluable cell published_rows == evaluable_rows, + # so this is a no-op there. A genuinely incomplete/out-of-scope cell has + # neither set and produces an empty (evidence-only) result origin. + rows_for_result = result.published_rows if result.published_rows else result.evaluable_rows + lineage_components = [] + for row in rows_for_result: + qid = row["question_id"] + raw_row = dict(validated.rows_by_question_id[qid].values) + component = make_lineage_component( + operation_type="external_import", + question_id=qid, + parent_digests=(), + source_digests=all_source_digests, + authorization_digest=None, + implementation=_external_import_adapter, + parameters={_ADAPTER_STRICT_KEY: strict}, + input_digest=integrity_digest({"question_id": qid, "raw_row": raw_row}), + preownership_output_digest=integrity_digest({"question_id": qid, "row": row}), + prediction_origin=row["prediction_origin"], + ) + lineage_components.append(component) + + row_assignments = [ + ( + component["identity"]["question_id"], + component["identity"]["prediction_origin"], + component["lineage_id"], + ) + for component in lineage_components + ] + result_origin = make_result_origin(derivation_origin="external_import", row_assignments=row_assignments) + + qualification_digest = integrity_digest(list(result.qualifications)) + limitation_digest = integrity_digest(list(result.limitations)) + defect_digest = integrity_digest(list(result.defect_question_ids)) + + runtime_importer = importer_implementation_identity( + adapter=_external_import_adapter, validator=_external_import_validator + ) + + realization_identity = { + "import_spec_digest": integrity_digest(canonicalize(spec.provenance)), + "sources": source_records, + "expected_dataset": { + "snapshot_digest": dataset.snapshot_digest, + "question_set_digest": dataset.question_set_digest, + "derivation_digest": dataset.derivation_digest, + "question_ids": list(dataset.selected_question_ids), + }, + "importer_implementation": runtime_importer, + "parsing_policy": { + "dialect": validate_csv_dialect_identity(asdict(primary_declaration.dialect)), + "mapping": dict(primary_declaration.columns), + "null_values": list(primary_declaration.null_values), + "numeric_columns": validate_numeric_columns_identity( + [asdict(item) for item in primary_declaration.numeric_columns] + ), + "option_mapping": validate_option_mapping_identity( + asdict(primary_declaration.option_mapping) + ), + "extra_field_policy": primary_declaration.extra_field_policy, + }, + "validation": {"findings_digest": validated.findings_digest}, + "evidence": { + "evidence_status": result.computed_evidence_status, + "qualification_digest": qualification_digest, + "limitation_digest": limitation_digest, + "defect_digest": defect_digest, + "scope_disposition": condition.scope_disposition, + }, + "lineage_components": lineage_components, + "result_origin": result_origin, + "parent_digests": {"realization_digests": [], "evidence_digests": [], "result_digests": []}, + "authorization_digest": None, + "overlay": None, + } + + lineage_runtime_callables = { + component["lineage_id"]: _external_import_adapter for component in lineage_components + } + realization = make_realization( + condition_id=semantic.condition["condition_id"], + condition_digest=semantic.condition["condition_digest"], + identity=realization_identity, + fields={}, + adapter=_external_import_adapter, + validator=_external_import_validator, + expected_dataset=dataset, + semantic_condition=semantic.condition, + lineage_runtime_callables=lineage_runtime_callables, + ) + return realization, semantic.condition, result, tuple(evidence_records) + + +def build_import_plan(request: ImportRequest) -> ImportPlan: + """Resolve, adapt, validate, and identity-bind every declared condition. + Performs every read/validation/identity step but writes nothing.""" + spec = request.spec + sources_by_id = {source.source_id: source for source in spec.sources} + containment_root = request.source_containment_root or request.workspace_root + if request.expected_datasets is not None: + expected_datasets: dict[str, ExpectedDataset] = dict(request.expected_datasets) + missing_datasets = sorted( + {condition.dataset_id for condition in spec.conditions} - set(expected_datasets) + ) + if missing_datasets: + raise ImportEngineError( + f"request.expected_datasets is missing dataset(s) {missing_datasets} " + "referenced by spec.conditions." + ) + else: + expected_datasets = {} + for declaration in spec.datasets: + opened = { + source_id: open_verified_source( + sources_by_id[source_id], containment_root=containment_root + ) + for source_id in declaration.source_ids + } + expected_datasets[declaration.dataset_id] = build_expected_dataset(declaration, opened) + + models_by_key = {model.model_key: model for model in spec.models} + methods_by_key = {method.method_key: method for method in spec.methods} + prompts_by_key = {prompt.prompt_key: prompt for prompt in spec.prompts} + + semantic_conditions: dict[str, dict[str, Any]] = {} + realizations: dict[str, dict[str, Any]] = {} + realization_validations: dict[str, Any] = {} + all_evidence_records: list[dict[str, Any]] = [] + all_findings: list[Any] = [] + + for condition in spec.conditions: + dataset = expected_datasets[condition.dataset_id] + model = models_by_key[condition.model_key] + method = methods_by_key[condition.method_key] + prompt = prompts_by_key[condition.prompt_key] + realization, condition_record, result, evidence_records = _build_realization_for_condition( + condition, spec=spec, dataset=dataset, model=model, method=method, prompt=prompt, + workspace_root=containment_root, strict=request.strict, + ) + semantic_conditions[condition_record["condition_id"]] = condition_record + realizations[realization["realization_id"]] = realization + realization_validations[realization["realization_id"]] = result + all_evidence_records.extend(evidence_records) + all_findings.extend(result.findings) + + payload = { + "protocol_version": PROTOCOL_V3_VERSION, + "canonicalization_version": CANONICALIZATION_VERSION, + "semantic_conditions": semantic_conditions, + "realizations": realizations, + } + manifest = make_manifest_v3( + payload, + audit={ + "source_location": str(request.workspace_root), + }, + ) + validate_manifest_v3(manifest) + + return ImportPlan( + manifest=manifest, + expected_datasets=expected_datasets, + realizations=realization_validations, + evidence_records=tuple(all_evidence_records), + findings=tuple(all_findings), + ) + + +def _counts_for_plan(plan: ImportPlan) -> ImportCounts: + conditions = {"declared": len(plan.manifest["payload"]["semantic_conditions"])} + evidence_status: dict[str, int] = {} + scope_disposition: dict[str, int] = {} + for realization_id, result in plan.realizations.items(): + evidence_status[result.computed_evidence_status] = ( + evidence_status.get(result.computed_evidence_status, 0) + 1 + ) + realization = plan.manifest["payload"]["realizations"][realization_id] + disposition = realization["identity"]["realization"]["evidence"]["scope_disposition"] + scope_disposition[disposition] = scope_disposition.get(disposition, 0) + 1 + return ImportCounts( + conditions=conditions, evidence_status=evidence_status, scope_disposition=scope_disposition, + ) + + +def _defects_for_plan(plan: ImportPlan) -> DefectReport: + by_code: dict[str, int] = {} + for finding in plan.findings: + by_code[finding.code] = by_code.get(finding.code, 0) + 1 + findings_payload = [finding.to_json() for finding in plan.findings] + return DefectReport( + total_findings=len(plan.findings), by_code=by_code, + findings_digest=integrity_digest(findings_payload), + ) + + +def execute_import(request: ImportRequest, *, dry_run: bool = False) -> ImportReport: + try: + plan = build_import_plan(request) + except (ImportIdentityError, ImportValidationError, ValueError) as exc: + return ImportReport( + schema_version="choicebench.import-report.v1", + import_state="failed", + wrote_artifacts=False, + idempotent_noop=False, + run_id=request.run_id, + experiment_id=None, + counts=ImportCounts(conditions={}, evidence_status={}, scope_disposition={}), + defects=DefectReport(total_findings=0, by_code={}, findings_digest=integrity_digest([])), + condition_digests={}, + realization_digests={}, + failures=({"code": "IMPORT_PLAN_FAILED", "sanitized_message": str(exc)},), + ) + + counts = _counts_for_plan(plan) + defects = _defects_for_plan(plan) + condition_digests = { + condition_id: record["condition_digest"] + for condition_id, record in plan.manifest["payload"]["semantic_conditions"].items() + } + realization_digests = { + realization_id: record["realization_digest"] + for realization_id, record in plan.manifest["payload"]["realizations"].items() + } + + if dry_run: + return ImportReport( + schema_version="choicebench.import-report.v1", + import_state="validated", + wrote_artifacts=False, + idempotent_noop=False, + run_id=request.run_id, + experiment_id=plan.manifest["experiment_id"], + counts=counts, + defects=defects, + condition_digests=condition_digests, + realization_digests=realization_digests, + failures=(), + ) + + runs_dir = Path(request.workspace_root) / "runs" + idempotent_noop = [False] + + def _validator(published_root: Path) -> None: + if published_root == runs_dir / request.run_id and (published_root / MANIFEST_FILENAME).is_file(): + verify_import_run(published_root, expected=plan) + idempotent_noop[0] = True + return + _write_staged_run(published_root, request, plan) + + with ImportTransaction(runs_dir=runs_dir, run_id=request.run_id) as txn: + final_path = txn.publish(_validator) + + return ImportReport( + schema_version="choicebench.import-report.v1", + import_state="imported", + wrote_artifacts=not idempotent_noop[0], + idempotent_noop=idempotent_noop[0], + run_id=request.run_id, + experiment_id=plan.manifest["experiment_id"], + counts=counts, + defects=defects, + condition_digests=condition_digests, + realization_digests=realization_digests, + failures=(), + ) + + +def _write_staged_run(staged_run: Path, request: ImportRequest, plan: ImportPlan) -> None: + staged_run.mkdir(parents=True, exist_ok=True) + sources_by_id = {source.source_id: source for source in request.spec.sources} + + written_digests: set[str] = set() + index_records: list[Mapping[str, Any]] = [] + for record in plan.evidence_records: + if record["sha256"] in written_digests: + continue + written_digests.add(record["sha256"]) + declaration = sources_by_id[record["source_id"]] + opened = open_verified_source( + declaration, + containment_root=request.source_containment_root or request.workspace_root, + ) + index_records.append(write_evidence_blob(staged_run, opened, references=record["references"])) + write_evidence_index(staged_run, index_records) + + atomic_write_json(staged_run / MANIFEST_FILENAME, plan.manifest) + + state = initial_run_state_v3(plan.manifest) + for realization_id, realization in plan.manifest["payload"]["realizations"].items(): + result = plan.realizations[realization_id] + realization_state = state["realizations"][realization_id] + # A structurally-complete, in-scope cell publishes its full canonical + # result CSV even when not evaluable ("published_unscored"): the rows + # are retained verbatim (some predictions may be unparseable) so an + # authorized overlay can later replace specific rows. A genuinely + # incomplete/out-of-scope cell stays evidence-only with no result CSV. + publishes_result = bool(result.published_rows) + if result.evaluable: + realization_state["status"] = "completed" + elif publishes_result: + realization_state["status"] = "published_unscored" + else: + realization_state["status"] = "evidence_only" + + prepared_validation = prepare_realization_validation_artifact( + result, realization=realization, evidence_records=plan.evidence_records + ) + write_realization_validation_artifact(staged_run, prepared_validation) + realization_state["validation_sha256"] = prepared_validation.file_sha256 + + if publishes_result: + rows_for_result = result.published_rows if result.published_rows else result.evaluable_rows + condition_id = realization["condition_id"] + condition_record = plan.manifest["payload"]["semantic_conditions"][condition_id] + benchmark = condition_record["identity"]["benchmark"] + identity_columns = { + "condition_id": condition_id, + "realization_id": realization_id, + "experiment_id": plan.manifest["experiment_id"], + "dataset_artifact_id": benchmark["artifact_id"], + "dataset_selection_id": benchmark["selection_id"], + "model_id": condition_record["identity"]["model_id"], + "method_id": condition_record["identity"]["method_id"], + "prompt_id": condition_record["identity"]["prompt_id"], + "benchmark_name": benchmark["name"], + "benchmark_split": benchmark["split"], + } + lineage_ids_by_qid = { + assignment["question_id"]: assignment["prediction_lineage_id"] + for assignment in realization["identity"]["realization"]["result_origin"]["row_assignments"] + } + result_rows = [ + { + **identity_columns, + **row, + "prediction_lineage_id": lineage_ids_by_qid[row["question_id"]], + } + for row in rows_for_result + ] + prepared_result = prepare_manifest_result( + result_rows, manifest=plan.manifest, realization_id=realization_id + ) + _, published_metadata, _ = publish_manifest_result(prepared_result, run_dir=staged_run) + realization_state["result_artifact_id"] = published_metadata["result_artifact_id"] + realization_state["result_artifact_digest"] = published_metadata["result_artifact_digest"] + realization_state["result_sha256"] = published_metadata["file_sha256"] + + atomic_write_json(staged_run / RUN_STATE_FILENAME, canonicalize(state, redact_secrets=False)) + validate_manifest_v3(plan.manifest) + + +def verify_import_run( + run_dir: Path, expected: ImportPlan | SimpleNamespace | None = None +) -> VerifiedImportRun: + """The single shared full-graph verifier: validates manifest, state, + validation artifacts, and results, and returns immutable trusted values + only after the whole graph passes. `expected` need only expose a + `.manifest` mapping (an ImportPlan, or a bare SimpleNamespace(manifest=...) + for the overlay path, which has no full ImportPlan of its own).""" + run_dir = Path(run_dir) + manifest = json.loads((run_dir / MANIFEST_FILENAME).read_text()) + validate_manifest_v3(manifest) + state = json.loads((run_dir / RUN_STATE_FILENAME).read_text()) + validate_run_state_v3(state, manifest) + + if expected is not None and manifest["experiment_digest"] != expected.manifest["experiment_digest"]: + raise ImportEngineError( + f"Run {run_dir} belongs to a different experiment than expected." + ) + + evidence_index_path = run_dir / "artifacts" / "imports" / "evidence" / "index.json" + evidence_index = json.loads(evidence_index_path.read_text()) + evidence_index_digest = evidence_index["evidence_index_digest"] + validate_evidence_index(run_dir, evidence_index_digest) + + realizations: dict[str, VerifiedBaseRealization] = {} + for realization_id, realization in manifest["payload"]["realizations"].items(): + realization_identity = realization["identity"]["realization"] + realization_state = state["realizations"][realization_id] + status = realization_state["status"] + + validation_path = run_dir / f"artifacts/imports/validation/{realization_id}.json" + validate_realization_validation_artifact( + validation_path, manifest=manifest, realization=realization, + expected_sha256=realization_state["validation_sha256"], + ) + + result_sha256 = None + rows_by_question_id: dict[str, Any] = {} + prediction_origins: dict[str, str] = {} + # Both a scored ("completed") and an unscored-but-structurally-complete + # ("published_unscored") realization publish a full result CSV; read + # the retained row set from either so an overlay can retain rows. + if status in ("completed", "published_unscored"): + result_path = run_dir / f"results/{realization_id}.csv" + metadata = validate_result_artifact(result_path, manifest=manifest, realization=realization) + if ( + metadata["result_artifact_id"] != realization_state["result_artifact_id"] + or metadata["result_artifact_digest"] != realization_state["result_artifact_digest"] + ): + raise ImportEngineError( + f"Realization {realization_id!r} result artifact does not match its " + "recorded run-state identity." + ) + result_sha256 = metadata["file_sha256"] + frame = pd.read_csv(result_path, dtype={"question_id": "string"}) + for _, row in frame.iterrows(): + qid = str(row["question_id"]) + # A "published_unscored" realization's result CSV can carry a + # genuinely blank/unparseable prediction (e.g. an unrelated + # parse-missing row retained verbatim through an overlay). + # Series.to_dict() re-coerces a blank cell to a float NaN even + # after DataFrame-level NA cleanup (confirmed empirically: + # frame.astype(object).where(frame.notna(), None) leaves a + # real None at the DataFrame level, but per-row Series dtype + # inference during iterrows() collapses it back to NaN). The + # manifest writer's canonicalize() rejects NaN outright + # ("cannot contain NaN or infinity"), so normalize per-value + # here instead. A "completed" realization's rows never + # contain NaN (evaluable requires a valid prediction in every + # row), so this is a no-op on the existing path. + row_dict = { + key: (None if isinstance(value, float) and math.isnan(value) else value) + for key, value in row.to_dict().items() + } + rows_by_question_id[qid] = row_dict + prediction_origins[qid] = str(row["prediction_origin"]) + + evidence_digests = { + source["source_id"]: source["sha256"] for source in realization_identity["sources"] + } + + realizations[realization_id] = VerifiedBaseRealization( + condition_id=realization["condition_id"], + condition_digest=realization["condition_digest"], + realization_id=realization_id, + realization_digest=realization["realization_digest"], + evidence_index_digest=evidence_index_digest, + evidence_source_digests=evidence_digests, + validation_artifact_sha256=realization_state["validation_sha256"], + result_sha256=result_sha256, + rows_by_question_id=rows_by_question_id, + prediction_origins=prediction_origins, + ) + + return VerifiedImportRun( + manifest=manifest, manifest_digest=manifest["experiment_digest"], realizations=realizations + ) + + +# --- Overlay merge-into-new-run orchestration ------------------------------- +# +# Scope boundary: requires an *overlayable* base realization -- one that +# published a full canonical result CSV whose row-identity is exactly the +# expected complete question set (every ID present exactly once), with its +# source hash verified by verify_import_run. Overlayability is a data-identity +# property, distinct from evaluability (scoreability): a +# malformed-but-structurally-complete base (all IDs present, but some rows +# carry an unparseable prediction) is NOT evaluable yet IS overlayable -- it +# has a full published row set ("published_unscored" status) from which the +# non-authorized rows are retained verbatim while the authorized rows are +# replaced. Only a genuinely broken base -- one with missing/duplicate/ +# unexpected rows or no published result at all (evidence-only) -- is refused, +# with the same fail-closed behavior as before. The derived realization's +# evidence_status is recomputed from the actual merged row content (it is NOT +# forced to the overlay's declared status); the base's own freeze-declared +# historical status remains represented separately (see below). + +def _recompute_overlay_evidence_status( + rows_by_qid: Mapping[str, Mapping[str, Any]], *, expected: ExpectedDataset, base_qualified: bool +) -> str: + """Recompute the derived realization's evidence_status from the ACTUAL + merged row content, never forcing it to the overlay's declared status. Any + row whose predicted_option is empty or is not one of that question's option + letters keeps the derived cell 'malformed' (e.g. a repair that fixed its 3 + authorized rows but left other unrelated parse-missing rows in place); a + fully-parseable merged set is 'complete', or 'qualified' when the base + itself was qualified. The merged set is always structurally complete here + (the overlayable gate guaranteed the base's full row-identity), so + 'partial'/'failed' cannot arise.""" + frame = expected.frame + for qid, row in rows_by_qid.items(): + matches = frame[frame["question_id"].astype(str) == str(qid)] + letters = _expected_choice_letters(matches.iloc[0].to_dict()) if not matches.empty else [] + pred = row.get("predicted_option") + pred = None if pred is None else str(pred).strip().upper() + if not pred or pred == "NAN" or pred not in letters: + return "malformed" + return "qualified" if base_qualified else "complete" + + +@dataclass(frozen=True) +class OverlayImportRequest: + base_run_dir: Path + base_realization_id: str + authorization: AuthorizationSpec + authorization_source: SourceArtifactSpec + overlay: OverlaySpec + overlay_source: SourceArtifactSpec + overlay_mapping: Mapping[str, str] + condition_digests: Mapping[str, str] + expected_dataset: ExpectedDataset + run_id: str + workspace_root: Path + source_containment_root: Path | None = None + """Root every declared source path (authorization_source, overlay_source) + must resolve under. Defaults to workspace_root; set separately when + those sources live under an independent, read-only root distinct from + where the derived run is written -- see ImportRequest.source_containment_root.""" + + +def _merge_overlay_realization( + request: OverlayImportRequest, +) -> tuple[dict[str, Any], dict[str, Any], list[dict[str, Any]], dict[str, Any]]: + """Returns (realization_record, semantic_condition_record, result_rows, + overlay_evidence_record). The evidence record covers only the new + overlay source -- base evidence remains immutably stored in the base + run's own directory and is reachable via parent_digests.""" + verified = verify_import_run(request.base_run_dir) + base = verified.realizations.get(request.base_realization_id) + if base is None: + raise ImportEngineError(f"Unknown base realization {request.base_realization_id!r}.") + # Overlayable gate: the base must have a published canonical row set whose + # question-ID identity is exactly the expected complete set. This is + # independent of the base's evidence_status -- a malformed base that is + # structurally complete IS overlayable; a base with no published result, or + # with missing/duplicate/unexpected rows, is refused fail-closed. + expected_ids = tuple(request.expected_dataset.selected_question_ids) + base_row_ids = list(base.rows_by_question_id) + if ( + base.result_sha256 is None + or set(base_row_ids) != set(expected_ids) + or len(base_row_ids) != len(expected_ids) + ): + raise ImportEngineError( + f"Realization {request.base_realization_id!r} is not overlayable: it lacks a " + "structurally complete published canonical row set (every expected question ID " + "present exactly once). A malformed-but-structurally-complete base is " + "overlayable, but an evidence-only base or one with missing/duplicate/unexpected " + "rows cannot be overlaid." + ) + + base_manifest = verified.manifest + base_realization_record = base_manifest["payload"]["realizations"][request.base_realization_id] + base_realization_identity = base_realization_record["identity"]["realization"] + condition_id = base_realization_record["condition_id"] + semantic_condition = base_manifest["payload"]["semantic_conditions"][condition_id] + + containment_root = request.source_containment_root or request.workspace_root + opened_auth_source = open_verified_source( + request.authorization_source, containment_root=containment_root + ) + bundle = validate_authorization_bundle( + request.authorization, opened_source=opened_auth_source, + condition_digests=request.condition_digests, + expected={base.condition_digest: request.expected_dataset}, + ) + authorization = authorization_for_condition(bundle, condition_digest=base.condition_digest) + + opened_overlay_source = open_verified_source( + request.overlay_source, containment_root=containment_root + ) + overlay_table = parse_csv_source(opened_overlay_source, request.overlay_source, strict=True) + + derived = derive_overlay( + base=base, overlay=request.overlay, authorization=authorization, + overlay_table=overlay_table, overlay_mapping=request.overlay_mapping, + expected=request.expected_dataset, + ) + + base_lineage_by_id = { + component["lineage_id"]: component + for component in base_realization_identity["lineage_components"] + } + base_row_assignments = base_realization_identity["result_origin"]["row_assignments"] + retained_assignments = [ + assignment for assignment in base_row_assignments + if assignment["question_id"] not in derived.replacement_question_ids + ] + retained_lineage_ids = {assignment["prediction_lineage_id"] for assignment in retained_assignments} + + # Realization identity requires row_assignments to be an ordered subset of + # the expected dataset's question order (identity.py), not "retained then + # replaced" -- reassemble in dataset order regardless of which side (base + # or overlay) supplies each question. lineage_components must follow the + # SAME deterministic order: canonicalize() hashes list order, so building + # it from set iteration (as an earlier version of this function did) made + # the realization/experiment digest nondeterministic across processes + # whenever more than one row was retained. + assignments_by_qid = { + assignment["question_id"]: ( + assignment["question_id"], assignment["prediction_origin"], assignment["prediction_lineage_id"] + ) + for assignment in retained_assignments + } + assignments_by_qid.update({ + component["identity"]["question_id"]: ( + component["identity"]["question_id"], + component["identity"]["prediction_origin"], + component["lineage_id"], + ) + for component in derived.lineage_components + }) + row_assignments = [ + assignments_by_qid[qid] + for qid in request.expected_dataset.selected_question_ids + if qid in assignments_by_qid + ] + + lineage_components_by_id = {**base_lineage_by_id} + lineage_components_by_id.update( + {component["lineage_id"]: component for component in derived.lineage_components} + ) + all_lineage_components = [ + lineage_components_by_id[lineage_id] + for _question_id, _origin, lineage_id in row_assignments + ] + derivation_origin = derived.lineage_components[0]["identity"]["operation_type"] + result_origin = make_result_origin(derivation_origin=derivation_origin, row_assignments=row_assignments) + + # Assemble the merged result rows now -- non-authorized base rows retained + # verbatim (numeric-tolerant, value-identical), authorized rows replaced -- + # so the derived realization's evidence_status is recomputed from the + # ACTUAL merged content below instead of being forced to the overlay's + # declared status. + _ROW_KEYS = { + "question_id", "question_text", "correct_option", "choices_json", + "prediction_origin", "predicted_option", + } + retained_rows_by_qid = { + assignment["question_id"]: { + key: value + for key, value in base.rows_by_question_id[assignment["question_id"]].items() + if key in _ROW_KEYS + } + for assignment in retained_assignments + } + derived_rows_by_qid = {row["question_id"]: dict(row) for row in derived.preownership_rows} + rows_by_qid = {**retained_rows_by_qid, **derived_rows_by_qid} + result_rows = [ + rows_by_qid[qid] for qid in request.expected_dataset.selected_question_ids if qid in rows_by_qid + ] + base_qualified = base_realization_identity["evidence"]["evidence_status"] == "qualified" + recomputed_evidence_status = _recompute_overlay_evidence_status( + rows_by_qid, expected=request.expected_dataset, base_qualified=base_qualified + ) + + replacement_digest = integrity_digest( + canonicalize({ + "replacement_question_ids": list(derived.replacement_question_ids), + "replacement_reasons": dict(request.overlay.replacement_reasons), + }) + ) + overlay_dict: dict[str, Any] = { + "source_sha256": overlay_table.source_sha256, + "replacement_digest": replacement_digest, + "transformation_input_digest": None, + "preownership_output_digest": None, + "implementation_digest": None, + } + if derivation_origin == "offline_transformation": + # The lineage/overlay digests MUST be computed from the derived + # components in the SAME canonical order the identity verifier uses + # (identity.py filters the realization's own dataset-ordered + # lineage_components by operation_type), NOT in derive_overlay's + # sorted(replacement_ids) order. integrity_digest hashes list order, so + # whenever the replaced IDs' dataset order differs from their sorted + # order the two digests diverge and identity validation raised a + # false-positive "digest conflicts with its lineage components". + ordered_derived_components = [ + component + for component in all_lineage_components + if component["identity"]["operation_type"] == derivation_origin + ] + overlay_dict["transformation_input_digest"] = integrity_digest( + [component["identity"]["input_digest"] for component in ordered_derived_components] + ) + overlay_dict["preownership_output_digest"] = integrity_digest( + [component["identity"]["preownership_output_digest"] for component in ordered_derived_components] + ) + overlay_dict["implementation_digest"] = integrity_digest( + ordered_derived_components[0]["identity"]["implementation"] + ) + + overlay_source_notes_digest = ( + integrity_digest(request.overlay_source.notes) if request.overlay_source.notes else None + ) + source_records = [ + *base_realization_identity["sources"], + _source_record(request.overlay_source, opened_overlay_source, notes_digest=overlay_source_notes_digest), + ] + + realization_identity = { + "import_spec_digest": base_realization_identity["import_spec_digest"], + "sources": source_records, + "expected_dataset": base_realization_identity["expected_dataset"], + "importer_implementation": base_realization_identity["importer_implementation"], + "parsing_policy": base_realization_identity["parsing_policy"], + "validation": base_realization_identity["validation"], + "evidence": { + **base_realization_identity["evidence"], + "evidence_status": recomputed_evidence_status, + }, + "lineage_components": all_lineage_components, + "result_origin": result_origin, + "parent_digests": { + "realization_digests": [base.realization_digest], + "evidence_digests": [base.validation_artifact_sha256], + "result_digests": [base.result_sha256] if base.result_sha256 else [], + }, + "authorization_digest": authorization.bundle_digest, + "overlay": overlay_dict, + } + realization = make_realization( + condition_id=condition_id, + condition_digest=base.condition_digest, + identity=realization_identity, + fields={}, + adapter=_external_import_adapter, + validator=_external_import_validator, + expected_dataset=request.expected_dataset, + semantic_condition=semantic_condition, + lineage_runtime_callables={lid: _external_import_adapter for lid in retained_lineage_ids}, + ) + + overlay_evidence_record = evidence_record( + opened_overlay_source, references=sorted(derived.replacement_question_ids) + ) + return realization, semantic_condition, result_rows, overlay_evidence_record + + +def execute_overlay_import(request: OverlayImportRequest, *, dry_run: bool = False) -> ImportReport: + try: + realization, semantic_condition, result_rows, overlay_evidence_record = ( + _merge_overlay_realization(request) + ) + except (ImportIdentityError, ImportEngineError, ValueError) as exc: + return ImportReport( + schema_version="choicebench.import-report.v1", + import_state="failed", + wrote_artifacts=False, + idempotent_noop=False, + run_id=request.run_id, + experiment_id=None, + counts=ImportCounts(conditions={}, evidence_status={}, scope_disposition={}), + defects=DefectReport(total_findings=0, by_code={}, findings_digest=integrity_digest([])), + condition_digests={}, + realization_digests={}, + failures=({"code": "OVERLAY_IMPORT_FAILED", "sanitized_message": str(exc)},), + ) + + condition_id = realization["condition_id"] + payload = { + "protocol_version": PROTOCOL_V3_VERSION, + "canonicalization_version": CANONICALIZATION_VERSION, + "semantic_conditions": {condition_id: semantic_condition}, + "realizations": {realization["realization_id"]: realization}, + } + manifest = make_manifest_v3(payload, audit={"source_location": str(request.workspace_root)}) + validate_manifest_v3(manifest) + + evidence_status = realization["identity"]["realization"]["evidence"]["evidence_status"] + scope_disposition = realization["identity"]["realization"]["evidence"]["scope_disposition"] + counts = ImportCounts( + conditions={"declared": 1}, + evidence_status={evidence_status: 1}, + scope_disposition={scope_disposition: 1}, + ) + condition_digests = {condition_id: semantic_condition["condition_digest"]} + realization_digests = {realization["realization_id"]: realization["realization_digest"]} + + if dry_run: + return ImportReport( + schema_version="choicebench.import-report.v1", + import_state="validated", + wrote_artifacts=False, + idempotent_noop=False, + run_id=request.run_id, + experiment_id=manifest["experiment_id"], + counts=counts, + defects=DefectReport(total_findings=0, by_code={}, findings_digest=integrity_digest([])), + condition_digests=condition_digests, + realization_digests=realization_digests, + failures=(), + ) + + runs_dir = Path(request.workspace_root) / "runs" + idempotent_noop = [False] + + def _validator(published_root: Path) -> None: + if published_root == runs_dir / request.run_id and (published_root / MANIFEST_FILENAME).is_file(): + verify_import_run(published_root, expected=SimpleNamespace(manifest=manifest)) + idempotent_noop[0] = True + return + _write_overlay_staged_run( + published_root, manifest, realization, result_rows, + overlay_source=request.overlay_source, + overlay_evidence_record=overlay_evidence_record, + workspace_root=request.source_containment_root or request.workspace_root, + declared_evidence_status=request.overlay.expected_evidence_status, + ) + + with ImportTransaction(runs_dir=runs_dir, run_id=request.run_id) as txn: + txn.publish(_validator) + + return ImportReport( + schema_version="choicebench.import-report.v1", + import_state="imported", + wrote_artifacts=not idempotent_noop[0], + idempotent_noop=idempotent_noop[0], + run_id=request.run_id, + experiment_id=manifest["experiment_id"], + counts=counts, + defects=DefectReport(total_findings=0, by_code={}, findings_digest=integrity_digest([])), + condition_digests=condition_digests, + realization_digests=realization_digests, + failures=(), + ) + + +def _write_overlay_staged_run( + staged_run: Path, manifest: Mapping[str, Any], realization: Mapping[str, Any], + result_rows: list[dict[str, Any]], + *, + overlay_source: SourceArtifactSpec, + overlay_evidence_record: Mapping[str, Any], + workspace_root: Path, + declared_evidence_status: str, +) -> None: + staged_run.mkdir(parents=True, exist_ok=True) + + opened_overlay_source = open_verified_source(overlay_source, containment_root=workspace_root) + written_record = write_evidence_blob( + staged_run, opened_overlay_source, references=overlay_evidence_record["references"] + ) + write_evidence_index(staged_run, [written_record]) + + atomic_write_json(staged_run / MANIFEST_FILENAME, manifest) + + realization_id = realization["realization_id"] + state = initial_run_state_v3(manifest) + evidence = realization["identity"]["realization"]["evidence"] + recomputed_status = evidence["evidence_status"] + evaluable = recomputed_status in ("complete", "qualified") and evidence["scope_disposition"] == "included" + # A derived overlay run always publishes a full canonical result CSV; its + # run status reflects the RECOMPUTED evidence_status ("completed" when the + # merged content is evaluable, else "published_unscored"). The overlay's + # declared/intended historical status is preserved separately here so it is + # never conflated with the observed evidence_status. + state["realizations"][realization_id]["status"] = "completed" if evaluable else "published_unscored" + state["realizations"][realization_id]["declared_evidence_status"] = declared_evidence_status + + condition_id = realization["condition_id"] + condition_record = manifest["payload"]["semantic_conditions"][condition_id] + benchmark = condition_record["identity"]["benchmark"] + identity_columns = { + "condition_id": condition_id, + "realization_id": realization_id, + "experiment_id": manifest["experiment_id"], + "dataset_artifact_id": benchmark["artifact_id"], + "dataset_selection_id": benchmark["selection_id"], + "model_id": condition_record["identity"]["model_id"], + "method_id": condition_record["identity"]["method_id"], + "prompt_id": condition_record["identity"]["prompt_id"], + "benchmark_name": benchmark["name"], + "benchmark_split": benchmark["split"], + } + lineage_ids_by_qid = { + assignment["question_id"]: assignment["prediction_lineage_id"] + for assignment in realization["identity"]["realization"]["result_origin"]["row_assignments"] + } + final_rows = [ + {**identity_columns, **row, "prediction_lineage_id": lineage_ids_by_qid[row["question_id"]]} + for row in result_rows + ] + prepared_result = prepare_manifest_result(final_rows, manifest=manifest, realization_id=realization_id) + _, published_metadata, _ = publish_manifest_result(prepared_result, run_dir=staged_run) + state["realizations"][realization_id]["result_artifact_id"] = published_metadata["result_artifact_id"] + state["realizations"][realization_id]["result_artifact_digest"] = published_metadata["result_artifact_digest"] + state["realizations"][realization_id]["result_sha256"] = published_metadata["file_sha256"] + + merged_validation = RealizationValidation( + evaluable_rows=tuple(result_rows) if evaluable else (), + findings=(), + validation_digest=integrity_digest({"evaluable_rows": canonicalize(result_rows)}), + computed_evidence_status=recomputed_status, + evaluable=evaluable, + qualifications=(), + limitations=(), + defect_question_ids=(), + published_rows=tuple(result_rows), + structurally_complete=True, + ) + prepared_validation = prepare_realization_validation_artifact( + merged_validation, realization=realization, evidence_records=() + ) + write_realization_validation_artifact(staged_run, prepared_validation) + state["realizations"][realization_id]["validation_sha256"] = prepared_validation.file_sha256 + + atomic_write_json(staged_run / RUN_STATE_FILENAME, canonicalize(state, redact_secrets=False)) + validate_manifest_v3(manifest) + + +def serialize_import_report(report: ImportReport) -> dict[str, Any]: + return { + "schema_version": report.schema_version, + "import_state": report.import_state, + "wrote_artifacts": report.wrote_artifacts, + "idempotent_noop": report.idempotent_noop, + "run_id": report.run_id, + "experiment_id": report.experiment_id, + "counts": { + "conditions": dict(report.counts.conditions), + "evidence_status": dict(report.counts.evidence_status), + "scope_disposition": dict(report.counts.scope_disposition), + }, + "defects": { + "total_findings": report.defects.total_findings, + "by_code": dict(report.defects.by_code), + "findings_digest": report.defects.findings_digest, + }, + "condition_digests": dict(report.condition_digests), + "realization_digests": dict(report.realization_digests), + "failures": [dict(failure) for failure in report.failures], + } diff --git a/src/choicebench/importing/evidence.py b/src/choicebench/importing/evidence.py new file mode 100644 index 0000000..fb6013e --- /dev/null +++ b/src/choicebench/importing/evidence.py @@ -0,0 +1,176 @@ +"""Open declared sources with a verified checksum and store referenced +row-level evidence bytes as run-local content-addressed blobs. + +Reduced scope: covers the concrete safety guarantees the repair-import +workflow needs -- read-once-hash-forward source opening, path containment, +symlink refusal, and a self-digested evidence index -- without the full +historical suffix-derivation/dedup-across-runs richness of the original plan. +""" + +from __future__ import annotations + +from hashlib import sha256 +import json +from pathlib import Path +from typing import Any, Mapping, Sequence + +from choicebench.identity import canonicalize, integrity_digest +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.schema import SourceArtifactSpec +from choicebench.infra.artifacts import atomic_write_bytes, atomic_write_json + +EVIDENCE_INDEX_SCHEMA_VERSION = "choicebench.evidence-index.v1" +_SAFE_SUFFIXES = {"csv": "csv"} + + +class EvidenceError(ValueError): + """Raised when a source or stored evidence fails a safety or integrity check.""" + + +def open_verified_source( + declaration: SourceArtifactSpec, *, containment_root: Path | None = None +) -> OpenedSource: + """Open exactly the declared regular file, read its bytes once, and hash + those same bytes -- never re-read or re-derive the digest from a + separately reopened handle.""" + path = Path(declaration.path) + if containment_root is not None: + root = Path(containment_root).resolve() + try: + path.resolve().relative_to(root) + except ValueError as exc: + raise EvidenceError( + f"Source {declaration.source_id!r} path {path} escapes containment " + f"root {root}." + ) from exc + if path.is_symlink(): + raise EvidenceError(f"Source {declaration.source_id!r} path {path} is a symlink.") + if not path.is_file(): + raise EvidenceError(f"Source {declaration.source_id!r} path {path} is not a regular file.") + data = path.read_bytes() + actual_sha256 = sha256(data).hexdigest() + if actual_sha256 != declaration.expected_sha256: + raise EvidenceError( + f"Source {declaration.source_id!r} checksum mismatch: expected " + f"{declaration.expected_sha256}, found {actual_sha256}." + ) + return OpenedSource( + source_id=declaration.source_id, + audit_path=path, + logical_path=declaration.logical_path, + data=data, + sha256=actual_sha256, + ) + + +def evidence_blob_path(staged_run: Path, sha256_digest: str, format_name: str) -> Path: + suffix = _SAFE_SUFFIXES.get(format_name) + if suffix is None: + raise EvidenceError(f"Unsupported evidence blob format {format_name!r}.") + if len(sha256_digest) != 64 or any(c not in "0123456789abcdef" for c in sha256_digest): + raise EvidenceError(f"Invalid evidence blob digest {sha256_digest!r}.") + return ( + Path(staged_run) + / "artifacts" / "imports" / "evidence" / "sha256" + / sha256_digest[:2] / f"{sha256_digest}.{suffix}" + ) + + +def evidence_record(source: OpenedSource, references: Sequence[str]) -> dict[str, Any]: + """The evidence blob record for a source, computed with no filesystem + I/O. Used both to write a blob (write_evidence_blob) and, during + dry-run/planning, to know what a real import would write without + actually writing it.""" + return canonicalize( + { + "source_id": source.source_id, + "logical_path": source.logical_path, + "sha256": source.sha256, + "size": len(source.data), + "format": "csv", + "references": sorted(set(references)), + } + ) + + +def write_evidence_blob( + staged_run: Path, source: OpenedSource, references: Sequence[str] +) -> dict[str, Any]: + """Store one referenced source's bytes as a run-local content-addressed + blob. Identical bytes already staged (e.g. two references to the same + source) are reused rather than rewritten; a divergent existing blob at + the same digest path is refused.""" + blob_path = evidence_blob_path(staged_run, source.sha256, "csv") + if blob_path.exists(): + if blob_path.read_bytes() != source.data: + raise EvidenceError( + f"Refusing to overwrite a divergent existing evidence blob: {blob_path}" + ) + else: + atomic_write_bytes(blob_path, source.data) + record = evidence_record(source, references) + sidecar_path = blob_path.with_suffix(blob_path.suffix + ".json") + if sidecar_path.exists(): + existing = json.loads(sidecar_path.read_text()) + if existing != record: + raise EvidenceError( + f"Refusing to overwrite a divergent evidence blob sidecar: {sidecar_path}" + ) + else: + atomic_write_json(sidecar_path, record) + return record + + +def validate_evidence_blob(run_dir: Path, record: Mapping[str, Any]) -> None: + blob_path = evidence_blob_path(run_dir, record["sha256"], record["format"]) + if not blob_path.is_file(): + raise EvidenceError(f"Missing evidence blob: {blob_path}") + actual = sha256(blob_path.read_bytes()).hexdigest() + if actual != record["sha256"]: + raise EvidenceError( + f"Evidence blob content integrity check failed: {blob_path}; expected " + f"{record['sha256']}, found {actual}." + ) + sidecar_path = blob_path.with_suffix(blob_path.suffix + ".json") + try: + stored = json.loads(sidecar_path.read_text()) + except (OSError, json.JSONDecodeError) as exc: + raise EvidenceError(f"Evidence blob sidecar is unreadable: {sidecar_path}: {exc}") from exc + if canonicalize(stored) != canonicalize(dict(record)): + raise EvidenceError(f"Evidence blob sidecar does not match its declared record: {sidecar_path}") + + +def write_evidence_index( + staged_run: Path, records: Sequence[Mapping[str, Any]] +) -> Mapping[str, Any]: + payload = canonicalize( + { + "schema_version": EVIDENCE_INDEX_SCHEMA_VERSION, + "records": list(records), + } + ) + evidence_index_digest = integrity_digest(payload) + index = {**payload, "evidence_index_digest": evidence_index_digest} + index_path = Path(staged_run) / "artifacts" / "imports" / "evidence" / "index.json" + atomic_write_json(index_path, index) + return index + + +def validate_evidence_index(run_dir: Path, expected_digest: str) -> Mapping[str, Any]: + index_path = Path(run_dir) / "artifacts" / "imports" / "evidence" / "index.json" + try: + index = json.loads(index_path.read_text()) + except (OSError, json.JSONDecodeError) as exc: + raise EvidenceError(f"Evidence index is unreadable: {index_path}: {exc}") from exc + if index.get("schema_version") != EVIDENCE_INDEX_SCHEMA_VERSION: + raise EvidenceError(f"Unsupported evidence index schema: {index_path}") + payload = {key: value for key, value in index.items() if key != "evidence_index_digest"} + if integrity_digest(canonicalize(payload)) != index.get("evidence_index_digest"): + raise EvidenceError(f"Evidence index integrity check failed: {index_path}") + if index["evidence_index_digest"] != expected_digest: + raise EvidenceError( + f"Evidence index digest does not match its expected binding: {index_path}" + ) + for record in index["records"]: + validate_evidence_blob(run_dir, record) + return index diff --git a/src/choicebench/importing/identity.py b/src/choicebench/importing/identity.py new file mode 100644 index 0000000..b8eba10 --- /dev/null +++ b/src/choicebench/importing/identity.py @@ -0,0 +1,2616 @@ +"""Layered semantic, realization, lineage, and prediction-origin identities.""" + +from __future__ import annotations + +from dataclasses import dataclass +import functools +import inspect +import math +from pathlib import PurePosixPath +import re +from typing import Any, Callable, Literal, Mapping, Sequence + +from choicebench import __version__ +from choicebench.datasets import dataset_content_digest, dataset_sample_identities +from choicebench.identity import ( + canonicalize, + integrity_digest, + is_credential_key, + short_id, +) +from choicebench.importing.dataset_reference import ExpectedDataset +from choicebench.importing.schema import ( + CsvDialectSpec, + ImportConditionSpec, + ImportSpecError, + ImportMethodSpec, + ImportModelSpec, + ImportPromptSpec, + validate_csv_dialect_identity, + validate_implementation_identity_record, + validate_native_method_payload, + validate_native_model_payload, + validate_numeric_columns_identity, + validate_option_mapping_identity, +) +from choicebench.provenance import implementation_identity +from choicebench.registry import METHOD_REGISTRY + + +class ImportIdentityError(ValueError): + """Raised when an importer identity is cyclic, unsafe, or inconsistent.""" + + +@dataclass(frozen=True) +class ImportSemanticRecords: + dataset_artifact: Mapping[str, Any] + selection: Mapping[str, Any] + model: Mapping[str, Any] + method: Mapping[str, Any] + prompt: Mapping[str, Any] + condition: Mapping[str, Any] + + +_CONDITION_KEYS = { + "benchmark", + "preflight", + "model_id", + "method_id", + "prompt_id", + "prompt_snapshot_path", + "seed", +} +_DERIVATION_ORIGINS = { + "native_execution", + "external_import", + "repair_overlay", + "offline_transformation", +} +_PREDICTION_ORIGINS = { + "native_inference", + "external_historical_inference", + "external_repair_inference", +} +_REALIZATION_KEYS = { + "import_spec_digest", + "sources", + "expected_dataset", + "importer_implementation", + "parsing_policy", + "validation", + "evidence", + "lineage_components", + "result_origin", + "parent_digests", + "authorization_digest", + "overlay", +} +_SOURCE_KEYS = { + "source_id", + "logical_path", + "classification", + "format", + "format_version", + "sha256", + "provenance", +} +_SOURCE_PROVENANCE_KEYS = { + "source_run_id", + "source_repository", + "source_commit", + "notes_digest", + "evidence_digest", + "unknown_reasons", +} +_SOURCE_PROVENANCE_UNKNOWN_FIELDS = { + "source_run_id", + "source_repository", + "source_commit", +} +_PARSING_POLICY_KEYS = { + "dialect", + "mapping", + "null_values", + "numeric_columns", + "option_mapping", + "extra_field_policy", +} +_DIALECT_KEYS = set(CsvDialectSpec.__dataclass_fields__) +_EVIDENCE_STATUSES = { + "complete", + "qualified", + "partial", + "malformed", + "recoverable", + "failed", +} +_SCOPE_DISPOSITIONS = { + "included", + "excluded_from_paper_matrix", + "held", + "superseded", +} +_LINEAGE_ID_RE = re.compile(r"^lin_[0-9a-f]{16}$") +_FORBIDDEN_IDENTITY_KEYS = { + "audit_path", + "absolute_path", + "choicebench_home", + "experiment_id", + "final_csv_sha256", + "final_result_digest", + "host", + "hostname", + "import_date", + "import_timestamp", + "output_path", + "output_root", + "realization_id", + "realization_digest", + "report_path", + "result_artifact_id", + "result_artifact_digest", + "result_sha256", + "source_path", + "temporary_path", + "timestamp", + "machine_id", + "machine", + "node", + "operator", + "platform", + "recorded_at", + "executed_at", + "run_date", + "cwd", + "child_lineage_id", +} +_PROTOCOL_SETTING_KEYS = {"pride_modal_k_threshold"} +_NATIVE_DATASET_SOURCE_KEYS = { + "generator", + "seed", + "hf_path", + "hf_subset", + "split", + "requested_revision", + "resolved_revision", + "hf_fingerprint", + "hf_dataset_info", + "revision", +} +_METHOD_RUNTIME_PARAMETERS = { + "self", + "args", + "kwargs", + "backend", + "method_name", + "split_name", + "prompt_version", + "prompts_dir", + "run_id", + "temperature", + "max_tokens", + "seed", + "perturbation_name", + "model_label", + "preflight_questions", + "calibration_questions", + "calibration_runs_dir", + "modal_k", + "gate_summary", + "condition_id", + "calibration_identity", +} +_DATASET_DERIVATION_KEYS = { + "schema_version", + "dataset_id", + "reference_kind", + "trust_label", + "source_chain", + "selection_source_id", + "columns", + "revision", + "fingerprint", + "declared_derivation", + "selected_question_ids", + "selection_semantics", + "selection_unknown_reasons", + "identity_mode", + "limitations", +} + + +def _unknown(reason: str | None, field: str) -> dict[str, Any]: + if not isinstance(reason, str) or not reason.strip(): + raise ImportIdentityError(f"Missing {field} requires an explicit unknown reason.") + return {"value": None, "reason": reason} + + +def _nullable_semantic(value: Any, reason: str | None, field: str) -> Any: + if value is not None: + return value + if isinstance(reason, str) and reason.strip().casefold() == "not applicable": + return None + return _unknown(reason, field) + + +def _validate_unknown_reason_contract( + values: Mapping[str, Any], reasons: Mapping[str, str], where: str +) -> None: + if not isinstance(reasons, Mapping): + raise ImportIdentityError(f"{where} unknown reasons must be a mapping.") + expected = {field for field, value in values.items() if value is None} + if set(reasons) != expected or not all( + isinstance(reason, str) and reason.strip() for reason in reasons.values() + ): + raise ImportIdentityError( + f"{where} unknown reasons contradict known and missing values." + ) + + +def _validate_direct_declarations( + condition: ImportConditionSpec, + model: ImportModelSpec, + method: ImportMethodSpec, + prompt: ImportPromptSpec, +) -> None: + _validate_stable_values( + { + "model_key": model.model_key, + "display_name": model.display_name, + "backend": model.backend, + "provider": model.provider, + "revision": model.revision, + "method_key": method.method_key, + "method_name": method.name, + "prompt_key": prompt.prompt_key, + "template_identity": prompt.template_identity, + }, + "Semantic child declarations", + ) + _validate_stable_parameters( + model.effective_parameters, "Model effective parameters" + ) + _validate_stable_parameters( + condition.generation_parameters, "Condition generation parameters" + ) + _validate_stable_parameters( + method.effective_parameters, "Method effective parameters" + ) + if condition.preflight_identity is not None: + _validate_stable_parameters( + condition.preflight_identity, "Condition preflight identity" + ) + if method.implementation is not None: + try: + validated_implementation = validate_implementation_identity_record( + method.implementation + ) + except ImportSpecError as exc: + raise ImportIdentityError(f"Invalid method implementation: {exc}") from exc + if canonicalize(validated_implementation) != canonicalize( + method.implementation + ): + raise ImportIdentityError( + "Method implementation does not match the closed code identity schema." + ) + _validate_stable_values(validated_implementation, "Method implementation") + _validate_unknown_reason_contract( + { + "backend": model.backend, + "provider": model.provider, + "revision": model.revision, + }, + model.unknown_reasons, + "Model declaration", + ) + _validate_unknown_reason_contract( + {"implementation": method.implementation}, + method.unknown_reasons, + "Method declaration", + ) + _validate_unknown_reason_contract( + { + "seed": condition.seed, + "calibration_identity": condition.calibration_identity, + "preflight_identity": condition.preflight_identity, + }, + condition.unknown_reasons, + "Condition declaration", + ) + prompt_values = ( + prompt.template_identity, + prompt.template_digest, + prompt.template_contents, + ) + has_unknown_prompt = any(value is None for value in prompt_values) + if has_unknown_prompt != (prompt.unknown_reason is not None) or ( + prompt.unknown_reason is not None + and (not isinstance(prompt.unknown_reason, str) or not prompt.unknown_reason.strip()) + ): + raise ImportIdentityError( + "Prompt declaration unknown reason contradicts known and missing values." + ) + if prompt.template_digest is not None: + _validate_digest(prompt.template_digest, "prompt.template_digest") + if ( + prompt.native_compatibility_identity is None + and prompt.template_digest is not None + and prompt.template_contents is not None + and prompt.template_digest != integrity_digest(prompt.template_contents) + ): + raise ImportIdentityError( + "Fallback prompt digest does not own the exact template contents." + ) + + +def _identity_record(prefix: str, payload: Mapping[str, Any]) -> tuple[str, str, dict]: + identity = canonicalize(payload) + return short_id(prefix, identity), integrity_digest(identity), identity + + +def _verify_claim( + *, prefix: str, payload: Mapping[str, Any], claimed_id: Any, claimed_digest: Any +) -> tuple[str, str, dict]: + if not isinstance(payload, Mapping): + raise ImportIdentityError(f"Invalid validated native {prefix} identity payload.") + identity_id, digest, identity = _identity_record(prefix, payload) + if claimed_id != identity_id or claimed_digest != digest: + raise ImportIdentityError(f"Invalid validated native {prefix} identity claim.") + return identity_id, digest, identity + + +def _reject_forbidden( + value: Any, where: str, *, allowed_path_keys: frozenset[str] = frozenset() +) -> None: + if isinstance(value, Mapping): + for key, item in value.items(): + if not isinstance(key, str): + raise ImportIdentityError(f"{where} contains a non-string field name.") + normalized = key.casefold() + unsafe_alias = ( + normalized in _FORBIDDEN_IDENTITY_KEYS + or (normalized not in allowed_path_keys and normalized.endswith("_path")) + or normalized in {"path", "created_at", "updated_at", "import_state"} + or "timestamp" in normalized + or normalized.endswith( + ("realization_id", "result_artifact_id", "experiment_id") + ) + or ("csv" in normalized and normalized.endswith(("sha256", "digest"))) + ) + if unsafe_alias: + raise ImportIdentityError(f"{where} contains forbidden field {key!r}.") + _reject_forbidden(item, where, allowed_path_keys=allowed_path_keys) + elif isinstance(value, (list, tuple)): + for item in value: + _reject_forbidden(item, where, allowed_path_keys=allowed_path_keys) + + +def _is_machine_path(value: str) -> bool: + return bool( + value.startswith(("/", "./", "../", "~/", "~\\", "\\\\")) + or re.match(r"^[A-Za-z]:[\\/]", value) + ) + + +def _validate_stable_values(value: Any, where: str) -> None: + _reject_forbidden(value, where) + if isinstance(value, Mapping): + for item in value.values(): + _validate_stable_values(item, where) + elif isinstance(value, (list, tuple)): + for item in value: + _validate_stable_values(item, where) + elif isinstance(value, str) and _is_machine_path(value): + raise ImportIdentityError(f"{where} contains a machine-local path value.") + + +def _validate_no_machine_path_values(value: Any, where: str) -> None: + if isinstance(value, Mapping): + for item in value.values(): + _validate_no_machine_path_values(item, where) + elif isinstance(value, (list, tuple)): + for item in value: + _validate_no_machine_path_values(item, where) + elif isinstance(value, str) and _is_machine_path(value): + raise ImportIdentityError(f"{where} contains a machine-local path value.") + + +def _validate_stable_parameters(value: Any, where: str) -> None: + _reject_forbidden(value, where) + if isinstance(value, Mapping): + for key, item in value.items(): + if not isinstance(key, str) or not key: + raise ImportIdentityError(f"{where} keys must be non-empty strings.") + normalized = key.casefold().replace("-", "_") + if ( + is_credential_key(key) + or normalized.endswith( + ( + "_id", + "_ids", + "_digest", + "_digests", + "_checksum", + "_hash", + "_sha256", + ) + ) + ): + raise ImportIdentityError( + f"{where} contains forbidden metadata field {key!r}." + ) + _validate_stable_parameters(item, where) + elif isinstance(value, (list, tuple)): + for item in value: + _validate_stable_parameters(item, where) + elif isinstance(value, str) and ( + _is_machine_path(value) or "/" in value or "\\" in value + ): + raise ImportIdentityError(f"{where} contains a path-like metadata value.") + else: + try: + canonicalize(value) + except (TypeError, ValueError) as exc: + raise ImportIdentityError(f"{where} contains unsafe parameter data.") from exc + + +def _validate_condition_identity(value: Mapping[str, Any]) -> None: + benchmark = _exact_mapping( + value["benchmark"], + {"name", "split", "selection_id", "artifact_id"}, + "Semantic condition benchmark", + ) + for field in ("name", "split"): + if not isinstance(benchmark[field], str) or not benchmark[field]: + raise ImportIdentityError( + f"Semantic condition benchmark {field} must be non-empty." + ) + for field, prefix in ( + ("selection_id", "sel"), + ("artifact_id", "ds"), + ("model_id", "model"), + ("method_id", "method"), + ("prompt_id", "prompt"), + ): + identifier = benchmark[field] if field in benchmark else value[field] + if not isinstance(identifier, str) or not re.fullmatch( + rf"{prefix}_[0-9a-f]{{16}}", identifier + ): + raise ImportIdentityError( + f"Semantic condition {field} is not a valid {prefix} identity." + ) + preflight = value["preflight"] + if preflight is not None: + if not isinstance(preflight, Mapping): + raise ImportIdentityError("Semantic condition preflight must be a mapping or null.") + if set(preflight) == {"value", "reason"}: + if preflight["value"] is not None or not isinstance( + preflight["reason"], str + ) or not preflight["reason"].strip(): + raise ImportIdentityError( + "Semantic condition unknown preflight must include a reason." + ) + else: + native = _exact_mapping( + preflight, + {"artifact_id", "selection_id", "split", "content_digest"}, + "Semantic condition preflight", + ) + if not re.fullmatch(r"ds_[0-9a-f]{16}", str(native["artifact_id"])): + raise ImportIdentityError("Semantic condition preflight artifact_id is invalid.") + if not re.fullmatch(r"sel_[0-9a-f]{16}", str(native["selection_id"])): + raise ImportIdentityError("Semantic condition preflight selection_id is invalid.") + if not isinstance(native["split"], str) or not native["split"]: + raise ImportIdentityError("Semantic condition preflight split is invalid.") + _validate_digest(native["content_digest"], "preflight.content_digest") + seed = value["seed"] + if not ( + isinstance(seed, int) + and not isinstance(seed, bool) + or ( + isinstance(seed, Mapping) + and set(seed) == {"value", "reason"} + and seed.get("value") is None + and isinstance(seed.get("reason"), str) + and bool(seed["reason"].strip()) + ) + ): + raise ImportIdentityError("Semantic condition seed is invalid.") + if "protocol_settings" in value: + if not isinstance(value["protocol_settings"], Mapping): + raise ImportIdentityError("Semantic condition protocol_settings must be a mapping.") + unexpected_protocol = sorted( + set(value["protocol_settings"]) - _PROTOCOL_SETTING_KEYS + ) + if unexpected_protocol: + raise ImportIdentityError( + "Semantic condition protocol_settings contains unsupported metadata " + f"or protocol fields {unexpected_protocol}." + ) + threshold = value["protocol_settings"].get("pride_modal_k_threshold") + if threshold is not None and ( + isinstance(threshold, bool) + or not isinstance(threshold, (int, float)) + or not math.isfinite(float(threshold)) + ): + raise ImportIdentityError( + "Semantic condition pride_modal_k_threshold must be finite numeric data." + ) + _validate_stable_parameters( + value["protocol_settings"], "Semantic condition protocol_settings" + ) + + +def _exact_mapping(value: Any, keys: set[str], where: str) -> Mapping[str, Any]: + if not isinstance(value, Mapping): + raise ImportIdentityError(f"{where} must be a mapping.") + missing = sorted(keys - set(value)) + unexpected = sorted(set(value) - keys) + if missing or unexpected: + raise ImportIdentityError( + f"{where} fields are invalid; missing={missing}, unexpected={unexpected}." + ) + return value + + +def _digest_sequence(value: Any, field: str) -> list[str]: + if not isinstance(value, (list, tuple)): + raise ImportIdentityError(f"{field} must be a digest sequence.") + result = [] + for index, digest in enumerate(value): + checked = _validate_digest(digest, f"{field}[{index}]") + assert checked is not None + result.append(checked) + if len(result) != len(set(result)): + raise ImportIdentityError(f"{field} contains duplicate digest edges.") + return sorted(result) + + +def _validate_expected_dataset_contract(dataset: ExpectedDataset) -> None: + artifact_payload = dataset.artifact_payload + if dataset.identity_mode == "imported_semantic_fallback": + artifact = _exact_mapping( + artifact_payload, + {"schema_version", "benchmark", "split", "content_digest"}, + "Expected dataset fallback artifact", + ) + if artifact["schema_version"] != "choicebench.semantic-dataset.v1": + raise ImportIdentityError("Expected dataset fallback schema is invalid.") + if ( + artifact["benchmark"] != dataset.benchmark_name + or artifact["split"] != dataset.split + ): + raise ImportIdentityError( + "Expected dataset fallback artifact conflicts with its declaration." + ) + else: + artifact = _exact_mapping( + artifact_payload, + {"spec", "content_digest", "source"}, + "Expected dataset native artifact", + ) + spec = _exact_mapping( + artifact["spec"], + { + "benchmark", + "split", + "hf_path", + "hf_subset", + "source_revision", + "normalization_version", + "transforms", + "output_name", + }, + "Expected dataset native artifact spec", + ) + if spec["benchmark"] != dataset.benchmark_name or spec["split"] != dataset.split: + raise ImportIdentityError( + "Expected dataset native artifact conflicts with its declaration." + ) + _validate_no_machine_path_values( + spec, "Expected dataset native artifact spec" + ) + if not isinstance(artifact["source"], Mapping): + raise ImportIdentityError("Expected dataset native source is invalid.") + source = artifact["source"] + unexpected_source = sorted(set(source) - _NATIVE_DATASET_SOURCE_KEYS) + if unexpected_source: + raise ImportIdentityError( + "Expected dataset native source contains unsupported audit or " + f"source fields {unexpected_source}." + ) + if "generator" in source and ( + not isinstance(source["generator"], str) or not source["generator"] + ): + raise ImportIdentityError("Expected dataset native source generator is invalid.") + if isinstance(source.get("generator"), str) and _is_machine_path( + source["generator"] + ): + raise ImportIdentityError( + "Expected dataset native source generator contains a machine-local path." + ) + if "seed" in source and ( + not isinstance(source["seed"], int) or isinstance(source["seed"], bool) + ): + raise ImportIdentityError("Expected dataset native source seed is invalid.") + for field in ( + "hf_path", + "hf_subset", + "split", + "requested_revision", + "resolved_revision", + "hf_fingerprint", + "revision", + ): + if field in source and source[field] is not None and not isinstance( + source[field], str + ): + raise ImportIdentityError( + f"Expected dataset native source {field} is invalid." + ) + if isinstance(source.get(field), str) and _is_machine_path(source[field]): + raise ImportIdentityError( + f"Expected dataset native source {field} contains a " + "machine-local path." + ) + if "hf_dataset_info" in source: + info = _exact_mapping( + source["hf_dataset_info"], + {"builder_name", "config_name", "version"}, + "Expected dataset native source hf_dataset_info", + ) + if any( + item is not None and not isinstance(item, str) + for item in info.values() + ): + raise ImportIdentityError( + "Expected dataset native source hf_dataset_info is invalid." + ) + _validate_no_machine_path_values( + info, "Expected dataset native source hf_dataset_info" + ) + if artifact["content_digest"] != dataset_content_digest(dataset.artifact_frame): + raise ImportIdentityError( + "Expected dataset artifact content digest does not own its frame." + ) + + semantics = _exact_mapping( + dataset.selection_semantics, + {"seed", "n_samples", "subject_filter"}, + "Expected dataset selection semantics", + ) + selection = _exact_mapping( + dataset.selection_payload, + { + "artifact_id", + "content_digest", + "sample_identities", + "seed", + "n_samples", + "subject_filter", + }, + "Expected dataset selection", + ) + expected_selection = canonicalize( + { + "artifact_id": dataset.artifact_id, + "content_digest": dataset_content_digest(dataset.frame), + "sample_identities": dataset_sample_identities(dataset.frame), + "seed": semantics["seed"], + "n_samples": semantics["n_samples"], + "subject_filter": sorted(semantics["subject_filter"]), + } + ) + if canonicalize(selection) != expected_selection: + raise ImportIdentityError( + "Expected dataset selection does not own its selected frame." + ) + frame_question_ids = tuple(dataset.frame["question_id"].astype(str)) + if dataset.selected_question_ids != frame_question_ids: + raise ImportIdentityError( + "Expected dataset selected question IDs do not match its frame." + ) + artifact_question_ids = tuple(dataset.artifact_frame["question_id"].astype(str)) + if len(artifact_question_ids) != len(set(artifact_question_ids)): + raise ImportIdentityError("Expected dataset artifact has duplicate question IDs.") + artifact_samples = dict( + zip( + artifact_question_ids, + dataset_sample_identities(dataset.artifact_frame), + strict=True, + ) + ) + selected_samples = dataset_sample_identities(dataset.frame) + for question_id, sample_identity in zip( + frame_question_ids, selected_samples, strict=True + ): + if artifact_samples.get(question_id) != sample_identity: + raise ImportIdentityError( + "Expected dataset selected content is not owned by its artifact." + ) + expected_unknowns = { + field + for field, value in ( + ("selection_seed", semantics["seed"]), + ("selection_n_samples", semantics["n_samples"]), + ) + if value is None + } + if set(dataset.selection_unknown_reasons) != expected_unknowns or not all( + isinstance(reason, str) and reason.strip() + for reason in dataset.selection_unknown_reasons.values() + ): + raise ImportIdentityError( + "Expected dataset selection unknown reasons are contradictory." + ) + + if dataset.reference_kind not in { + "independent_input_snapshot", + "profile_derived_reference_snapshot", + }: + raise ImportIdentityError("Expected dataset reference kind is invalid.") + if not isinstance(dataset.trust_label, str) or not dataset.trust_label.strip(): + raise ImportIdentityError("Expected dataset trust label is invalid.") + if dataset.question_set_digest != integrity_digest(list(frame_question_ids)): + raise ImportIdentityError( + "Expected dataset question-set digest does not own its selected IDs." + ) + derivation = _exact_mapping( + dataset.derivation, + _DATASET_DERIVATION_KEYS, + "Expected dataset derivation", + ) + expected_derivation_fields = { + "schema_version": "choicebench.dataset-reference.v1", + "dataset_id": dataset.dataset_id, + "reference_kind": dataset.reference_kind, + "trust_label": dataset.trust_label, + "selected_question_ids": list(frame_question_ids), + "selection_semantics": semantics, + "selection_unknown_reasons": dict(dataset.selection_unknown_reasons), + "identity_mode": dataset.identity_mode, + "limitations": list(dataset.limitations), + } + for field, expected_value in expected_derivation_fields.items(): + if canonicalize(derivation[field]) != canonicalize(expected_value): + raise ImportIdentityError( + f"Expected dataset derivation {field} conflicts with its snapshot." + ) + source_chain = derivation["source_chain"] + if not isinstance(source_chain, (list, tuple)) or not source_chain: + raise ImportIdentityError("Expected dataset derivation source chain is invalid.") + seen_source_ids: set[str] = set() + for index, item in enumerate(source_chain): + source = _exact_mapping( + item, + {"source_id", "logical_path", "sha256"}, + f"Expected dataset derivation source_chain[{index}]", + ) + if ( + not isinstance(source["source_id"], str) + or not source["source_id"] + or source["source_id"] in seen_source_ids + ): + raise ImportIdentityError( + "Expected dataset derivation source IDs must be non-empty and unique." + ) + logical_path = source["logical_path"] + if not isinstance(logical_path, str): + raise ImportIdentityError( + "Expected dataset derivation logical path is invalid." + ) + pure_path = PurePosixPath(logical_path) + if ( + pure_path.is_absolute() + or ".." in pure_path.parts + or not pure_path.parts + or _is_machine_path(logical_path) + ): + raise ImportIdentityError( + "Expected dataset derivation logical path is unsafe." + ) + _validate_digest( + source["sha256"], + f"Expected dataset derivation source_chain[{index}].sha256", + ) + seen_source_ids.add(source["source_id"]) + if ( + not isinstance(derivation["selection_source_id"], str) + or derivation["selection_source_id"] not in seen_source_ids + ): + raise ImportIdentityError( + "Expected dataset derivation selection source is not in its source chain." + ) + if not isinstance(derivation["columns"], Mapping) or not all( + isinstance(key, str) and isinstance(value, str) + for key, value in derivation["columns"].items() + ): + raise ImportIdentityError("Expected dataset derivation columns are invalid.") + _validate_stable_values( + derivation["revision"], "Expected dataset derivation revision" + ) + _validate_stable_values( + derivation["fingerprint"], "Expected dataset derivation fingerprint" + ) + _validate_stable_values( + derivation["declared_derivation"], + "Expected dataset declared derivation", + ) + if dataset.reference_kind == "profile_derived_reference_snapshot": + declared_derivation = derivation["declared_derivation"] + if not isinstance(declared_derivation, Mapping): + raise ImportIdentityError( + "Profile-derived reference declared derivation is invalid." + ) + raw_groups = declared_derivation.get("independent_source_groups") + if not isinstance(raw_groups, (list, tuple)): + raise ImportIdentityError( + "Profile-derived reference requires at least two independent " + "source groups." + ) + groups: list[tuple[str, ...]] = [] + for item in raw_groups: + if not isinstance(item, (list, tuple)) or not item: + raise ImportIdentityError( + "Profile-derived reference requires at least two independent " + "source groups." + ) + groups.append(tuple(str(source_id) for source_id in item)) + if len(groups) < 2: + raise ImportIdentityError( + "Profile-derived reference requires at least two independent " + "source groups." + ) + flattened = [source_id for group in groups for source_id in group] + if ( + len(flattened) != len(set(flattened)) + or set(flattened) != seen_source_ids + ): + raise ImportIdentityError( + "Profile-derived reference independent source groups must be " + "disjoint and cover every declared source." + ) + expected_derivation_digest = integrity_digest(derivation) + if dataset.derivation_digest != expected_derivation_digest: + raise ImportIdentityError("Expected dataset derivation digest is inconsistent.") + expected_snapshot_digest = integrity_digest( + { + "artifact_digest": dataset.artifact_digest, + "selection_digest": dataset.selection_digest, + "question_set_digest": dataset.question_set_digest, + "derivation_digest": expected_derivation_digest, + "reference_kind": dataset.reference_kind, + "trust_label": dataset.trust_label, + } + ) + if dataset.snapshot_digest != expected_snapshot_digest: + raise ImportIdentityError("Expected dataset snapshot digest is inconsistent.") + + +def _semantic_child_records( + *, + condition: ImportConditionSpec, + dataset: ExpectedDataset, + model: ImportModelSpec, + method: ImportMethodSpec, + prompt: ImportPromptSpec, +) -> tuple[dict, dict, dict, dict, dict]: + dataset_record = { + "artifact_id": dataset.artifact_id, + "artifact_digest": dataset.artifact_digest, + "identity": canonicalize(dataset.artifact_payload), + "identity_mode": dataset.identity_mode, + } + if dataset.artifact_digest != integrity_digest(dataset_record["identity"]): + raise ImportIdentityError("Expected dataset artifact digest is inconsistent.") + if dataset.artifact_id != short_id("ds", dataset_record["identity"]): + raise ImportIdentityError("Expected dataset artifact ID is inconsistent.") + selection_record = { + "selection_id": dataset.selection_id, + "selection_digest": dataset.selection_digest, + "identity": canonicalize(dataset.selection_payload), + } + if dataset.selection_digest != integrity_digest(selection_record["identity"]): + raise ImportIdentityError("Expected dataset selection digest is inconsistent.") + if dataset.selection_id != short_id("sel", selection_record["identity"]): + raise ImportIdentityError("Expected dataset selection ID is inconsistent.") + + if model.native_compatibility_identity is not None: + native = model.native_compatibility_identity + if not isinstance(native, Mapping) or set(native) != { + "payload", + "digest", + "model_id", + }: + raise ImportIdentityError("Invalid validated native model identity fields.") + model_id, model_digest, model_payload = _verify_claim( + prefix="model", + payload=native["payload"], + claimed_id=native["model_id"], + claimed_digest=native["digest"], + ) + if model.backend is None: + raise ImportIdentityError( + "An unknown model backend cannot be upgraded by a native identity claim." + ) + try: + validated_model_payload = validate_native_model_payload(model_payload) + except ImportSpecError as exc: + raise ImportIdentityError(f"Invalid validated native model payload: {exc}") from exc + if canonicalize(validated_model_payload) != canonicalize(model_payload): + raise ImportIdentityError( + "Validated native model payload does not match the current backend schema." + ) + model_payload = validated_model_payload + for nullable_field in ("provider", "revision"): + value = getattr(model, nullable_field) + reason = model.unknown_reasons.get(nullable_field) + if value is None and ( + not isinstance(reason, str) + or reason.strip().casefold() != "not applicable" + ): + raise ImportIdentityError( + f"Unknown model {nullable_field} cannot be upgraded by a native " + "identity claim." + ) + if model.backend is not None and model_payload.get("backend") != model.backend: + raise ImportIdentityError( + "Validated native model backend conflicts with its declaration." + ) + if model.provider is not None and model_payload.get("provider") != model.provider: + raise ImportIdentityError( + "Validated native model provider conflicts with its declaration." + ) + if ( + "model_name_or_path" in model_payload + and model_payload["model_name_or_path"] != model.display_name + ): + raise ImportIdentityError( + "Validated native model name conflicts with its declaration." + ) + if model.backend == "huggingface": + resolved_model = model_payload["model"] + declared_name = ( + resolved_model["repo_id"] + if resolved_model["kind"] == "huggingface-hub" + else resolved_model["logical_name"] + ) + if declared_name != model.display_name: + raise ImportIdentityError( + "Validated native model name conflicts with its declaration." + ) + if resolved_model.get("requested_revision") != model.revision: + raise ImportIdentityError( + "Validated native model revision conflicts with its declaration." + ) + elif model.revision is not None: + raise ImportIdentityError( + "The current native model identity cannot represent a declared " + "revision; use the imported semantic fallback." + ) + expected_generation = dict(model.effective_parameters) + for name, value in condition.generation_parameters.items(): + if name in expected_generation and expected_generation[name] != value: + raise ImportIdentityError( + f"Generation parameter {name!r} conflicts with model effective parameters." + ) + expected_generation[name] = value + native_generation = model_payload.get("generation_kwargs", {}) + if canonicalize(native_generation) != canonicalize(expected_generation): + raise ImportIdentityError( + "Condition generation parameters are not bound by the validated " + "native model identity." + ) + model_mode = "native_compatibility" + else: + effective_parameters = dict(model.effective_parameters) + for name, value in condition.generation_parameters.items(): + if name in effective_parameters and effective_parameters[name] != value: + raise ImportIdentityError( + f"Generation parameter {name!r} conflicts with model effective parameters." + ) + effective_parameters[name] = value + model_payload = { + "schema_version": "choicebench.semantic-model.v1", + "backend": ( + model.backend + if model.backend is not None + else _unknown(model.unknown_reasons.get("backend"), "model backend") + ), + "provider": ( + model.provider + if model.provider is not None + else _unknown(model.unknown_reasons.get("provider"), "model provider") + ), + "model": model.display_name, + "revision": ( + model.revision + if model.revision is not None + else _unknown(model.unknown_reasons.get("revision"), "model revision") + ), + "effective_parameters": effective_parameters, + } + model_id, model_digest, model_payload = _identity_record("model", model_payload) + model_mode = "imported_semantic_fallback" + model_record = { + "model_id": model_id, + "model_digest": model_digest, + "identity": model_payload, + "identity_mode": model_mode, + } + + if method.native_compatibility_identity is not None: + native = method.native_compatibility_identity + if not isinstance(native, Mapping) or set(native) != { + "payload", + "digest", + "method_id", + }: + raise ImportIdentityError("Invalid validated native method identity fields.") + method_id, method_digest, method_payload = _verify_claim( + prefix="method", + payload=native["payload"], + claimed_id=native["method_id"], + claimed_digest=native["digest"], + ) + try: + validated_method_payload = validate_native_method_payload(method_payload) + except ImportSpecError as exc: + raise ImportIdentityError( + f"Invalid validated native method payload: {exc}" + ) from exc + if canonicalize(validated_method_payload) != canonicalize(method_payload): + raise ImportIdentityError( + "Validated native method payload does not match the current schema." + ) + method_payload = validated_method_payload + runner_cls = METHOD_REGISTRY.get(method.name) + if runner_cls is None: + raise ImportIdentityError( + "Validated native method does not name a registered current runner." + ) + current_effective_params: dict[str, Any] = {} + accepted_parameters: set[str] = set() + for name, parameter in inspect.signature(runner_cls.__init__).parameters.items(): + if name in _METHOD_RUNTIME_PARAMETERS: + continue + accepted_parameters.add(name) + if parameter.default is not inspect.Parameter.empty: + current_effective_params[name] = parameter.default + unexpected_effective = sorted( + set(method.effective_parameters) - accepted_parameters + ) + if unexpected_effective: + raise ImportIdentityError( + "Validated native method declares parameters not accepted by its " + f"registered runner: {unexpected_effective}." + ) + current_effective_params.update(method.effective_parameters) + declared_preflight = _nullable_semantic( + condition.preflight_identity, + condition.unknown_reasons.get("preflight_identity"), + "method preflight identity", + ) + expected_current_payload = canonicalize( + { + "name": method.name, + "effective_params": current_effective_params, + "preflight": declared_preflight, + "implementation": implementation_identity(runner_cls), + } + ) + if canonicalize(method_payload) != expected_current_payload: + raise ImportIdentityError( + "Validated native method does not match its registered current runner." + ) + if method.implementation is None: + raise ImportIdentityError( + "An unknown method implementation cannot be upgraded by a native " + "identity claim." + ) + if ( + method_payload.get("name") != method.name + or canonicalize(method_payload.get("effective_params")) + != canonicalize(method.effective_parameters) + or canonicalize(method_payload.get("preflight")) + != canonicalize(declared_preflight) + or ( + method.implementation is not None + and canonicalize(method_payload.get("implementation")) + != canonicalize(method.implementation) + ) + ): + raise ImportIdentityError( + "Validated native method identity conflicts with its semantic declaration." + ) + method_mode = "native_compatibility" + else: + method_payload = { + "schema_version": "choicebench.semantic-method.v1", + "name": method.name, + "effective_params": dict(method.effective_parameters), + "preflight": _nullable_semantic( + condition.preflight_identity, + condition.unknown_reasons.get("preflight_identity"), + "method preflight identity", + ), + "implementation": ( + method.implementation + if method.implementation is not None + else _unknown( + method.unknown_reasons.get("implementation"), + "method implementation", + ) + ), + } + method_id, method_digest, method_payload = _identity_record( + "method", method_payload + ) + method_mode = "imported_semantic_fallback" + method_record = { + "method_id": method_id, + "method_digest": method_digest, + "identity": method_payload, + "identity_mode": method_mode, + } + + if prompt.native_compatibility_identity is not None: + native = prompt.native_compatibility_identity + if not isinstance(native, Mapping) or set(native) != { + "prompt_id", + "version", + "files", + }: + raise ImportIdentityError("Invalid validated native prompt identity fields.") + prompt_payload = canonicalize( + {"version": native["version"], "files": native["files"]}, + redact_secrets=False, + ) + if any( + value is None + for value in ( + prompt.template_identity, + prompt.template_digest, + prompt.template_contents, + ) + ): + raise ImportIdentityError( + "An unknown prompt cannot be upgraded by a native identity claim." + ) + native_files_raw = _exact_mapping( + prompt_payload.get("files"), + {"direct_mcq", "free_text", "option_matching"}, + "Validated native prompt files", + ) + for name, raw_item in native_files_raw.items(): + item = _exact_mapping( + raw_item, {"sha256", "content"}, f"Validated native prompt file {name}" + ) + if not isinstance(item["content"], str) or item["sha256"] != integrity_digest( + item["content"] + ): + raise ImportIdentityError( + f"Validated native prompt file {name!r} has an invalid content hash." + ) + prompt_id = f"prompt_{integrity_digest(prompt_payload)[:16]}" + if native["prompt_id"] != prompt_id: + raise ImportIdentityError("Invalid validated native prompt identity claim.") + declared_contents = prompt.template_contents + native_files = prompt_payload.get("files") + native_contents = ( + { + name: item.get("content") + for name, item in native_files.items() + if isinstance(item, Mapping) + } + if isinstance(native_files, Mapping) + else None + ) + if ( + declared_contents is not None + and canonicalize(declared_contents, redact_secrets=False) + != canonicalize(native_contents, redact_secrets=False) + ): + raise ImportIdentityError( + "Validated native prompt contents conflict with their declaration." + ) + if ( + prompt.template_identity is not None + and prompt.template_identity != prompt_payload.get("version") + ): + raise ImportIdentityError( + "Validated native prompt identity conflicts with its declaration." + ) + if ( + prompt.template_digest is not None + and prompt.template_digest != integrity_digest(native_files) + ): + raise ImportIdentityError( + "Validated native prompt digest conflicts with its declaration." + ) + prompt_digest = integrity_digest(prompt_payload) + prompt_mode = "native_compatibility" + else: + unknown = lambda field: _unknown(prompt.unknown_reason, f"prompt {field}") + prompt_payload = { + "schema_version": "choicebench.semantic-prompt.v1", + "template_identity": ( + prompt.template_identity + if prompt.template_identity is not None + else unknown("template identity") + ), + "template_digest": ( + prompt.template_digest + if prompt.template_digest is not None + else unknown("template digest") + ), + "template_contents": ( + prompt.template_contents + if prompt.template_contents is not None + else unknown("template contents") + ), + } + prompt_payload = canonicalize(prompt_payload, redact_secrets=False) + prompt_digest = integrity_digest(prompt_payload) + prompt_id = f"prompt_{prompt_digest[:16]}" + prompt_mode = "imported_semantic_fallback" + prompt_record = { + "prompt_id": prompt_id, + "prompt_digest": prompt_digest, + "identity": prompt_payload, + "identity_mode": prompt_mode, + } + return dataset_record, selection_record, model_record, method_record, prompt_record + + +def make_semantic_condition( + *, identity: Mapping[str, Any], fields: Mapping[str, Any] +) -> dict[str, Any]: + """Create one scientific condition without realization provenance.""" + allowed = set(_CONDITION_KEYS) + if "protocol_settings" in identity: + allowed.add("protocol_settings") + if set(identity) != allowed: + raise ImportIdentityError("Semantic condition identity fields are invalid.") + prompt_id = identity.get("prompt_id") + expected_prompt_path = f"artifacts/prompts/{prompt_id}" + if identity.get("prompt_snapshot_path") != expected_prompt_path: + raise ImportIdentityError( + "Semantic condition prompt snapshot path must be derived from prompt_id." + ) + _validate_condition_identity(identity) + _reject_forbidden( + identity, + "Semantic condition identity", + allowed_path_keys=frozenset({"prompt_snapshot_path"}), + ) + payload = canonicalize(identity) + condition_id = short_id("cond", payload) + condition_digest = integrity_digest(payload) + if not isinstance(fields, Mapping) or not set(fields) <= { + "condition_key", + "expected_question_ids", + }: + raise ImportIdentityError( + "Semantic condition fields may contain only condition_key and " + "expected_question_ids." + ) + extra = canonicalize(fields) + return { + "condition_id": condition_id, + "condition_digest": condition_digest, + "identity": payload, + **extra, + } + + +def build_import_semantic_identity( + *, + condition: ImportConditionSpec, + dataset: ExpectedDataset, + model: ImportModelSpec, + method: ImportMethodSpec, + prompt: ImportPromptSpec, +) -> ImportSemanticRecords: + """Build semantic children and a current-shape scientific condition.""" + _validate_direct_declarations(condition, model, method, prompt) + if dataset.identity_mode not in { + "imported_semantic_fallback", + "native_compatibility", + }: + raise ImportIdentityError( + f"Unsupported expected dataset identity mode {dataset.identity_mode!r}." + ) + _validate_expected_dataset_contract(dataset) + references = ( + ("dataset_id", condition.dataset_id, dataset.dataset_id), + ("model_key", condition.model_key, model.model_key), + ("method_key", condition.method_key, method.method_key), + ("prompt_key", condition.prompt_key, prompt.prompt_key), + ) + for field, declared, supplied in references: + if declared != supplied: + raise ImportIdentityError( + f"Condition {field}={declared!r} does not match supplied child {supplied!r}." + ) + if tuple(condition.expected_question_ids) != dataset.selected_question_ids: + raise ImportIdentityError("Condition question set does not match dataset selection.") + artifact_identity = dataset.artifact_payload + if dataset.identity_mode == "imported_semantic_fallback": + if ( + artifact_identity.get("benchmark") != dataset.benchmark_name + or artifact_identity.get("split") != dataset.split + ): + raise ImportIdentityError( + "Expected dataset description conflicts with artifact identity." + ) + else: + artifact_spec = artifact_identity.get("spec") + if ( + not isinstance(artifact_spec, Mapping) + or artifact_spec.get("benchmark") != dataset.benchmark_name + or artifact_spec.get("split") != dataset.split + ): + raise ImportIdentityError( + "Expected dataset description conflicts with native artifact identity." + ) + if dataset.selection_payload.get("artifact_id") != dataset.artifact_id: + raise ImportIdentityError("Dataset selection points to a different artifact.") + dataset_record, selection, model_record, method_record, prompt_record = ( + _semantic_child_records( + condition=condition, + dataset=dataset, + model=model, + method=method, + prompt=prompt, + ) + ) + seed = ( + condition.seed + if condition.seed is not None + else _unknown(condition.unknown_reasons.get("seed"), "condition seed") + ) + preflight = _nullable_semantic( + condition.calibration_identity, + condition.unknown_reasons.get("calibration_identity"), + "condition calibration identity", + ) + condition_payload: dict[str, Any] = { + "benchmark": { + "name": dataset.benchmark_name, + "split": dataset.split, + "selection_id": selection["selection_id"], + "artifact_id": dataset_record["artifact_id"], + }, + "preflight": preflight, + "model_id": model_record["model_id"], + "method_id": method_record["method_id"], + "prompt_id": prompt_record["prompt_id"], + "prompt_snapshot_path": f"artifacts/prompts/{prompt_record['prompt_id']}", + "seed": seed, + } + if condition.protocol_settings: + condition_payload["protocol_settings"] = condition.protocol_settings + condition_record = make_semantic_condition( + identity=condition_payload, + fields={ + "condition_key": condition.condition_key, + "expected_question_ids": list(condition.expected_question_ids), + }, + ) + return ImportSemanticRecords( + dataset_artifact=dataset_record, + selection=selection, + model=model_record, + method=method_record, + prompt=prompt_record, + condition=condition_record, + ) + + +def _validate_digest(value: Any, field: str, *, optional: bool = False) -> str | None: + if optional and value is None: + return None + if ( + not isinstance(value, str) + or len(value) != 64 + or any(character not in "0123456789abcdef" for character in value) + ): + raise ImportIdentityError(f"{field} must be a lowercase SHA-256 digest.") + return value + + +@functools.lru_cache(maxsize=None) +def _cached_global_implementation_identity(value: Any) -> dict[str, Any]: + """Cache `implementation_identity` per referenced global object. + + `implementation_identity` walks and hashes a non-`choicebench` package's + entire file tree; without caching, every callable that references the + same third-party helper (e.g. a bare `from pandas import isna`) would + repeat that walk on every identity computation. The referenced object's + defining files do not change within a process, so caching by object + identity (functions/classes hash by identity) is safe and keeps repeated + identity computations cheap. + """ + return validate_implementation_identity_record(implementation_identity(value)) + + +def _resolved_global_bindings(target: Callable[..., Any], where: str) -> dict[str, Any]: + """Bind behavior-affecting global values the callable's code resolves by name. + + Function/class globals (helpers) are bound by their own defining-file + identity rather than skipped: `implementation_identity` only hashes a + callable's own defining file, so a helper imported from another module is + never covered by the calling callable's own source digest. Binding it + shallowly (not expanding its own resolved globals) closes that gap without + risking unbounded or cyclic expansion through mutually referencing + helpers. Modules and dunder names are excluded as irrelevant runtime + state; anything else that cannot be represented stably fails closed + rather than being silently dropped. + """ + code = getattr(target, "__code__", None) + global_ns = getattr(target, "__globals__", None) + if code is None or global_ns is None: + return {} + bindings: dict[str, Any] = {} + for name in sorted(set(code.co_names)): + if name.startswith("__") or name not in global_ns: + continue + value = global_ns[name] + if value is target or inspect.ismodule(value): + continue + if inspect.isfunction(value) or inspect.isclass(value): + try: + bindings[name] = _cached_global_implementation_identity(value) + except (ImportSpecError, OSError, TypeError, ValueError) as exc: + raise ImportIdentityError( + f"{where} references global {name!r} with no inspectable " + "implementation identity." + ) from exc + continue + if callable(value): + raise ImportIdentityError( + f"{where} references global {name!r} that is an uninspectable " + "stateful callable." + ) + try: + bindings[name] = canonicalize(value) + except (TypeError, ValueError) as exc: + raise ImportIdentityError( + f"{where} references global {name!r} with an unsupported value." + ) from exc + return bindings + + +def _runtime_callable_record(target: Callable[..., Any], where: str) -> dict[str, Any]: + if inspect.ismethod(target) or not ( + inspect.isfunction(target) or inspect.isclass(target) + ): + raise ImportIdentityError( + f"{where} must be a stateless module function or class, not a bound " + "method or stateful callable instance." + ) + if inspect.isfunction(target) and target.__closure__: + raise ImportIdentityError(f"{where} cannot be a stateful closure.") + qualified_name = getattr(target, "__qualname__", "") + if "" in qualified_name or "" in qualified_name: + raise ImportIdentityError( + f"{where} must be a uniquely addressable named module callable." + ) + try: + callable_source = inspect.getsource(target) + except (OSError, TypeError) as exc: + raise ImportIdentityError( + f"{where} lacks uniquely inspectable callable source." + ) from exc + callable_payload = { + "source": callable_source, + "defaults": getattr(target, "__defaults__", None), + "keyword_defaults": getattr(target, "__kwdefaults__", None), + "resolved_globals": _resolved_global_bindings(target, where), + } + try: + callable_digest = integrity_digest(callable_payload) + except (TypeError, ValueError) as exc: + raise ImportIdentityError( + f"{where} has unsupported callable defaults." + ) from exc + try: + record = validate_implementation_identity_record( + implementation_identity(target) + ) + except (ImportSpecError, OSError, TypeError, ValueError) as exc: + raise ImportIdentityError( + f"{where} lacks an inspectable runtime code identity." + ) from exc + if "source_digest" not in record: + raise ImportIdentityError(f"{where} lacks an inspectable source digest.") + record["callable_digest"] = callable_digest + return record + + +def _runtime_parameter_schema( + target: Callable[..., Any], +) -> tuple[set[str], set[str]]: + try: + signature = inspect.signature(target) + except (TypeError, ValueError) as exc: + raise ImportIdentityError( + "Lineage runtime implementation has no inspectable parameter schema." + ) from exc + allowed = { + name + for name, parameter in signature.parameters.items() + if parameter.kind is inspect.Parameter.KEYWORD_ONLY + } + required = { + name + for name, parameter in signature.parameters.items() + if parameter.kind is inspect.Parameter.KEYWORD_ONLY + and parameter.default is inspect.Parameter.empty + } + return allowed, required + + +def make_lineage_component( + *, + operation_type: str, + question_id: str, + parent_digests: Sequence[str], + source_digests: Sequence[str], + authorization_digest: str | None, + implementation: Callable[..., Any] | Mapping[str, Any], + implementation_mode: Literal["runtime_callable", "declared_external"] = ( + "runtime_callable" + ), + parameters: Mapping[str, Any], + input_digest: str, + preownership_output_digest: str, + prediction_origin: str, +) -> dict[str, Any]: + """Build a parent-derived lineage node without any child/self edge.""" + if not isinstance(operation_type, str) or not operation_type: + raise ImportIdentityError("Lineage operation_type must be non-empty.") + if not isinstance(question_id, str) or not question_id: + raise ImportIdentityError("Lineage question_id must be non-empty.") + if prediction_origin not in _PREDICTION_ORIGINS: + raise ImportIdentityError(f"Invalid prediction origin {prediction_origin!r}.") + for index, digest in enumerate(parent_digests): + _validate_digest(digest, f"parent_digests[{index}]") + for index, digest in enumerate(source_digests): + _validate_digest(digest, f"source_digests[{index}]") + if len(parent_digests) != len(set(parent_digests)): + raise ImportIdentityError("Lineage parent_digests contains duplicate edges.") + if len(source_digests) != len(set(source_digests)): + raise ImportIdentityError("Lineage source_digests contains duplicate edges.") + _validate_digest(authorization_digest, "authorization_digest", optional=True) + _validate_digest(input_digest, "input_digest") + _validate_digest(preownership_output_digest, "preownership_output_digest") + if implementation_mode == "runtime_callable": + if not callable(implementation): + raise ImportIdentityError( + "Lineage runtime implementation must be an inspectable callable." + ) + validated_implementation = _runtime_callable_record( + implementation, "Lineage runtime implementation" + ) + allowed_parameters, required_parameters = _runtime_parameter_schema( + implementation + ) + unexpected_parameters = sorted(set(parameters) - allowed_parameters) + if unexpected_parameters: + raise ImportIdentityError( + "Lineage parameters contain undeclared field(s) " + f"{unexpected_parameters}." + ) + missing_parameters = sorted(required_parameters - set(parameters)) + if missing_parameters: + raise ImportIdentityError( + "Lineage parameters are missing required field(s) " + f"{missing_parameters}." + ) + elif implementation_mode == "declared_external": + try: + validated_implementation = validate_implementation_identity_record( + implementation + ) + except ImportSpecError as exc: + raise ImportIdentityError( + f"Invalid declared external lineage implementation: {exc}" + ) from exc + source_digest = validated_implementation.get("source_digest") + if source_digest is None or source_digest not in source_digests: + raise ImportIdentityError( + "Declared external lineage implementation source digest must be " + "present in source_digests." + ) + if parameters: + raise ImportIdentityError( + "Declared external lineage implementations cannot self-declare " + "unverified parameters." + ) + else: + raise ImportIdentityError( + f"Invalid lineage implementation_mode {implementation_mode!r}." + ) + implementation_record = { + "identity_mode": implementation_mode, + "identity": validated_implementation, + } + _validate_stable_values(implementation_record, "Lineage implementation") + _validate_stable_parameters(parameters, "Lineage parameters") + payload = canonicalize( + { + "schema_version": "choicebench.lineage-component.v1", + "operation_type": operation_type, + "question_id": question_id, + "parent_digests": sorted(parent_digests), + "source_digests": sorted(source_digests), + "authorization_digest": authorization_digest, + "implementation": implementation_record, + "parameters": parameters, + "input_digest": input_digest, + "preownership_output_digest": preownership_output_digest, + "prediction_origin": prediction_origin, + } + ) + return { + "lineage_id": short_id("lin", payload), + "lineage_digest": integrity_digest(payload), + "identity": payload, + } + + +def make_result_origin( + *, + derivation_origin: Literal[ + "native_execution", "external_import", "repair_overlay", "offline_transformation" + ], + row_assignments: Sequence[tuple[str, str, str]], +) -> dict[str, Any]: + """Record constituent and ordered per-row prediction origins.""" + if derivation_origin not in _DERIVATION_ORIGINS: + raise ImportIdentityError(f"Invalid derivation origin {derivation_origin!r}.") + rows: list[dict[str, str]] = [] + seen: set[str] = set() + counts: dict[str, int] = {} + for index, assignment in enumerate(row_assignments): + if not isinstance(assignment, (list, tuple)) or len(assignment) != 3: + raise ImportIdentityError( + f"Result origin row assignment {index} must contain exactly " + "question_id, prediction_origin, and prediction_lineage_id." + ) + question_id, prediction_origin, lineage_id = assignment + if not isinstance(question_id, str) or not question_id or question_id in seen: + raise ImportIdentityError("Result origin question IDs must be non-empty and unique.") + if prediction_origin not in _PREDICTION_ORIGINS: + raise ImportIdentityError(f"Invalid prediction origin {prediction_origin!r}.") + if not isinstance(lineage_id, str) or not _LINEAGE_ID_RE.fullmatch(lineage_id): + raise ImportIdentityError( + "Result origin prediction_lineage_id must be a lineage-component ID." + ) + seen.add(question_id) + counts[prediction_origin] = counts.get(prediction_origin, 0) + 1 + rows.append( + { + "question_id": question_id, + "prediction_origin": prediction_origin, + "prediction_lineage_id": lineage_id, + } + ) + rows = canonicalize(rows) + lineage_projection = [ + { + "prediction_origin": row["prediction_origin"], + "prediction_lineage_id": row["prediction_lineage_id"], + } + for row in rows + ] + identity = { + "derivation_origin": derivation_origin, + "prediction_origins": sorted(counts), + "prediction_origin_counts": {key: counts[key] for key in sorted(counts)}, + "row_assignments": rows, + "ordered_row_origin_digest": integrity_digest(rows), + "lineage_component_origin_digest": integrity_digest(lineage_projection), + } + return { + "origin_id": short_id("origin", identity), + "origin_digest": integrity_digest(identity), + **identity, + } + + +def _validate_lineage_component_record( + value: Any, *, runtime_callable: Callable[..., Any] | None +) -> dict[str, Any]: + record = _exact_mapping( + value, + {"lineage_id", "lineage_digest", "identity"}, + "Realization lineage component", + ) + identity = _exact_mapping( + record["identity"], + { + "schema_version", + "operation_type", + "question_id", + "parent_digests", + "source_digests", + "authorization_digest", + "implementation", + "parameters", + "input_digest", + "preownership_output_digest", + "prediction_origin", + }, + "Realization lineage component identity", + ) + if identity["schema_version"] != "choicebench.lineage-component.v1": + raise ImportIdentityError("Realization lineage component schema is invalid.") + for field in ("operation_type", "question_id"): + if not isinstance(identity[field], str) or not identity[field]: + raise ImportIdentityError( + f"Realization lineage component {field} must be non-empty." + ) + normalized_parents = _digest_sequence( + identity["parent_digests"], "lineage.parent_digests" + ) + normalized_sources = _digest_sequence( + identity["source_digests"], "lineage.source_digests" + ) + authorization_digest = _validate_digest( + identity["authorization_digest"], + "lineage.authorization_digest", + optional=True, + ) + implementation = _exact_mapping( + identity["implementation"], + {"identity_mode", "identity"}, + "Realization lineage implementation", + ) + if implementation["identity_mode"] not in { + "runtime_callable", + "declared_external", + }: + raise ImportIdentityError( + "Realization lineage implementation identity mode is invalid." + ) + try: + implementation_identity_record = validate_implementation_identity_record( + implementation["identity"] + ) + except ImportSpecError as exc: + raise ImportIdentityError( + f"Realization lineage implementation is invalid: {exc}" + ) from exc + source_digest = implementation_identity_record.get("source_digest") + if source_digest is None: + raise ImportIdentityError( + "Realization lineage implementation lacks a source digest." + ) + if implementation["identity_mode"] == "runtime_callable": + if runtime_callable is None: + raise ImportIdentityError( + "Realization runtime lineage requires its registered callable." + ) + runtime_identity = _runtime_callable_record( + runtime_callable, "Realization lineage runtime implementation" + ) + if canonicalize(runtime_identity) != canonicalize( + implementation_identity_record + ): + raise ImportIdentityError( + "Realization runtime lineage implementation does not match its " + "registered callable." + ) + elif runtime_callable is not None: + raise ImportIdentityError( + "Declared external lineage cannot claim a runtime callable." + ) + if ( + implementation["identity_mode"] == "declared_external" + and source_digest not in normalized_sources + ): + raise ImportIdentityError( + "Declared external lineage implementation is not owned by its source edge." + ) + parameters = identity["parameters"] + if not isinstance(parameters, Mapping): + raise ImportIdentityError("Realization lineage parameters must be a mapping.") + _validate_stable_parameters(parameters, "Realization lineage parameters") + input_digest = _validate_digest(identity["input_digest"], "lineage.input_digest") + output_digest = _validate_digest( + identity["preownership_output_digest"], + "lineage.preownership_output_digest", + ) + prediction_origin = identity["prediction_origin"] + if prediction_origin not in _PREDICTION_ORIGINS: + raise ImportIdentityError("Realization lineage prediction origin is invalid.") + normalized_identity = canonicalize( + { + "schema_version": "choicebench.lineage-component.v1", + "operation_type": identity["operation_type"], + "question_id": identity["question_id"], + "parent_digests": normalized_parents, + "source_digests": normalized_sources, + "authorization_digest": authorization_digest, + "implementation": { + "identity_mode": implementation["identity_mode"], + "identity": implementation_identity_record, + }, + "parameters": parameters, + "input_digest": input_digest, + "preownership_output_digest": output_digest, + "prediction_origin": prediction_origin, + } + ) + expected_digest = integrity_digest(normalized_identity) + expected_id = short_id("lin", normalized_identity) + if ( + record["lineage_digest"] != expected_digest + or record["lineage_id"] != expected_id + or canonicalize(record["identity"]) != normalized_identity + ): + raise ImportIdentityError( + "Realization lineage component identity or digest is inconsistent." + ) + return { + "lineage_id": expected_id, + "lineage_digest": expected_digest, + "identity": normalized_identity, + } + + +def importer_implementation_identity( + *, adapter: Callable[..., Any], validator: Callable[..., Any] +) -> dict[str, Any]: + """Bind installed ChoiceBench plus importer-core, adapter, and validator code.""" + records: dict[str, Mapping[str, Any]] = {} + for role, target in (("adapter", adapter), ("validator", validator)): + records[role] = _runtime_callable_record( + target, f"Importer {role} implementation" + ) + importer_record = _runtime_callable_record( + _identity_record, "Importer core implementation" + ) + return canonicalize( + { + "schema_version": "choicebench.importer-implementation.v1", + "package": {"name": "choicebench", "version": __version__}, + "importer": importer_record, + "adapter": records["adapter"], + "validator": records["validator"], + } + ) + + +def _validate_result_origin_record(value: Any) -> dict[str, Any]: + record = _exact_mapping( + value, + { + "origin_id", + "origin_digest", + "derivation_origin", + "prediction_origins", + "prediction_origin_counts", + "row_assignments", + "ordered_row_origin_digest", + "lineage_component_origin_digest", + }, + "Realization result_origin", + ) + raw_rows = record["row_assignments"] + if not isinstance(raw_rows, (list, tuple)): + raise ImportIdentityError("Realization result_origin row_assignments must be a list.") + assignments = [] + for index, raw_row in enumerate(raw_rows): + row = _exact_mapping( + raw_row, + {"question_id", "prediction_origin", "prediction_lineage_id"}, + f"Realization result_origin row_assignments[{index}]", + ) + assignments.append( + ( + row["question_id"], + row["prediction_origin"], + row["prediction_lineage_id"], + ) + ) + rebuilt = make_result_origin( + derivation_origin=record["derivation_origin"], + row_assignments=assignments, + ) + if canonicalize(record) != canonicalize(rebuilt): + raise ImportIdentityError("Realization result_origin is internally inconsistent.") + return rebuilt + + +def _validate_realization_identity( + value: Mapping[str, Any], + *, + runtime_importer: Mapping[str, Any], + expected_dataset: ExpectedDataset, + lineage_runtime_callables: Mapping[str, Callable[..., Any]], +) -> dict[str, Any]: + raw = _exact_mapping(value, _REALIZATION_KEYS, "Realization identity") + _validate_digest(raw["import_spec_digest"], "import_spec_digest") + + sources = raw["sources"] + if not isinstance(sources, (list, tuple)) or not sources: + raise ImportIdentityError("Realization sources must be a non-empty list.") + normalized_sources = [] + seen_source_ids: set[str] = set() + for index, raw_source in enumerate(sources): + source = _exact_mapping(raw_source, _SOURCE_KEYS, f"Realization sources[{index}]") + source_id = source["source_id"] + logical_path = source["logical_path"] + if ( + not isinstance(source_id, str) + or not source_id + or source_id in seen_source_ids + ): + raise ImportIdentityError("Realization source IDs must be non-empty and unique.") + if not isinstance(logical_path, str): + raise ImportIdentityError("Realization source logical_path must be a string.") + pure_path = PurePosixPath(logical_path) + if ( + pure_path.is_absolute() + or ".." in pure_path.parts + or not pure_path.parts + or _is_machine_path(logical_path) + ): + raise ImportIdentityError("Realization source logical_path is unsafe.") + if source["classification"] not in { + "raw", + "canonical", + "derived", + "repaired", + "aggregate_only", + }: + raise ImportIdentityError("Realization source classification is invalid.") + if source["format"] != "csv": + raise ImportIdentityError("Realization source format is unsupported.") + if not isinstance(source["format_version"], str) or not source["format_version"]: + raise ImportIdentityError("Realization source format_version is invalid.") + _validate_digest(source["sha256"], f"sources[{index}].sha256") + provenance = _exact_mapping( + source["provenance"], + _SOURCE_PROVENANCE_KEYS, + f"Realization sources[{index}].provenance", + ) + for field in ("source_run_id", "source_repository", "source_commit"): + value = provenance[field] + if value is not None and (not isinstance(value, str) or not value.strip()): + raise ImportIdentityError( + f"Realization source provenance {field} must be null or non-empty." + ) + if isinstance(value, str) and _is_machine_path(value): + raise ImportIdentityError( + f"Realization source provenance {field} contains a machine-local path." + ) + for field in ("notes_digest", "evidence_digest"): + _validate_digest( + provenance[field], f"sources[{index}].provenance.{field}", optional=True + ) + _validate_unknown_reason_contract( + {field: provenance[field] for field in _SOURCE_PROVENANCE_UNKNOWN_FIELDS}, + provenance["unknown_reasons"], + f"Realization sources[{index}].provenance", + ) + source = {**source, "provenance": canonicalize(provenance)} + seen_source_ids.add(source_id) + normalized_sources.append(canonicalize(source)) + + expected = _exact_mapping( + raw["expected_dataset"], + { + "snapshot_digest", + "question_set_digest", + "derivation_digest", + "question_ids", + }, + "Realization expected_dataset", + ) + for field in ("snapshot_digest", "question_set_digest", "derivation_digest"): + digest = expected[field] + _validate_digest(digest, f"expected_dataset.{field}") + expected_question_ids = expected["question_ids"] + if not isinstance(expected_question_ids, (list, tuple)) or not all( + isinstance(question_id, str) and question_id + for question_id in expected_question_ids + ): + raise ImportIdentityError( + "Realization expected dataset question_ids must be a non-empty string list." + ) + if not expected_question_ids or len(expected_question_ids) != len( + set(expected_question_ids) + ): + raise ImportIdentityError( + "Realization expected dataset question_ids must be non-empty and unique." + ) + if integrity_digest(list(expected_question_ids)) != expected["question_set_digest"]: + raise ImportIdentityError( + "Realization expected dataset question_ids do not match question_set_digest." + ) + _validate_expected_dataset_contract(expected_dataset) + authoritative_expected = canonicalize( + { + "snapshot_digest": expected_dataset.snapshot_digest, + "question_set_digest": expected_dataset.question_set_digest, + "derivation_digest": expected_dataset.derivation_digest, + "question_ids": list(expected_dataset.selected_question_ids), + } + ) + if canonicalize(expected) != authoritative_expected: + raise ImportIdentityError( + "Realization expected dataset does not match the validated dataset snapshot." + ) + + raw_importer = raw["importer_implementation"] + if canonicalize(raw_importer) != canonicalize(runtime_importer): + raise ImportIdentityError( + "Realization importer implementation does not match the supplied runtime " + "adapter and validator callables." + ) + importer = _exact_mapping( + raw_importer, + {"schema_version", "package", "importer", "adapter", "validator"}, + "Realization importer_implementation", + ) + if importer["schema_version"] != "choicebench.importer-implementation.v1": + raise ImportIdentityError("Realization importer implementation schema is invalid.") + package = _exact_mapping( + importer["package"], {"name", "version"}, "Realization importer package" + ) + if package != {"name": "choicebench", "version": __version__}: + raise ImportIdentityError("Realization importer package identity is invalid.") + for role in ("importer", "adapter", "validator"): + try: + component = validate_implementation_identity_record(importer[role]) + except ImportSpecError as exc: + raise ImportIdentityError( + f"Realization importer {role} identity is invalid: {exc}" + ) from exc + if not isinstance(component.get("source_digest"), str): + raise ImportIdentityError( + f"Realization importer {role} lacks a source identity." + ) + _validate_digest(component["source_digest"], f"importer.{role}.source_digest") + + parsing = _exact_mapping( + raw["parsing_policy"], _PARSING_POLICY_KEYS, "Realization parsing_policy" + ) + _exact_mapping( + parsing["dialect"], _DIALECT_KEYS, "Realization parsing_policy.dialect" + ) + try: + dialect = validate_csv_dialect_identity(parsing["dialect"]) + except ImportSpecError as exc: + raise ImportIdentityError(f"Invalid realization parsing dialect: {exc}") from exc + mapping = parsing["mapping"] + if not isinstance(mapping, Mapping): + raise ImportIdentityError("Realization parsing mapping must be a mapping.") + normalized_mapping: dict[str, str] = {} + for semantic_field, source_column in mapping.items(): + if not isinstance(semantic_field, str) or not semantic_field.strip(): + raise ImportIdentityError( + "Realization parsing mapping keys must be non-empty strings." + ) + if not isinstance(source_column, str) or not source_column.strip(): + raise ImportIdentityError( + "Realization parsing mapping values must be non-empty source columns." + ) + normalized_mapping[semantic_field] = source_column + if len(normalized_mapping.values()) != len(set(normalized_mapping.values())): + raise ImportIdentityError("Realization parsing mapping reuses a source column.") + null_values = parsing["null_values"] + if not isinstance(null_values, (list, tuple)) or not all( + isinstance(item, str) for item in null_values + ): + raise ImportIdentityError( + "Realization parsing null_values must be a string list." + ) + if len(null_values) != len(set(null_values)): + raise ImportIdentityError("Realization parsing null_values contains duplicates.") + try: + numeric_columns = validate_numeric_columns_identity( + parsing["numeric_columns"] + ) + option_mapping = validate_option_mapping_identity(parsing["option_mapping"]) + except ImportSpecError as exc: + raise ImportIdentityError(f"Invalid realization parsing policy: {exc}") from exc + if parsing["extra_field_policy"] not in { + "preserve_unmapped", + "reject_unmapped", + }: + raise ImportIdentityError("Realization extra_field_policy is invalid.") + + validation = _exact_mapping( + raw["validation"], {"findings_digest"}, "Realization validation" + ) + _validate_digest(validation["findings_digest"], "validation.findings_digest") + evidence = _exact_mapping( + raw["evidence"], + { + "evidence_status", + "qualification_digest", + "limitation_digest", + "defect_digest", + "scope_disposition", + }, + "Realization evidence", + ) + if evidence["evidence_status"] not in _EVIDENCE_STATUSES: + raise ImportIdentityError("Realization evidence_status is invalid.") + if evidence["scope_disposition"] not in _SCOPE_DISPOSITIONS: + raise ImportIdentityError("Realization scope_disposition is invalid.") + for field in ("qualification_digest", "limitation_digest", "defect_digest"): + _validate_digest(evidence[field], f"evidence.{field}") + + raw_lineage_components = raw["lineage_components"] + if not isinstance(raw_lineage_components, (list, tuple)): + raise ImportIdentityError( + "Realization lineage_components must be a list." + ) + if not isinstance(lineage_runtime_callables, Mapping) or not all( + isinstance(lineage_id, str) and callable(target) + for lineage_id, target in lineage_runtime_callables.items() + ): + raise ImportIdentityError( + "Realization lineage runtime callable registry is invalid." + ) + lineage_components = [] + for component in raw_lineage_components: + claimed_lineage_id = ( + component.get("lineage_id") if isinstance(component, Mapping) else None + ) + lineage_components.append( + _validate_lineage_component_record( + component, + runtime_callable=lineage_runtime_callables.get(claimed_lineage_id), + ) + ) + lineage_by_id: dict[str, dict[str, Any]] = {} + lineage_digests: set[str] = set() + for component in lineage_components: + lineage_id = component["lineage_id"] + lineage_digest = component["lineage_digest"] + if lineage_id in lineage_by_id or lineage_digest in lineage_digests: + raise ImportIdentityError( + "Realization lineage components contain duplicate identities." + ) + lineage_by_id[lineage_id] = component + lineage_digests.add(lineage_digest) + runtime_lineage_ids = { + component["lineage_id"] + for component in lineage_components + if component["identity"]["implementation"]["identity_mode"] + == "runtime_callable" + } + if set(lineage_runtime_callables) != runtime_lineage_ids: + raise ImportIdentityError( + "Realization lineage runtime callable registry does not exactly own " + "the runtime lineage components." + ) + + result_origin = _validate_result_origin_record(raw["result_origin"]) + origin_question_ids = [ + row["question_id"] for row in result_origin["row_assignments"] + ] + evidence_status = evidence["evidence_status"] + origin_question_id_set = set(origin_question_ids) + unexpected_origin_ids = sorted( + origin_question_id_set - set(expected_question_ids) + ) + ordered_subset = [ + question_id + for question_id in expected_question_ids + if question_id in origin_question_id_set + ] + if unexpected_origin_ids or ordered_subset != origin_question_ids: + raise ImportIdentityError( + "Realization result origin question IDs are not an ordered subset of " + "the expected dataset question set." + ) + if ( + evidence_status in {"complete", "qualified"} + and evidence["scope_disposition"] == "included" + and origin_question_ids != list(expected_question_ids) + ): + raise ImportIdentityError( + "Complete or qualified realization result origin question IDs do not " + "match the full expected dataset question set." + ) + for assignment in result_origin["row_assignments"]: + lineage = lineage_by_id.get(assignment["prediction_lineage_id"]) + if lineage is None: + raise ImportIdentityError( + "Realization result origin references an unowned lineage component." + ) + lineage_identity = lineage["identity"] + if ( + lineage_identity["question_id"] != assignment["question_id"] + or lineage_identity["prediction_origin"] + != assignment["prediction_origin"] + ): + raise ImportIdentityError( + "Realization result origin conflicts with its lineage component." + ) + assigned_lineage_ids = { + assignment["prediction_lineage_id"] + for assignment in result_origin["row_assignments"] + } + unreachable_lineage_ids = sorted(set(lineage_by_id) - assigned_lineage_ids) + if unreachable_lineage_ids: + raise ImportIdentityError( + f"Realization lineage component(s) {unreachable_lineage_ids} are not " + "reachable from any result-origin row." + ) + parents = _exact_mapping( + raw["parent_digests"], + {"realization_digests", "evidence_digests", "result_digests"}, + "Realization parent_digests", + ) + normalized_parents = { + field: _digest_sequence(items, f"parent_digests.{field}") + for field, items in parents.items() + } + authorization_digest = _validate_digest( + raw["authorization_digest"], "authorization_digest", optional=True + ) + lineage_authorizations = { + component["identity"]["authorization_digest"] + for component in lineage_components + if component["identity"]["authorization_digest"] is not None + } + if any(digest != authorization_digest for digest in lineage_authorizations): + raise ImportIdentityError( + "Realization lineage authorization is not owned by the realization." + ) + declared_parent_edges = { + digest + for digests in normalized_parents.values() + for digest in digests + } + lineage_parent_edges = { + digest + for component in lineage_components + for digest in component["identity"]["parent_digests"] + } + if not lineage_parent_edges <= declared_parent_edges: + raise ImportIdentityError( + "Realization lineage parent edge is not owned by the realization." + ) + overlay = raw["overlay"] + normalized_overlay = None + if overlay is not None: + overlay = _exact_mapping( + overlay, + { + "source_sha256", + "replacement_digest", + "transformation_input_digest", + "preownership_output_digest", + "implementation_digest", + }, + "Realization overlay", + ) + normalized_overlay = {} + for field, digest in overlay.items(): + normalized_overlay[field] = _validate_digest( + digest, + f"overlay.{field}", + optional=(field not in {"source_sha256", "replacement_digest"}), + ) + declared_source_digests = {source["sha256"] for source in normalized_sources} + if normalized_overlay is not None: + declared_source_digests.add(normalized_overlay["source_sha256"]) + lineage_source_edges = { + digest + for component in lineage_components + for digest in component["identity"]["source_digests"] + } + if not lineage_source_edges <= declared_source_digests: + raise ImportIdentityError( + "Realization lineage source edge is not owned by the realization." + ) + unreachable_source_digests = ( + sorted(declared_source_digests - lineage_source_edges) + if lineage_components + else [] + ) + if unreachable_source_digests: + raise ImportIdentityError( + f"Realization source(s) {unreachable_source_digests} are not reachable " + "from any lineage component." + ) + + derivation_origin = result_origin["derivation_origin"] + derived = derivation_origin in {"repair_overlay", "offline_transformation"} + if derived: + if authorization_digest is None: + raise ImportIdentityError( + f"{derivation_origin} realization requires an authorization digest." + ) + if not normalized_parents["realization_digests"]: + raise ImportIdentityError( + f"{derivation_origin} realization requires a parent realization digest." + ) + if not normalized_parents["evidence_digests"]: + raise ImportIdentityError( + f"{derivation_origin} realization requires a parent evidence digest." + ) + if normalized_overlay is None: + raise ImportIdentityError( + f"{derivation_origin} realization requires an overlay identity." + ) + if authorization_digest not in lineage_authorizations: + raise ImportIdentityError( + f"{derivation_origin} realization requires an authorized lineage " + "component." + ) + if not lineage_parent_edges: + raise ImportIdentityError( + f"{derivation_origin} realization requires a parent-derived lineage " + "component." + ) + derived_components = [ + component + for component in lineage_components + if component["identity"]["operation_type"] == derivation_origin + ] + if not derived_components: + raise ImportIdentityError( + f"{derivation_origin} realization requires a matching derived " + "lineage component." + ) + unassigned_derived_ids = sorted( + component["lineage_id"] + for component in derived_components + if component["lineage_id"] not in assigned_lineage_ids + ) + if unassigned_derived_ids: + raise ImportIdentityError( + f"{derivation_origin} lineage component(s) {unassigned_derived_ids} " + "do not own a result-origin row assignment." + ) + for component in derived_components: + component_identity = component["identity"] + if component_identity["authorization_digest"] != authorization_digest: + raise ImportIdentityError( + f"{derivation_origin} lineage component lacks the realization " + "authorization." + ) + if not component_identity["parent_digests"]: + raise ImportIdentityError( + f"{derivation_origin} lineage component lacks a parent edge." + ) + if not ( + set(component_identity["parent_digests"]) + & set(normalized_parents["realization_digests"]) + ): + raise ImportIdentityError( + f"{derivation_origin} lineage component lacks the declared base " + "realization edge." + ) + if normalized_overlay["source_sha256"] not in component_identity[ + "source_digests" + ]: + raise ImportIdentityError( + f"{derivation_origin} lineage component lacks the overlay source " + "edge." + ) + if derivation_origin == "repair_overlay": + for assignment in result_origin["row_assignments"]: + if assignment["prediction_origin"] in { + "native_inference", + "external_repair_inference", + }: + component = lineage_by_id[assignment["prediction_lineage_id"]] + if component["identity"]["operation_type"] != "repair_overlay": + raise ImportIdentityError( + "Repair prediction row lacks repair_overlay lineage." + ) + if derivation_origin == "offline_transformation" and any( + normalized_overlay[field] is None + for field in ( + "transformation_input_digest", + "preownership_output_digest", + "implementation_digest", + ) + ): + raise ImportIdentityError( + "offline_transformation overlay requires input, output, and " + "implementation digests." + ) + if derivation_origin == "offline_transformation": + expected_input_digest = integrity_digest( + [ + component["identity"]["input_digest"] + for component in derived_components + ] + ) + expected_output_digest = integrity_digest( + [ + component["identity"]["preownership_output_digest"] + for component in derived_components + ] + ) + implementation_identities = { + integrity_digest(component["identity"]["implementation"]): component[ + "identity" + ]["implementation"] + for component in derived_components + } + if len(implementation_identities) != 1: + raise ImportIdentityError( + "offline_transformation lineage components use conflicting " + "implementations." + ) + expected_implementation_digest = next(iter(implementation_identities)) + if ( + normalized_overlay["transformation_input_digest"] + != expected_input_digest + or normalized_overlay["preownership_output_digest"] + != expected_output_digest + or normalized_overlay["implementation_digest"] + != expected_implementation_digest + ): + raise ImportIdentityError( + "offline_transformation overlay input, output, or implementation " + "digest conflicts with its lineage components." + ) + if derivation_origin == "repair_overlay" and not ( + {"native_inference", "external_repair_inference"} + & set(result_origin["prediction_origins"]) + ): + raise ImportIdentityError( + "repair_overlay requires at least one repair prediction origin." + ) + else: + if authorization_digest is not None or normalized_overlay is not None: + raise ImportIdentityError( + f"{derivation_origin} realization cannot claim repair authorization " + "or overlay." + ) + incompatible_operations = sorted( + { + component["identity"]["operation_type"] + for component in lineage_components + if component["identity"]["operation_type"] != derivation_origin + } + ) + if incompatible_operations: + raise ImportIdentityError( + f"{derivation_origin} realization lineage components have " + f"incompatible operation type(s) {incompatible_operations}." + ) + + normalized = { + "import_spec_digest": raw["import_spec_digest"], + "sources": normalized_sources, + "expected_dataset": canonicalize(expected), + "importer_implementation": canonicalize(importer), + "parsing_policy": canonicalize( + { + "dialect": dialect, + "mapping": normalized_mapping, + "null_values": list(null_values), + "numeric_columns": numeric_columns, + "option_mapping": option_mapping, + "extra_field_policy": parsing["extra_field_policy"], + } + ), + "validation": canonicalize(validation), + "evidence": canonicalize(evidence), + "lineage_components": canonicalize(lineage_components), + "result_origin": result_origin, + "parent_digests": normalized_parents, + "authorization_digest": authorization_digest, + "overlay": normalized_overlay, + } + _reject_forbidden( + normalized, + "Realization identity", + allowed_path_keys=frozenset({"logical_path"}), + ) + return canonicalize(normalized) + + +def make_realization( + *, + condition_id: str, + condition_digest: str, + identity: Mapping[str, Any], + fields: Mapping[str, Any], + adapter: Callable[..., Any], + validator: Callable[..., Any], + expected_dataset: ExpectedDataset, + semantic_condition: Mapping[str, Any], + lineage_runtime_callables: Mapping[str, Callable[..., Any]], +) -> dict[str, Any]: + """Create an immutable realization identity distinct from its condition.""" + if not isinstance(condition_id, str) or not re.fullmatch(r"cond_[0-9a-f]{16}", condition_id): + raise ImportIdentityError("Realization condition_id is invalid.") + _validate_digest(condition_digest, "condition_digest") + if condition_id != f"cond_{condition_digest[:16]}": + raise ImportIdentityError( + "Realization condition_id does not match the condition_digest." + ) + if not isinstance(identity, Mapping) or not identity: + raise ImportIdentityError("Realization identity must be a non-empty mapping.") + condition_record = _exact_mapping( + semantic_condition, + { + "condition_id", + "condition_digest", + "identity", + "condition_key", + "expected_question_ids", + }, + "Realization semantic condition", + ) + rebuilt_condition = make_semantic_condition( + identity=condition_record["identity"], + fields={ + "condition_key": condition_record["condition_key"], + "expected_question_ids": condition_record["expected_question_ids"], + }, + ) + if canonicalize(rebuilt_condition) != canonicalize(condition_record): + raise ImportIdentityError( + "Realization semantic condition identity is inconsistent." + ) + if ( + condition_record["condition_id"] != condition_id + or condition_record["condition_digest"] != condition_digest + ): + raise ImportIdentityError( + "Realization condition ID/digest do not match its semantic condition." + ) + benchmark = condition_record["identity"]["benchmark"] + if ( + benchmark["name"] != expected_dataset.benchmark_name + or benchmark["split"] != expected_dataset.split + or benchmark["artifact_id"] != expected_dataset.artifact_id + or benchmark["selection_id"] != expected_dataset.selection_id + or tuple(condition_record["expected_question_ids"]) + != expected_dataset.selected_question_ids + ): + raise ImportIdentityError( + "Realization semantic condition is not owned by its expected dataset " + "artifact and selection." + ) + runtime_importer = importer_implementation_identity( + adapter=adapter, validator=validator + ) + validated_identity = _validate_realization_identity( + identity, + runtime_importer=runtime_importer, + expected_dataset=expected_dataset, + lineage_runtime_callables=lineage_runtime_callables, + ) + payload = canonicalize( + { + "schema_version": "choicebench.realization.v1", + "condition_id": condition_id, + "condition_digest": condition_digest, + "realization": validated_identity, + } + ) + if not isinstance(fields, Mapping) or not set(fields) <= {"audit", "realization_key"}: + raise ImportIdentityError( + "Realization fields may contain only audit and realization_key." + ) + if "audit" in fields and not isinstance(fields["audit"], Mapping): + raise ImportIdentityError("Realization audit field must be a mapping.") + extra = canonicalize(fields) + return { + "realization_id": short_id("real", payload), + "realization_digest": integrity_digest(payload), + "condition_id": condition_id, + "condition_digest": condition_digest, + "identity": payload, + **extra, + } diff --git a/src/choicebench/importing/overlays.py b/src/choicebench/importing/overlays.py new file mode 100644 index 0000000..eacd484 --- /dev/null +++ b/src/choicebench/importing/overlays.py @@ -0,0 +1,266 @@ +"""Pure derivation of a repair/offline-transformation overlay from an +already-verified base realization. + +Scope boundary: this module never opens a filesystem path, runs inference, or +merges the derived rows into a full realization -- it only validates the +overlay's binding to its base/authorization and produces the lineage/row +pieces for the *replaced* questions. Merging those with the base's retained +rows into a complete result set and constructing the final realization is +the import engine's job (it alone has the full base row set, the base run +directory, and the destination for the new run). +""" + +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any, Mapping + +from choicebench.identity import integrity_digest +from choicebench.importing.authorization import ValidatedAuthorization +from choicebench.importing.csv_adapter import AdaptedTable +from choicebench.importing.dataset_reference import ExpectedDataset +from choicebench.importing.identity import make_lineage_component, make_result_origin +from choicebench.importing.schema import OverlaySpec + +_EVIDENCE_STATUSES = { + "complete", "qualified", "partial", "malformed", "recoverable", "failed", +} +_REPAIR_ORIGINS = {"native_inference", "external_repair_inference"} + + +class OverlayError(ValueError): + """Raised when an overlay declaration or its derivation is invalid or unsafe.""" + + +@dataclass(frozen=True) +class VerifiedBaseRealization: + condition_id: str + condition_digest: str + realization_id: str + realization_digest: str + evidence_index_digest: str + evidence_source_digests: Mapping[str, str] + validation_artifact_sha256: str + result_sha256: str | None + rows_by_question_id: Mapping[str, Mapping[str, Any]] + prediction_origins: Mapping[str, str] + + +@dataclass(frozen=True) +class DerivedRealizationPayload: + condition_id: str + condition_digest: str + preownership_rows: tuple[Mapping[str, Any], ...] + lineage_components: tuple[Mapping[str, Any], ...] + result_origin: Mapping[str, Any] + evidence_status: str + replacement_question_ids: tuple[str, ...] + + +def _rows_by_question_id(table: AdaptedTable, mapping: Mapping[str, str]) -> dict[str, Any]: + question_id_column = mapping["question_id"] + rows_by_qid: dict[str, list] = {} + for row in table.rows: + qid = row.values.get(question_id_column) + if qid is None: + continue + rows_by_qid.setdefault(str(qid), []).append(row) + return rows_by_qid + + +def derive_overlay( + *, + base: VerifiedBaseRealization, + overlay: OverlaySpec, + authorization: ValidatedAuthorization, + overlay_table: AdaptedTable, + overlay_mapping: Mapping[str, str], + expected: ExpectedDataset, +) -> DerivedRealizationPayload: + """Derive the replaced-row lineage/content for one authorized overlay. + + `base` must be a value the caller obtained by independently verifying the + full base run graph (never a bare path or unverified dict) -- this + function performs no filesystem or verification work of its own, only + binding checks against the values it is given. + """ + if authorization.condition_digest != base.condition_digest: + raise OverlayError( + "Authorization condition does not match the base realization's " + "condition; an authorization cannot be borrowed across conditions." + ) + if ( + overlay.base_condition_digest != base.condition_digest + or overlay.base_realization_id != base.realization_id + or overlay.base_realization_digest != base.realization_digest + ): + raise OverlayError( + "Overlay base condition/realization binding does not match the " + "verified base realization." + ) + if overlay.base_validation_artifact_sha256 != base.validation_artifact_sha256: + raise OverlayError("Overlay base validation-artifact checksum does not match.") + if overlay.base_result_sha256 != base.result_sha256: + raise OverlayError("Overlay base result checksum does not match.") + if overlay.authorization_id != authorization.bundle_id: + raise OverlayError( + "Overlay declares a different authorization bundle than the one supplied." + ) + + if set(overlay.base_evidence_digests) != set(authorization.input_evidence_digests): + raise OverlayError( + "Overlay base_evidence_digests do not exactly match the authorization's " + "input evidence digests." + ) + for key, digest in authorization.input_evidence_digests.items(): + if ( + overlay.base_evidence_digests.get(key) != digest + or base.evidence_source_digests.get(key) != digest + ): + raise OverlayError( + f"Overlay evidence digest {key!r} is not owned by the base realization's " + "own evidence sources." + ) + + replacement_ids = set(overlay.replacement_reasons) + if not replacement_ids: + raise OverlayError("Overlay declares no replacement question IDs.") + unauthorized = replacement_ids - set(authorization.question_reasons) + if unauthorized: + raise OverlayError( + f"Overlay replaces question(s) {sorted(unauthorized)} that are not granted " + "by its authorization." + ) + unexpected = replacement_ids - set(expected.selected_question_ids) + if unexpected: + raise OverlayError( + f"Overlay replaces question(s) {sorted(unexpected)} outside the expected " + "dataset selection." + ) + + per_question_origin = overlay.result_origin.per_question_prediction_origins + default_origin = overlay.result_origin.default_prediction_origin + + if authorization.authorization_type == "inference_repair": + operation_type = "repair_overlay" + if overlay.result_origin.derivation_origin != operation_type: + raise OverlayError( + f"inference_repair overlay must declare derivation_origin=" + f"{operation_type!r}, found {overlay.result_origin.derivation_origin!r}." + ) + extra_origins = set(per_question_origin) - replacement_ids + if extra_origins: + raise OverlayError( + f"inference_repair origin assignment(s) for unreplaced question(s) " + f"{sorted(extra_origins)} are not permitted." + ) + resolved_origins: dict[str, str] = {} + for qid in replacement_ids: + origin = per_question_origin.get(qid, default_origin) + if origin is None: + raise OverlayError( + f"inference_repair overlay has no declared origin assignment for " + f"replaced question {qid!r}." + ) + if origin not in _REPAIR_ORIGINS: + raise OverlayError( + f"inference_repair origin {origin!r} for question {qid!r} is not a " + f"valid repair prediction origin {sorted(_REPAIR_ORIGINS)}." + ) + resolved_origins[qid] = origin + elif authorization.authorization_type == "offline_transformation": + operation_type = "offline_transformation" + if overlay.result_origin.derivation_origin != operation_type: + raise OverlayError( + f"offline_transformation overlay must declare derivation_origin=" + f"{operation_type!r}, found {overlay.result_origin.derivation_origin!r}." + ) + if default_origin is not None: + raise OverlayError( + "offline_transformation cannot declare a default prediction origin; it " + "retains each underlying response's own origin." + ) + resolved_origins = {} + for qid in replacement_ids: + origin = per_question_origin.get(qid) or base.prediction_origins.get(qid) + if origin is None: + raise OverlayError( + f"offline_transformation has no underlying prediction origin for " + f"question {qid!r}." + ) + resolved_origins[qid] = origin + else: + raise OverlayError( + f"Unsupported authorization_type {authorization.authorization_type!r}." + ) + + overlay_rows_by_qid = _rows_by_question_id(overlay_table, overlay_mapping) + prediction_column = overlay_mapping.get("prediction") + preownership_rows: list[dict[str, Any]] = [] + lineage_components: list[Mapping[str, Any]] = [] + for qid in sorted(replacement_ids): + rows = overlay_rows_by_qid.get(qid, []) + if len(rows) != 1: + raise OverlayError( + f"Overlay source has {len(rows)} row(s) for replaced question " + f"{qid!r}; exactly one is required." + ) + overlay_row = rows[0] + expected_matches = expected.frame[expected.frame["question_id"].astype(str) == qid] + if expected_matches.empty: + raise OverlayError(f"Replaced question {qid!r} is not in the expected dataset.") + expected_row = expected_matches.iloc[0].to_dict() + predicted_option = None + if prediction_column is not None: + raw_prediction = overlay_row.values.get(prediction_column) + predicted_option = None if raw_prediction is None else str(raw_prediction).strip().upper() + origin = resolved_origins[qid] + row_payload = { + "question_id": qid, + "question_text": expected_row.get("question_text"), + "correct_option": expected_row.get("correct_option"), + "choices_json": expected_row.get("choices_json"), + "prediction_origin": origin, + "predicted_option": predicted_option, + } + preownership_rows.append(row_payload) + + input_digest = integrity_digest( + {"question_id": qid, "base_row": base.rows_by_question_id.get(qid)} + ) + preownership_output_digest = integrity_digest({"question_id": qid, "row": row_payload}) + component = make_lineage_component( + operation_type=operation_type, + question_id=qid, + parent_digests=(base.realization_digest,), + source_digests=(overlay_table.source_sha256,), + authorization_digest=authorization.bundle_digest, + implementation=dict(overlay.implementation), + implementation_mode="declared_external", + parameters={}, + input_digest=input_digest, + preownership_output_digest=preownership_output_digest, + prediction_origin=origin, + ) + lineage_components.append(component) + + row_assignments = [ + (component["identity"]["question_id"], component["identity"]["prediction_origin"], component["lineage_id"]) + for component in lineage_components + ] + result_origin = make_result_origin(derivation_origin=operation_type, row_assignments=row_assignments) + + if overlay.expected_evidence_status not in _EVIDENCE_STATUSES: + raise OverlayError( + f"Overlay expected_evidence_status {overlay.expected_evidence_status!r} is invalid." + ) + + return DerivedRealizationPayload( + condition_id=base.condition_id, + condition_digest=base.condition_digest, + preownership_rows=tuple(preownership_rows), + lineage_components=tuple(lineage_components), + result_origin=result_origin, + evidence_status=overlay.expected_evidence_status, + replacement_question_ids=tuple(sorted(replacement_ids)), + ) diff --git a/src/choicebench/importing/profiles/__init__.py b/src/choicebench/importing/profiles/__init__.py new file mode 100644 index 0000000..5c3f53e --- /dev/null +++ b/src/choicebench/importing/profiles/__init__.py @@ -0,0 +1,6 @@ +"""Paper-specific import profiles. + +No generic importer module may import anything from this package, and no +paper-specific constant (cell IDs, method names, queue file names) may leak +into src/choicebench/importing/*.py outside this directory. +""" diff --git a/src/choicebench/importing/profiles/stage1_paper_freeze.py b/src/choicebench/importing/profiles/stage1_paper_freeze.py new file mode 100644 index 0000000..8f97345 --- /dev/null +++ b/src/choicebench/importing/profiles/stage1_paper_freeze.py @@ -0,0 +1,984 @@ +"""Translate the read-only Stage 1 paper-freeze into verified queue/status/ +matrix accounting. + +Reduced scope: this verifies checksums, parses the checksum-covered +canonical manifest and the four queue files, explodes them into +(cell_id, question_id) pairs, and reproduces the queue/status/matrix counts +the research objective needs to confirm (exact 138/687/3 approved/held/ +excluded pairs, all pairwise disjoint; rerun_queue.csv's 776 rerun-candidate +pairs verified as a subset of the 828 classified pairs, not a fourth +authority tier -- rerun_queue.csv carries no queue_disposition/ +execution_authority/executable columns at all, unlike the three +classification files; the 100-cell paper matrix breakdown; and the +6-cell/18-pair offline recoverable authority). It does NOT yet build +ImportSpec/AuthorizationSpec +objects ready to feed into the generic engine -- that requires the ARC/MMLU +ExpectedDataset trust chain (to validate authorization grants against real +selected question IDs, i.e. Task 16/Unit J) plus full per-cell +SourceArtifactSpec construction across the freeze's distinct method CSV +header schemas, neither of which is done here. translate_stage1_paper_freeze +is the verified foundation both of those build on: checksum-verified cell +records and verified, pairwise-disjoint, correctly-typed queue pairs. + +This module is the only place in the importer allowed to know Stage 1 +paper-specific facts (cell ID format, queue file names, method names). No +generic importer module imports from here. +""" + +from __future__ import annotations + +import ast +from dataclasses import dataclass, replace +from hashlib import sha256 +import csv +import io +import json +from pathlib import Path +from typing import Any, Callable, Mapping + +from choicebench.identity import short_id +from choicebench.importing.csv_adapter import OpenedSource, parse_csv_source +from choicebench.importing.dataset_reference import ExpectedDataset, build_expected_dataset +from choicebench.importing.evidence import open_verified_source +from choicebench.importing.validation import compute_evidence_status, validate_source_rows +from choicebench.importing.schema import ( + AuthorizationSpec, + CsvDialectSpec, + DatasetReferenceSpec, + ImportConditionSpec, + ImportMethodSpec, + ImportModelSpec, + ImportPromptSpec, + ImportSpec, + OptionMappingSpec, + ResultOriginSpec, + SourceArtifactSpec, +) + +STATUS_MAP = { + "canonical_complete": "complete", + "canonical_qualified": "qualified", + "recoverable_from_existing_artifacts": "recoverable", + "incomplete_requires_inference": "partial", + "malformed_requires_inference": "malformed", +} +_EXCLUDED_STATUS = "excluded_from_paper_matrix" +_PAPER_SCOPE_EXEMPT_METHOD = "pride" +_IHS_METHOD = "independent_hypothesis" + +_VERIFIED_RELATIVE_PATHS = ( + "manifests/canonical_results_manifest.json", + "manifests/canonical_results_manifest.csv", + "manifests/cell_status_matrix.csv", + "manifests/expected_matrix.csv", + "manifests/approved_rerun_queue.csv", + "manifests/held_or_declined_reruns.csv", + "manifests/paper_scope_excluded_reruns.csv", + "manifests/rerun_queue.csv", + "reports/canonical_freeze_report.md", +) + + +class Stage1ProfileError(ValueError): + """Raised when the Stage 1 freeze fails checksum or consistency validation.""" + + +def _load_checksum_ledger(freeze_root: Path) -> dict[str, str]: + ledger: dict[str, str] = {} + path = freeze_root / "checksums" / "checksums.sha256" + with path.open("r", encoding="utf-8") as handle: + for line in handle: + line = line.rstrip("\n") + if not line: + continue + digest, _, relative_path = line.partition(" ") + ledger[relative_path] = digest + return ledger + + +def _verify_checksummed_file(freeze_root: Path, ledger: Mapping[str, str], relative_path: str) -> bytes: + path = freeze_root / relative_path + try: + data = path.read_bytes() + except OSError as exc: + raise Stage1ProfileError(f"{relative_path} is unreadable: {exc}") from exc + expected = ledger.get(relative_path) + if expected is None: + raise Stage1ProfileError(f"{relative_path} has no checksum ledger entry.") + actual = sha256(data).hexdigest() + if actual != expected: + raise Stage1ProfileError( + f"{relative_path} checksum mismatch: expected {expected}, found {actual}." + ) + return data + + +def _read_csv_rows(path: Path) -> list[dict[str, str]]: + with path.open(newline="", encoding="utf-8") as handle: + return list(csv.DictReader(handle)) + + +def _explode_question_pairs( + rows: list[dict[str, str]], *, source_label: str +) -> tuple[dict[tuple[str, str], dict[str, str]], dict[tuple[str, str], str]]: + pairs: dict[tuple[str, str], dict[str, str]] = {} + reasons: dict[tuple[str, str], str] = {} + for row in rows: + cell_id = row["cell_id"] + question_ids = json.loads(row["exact_question_ids"]) + question_reasons = json.loads(row["question_reasons"]) if row.get("question_reasons") else {} + for question_id in question_ids: + pair = (cell_id, question_id) + if pair in pairs: + raise Stage1ProfileError( + f"{source_label} declares duplicate (cell_id, question_id) pair {pair}." + ) + pairs[pair] = row + reasons[pair] = question_reasons.get(question_id, "") + return pairs, reasons + + +def _require_all(rows: list[dict[str, str]], *, field: str, expected: str, source_label: str) -> None: + bad = [row for row in rows if row.get(field) != expected] + if bad: + raise Stage1ProfileError( + f"{source_label} has {len(bad)} row(s) with {field} != {expected!r}." + ) + + +def _require_all_bool(rows: list[dict[str, str]], *, field: str, expected: bool, source_label: str) -> None: + expected_text = "true" if expected else "false" + bad = [row for row in rows if row.get(field, "").strip().lower() != expected_text] + if bad: + raise Stage1ProfileError( + f"{source_label} has {len(bad)} row(s) with {field} != {expected_text!r}." + ) + + +@dataclass(frozen=True) +class Stage1MatrixCounts: + intended_cells: int + core_method_cells: int + local_pride_cells: int + ihs_cells: int + excluded_preserved_cells: int + + +@dataclass(frozen=True) +class Stage1QueueCounts: + approved_executable_question_cells: int + held_nonexecutable_question_cells: int + excluded_nonexecutable_question_cells: int + rerun_candidate_question_cells: int + classified_question_cells: int + pairwise_disjoint: bool + approved_authoritative_executable: bool + held_executable: bool + excluded_executable: bool + rerun_candidates_are_classified_subset: bool + + +@dataclass(frozen=True) +class Stage1QueuePairs: + approved: Mapping[tuple[str, str], str] + held: Mapping[tuple[str, str], str] + excluded: Mapping[tuple[str, str], str] + rerun_candidates: frozenset[tuple[str, str]] + + +@dataclass(frozen=True) +class Stage1OfflineAuthority: + condition_count: int + question_cell_count: int + recoverable_question_ids: tuple[str, ...] + cell_ids: tuple[str, ...] + inference_executable: bool + + +@dataclass(frozen=True) +class Stage1Translation: + cell_records: Mapping[str, Mapping[str, Any]] + matrix: Stage1MatrixCounts + queue_counts: Stage1QueueCounts + queue_pairs: Stage1QueuePairs + offline_authority: Stage1OfflineAuthority + verified_relative_paths: tuple[str, ...] + + +def _matrix_counts(cell_records: Mapping[str, Mapping[str, Any]]) -> Stage1MatrixCounts: + method_counts: dict[str, int] = {} + excluded_cells = 0 + for record in cell_records.values(): + method_counts[record["method"]] = method_counts.get(record["method"], 0) + 1 + if record["final_status"] == _EXCLUDED_STATUS: + excluded_cells += 1 + ihs_cells = method_counts.get(_IHS_METHOD, 0) + pride_cells_total = method_counts.get(_PAPER_SCOPE_EXEMPT_METHOD, 0) + local_pride_cells = pride_cells_total - excluded_cells + core_method_cells = sum( + count + for method, count in method_counts.items() + if method not in {_IHS_METHOD, _PAPER_SCOPE_EXEMPT_METHOD} + ) + return Stage1MatrixCounts( + intended_cells=core_method_cells + ihs_cells + local_pride_cells, + core_method_cells=core_method_cells, + local_pride_cells=local_pride_cells, + ihs_cells=ihs_cells, + excluded_preserved_cells=excluded_cells, + ) + + +def _offline_authority(cell_records: Mapping[str, Mapping[str, Any]]) -> Stage1OfflineAuthority: + recoverable_cells = { + cell_id: tuple(sorted(record["recoverable_question_ids"])) + for cell_id, record in cell_records.items() + if record["recoverable_question_ids"] + } + distinct_id_sets = set(recoverable_cells.values()) + if len(distinct_id_sets) != 1: + raise Stage1ProfileError( + "Recoverable cells do not all authorize the exact same question IDs: " + f"{sorted(distinct_id_sets)}." + ) + (recoverable_question_ids,) = distinct_id_sets + cell_ids = tuple(sorted(recoverable_cells)) + return Stage1OfflineAuthority( + condition_count=len(cell_ids), + question_cell_count=len(cell_ids) * len(recoverable_question_ids), + recoverable_question_ids=recoverable_question_ids, + cell_ids=cell_ids, + inference_executable=False, + ) + + +def translate_stage1_paper_freeze(freeze_root: Path) -> Stage1Translation: + freeze_root = Path(freeze_root) + ledger = _load_checksum_ledger(freeze_root) + for relative_path in _VERIFIED_RELATIVE_PATHS: + _verify_checksummed_file(freeze_root, ledger, relative_path) + + manifest_json = json.loads( + (freeze_root / "manifests" / "canonical_results_manifest.json").read_text() + ) + cell_records: dict[str, Mapping[str, Any]] = {} + for record in manifest_json: + cell_id = record["cell_id"] + if cell_id in cell_records: + raise Stage1ProfileError(f"Canonical manifest declares duplicate cell_id {cell_id!r}.") + if record["status"] not in STATUS_MAP and record["final_status"] != _EXCLUDED_STATUS: + raise Stage1ProfileError( + f"Cell {cell_id!r} has an unrecognized status {record['status']!r}." + ) + cell_records[cell_id] = record + + matrix = _matrix_counts(cell_records) + offline_authority = _offline_authority(cell_records) + + approved_rows = _read_csv_rows(freeze_root / "manifests" / "approved_rerun_queue.csv") + held_rows = _read_csv_rows(freeze_root / "manifests" / "held_or_declined_reruns.csv") + excluded_rows = _read_csv_rows(freeze_root / "manifests" / "paper_scope_excluded_reruns.csv") + # rerun_queue.csv carries no queue_disposition/execution_authority/executable + # columns at all (unlike the three files above): it is the technical rerun + # working-queue, not an authority classification. Verified below to be a + # subset of the classified pairs, never treated as its own authority tier. + rerun_candidate_rows = _read_csv_rows(freeze_root / "manifests" / "rerun_queue.csv") + + _require_all(approved_rows, field="queue_disposition", expected="approved", source_label="approved_rerun_queue.csv") + _require_all(approved_rows, field="execution_authority", expected="authoritative", source_label="approved_rerun_queue.csv") + _require_all_bool(approved_rows, field="executable", expected=True, source_label="approved_rerun_queue.csv") + _require_all_bool(held_rows, field="executable", expected=False, source_label="held_or_declined_reruns.csv") + _require_all_bool(excluded_rows, field="executable", expected=False, source_label="paper_scope_excluded_reruns.csv") + + approved_pairs, approved_reasons = _explode_question_pairs(approved_rows, source_label="approved_rerun_queue.csv") + held_pairs, held_reasons = _explode_question_pairs(held_rows, source_label="held_or_declined_reruns.csv") + excluded_pairs, excluded_reasons = _explode_question_pairs(excluded_rows, source_label="paper_scope_excluded_reruns.csv") + rerun_candidate_pairs, _ = _explode_question_pairs(rerun_candidate_rows, source_label="rerun_queue.csv") + + classified = set(approved_pairs) | set(held_pairs) | set(excluded_pairs) + total_classified = len(approved_pairs) + len(held_pairs) + len(excluded_pairs) + if len(classified) != total_classified: + raise Stage1ProfileError( + "approved/held/excluded queues are not pairwise disjoint: " + f"{total_classified} declared pairs but only {len(classified)} distinct." + ) + unclassified_candidates = set(rerun_candidate_pairs) - classified + if unclassified_candidates: + raise Stage1ProfileError( + f"{len(unclassified_candidates)} rerun-queue candidate pair(s) were never " + "classified as approved, held, or excluded." + ) + + queue_counts = Stage1QueueCounts( + approved_executable_question_cells=len(approved_pairs), + held_nonexecutable_question_cells=len(held_pairs), + excluded_nonexecutable_question_cells=len(excluded_pairs), + rerun_candidate_question_cells=len(rerun_candidate_pairs), + classified_question_cells=len(classified), + pairwise_disjoint=True, + approved_authoritative_executable=True, + held_executable=False, + excluded_executable=False, + rerun_candidates_are_classified_subset=True, + ) + queue_pairs = Stage1QueuePairs( + approved=approved_reasons, held=held_reasons, excluded=excluded_reasons, + rerun_candidates=frozenset(rerun_candidate_pairs), + ) + + return Stage1Translation( + cell_records=cell_records, + matrix=matrix, + queue_counts=queue_counts, + queue_pairs=queue_pairs, + offline_authority=offline_authority, + verified_relative_paths=_VERIFIED_RELATIVE_PATHS, + ) + + +# --- ARC/MMLU expected-dataset trust chain (Task 16) ----------------------- +# +# Reduced scope, verified directly against the real freeze: this recomputes +# question_text/correct_option/correct_answer_text/choices from each raw row +# and compares them fieldwise against the archived normalized row at the +# SAME position -- raw and normalized files have identical row counts and +# the same row order, so a positional join is sufficient and avoids needing +# to reproduce the archived normalizer's own question_id hash algorithm. +# This is a ChoiceBench revalidation of internal freeze consistency, not +# proof that the archived normalizer (or the raw upstream publisher) was +# authentic -- recorded as a limitation on the returned ExpectedDataset. + +_DATASET_RELATIVE_PATHS = { + "arc_challenge": { + "raw": "raw/local_model_generalization/data/raw/arc_challenge_raw.csv", + "normalized": "raw/local_model_generalization/data/processed/arc_challenge_normalized.csv", + "ids": "raw/local_model_generalization/data/splits/arc_challenge/robustness_ids.json", + "metadata": "raw/local_model_generalization/data/splits/arc_challenge/robustness_metadata.json", + }, + "mmlu": { + "raw": "raw/local_model_generalization/data/raw/mmlu_raw.csv", + "normalized": "raw/local_model_generalization/data/processed/mmlu_normalized.csv", + "ids": "raw/local_model_generalization/data/splits/benchmark/robustness_ids.json", + "metadata": "raw/local_model_generalization/data/splits/benchmark/robustness_metadata.json", + }, +} + + +_ARC_MAX_OPTIONS = 4 + + +def _recompute_arc_fields(raw_row: Mapping[str, str]) -> dict[str, str]: + """The archived normalizer caps ARC options at _ARC_MAX_OPTIONS, silently + dropping any raw option beyond it -- verified against the real freeze: + 3 raw ARC questions have a 5th option, and all 3 have their answerKey + within the first _ARC_MAX_OPTIONS, so truncation never drops the correct + option. If a future/other freeze ever did drop the correct option, the + index check below still fails closed rather than silently truncating it. + """ + try: + parsed = json.loads(raw_row["choices"]) + texts = list(parsed["text"]) + labels = list(parsed["label"]) + index = labels.index(raw_row["answerKey"]) + except (KeyError, ValueError, TypeError) as exc: + raise Stage1ProfileError( + f"ARC raw row {raw_row.get('id')!r} has malformed choices/answerKey: {exc}." + ) from exc + if index >= _ARC_MAX_OPTIONS: + raise Stage1ProfileError( + f"ARC raw row {raw_row.get('id')!r} answerKey index {index} falls outside the " + f"archived normalizer's first {_ARC_MAX_OPTIONS} options; truncation would drop " + "the correct option, so this cannot be safely revalidated." + ) + texts = texts[:_ARC_MAX_OPTIONS] + fields = { + "question_text": raw_row["question"], + "correct_option": chr(ord("A") + index), + "correct_answer_text": texts[index], + } + for offset in range(_ARC_MAX_OPTIONS): + fields[f"choice_{chr(ord('a') + offset)}"] = texts[offset] if offset < len(texts) else "" + return fields + + +def _recompute_mmlu_fields(raw_row: Mapping[str, str]) -> dict[str, str]: + try: + choices = list(ast.literal_eval(raw_row["choices"])) + index = int(raw_row["answer"]) + except (SyntaxError, ValueError, TypeError) as exc: + raise Stage1ProfileError( + f"MMLU raw row has malformed choices/answer: {exc}." + ) from exc + if not (0 <= index < len(choices)): + raise Stage1ProfileError( + f"MMLU raw row answer index {index} is out of range for choices {choices!r}." + ) + fields = { + "question_text": raw_row["question"], + "correct_option": chr(ord("A") + index), + "correct_answer_text": str(choices[index]), + } + for offset, text in enumerate(choices): + fields[f"choice_{chr(ord('a') + offset)}"] = str(text) + return fields + + +def _revalidate_positional( + raw_rows: list[dict[str, str]], + normalized_rows: list[dict[str, str]], + *, + recompute: Callable[[Mapping[str, str]], dict[str, str]], + benchmark_name: str, +) -> None: + if len(raw_rows) != len(normalized_rows): + raise Stage1ProfileError( + f"{benchmark_name}: raw ({len(raw_rows)}) and normalized ({len(normalized_rows)}) " + "row counts disagree; this revalidation assumes positional 1:1 correspondence." + ) + for index, (raw_row, normalized_row) in enumerate(zip(raw_rows, normalized_rows)): + recomputed = recompute(raw_row) + for field, expected_value in recomputed.items(): + actual_value = (normalized_row.get(field) or "").strip() + if actual_value != expected_value.strip(): + raise Stage1ProfileError( + f"{benchmark_name} row {index} ({normalized_row.get('question_id')!r}): " + f"recomputed {field}={expected_value!r} does not match the archived " + f"normalized value {actual_value!r} (ChoiceBench revalidation failure)." + ) + + +def _drop_duplicate_questions( + normalized_rows: list[dict[str, str]], *, benchmark_name: str +) -> list[dict[str, str]]: + """pandas.DataFrame.drop_duplicates(subset="question_id", keep="first") + semantics -- but only after requiring every duplicate occurrence to agree + on all parsed fields; a disagreeing duplicate fails closed rather than + silently picking one.""" + seen: dict[str, dict[str, str]] = {} + order: list[str] = [] + for row in normalized_rows: + question_id = row["question_id"] + if question_id in seen: + if row != seen[question_id]: + raise Stage1ProfileError( + f"{benchmark_name}: duplicate question_id {question_id!r} occurs with " + "disagreeing field values; cannot apply keep-first deduplication." + ) + continue + seen[question_id] = row + order.append(question_id) + return [seen[question_id] for question_id in order] + + +def _build_expected_dataset_for_benchmark( + freeze_root: Path, + ledger: Mapping[str, str], + *, + benchmark_name: str, + raw_recompute: Callable[[Mapping[str, str]], dict[str, str]], +) -> ExpectedDataset: + paths = _DATASET_RELATIVE_PATHS[benchmark_name] + raw_bytes = _verify_checksummed_file(freeze_root, ledger, paths["raw"]) + normalized_bytes = _verify_checksummed_file(freeze_root, ledger, paths["normalized"]) + ids_bytes = _verify_checksummed_file(freeze_root, ledger, paths["ids"]) + metadata_bytes = _verify_checksummed_file(freeze_root, ledger, paths["metadata"]) + + raw_rows = list(csv.DictReader(io.StringIO(raw_bytes.decode("utf-8")))) + normalized_rows = list(csv.DictReader(io.StringIO(normalized_bytes.decode("utf-8")))) + _revalidate_positional( + raw_rows, normalized_rows, recompute=raw_recompute, benchmark_name=benchmark_name + ) + + deduplicated_rows = _drop_duplicate_questions(normalized_rows, benchmark_name=benchmark_name) + selected_ids = json.loads(ids_bytes) + metadata = json.loads(metadata_bytes) + + available_ids = {row["question_id"] for row in deduplicated_rows} + missing = [question_id for question_id in selected_ids if question_id not in available_ids] + if missing: + raise Stage1ProfileError( + f"{benchmark_name}: {len(missing)} selected question ID(s) are absent from the " + f"normalized/deduplicated source, e.g. {missing[:3]}." + ) + + fieldnames = list(deduplicated_rows[0].keys()) + buffer = io.StringIO() + writer = csv.DictWriter(buffer, fieldnames=fieldnames) + writer.writeheader() + writer.writerows(deduplicated_rows) + deduplicated_bytes = buffer.getvalue().encode("utf-8") + opened_source = OpenedSource( + source_id="normalized", + audit_path=freeze_root / paths["normalized"], + logical_path=paths["normalized"], + data=deduplicated_bytes, + sha256=sha256(deduplicated_bytes).hexdigest(), + ) + + choice_columns = { + key: key for key in fieldnames if key.startswith("choice_") and len(key) == 8 + } + declaration = DatasetReferenceSpec( + dataset_id=benchmark_name, + benchmark_name=benchmark_name, + split="robustness", + reference_kind="independent_input_snapshot", + trust_label="checksum_verified_freeze_internal", + source_ids=("normalized",), + selection_source_id="normalized", + expected_question_ids=tuple(selected_ids), + selection_seed=metadata.get("seed"), + selection_n_samples=metadata.get("actual_size"), + subject_filter=(), + selection_unknown_reasons={}, + columns={ + "question_id": "question_id", + "question_text": "question_text", + "correct_option": "correct_option", + **choice_columns, + }, + revision=None, + fingerprint=None, + derivation={"source_role": f"{benchmark_name}_stage1_freeze_normalized"}, + limitations=( + "Stage 1 freeze: field content independently revalidated against the raw " + "source by ChoiceBench; this proves internal freeze consistency, not that " + "the archived normalizer or the upstream publisher was authentic.", + ), + native_compatibility_identity=None, + ) + return build_expected_dataset(declaration, {"normalized": opened_source}) + + +def build_stage1_expected_datasets(freeze_root: Path) -> dict[str, ExpectedDataset]: + freeze_root = Path(freeze_root) + ledger = _load_checksum_ledger(freeze_root) + return { + "arc_challenge": _build_expected_dataset_for_benchmark( + freeze_root, ledger, benchmark_name="arc_challenge", raw_recompute=_recompute_arc_fields, + ), + "mmlu": _build_expected_dataset_for_benchmark( + freeze_root, ledger, benchmark_name="mmlu", raw_recompute=_recompute_mmlu_fields, + ), + } + + +# --- Full ImportSpec assembly (Task 16/Unit K remainder) -------------------- +# +# Reduced scope, verified directly against the real freeze: all 102 cells +# have observed_unique_count==1000 with zero missing/unexpected/duplicate +# question IDs, so compute_evidence_status's "partial"/"recoverable" +# branches (which need missing_ids) never trigger here -- every cell +# reconciles to either "malformed" or "complete"/"qualified". The freeze's +# own STATUS_MAP-mapped status is a scientific/methodology classification, +# not always what ChoiceBench's generic row validator independently +# computes: the 6 "recoverable" cells and the "mmlu independent_hypothesis" +# "incomplete" cell have ZERO row-level defects (their defect -- option- +# hiding, or a protocol-invalid option-A score after MAX_TOKENS termination +# -- is invisible to generic validation, documented only via +# damaged_question_ids/qualification narrative). So every condition's +# declared evidence_status is reconciled via an actual probe +# (validate_source_rows + compute_evidence_status), never trusted blindly +# from STATUS_MAP. qualifications=({"reason": ...},) is pre-populated +# whenever record["qualification"] is non-empty, so a cell with a +# documented caveat but zero generic defects reconciles to "qualified" +# rather than bare "complete". +# +# question_id/correct_option/prediction are the only source->identity +# column mappings (question_text is deliberately NOT mapped), avoiding +# noise from formatting differences between each run's own recorded +# question_text and Unit J's normalized ExpectedDataset text -- +# correct_option is the meaningful invariant (a categorical answer key, +# not prone to whitespace/formatting variance). +# +# Real queue reconciliation (verified directly against the freeze): 102 +# cells = 62 needing no repair (57 complete + 5 qualified) + 6 recoverable +# (offline authority, 18 pairs, all arc_challenge/semantic_matching_v1) + +# 32 malformed/incomplete, of which 31 cells are fully approved (138 pairs +# total) and exactly 1 cell +# (cbp__gemini-2-5-flash__arc_challenge__independent_hypothesis) is fully +# HELD (687 questions, supervisor-stopped) + 2 excluded_from_paper_matrix +# cells (both "pride"/qwen-2.5-7b-instruct-turbo), one of which also +# carries 3 explicitly-excluded damaged questions in +# paper_scope_excluded_reruns.csv. Held and excluded pairs must NOT +# receive any authorization/overlay. + +_STAGE1_COLUMNS = { + "question_id": "question_id", + "correct_option": "correct_option", + "prediction": "parsed_choice", +} +_STAGE1_OPTION_COLUMNS = ("choice_a", "choice_b", "choice_c", "choice_d") +_STAGE1_PROMPT_KEY = "stage1_freeze_v1" + + +def _stage1_prompt() -> ImportPromptSpec: + return ImportPromptSpec( + prompt_key=_STAGE1_PROMPT_KEY, + template_identity=None, + template_digest=None, + template_contents=None, + unknown_reason=( + "Stage 1 freeze: each canonical CSV row preserves its own literal rendered " + "'prompt' text, but no single stable prompt-template identity/digest was " + "recorded across methods at the paper's authoring time; method_key already " + "captures the distinct prompting strategy." + ), + native_compatibility_identity=None, + ) + + +def _stage1_model(model_name: str, provider_backend: str) -> ImportModelSpec: + return ImportModelSpec( + model_key=model_name, + display_name=model_name, + backend=provider_backend, + provider=provider_backend, + revision=None, + effective_parameters={}, + unknown_reasons={"revision": "not recorded by paper_data_freeze canonical manifest"}, + native_compatibility_identity=None, + ) + + +def _stage1_method(method_name: str) -> ImportMethodSpec: + return ImportMethodSpec( + method_key=method_name, + name=method_name, + effective_parameters={}, + implementation=None, + unknown_reasons={"implementation": "not recorded by paper_data_freeze canonical manifest"}, + native_compatibility_identity=None, + ) + + +def _stage1_cell_source( + freeze_root: Path, ledger: Mapping[str, str], cell_id: str, record: Mapping[str, Any] +) -> SourceArtifactSpec: + relative_path = record["canonical_path"] + ledger_digest = ledger.get(relative_path) + if ledger_digest is None: + raise Stage1ProfileError(f"{relative_path} has no checksum ledger entry.") + if ledger_digest != record["canonical_sha256"]: + raise Stage1ProfileError( + f"Cell {cell_id!r} canonical_sha256 disagrees with the checksum ledger for " + f"{relative_path}." + ) + path = freeze_root / relative_path + with path.open(newline="", encoding="utf-8") as handle: + header = tuple(next(csv.reader(handle))) + required = set(_STAGE1_COLUMNS.values()) | set(_STAGE1_OPTION_COLUMNS) + missing = required - set(header) + if missing: + raise Stage1ProfileError( + f"Cell {cell_id!r} canonical CSV is missing column(s) {sorted(missing)}." + ) + return SourceArtifactSpec( + source_id=cell_id, + path=path, + logical_path=relative_path, + expected_sha256=record["canonical_sha256"], + format="csv", + format_version="paper_data_freeze.v1", + classification="raw", + dialect=CsvDialectSpec(), + columns=dict(_STAGE1_COLUMNS), + expected_columns=header, + ignored_columns={}, + null_values=("",), + numeric_columns=(), + option_mapping=OptionMappingSpec( + mode="ordered_columns", + ordered_columns=_STAGE1_OPTION_COLUMNS, + structured_column=None, + structured_label_key=None, + structured_text_key=None, + ), + extra_field_policy="preserve_unmapped", + preserve_namespace="stage1_paper_freeze", + source_run_id=None, + source_repository=None, + source_commit=None, + notes={ + "cell_id": cell_id, + "declared_status": record["status"], + "final_status": record["final_status"], + "qualification": record.get("qualification") or "", + }, + ) + + +def _stage1_normalized_dataset_source( + freeze_root: Path, ledger: Mapping[str, str], *, benchmark_name: str, source_id: str +) -> SourceArtifactSpec: + """Documentation-only source declaration for the ARC/MMLU normalized + reference file: real path/checksum, but never opened by the generic + engine, since ImportRequest.expected_datasets overrides the per-source + ExpectedDataset rebuild with build_stage1_expected_datasets' own + revalidation/dedup chain (needed for MMLU's real duplicate question_ids, + which the generic rebuild would reject).""" + relative_path = _DATASET_RELATIVE_PATHS[benchmark_name]["normalized"] + expected_sha256 = ledger.get(relative_path) + if expected_sha256 is None: + raise Stage1ProfileError(f"{relative_path} has no checksum ledger entry.") + return SourceArtifactSpec( + source_id=source_id, + path=freeze_root / relative_path, + logical_path=relative_path, + expected_sha256=expected_sha256, + format="csv", + format_version="paper_data_freeze.v1", + classification="canonical", + dialect=CsvDialectSpec(), + columns={}, + expected_columns=(), + ignored_columns={}, + null_values=(), + numeric_columns=(), + option_mapping=OptionMappingSpec( + mode="ordered_columns", ordered_columns=(), structured_column=None, + structured_label_key=None, structured_text_key=None, + ), + extra_field_policy="preserve_unmapped", + preserve_namespace="stage1_paper_freeze", + source_run_id=None, source_repository=None, source_commit=None, + notes={"role": f"{benchmark_name}_expected_dataset_reference"}, + ) + + +def _reconcile_evidence_status( + condition: ImportConditionSpec, + *, + source: SourceArtifactSpec, + freeze_root: Path, + dataset: ExpectedDataset, +) -> ImportConditionSpec: + opened = open_verified_source(source, containment_root=freeze_root) + table = parse_csv_source(opened, source, strict=True) + validated = validate_source_rows( + table, source_id=source.source_id, mapping=source.columns, condition=condition, expected=dataset + ) + computed, _ = compute_evidence_status(validated, condition=condition) + if computed != condition.evidence_status: + condition = replace(condition, evidence_status=computed) + return condition + + +def build_stage1_import_spec( + freeze_root: Path, + translation: Stage1Translation, + expected_datasets: Mapping[str, ExpectedDataset], +) -> ImportSpec: + freeze_root = Path(freeze_root) + ledger = _load_checksum_ledger(freeze_root) + + dataset_source_ids = { + "arc_challenge": "arc_challenge_normalized_reference", + "mmlu": "mmlu_normalized_reference", + } + sources: list[SourceArtifactSpec] = [ + _stage1_normalized_dataset_source( + freeze_root, ledger, benchmark_name=benchmark_name, source_id=source_id + ) + for benchmark_name, source_id in dataset_source_ids.items() + ] + + conditions: list[ImportConditionSpec] = [] + models_by_key: dict[str, ImportModelSpec] = {} + methods_by_key: dict[str, ImportMethodSpec] = {} + prompt = _stage1_prompt() + + for cell_id in sorted(translation.cell_records): + record = translation.cell_records[cell_id] + source = _stage1_cell_source(freeze_root, ledger, cell_id, record) + sources.append(source) + + model_name = record["model"] + if model_name not in models_by_key: + models_by_key[model_name] = _stage1_model(model_name, record["provider_backend"]) + method_name = record["method"] + if method_name not in methods_by_key: + methods_by_key[method_name] = _stage1_method(method_name) + + benchmark = record["benchmark"] + dataset = expected_datasets[benchmark] + scope_disposition = ( + "excluded_from_paper_matrix" if record["final_status"] == _EXCLUDED_STATUS else "included" + ) + qualification_text = record.get("qualification") or "" + qualifications = ({"reason": qualification_text},) if qualification_text else () + + condition = ImportConditionSpec( + condition_key=cell_id, + source_ids=(cell_id,), + dataset_id=benchmark, + model_key=model_name, + method_key=method_name, + prompt_key=prompt.prompt_key, + seed=42, + calibration_identity=None, + preflight_identity=None, + protocol_settings={}, + generation_parameters={"temperature": 0.0}, + unknown_reasons={ + "calibration_identity": "not recorded by paper_data_freeze canonical manifest", + "preflight_identity": "not recorded by paper_data_freeze canonical manifest", + }, + expected_question_ids=dataset.selected_question_ids, + evidence_status=STATUS_MAP.get(record["status"], "complete"), + scope_disposition=scope_disposition, + executable=None, + qualifications=qualifications, + limitations=(), + damaged_question_ids=tuple(record["damaged_question_ids"]), + recoverable_question_ids=tuple(record["recoverable_question_ids"]), + result_origin=ResultOriginSpec( + derivation_origin="external_import", + default_prediction_origin="external_historical_inference", + per_question_prediction_origins={}, + ), + ) + condition = _reconcile_evidence_status( + condition, source=source, freeze_root=freeze_root, dataset=dataset + ) + conditions.append(condition) + + datasets = tuple( + DatasetReferenceSpec( + dataset_id=benchmark_name, + benchmark_name=benchmark_name, + split="robustness", + reference_kind="independent_input_snapshot", + trust_label="checksum_verified_freeze_internal", + source_ids=(dataset_source_ids[benchmark_name],), + selection_source_id=dataset_source_ids[benchmark_name], + expected_question_ids=expected_datasets[benchmark_name].selected_question_ids, + selection_seed=None, + selection_n_samples=None, + subject_filter=(), + selection_unknown_reasons={}, + columns={}, + revision=None, + fingerprint=None, + derivation={}, + limitations=(), + native_compatibility_identity=None, + ) + for benchmark_name in ("arc_challenge", "mmlu") + ) + # This DatasetReferenceSpec is documentation-only -- the real + # ExpectedDataset comes from ImportRequest.expected_datasets, built by + # build_stage1_expected_datasets' own revalidation/dedup chain, which + # build_import_plan's generic per-source rebuild cannot reproduce for + # MMLU's real duplicate question_ids (build_import_plan skips its own + # per-declaration rebuild entirely whenever expected_datasets is + # supplied, so this declaration's source_ids are never actually opened). + + return ImportSpec( + schema_version="choicebench.import-spec.v1", + import_name="Stage 1 paper data freeze", + sources=tuple(sources), + datasets=datasets, + models=tuple(models_by_key.values()), + methods=tuple(methods_by_key.values()), + prompts=(prompt,), + conditions=tuple(conditions), + authorizations=(), + overlays=(), + metrics=["accuracy"], + provenance={"producer_request_id": {"value": None, "reason": "paper_data_freeze import"}}, + audit={"source_location": str(freeze_root)}, + ) + + +# --- Per-cell authorization bundle ------------------------------------- +# +# IMPORTANT, verified by reading derive_overlay's own cross-check +# (overlays.py): a ValidatedAuthorization's input_evidence_digests is the +# WHOLE bundle's input_evidence_digests, unsliced by +# authorization_for_condition -- and derive_overlay requires +# overlay.base_evidence_digests to match it key-for-key AND every one of +# those keys to also appear in the base realization's OWN +# evidence_source_digests. Since every Stage 1 cell has exactly one +# source, a bundle whose input_evidence_digests spans more than one cell +# can never satisfy that check for any single cell's overlay. So each +# AuthorizationSpec here is deliberately scoped to exactly ONE cell/ +# condition (one grants key, one evidence-digest key) -- NOT one bundle +# for all 31 approved cells or all 6 recoverable cells at once. + +def build_stage1_cell_authorization( + freeze_root: Path, + ledger: Mapping[str, str], + *, + cell_id: str, + condition_digest: str, + authorization_type: str, + question_reasons: Mapping[str, str], + dataset_snapshot_digest: str, + cell_evidence_sha256: str, +) -> tuple[AuthorizationSpec, SourceArtifactSpec]: + if authorization_type == "inference_repair": + relative_path = "manifests/approved_rerun_queue.csv" + executable = True + elif authorization_type == "offline_transformation": + relative_path = "manifests/canonical_results_manifest.csv" + executable = False + else: + raise Stage1ProfileError(f"Unsupported authorization_type {authorization_type!r}.") + + expected_sha256 = ledger.get(relative_path) + if expected_sha256 is None: + raise Stage1ProfileError(f"{relative_path} has no checksum ledger entry.") + source_id = f"{cell_id}__{authorization_type}_authorization_source" + source = SourceArtifactSpec( + source_id=source_id, + path=freeze_root / relative_path, + logical_path=relative_path, + expected_sha256=expected_sha256, + format="csv", + format_version="paper_data_freeze.v1", + classification="canonical", + dialect=CsvDialectSpec(), + columns={}, + expected_columns=(), + ignored_columns={}, + null_values=(), + numeric_columns=(), + option_mapping=OptionMappingSpec( + mode="ordered_columns", ordered_columns=(), structured_column=None, + structured_label_key=None, structured_text_key=None, + ), + extra_field_policy="preserve_unmapped", + preserve_namespace="stage1_paper_freeze", + source_run_id=None, source_repository=None, source_commit=None, + notes={"role": f"{authorization_type}_authorization_evidence", "cell_id": cell_id}, + ) + + purpose = ( + "Stage 1 approved rerun repair" if authorization_type == "inference_repair" + else "Stage 1 offline-recoverable semantic rematch" + ) + payload = { + "schema_version": "choicebench.authorization-bundle.v1", + "authorization_type": authorization_type, + "grants": {condition_digest: dict(question_reasons)}, + "authority": "paper_data_freeze.manifests", + "purpose": purpose, + "executable": executable, + "source_sha256": expected_sha256, + "input_evidence_digests": {cell_id: cell_evidence_sha256}, + "expected_snapshot_digests": {condition_digest: dataset_snapshot_digest}, + } + authorization_id = short_id("auth", payload) + authorization = AuthorizationSpec( + authorization_id=authorization_id, + authorization_type=authorization_type, + source_id=source_id, + condition_question_reasons={condition_digest: dict(question_reasons)}, + authority="paper_data_freeze.manifests", + purpose=purpose, + executable=executable, + input_evidence_digests={cell_id: cell_evidence_sha256}, + expected_snapshot_digests={condition_digest: dataset_snapshot_digest}, + ) + return authorization, source diff --git a/src/choicebench/importing/schema.py b/src/choicebench/importing/schema.py new file mode 100644 index 0000000..8d7967a --- /dev/null +++ b/src/choicebench/importing/schema.py @@ -0,0 +1,1749 @@ +"""Strict, paper-agnostic declarations for importing external results.""" + +from __future__ import annotations + +from dataclasses import asdict, dataclass +import math +from numbers import Real +from pathlib import Path, PurePosixPath +import re +from typing import Any, Literal, Mapping, TypeAlias + +import yaml + +from choicebench.identity import canonicalize, integrity_digest, is_credential_key, short_id +from choicebench.metrics import BUILTIN_METRICS + + +ImportState: TypeAlias = Literal["validated", "imported", "failed"] +EvidenceStatus: TypeAlias = Literal[ + "complete", "qualified", "partial", "malformed", "recoverable", "failed" +] +ScopeDisposition: TypeAlias = Literal[ + "included", "excluded_from_paper_matrix", "held", "superseded" +] +PredictionOrigin: TypeAlias = Literal[ + "native_inference", + "external_historical_inference", + "external_repair_inference", +] +DerivationOrigin: TypeAlias = Literal[ + "native_execution", "external_import", "repair_overlay", "offline_transformation" +] + + +class ImportSpecError(ValueError): + """Raised when an external import declaration is unsafe or inconsistent.""" + + +@dataclass(frozen=True) +class CsvDialectSpec: + encoding: Literal["utf-8"] = "utf-8" + bom_policy: Literal["forbid", "strip_utf8_bom"] = "forbid" + decoding_errors: Literal["strict"] = "strict" + delimiter: str = "," + quote_character: str = '"' + escape_character: str | None = None + double_quote: bool = True + line_terminators: tuple[str, ...] = ("crlf", "lf", "cr") + mixed_line_terminators: Literal["allow", "forbid"] = "allow" + final_record_without_terminator: Literal["allow", "forbid"] = "allow" + blank_record_policy: Literal["reject"] = "reject" + skip_initial_space: bool = False + header: Literal["first_logical_record"] = "first_logical_record" + strict_syntax: bool = True + + +@dataclass(frozen=True) +class NumericColumnSpec: + source_column: str + value_type: Literal["integer", "float"] + null_allowed: bool + finite_only: bool = True + + +@dataclass(frozen=True) +class OptionMappingSpec: + mode: Literal["ordered_columns", "structured_json"] + ordered_columns: tuple[str, ...] + structured_column: str | None + structured_label_key: str | None + structured_text_key: str | None + + +@dataclass(frozen=True) +class SourceArtifactSpec: + source_id: str + path: Path + logical_path: str + expected_sha256: str + format: Literal["csv"] + format_version: str + classification: Literal["raw", "canonical", "derived", "repaired", "aggregate_only"] + dialect: CsvDialectSpec + columns: Mapping[str, str] + expected_columns: tuple[str, ...] + ignored_columns: Mapping[str, str] + null_values: tuple[str, ...] + numeric_columns: tuple[NumericColumnSpec, ...] + option_mapping: OptionMappingSpec + extra_field_policy: Literal["preserve_unmapped", "reject_unmapped"] + preserve_namespace: str + source_run_id: str | None + source_repository: str | None + source_commit: str | None + notes: Mapping[str, Any] + + +@dataclass(frozen=True) +class DatasetReferenceSpec: + dataset_id: str + benchmark_name: str + split: str + reference_kind: Literal[ + "independent_input_snapshot", "profile_derived_reference_snapshot" + ] + trust_label: str + source_ids: tuple[str, ...] + selection_source_id: str + expected_question_ids: tuple[str, ...] + selection_seed: int | None + selection_n_samples: int | None + subject_filter: tuple[str, ...] + selection_unknown_reasons: Mapping[str, str] + columns: Mapping[str, str] + revision: str | None + fingerprint: str | None + derivation: Mapping[str, Any] + limitations: tuple[str, ...] + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ImportModelSpec: + model_key: str + display_name: str + backend: str | None + provider: str | None + revision: str | None + effective_parameters: Mapping[str, Any] + unknown_reasons: Mapping[str, str] + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ImportMethodSpec: + method_key: str + name: str + effective_parameters: Mapping[str, Any] + implementation: Mapping[str, Any] | None + unknown_reasons: Mapping[str, str] + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ImportPromptSpec: + prompt_key: str + template_identity: str | None + template_digest: str | None + template_contents: Mapping[str, str] | None + unknown_reason: str | None + native_compatibility_identity: Mapping[str, Any] | None + + +@dataclass(frozen=True) +class ResultOriginSpec: + derivation_origin: DerivationOrigin + default_prediction_origin: PredictionOrigin | None + per_question_prediction_origins: Mapping[str, PredictionOrigin] + + +@dataclass(frozen=True) +class ImportConditionSpec: + condition_key: str + source_ids: tuple[str, ...] + dataset_id: str + model_key: str + method_key: str + prompt_key: str + seed: int | None + calibration_identity: Mapping[str, Any] | None + preflight_identity: Mapping[str, Any] | None + protocol_settings: Mapping[str, Any] + generation_parameters: Mapping[str, Any] + unknown_reasons: Mapping[str, str] + expected_question_ids: tuple[str, ...] + evidence_status: EvidenceStatus + scope_disposition: ScopeDisposition + executable: bool | None + qualifications: tuple[Mapping[str, Any], ...] + limitations: tuple[Mapping[str, Any], ...] + damaged_question_ids: tuple[str, ...] + recoverable_question_ids: tuple[str, ...] + result_origin: ResultOriginSpec + + +@dataclass(frozen=True) +class AuthorizationSpec: + authorization_id: str + authorization_type: Literal["inference_repair", "offline_transformation"] + source_id: str + condition_question_reasons: Mapping[str, Mapping[str, str]] + authority: str + purpose: str + executable: bool + input_evidence_digests: Mapping[str, str] + expected_snapshot_digests: Mapping[str, str] + + +@dataclass(frozen=True) +class OverlaySpec: + overlay_id: str + base_run_path: Path + base_condition_digest: str + base_realization_id: str + base_realization_digest: str + base_evidence_digests: Mapping[str, str] + base_validation_artifact_sha256: str + base_result_sha256: str | None + source_id: str + authorization_id: str + replacement_reasons: Mapping[str, str] + result_origin: ResultOriginSpec + lineage_notes: Mapping[str, Any] + implementation: Mapping[str, Any] + input_digest: str + preownership_output_digest: str + expected_evidence_status: EvidenceStatus + + +@dataclass(frozen=True) +class ImportSpec: + schema_version: Literal["choicebench.import-spec.v1"] + import_name: str + sources: tuple[SourceArtifactSpec, ...] + datasets: tuple[DatasetReferenceSpec, ...] + models: tuple[ImportModelSpec, ...] + methods: tuple[ImportMethodSpec, ...] + prompts: tuple[ImportPromptSpec, ...] + conditions: tuple[ImportConditionSpec, ...] + authorizations: tuple[AuthorizationSpec, ...] + overlays: tuple[OverlaySpec, ...] + metrics: tuple[str, ...] + provenance: Mapping[str, Any] + audit: Mapping[str, Any] + + +_TOP_LEVEL_KEYS = { + "schema_version", + "import_name", + "sources", + "datasets", + "models", + "methods", + "prompts", + "conditions", + "authorizations", + "overlays", + "metrics", + "provenance", + "audit", +} +_SOURCE_KEYS = { + "source_id", + "path", + "logical_path", + "expected_sha256", + "format", + "format_version", + "classification", + "dialect", + "columns", + "expected_columns", + "ignored_columns", + "null_values", + "numeric_columns", + "option_mapping", + "extra_field_policy", + "preserve_namespace", + "source_run_id", + "source_repository", + "source_commit", + "notes", +} +_DIALECT_KEYS = { + "encoding", + "bom_policy", + "decoding_errors", + "delimiter", + "quote_character", + "escape_character", + "double_quote", + "line_terminators", + "mixed_line_terminators", + "final_record_without_terminator", + "blank_record_policy", + "skip_initial_space", + "header", + "strict_syntax", +} +_NUMERIC_KEYS = {"source_column", "value_type", "null_allowed", "finite_only"} +_OPTION_KEYS = { + "mode", + "ordered_columns", + "structured_column", + "structured_label_key", + "structured_text_key", +} +_DATASET_KEYS = { + "dataset_id", + "benchmark_name", + "split", + "reference_kind", + "trust_label", + "source_ids", + "selection_source_id", + "expected_question_ids", + "selection_seed", + "selection_n_samples", + "subject_filter", + "selection_unknown_reasons", + "columns", + "revision", + "fingerprint", + "derivation", + "limitations", + "native_compatibility_identity", +} +_MODEL_KEYS = { + "model_key", + "display_name", + "backend", + "provider", + "revision", + "effective_parameters", + "unknown_reasons", + "native_compatibility_identity", +} +_METHOD_KEYS = { + "method_key", + "name", + "effective_parameters", + "implementation", + "unknown_reasons", + "native_compatibility_identity", +} +_PROMPT_KEYS = { + "prompt_key", + "template_identity", + "template_digest", + "template_contents", + "unknown_reason", + "native_compatibility_identity", +} +_CONDITION_KEYS = { + "condition_key", + "source_ids", + "dataset_id", + "model_key", + "method_key", + "prompt_key", + "seed", + "calibration_identity", + "preflight_identity", + "protocol_settings", + "generation_parameters", + "unknown_reasons", + "expected_question_ids", + "evidence_status", + "scope_disposition", + "executable", + "qualifications", + "limitations", + "damaged_question_ids", + "recoverable_question_ids", + "result_origin", +} +_PROTOCOL_SETTING_KEYS = {"pride_modal_k_threshold"} +_RESULT_ORIGIN_KEYS = { + "derivation_origin", + "default_prediction_origin", + "per_question_prediction_origins", +} +_AUTHORIZATION_KEYS = { + "authorization_id", + "authorization_type", + "source_id", + "condition_question_reasons", + "authority", + "purpose", + "executable", + "input_evidence_digests", + "expected_snapshot_digests", +} +_OVERLAY_KEYS = { + "overlay_id", + "base_run_path", + "base_condition_digest", + "base_realization_id", + "base_realization_digest", + "base_evidence_digests", + "base_validation_artifact_sha256", + "base_result_sha256", + "source_id", + "authorization_id", + "replacement_reasons", + "result_origin", + "lineage_notes", + "implementation", + "input_digest", + "preownership_output_digest", + "expected_evidence_status", +} + +_EVIDENCE_STATUSES = { + "complete", "qualified", "partial", "malformed", "recoverable", "failed" +} +_SCOPE_DISPOSITIONS = { + "included", "excluded_from_paper_matrix", "held", "superseded" +} +_PREDICTION_ORIGINS = { + "native_inference", "external_historical_inference", "external_repair_inference" +} +_DERIVATION_ORIGINS = { + "native_execution", "external_import", "repair_overlay", "offline_transformation" +} +_SHA256_RE = re.compile(r"^[0-9a-f]{64}$") +_WINDOWS_ABSOLUTE_RE = re.compile(r"^[A-Za-z]:[\\/]") + + +def _mapping(value: Any, where: str) -> Mapping[str, Any]: + if not isinstance(value, Mapping): + raise ImportSpecError(f"{where} must be a YAML mapping; got {value!r}.") + bad_key = next((key for key in value if not isinstance(key, str)), None) + if bad_key is not None: + raise ImportSpecError( + f"{where} mapping keys must be strings; got {bad_key!r}." + ) + return value + + +def _exact_keys(value: Any, allowed: set[str], where: str) -> Mapping[str, Any]: + raw = _mapping(value, where) + unknown = sorted(set(raw) - allowed) + if unknown: + raise ImportSpecError(f"Unknown field(s) in {where}: {unknown}.") + missing = sorted(allowed - set(raw)) + if missing: + raise ImportSpecError(f"Missing required field(s) in {where}: {missing}.") + return raw + + +def _allowed_keys(value: Any, allowed: set[str], where: str) -> Mapping[str, Any]: + raw = _mapping(value, where) + unknown = sorted(set(raw) - allowed) + if unknown: + raise ImportSpecError(f"Unknown field(s) in {where}: {unknown}.") + return raw + + +def _nonempty(value: Any, where: str) -> str: + if not isinstance(value, str) or not value.strip(): + raise ImportSpecError(f"{where} must be a non-empty string; got {value!r}.") + return value.strip() + + +def _optional_string(value: Any, where: str) -> str | None: + if value is None: + return None + return _nonempty(value, where) + + +def _strict_bool(value: Any, where: str) -> bool: + if not isinstance(value, bool): + raise ImportSpecError(f"{where} must be true or false; got {value!r}.") + return value + + +def _optional_bool(value: Any, where: str) -> bool | None: + if value is None: + return None + return _strict_bool(value, where) + + +def _strict_int(value: Any, where: str, *, minimum: int | None = None) -> int: + if isinstance(value, bool) or not isinstance(value, int): + raise ImportSpecError(f"{where} must be an integer; got {value!r}.") + if minimum is not None and value < minimum: + raise ImportSpecError(f"{where} must be >= {minimum}; got {value!r}.") + return value + + +def _strict_number(value: Any, where: str) -> float: + if isinstance(value, bool) or not isinstance(value, Real): + raise ImportSpecError(f"{where} must be a finite number; got {value!r}.") + result = float(value) + if not math.isfinite(result): + raise ImportSpecError(f"{where} must be a finite number; got {value!r}.") + return result + + +def _optional_int(value: Any, where: str, *, minimum: int | None = None) -> int | None: + if value is None: + return None + return _strict_int(value, where, minimum=minimum) + + +def _enum(value: Any, values: set[str], where: str) -> str: + result = _nonempty(value, where) + if result not in values: + raise ImportSpecError( + f"{where} must be one of {sorted(values)}; got {result!r}." + ) + return result + + +def _sha256(value: Any, where: str) -> str: + result = _nonempty(value, where) + if not _SHA256_RE.fullmatch(result): + raise ImportSpecError(f"{where} must be a lowercase SHA-256 digest.") + return result + + +def _optional_sha256(value: Any, where: str) -> str | None: + if value is None: + return None + return _sha256(value, where) + + +def _sequence(value: Any, where: str) -> list[Any]: + if not isinstance(value, list): + raise ImportSpecError(f"{where} must be a YAML list; got {value!r}.") + return value + + +def _strings( + value: Any, + where: str, + *, + nonempty: bool = False, + unique: bool = False, + allow_empty_items: bool = False, +) -> tuple[str, ...]: + items = _sequence(value, where) + if allow_empty_items: + if not all(isinstance(item, str) for item in items): + bad = next(item for item in items if not isinstance(item, str)) + raise ImportSpecError(f"{where} entries must be strings; got {bad!r}.") + result = tuple(items) + else: + result = tuple( + _nonempty(item, f"{where}[{index}]") + for index, item in enumerate(items) + ) + if nonempty and not result: + raise ImportSpecError(f"{where} must be a non-empty list.") + if unique and len(result) != len(set(result)): + raise ImportSpecError(f"{where} contains duplicate values.") + return result + + +def _string_mapping(value: Any, where: str) -> dict[str, str]: + raw = _mapping(value, where) + return { + _nonempty(key, f"{where} key"): _nonempty(item, f"{where}.{key}") + for key, item in raw.items() + } + + +def _validate_unknown_reasons( + values: Mapping[str, Any], reasons: Mapping[str, str], where: str +) -> None: + unknown = sorted(set(reasons) - set(values)) + if unknown: + raise ImportSpecError(f"{where} contains unknown reason key(s) {unknown}.") + missing = sorted(key for key, value in values.items() if value is None and key not in reasons) + if missing: + raise ImportSpecError( + f"{where} must document every unknown nullable field; missing {missing}." + ) + contradictory = sorted( + key for key, value in values.items() if value is not None and key in reasons + ) + if contradictory: + raise ImportSpecError( + f"{where} gives unknown reasons for known field(s) {contradictory}." + ) + + +def _sha_mapping(value: Any, where: str) -> dict[str, str]: + raw = _mapping(value, where) + return { + _nonempty(key, f"{where} key"): _sha256(item, f"{where}.{key}") + for key, item in raw.items() + } + + +def _credential_keys_in(value: Any) -> set[str]: + found: set[str] = set() + if isinstance(value, Mapping): + for key, item in value.items(): + if isinstance(key, str) and is_credential_key(key): + found.add(key) + found |= _credential_keys_in(item) + elif isinstance(value, (list, tuple)): + for item in value: + found |= _credential_keys_in(item) + return found + + +def _canonical_mapping(value: Any, where: str) -> dict[str, Any]: + raw = _mapping(value, where) + try: + result = canonicalize(raw) + except (TypeError, ValueError) as exc: + raise ImportSpecError( + f"{where} must be finite, safe canonical data: {exc}" + ) from exc + return dict(result) + + +def _canonical_mapping_or_none(value: Any, where: str) -> dict[str, Any] | None: + if value is None: + return None + return _canonical_mapping(value, where) + + +def _mapping_sequence(value: Any, where: str) -> tuple[Mapping[str, Any], ...]: + return tuple( + _canonical_mapping(item, f"{where}[{index}]") + for index, item in enumerate(_sequence(value, where)) + ) + + +def _is_path_like(value: str) -> bool: + return ( + value.startswith(("/", "./", "../", "~/", "~\\")) + or _WINDOWS_ABSOLUTE_RE.match(value) is not None + ) + + +def _contains_path_like(value: Any) -> bool: + if isinstance(value, str): + return _is_path_like(value) + if isinstance(value, Mapping): + return any(_contains_path_like(item) for item in value.values()) + if isinstance(value, (list, tuple)): + return any(_contains_path_like(item) for item in value) + return False + + +def _build_provenance(value: Any) -> dict[str, Any]: + raw = _mapping(value, "provenance") + result: dict[str, Any] = {} + for key, item in raw.items(): + where = f"provenance.{key}" + record = _exact_keys(item, {"value", "reason"}, where) + reason = record["reason"] + if record["value"] is None: + reason = _nonempty(reason, f"{where}.reason") + elif reason is not None: + reason = _nonempty(reason, f"{where}.reason") + if _contains_path_like(record["value"]): + raise ImportSpecError( + f"{where}.value contains a machine-local path; put it in audit instead." + ) + canonical = _canonical_mapping( + {"value": record["value"], "reason": reason}, where + ) + result[_nonempty(key, "provenance key")] = canonical + return result + + +def _ascii_byte(value: Any, where: str, *, optional: bool = False) -> str | None: + if value is None and optional: + return None + if not isinstance(value, str) or len(value.encode("utf-8")) != 1: + raise ImportSpecError(f"{where} must be one ASCII byte.") + if ord(value) >= 128 or value in {"\x00", "\r", "\n"}: + raise ImportSpecError(f"{where} must be one usable ASCII byte.") + return value + + +def _build_dialect(value: Any, where: str) -> CsvDialectSpec: + raw = _allowed_keys(value, _DIALECT_KEYS, where) + encoding = _enum(raw.get("encoding", "utf-8"), {"utf-8"}, f"{where}.encoding") + decoding_errors = _enum( + raw.get("decoding_errors", "strict"), {"strict"}, f"{where}.decoding_errors" + ) + delimiter = _ascii_byte(raw.get("delimiter", ","), f"{where}.delimiter") + quote = _ascii_byte(raw.get("quote_character", '"'), f"{where}.quote_character") + escape = _ascii_byte( + raw.get("escape_character"), f"{where}.escape_character", optional=True + ) + characters = [item for item in (delimiter, quote, escape) if item is not None] + if len(characters) != len(set(characters)): + raise ImportSpecError( + f"{where} delimiter, quote_character, and escape_character must be distinct." + ) + line_terminators = _strings( + raw.get("line_terminators", ["crlf", "lf", "cr"]), + f"{where}.line_terminators", + nonempty=True, + unique=True, + ) + if not set(line_terminators) <= {"crlf", "lf", "cr"}: + raise ImportSpecError( + f"{where}.line_terminators may contain only crlf, lf, and cr." + ) + return CsvDialectSpec( + encoding=encoding, + bom_policy=_enum( + raw.get("bom_policy", "forbid"), + {"forbid", "strip_utf8_bom"}, + f"{where}.bom_policy", + ), + decoding_errors=decoding_errors, + delimiter=delimiter, + quote_character=quote, + escape_character=escape, + double_quote=_strict_bool(raw.get("double_quote", True), f"{where}.double_quote"), + line_terminators=line_terminators, + mixed_line_terminators=_enum( + raw.get("mixed_line_terminators", "allow"), + {"allow", "forbid"}, + f"{where}.mixed_line_terminators", + ), + final_record_without_terminator=_enum( + raw.get("final_record_without_terminator", "allow"), + {"allow", "forbid"}, + f"{where}.final_record_without_terminator", + ), + blank_record_policy=_enum( + raw.get("blank_record_policy", "reject"), + {"reject"}, + f"{where}.blank_record_policy", + ), + skip_initial_space=_strict_bool( + raw.get("skip_initial_space", False), f"{where}.skip_initial_space" + ), + header=_enum( + raw.get("header", "first_logical_record"), + {"first_logical_record"}, + f"{where}.header", + ), + strict_syntax=_strict_bool( + raw.get("strict_syntax", True), f"{where}.strict_syntax" + ), + ) + + +def _build_numeric(value: Any, where: str) -> NumericColumnSpec: + raw = _allowed_keys(value, _NUMERIC_KEYS, where) + missing = sorted({"source_column", "value_type", "null_allowed"} - set(raw)) + if missing: + raise ImportSpecError(f"Missing required field(s) in {where}: {missing}.") + return NumericColumnSpec( + source_column=_nonempty(raw["source_column"], f"{where}.source_column"), + value_type=_enum( + raw["value_type"], {"integer", "float"}, f"{where}.value_type" + ), + null_allowed=_strict_bool(raw["null_allowed"], f"{where}.null_allowed"), + finite_only=_strict_bool(raw.get("finite_only", True), f"{where}.finite_only"), + ) + + +def _build_option(value: Any, where: str) -> OptionMappingSpec: + raw = _exact_keys(value, _OPTION_KEYS, where) + mode = _enum(raw["mode"], {"ordered_columns", "structured_json"}, f"{where}.mode") + ordered = _strings( + raw["ordered_columns"], f"{where}.ordered_columns", unique=True + ) + structured_column = _optional_string( + raw["structured_column"], f"{where}.structured_column" + ) + label_key = _optional_string( + raw["structured_label_key"], f"{where}.structured_label_key" + ) + text_key = _optional_string( + raw["structured_text_key"], f"{where}.structured_text_key" + ) + if mode == "ordered_columns": + if not ordered or any( + item is not None for item in (structured_column, label_key, text_key) + ): + raise ImportSpecError( + f"{where} ordered_columns mode requires ordered columns and null structured fields." + ) + elif ordered or any(item is None for item in (structured_column, label_key, text_key)): + raise ImportSpecError( + f"{where} structured_json mode requires an empty ordered list and all structured fields." + ) + return OptionMappingSpec(mode, ordered, structured_column, label_key, text_key) + + +def _build_source(value: Any, index: int) -> SourceArtifactSpec: + where = f"sources[{index}]" + raw = _exact_keys(value, _SOURCE_KEYS, where) + logical_path = _nonempty(raw["logical_path"], f"{where}.logical_path") + pure_path = PurePosixPath(logical_path) + if pure_path.is_absolute() or ".." in pure_path.parts: + raise ImportSpecError(f"{where}.logical_path must be a safe relative logical path.") + expected_columns = _strings( + raw["expected_columns"], f"{where}.expected_columns", nonempty=True, unique=True + ) + numeric_columns = tuple( + _build_numeric(item, f"{where}.numeric_columns[{item_index}]") + for item_index, item in enumerate(_sequence(raw["numeric_columns"], f"{where}.numeric_columns")) + ) + numeric_names = [item.source_column for item in numeric_columns] + if len(numeric_names) != len(set(numeric_names)): + raise ImportSpecError(f"Duplicate numeric-column rules in {where}.numeric_columns.") + missing_numeric = sorted(set(numeric_names) - set(expected_columns)) + if missing_numeric: + raise ImportSpecError( + f"{where}.numeric_columns reference unknown expected columns {missing_numeric}." + ) + option_mapping = _build_option(raw["option_mapping"], f"{where}.option_mapping") + option_columns = ( + set(option_mapping.ordered_columns) + if option_mapping.mode == "ordered_columns" + else {option_mapping.structured_column} + ) + missing_options = sorted(item for item in option_columns if item not in expected_columns) + if missing_options: + raise ImportSpecError( + f"{where}.option_mapping references unknown expected columns {missing_options}." + ) + columns = _string_mapping(raw["columns"], f"{where}.columns") + missing_mappings = sorted(set(columns.values()) - set(expected_columns)) + if missing_mappings: + raise ImportSpecError( + f"{where}.columns references unknown expected columns {missing_mappings}." + ) + ignored_columns = _string_mapping(raw["ignored_columns"], f"{where}.ignored_columns") + missing_ignored = sorted(set(ignored_columns) - set(expected_columns)) + if missing_ignored: + raise ImportSpecError( + f"{where}.ignored_columns references unknown expected columns {missing_ignored}." + ) + mapped_values = list(columns.values()) + if len(mapped_values) != len(set(mapped_values)): + raise ImportSpecError( + f"{where}.columns assigns one source column to multiple mapped dispositions." + ) + mapped_columns = set(mapped_values) + ignored_set = set(ignored_columns) + conflicts = sorted( + (mapped_columns & option_columns) + | (mapped_columns & ignored_set) + | (option_columns & ignored_set) + ) + if conflicts: + raise ImportSpecError( + f"{where} source column disposition conflicts for {conflicts}; mapped, " + "option, and ignored columns must be pairwise disjoint." + ) + extra_field_policy = _enum( + raw["extra_field_policy"], + {"preserve_unmapped", "reject_unmapped"}, + f"{where}.extra_field_policy", + ) + disposed_columns = mapped_columns | option_columns | ignored_set + if extra_field_policy == "reject_unmapped": + undisposed = sorted(set(expected_columns) - disposed_columns) + if undisposed: + raise ImportSpecError( + f"{where} reject_unmapped requires one explicit disposition per source " + f"column; missing {undisposed}." + ) + path_text = _nonempty(raw["path"], f"{where}.path") + return SourceArtifactSpec( + source_id=_nonempty(raw["source_id"], f"{where}.source_id"), + path=Path(path_text), + logical_path=logical_path, + expected_sha256=_sha256(raw["expected_sha256"], f"{where}.expected_sha256"), + format=_enum(raw["format"], {"csv"}, f"{where}.format"), + format_version=_nonempty(raw["format_version"], f"{where}.format_version"), + classification=_enum( + raw["classification"], + {"raw", "canonical", "derived", "repaired", "aggregate_only"}, + f"{where}.classification", + ), + dialect=_build_dialect(raw["dialect"], f"{where}.dialect"), + columns=columns, + expected_columns=expected_columns, + ignored_columns=ignored_columns, + null_values=_strings( + raw["null_values"], + f"{where}.null_values", + unique=True, + allow_empty_items=True, + ), + numeric_columns=numeric_columns, + option_mapping=option_mapping, + extra_field_policy=extra_field_policy, + preserve_namespace=_nonempty( + raw["preserve_namespace"], f"{where}.preserve_namespace" + ), + source_run_id=_optional_string(raw["source_run_id"], f"{where}.source_run_id"), + source_repository=_optional_string( + raw["source_repository"], f"{where}.source_repository" + ), + source_commit=_optional_string(raw["source_commit"], f"{where}.source_commit"), + notes=_canonical_mapping(raw["notes"], f"{where}.notes"), + ) + + +def _validate_identity_claim( + raw: Mapping[str, Any], *, payload_key: str, digest_key: str, id_key: str, + prefix: str, where: str +) -> None: + payload = raw[payload_key] + if raw[digest_key] != integrity_digest(payload): + raise ImportSpecError(f"{where}.{digest_key} does not match its payload.") + if raw[id_key] != short_id(prefix, payload): + raise ImportSpecError(f"{where}.{id_key} does not match its payload.") + + +def _validate_dataset_native(value: Any, where: str) -> dict[str, Any]: + keys = { + "artifact_payload", "artifact_digest", "artifact_id", + "selection_payload", "selection_digest", "selection_id", + } + raw = _exact_keys(value, keys, where) + artifact = _exact_keys( + raw["artifact_payload"], {"spec", "content_digest", "source"}, + f"{where}.artifact_payload", + ) + spec = _exact_keys( + artifact["spec"], + { + "benchmark", "split", "hf_path", "hf_subset", "source_revision", + "normalization_version", "transforms", "output_name", + }, + f"{where}.artifact_payload.spec", + ) + normalized_artifact = { + "spec": { + "benchmark": _nonempty(spec["benchmark"], f"{where}.artifact_payload.spec.benchmark"), + "split": _nonempty(spec["split"], f"{where}.artifact_payload.spec.split"), + "hf_path": _optional_string(spec["hf_path"], f"{where}.artifact_payload.spec.hf_path"), + "hf_subset": _optional_string(spec["hf_subset"], f"{where}.artifact_payload.spec.hf_subset"), + "source_revision": _optional_string(spec["source_revision"], f"{where}.artifact_payload.spec.source_revision"), + "normalization_version": _nonempty(spec["normalization_version"], f"{where}.artifact_payload.spec.normalization_version"), + "transforms": list(_strings(spec["transforms"], f"{where}.artifact_payload.spec.transforms", unique=True)), + "output_name": _optional_string(spec["output_name"], f"{where}.artifact_payload.spec.output_name"), + }, + "content_digest": _sha256(artifact["content_digest"], f"{where}.artifact_payload.content_digest"), + "source": _canonical_mapping(artifact["source"], f"{where}.artifact_payload.source"), + } + selection = _exact_keys( + raw["selection_payload"], + {"artifact_id", "content_digest", "sample_identities", "seed", "n_samples", "subject_filter"}, + f"{where}.selection_payload", + ) + normalized_selection = { + "artifact_id": _nonempty(selection["artifact_id"], f"{where}.selection_payload.artifact_id"), + "content_digest": _sha256(selection["content_digest"], f"{where}.selection_payload.content_digest"), + "sample_identities": list(_strings(selection["sample_identities"], f"{where}.selection_payload.sample_identities", unique=False)), + "seed": _strict_int(selection["seed"], f"{where}.selection_payload.seed"), + "n_samples": _optional_int(selection["n_samples"], f"{where}.selection_payload.n_samples", minimum=1), + "subject_filter": list(_strings(selection["subject_filter"], f"{where}.selection_payload.subject_filter", unique=True)), + } + for index, digest in enumerate(normalized_selection["sample_identities"]): + _sha256(digest, f"{where}.selection_payload.sample_identities[{index}]") + if normalized_selection["subject_filter"] != sorted(normalized_selection["subject_filter"]): + raise ImportSpecError(f"{where}.selection_payload.subject_filter must be sorted.") + normalized = { + "artifact_payload": normalized_artifact, + "artifact_digest": _sha256(raw["artifact_digest"], f"{where}.artifact_digest"), + "artifact_id": _nonempty(raw["artifact_id"], f"{where}.artifact_id"), + "selection_payload": normalized_selection, + "selection_digest": _sha256(raw["selection_digest"], f"{where}.selection_digest"), + "selection_id": _nonempty(raw["selection_id"], f"{where}.selection_id"), + } + _validate_identity_claim( + normalized, payload_key="artifact_payload", digest_key="artifact_digest", + id_key="artifact_id", prefix="ds", where=where, + ) + if normalized_selection["artifact_id"] != normalized["artifact_id"]: + raise ImportSpecError(f"{where}.selection_payload.artifact_id is inconsistent.") + _validate_identity_claim( + normalized, payload_key="selection_payload", digest_key="selection_digest", + id_key="selection_id", prefix="sel", where=where, + ) + return normalized + + +def _validate_model_payload(value: Any, where: str) -> dict[str, Any]: + raw = _mapping(value, where) + backend = _enum(raw.get("backend"), {"dummy", "api", "huggingface"}, f"{where}.backend") + allowed = { + "dummy": {"backend", "model_name_or_path"}, + "api": {"backend", "provider", "model_name_or_path", "base_url", "generation_kwargs"}, + "huggingface": {"backend", "model", "device", "add_bos_token", "generation_kwargs", "loader"}, + }[backend] + _exact_keys(raw, allowed, where) + if backend == "dummy": + return { + "backend": backend, + "model_name_or_path": _nonempty( + raw["model_name_or_path"], f"{where}.model_name_or_path" + ), + } + generation_keys = ( + {"max_new_tokens", "temperature"} + if backend == "api" + else {"max_new_tokens", "temperature", "do_sample"} + ) + generation = _exact_keys( + raw["generation_kwargs"], generation_keys, f"{where}.generation_kwargs" + ) + normalized_generation: dict[str, Any] = { + "max_new_tokens": _strict_int( + generation["max_new_tokens"], + f"{where}.generation_kwargs.max_new_tokens", + minimum=1, + ), + "temperature": _strict_number( + generation["temperature"], f"{where}.generation_kwargs.temperature" + ), + } + if backend == "api": + return { + "backend": backend, + "provider": _nonempty(raw["provider"], f"{where}.provider"), + "model_name_or_path": _nonempty( + raw["model_name_or_path"], f"{where}.model_name_or_path" + ), + "base_url": _optional_string(raw["base_url"], f"{where}.base_url"), + "generation_kwargs": normalized_generation, + } + normalized_generation["do_sample"] = _strict_bool( + generation["do_sample"], f"{where}.generation_kwargs.do_sample" + ) + resolved = _mapping(raw["model"], f"{where}.model") + kind = _enum( + resolved.get("kind"), {"local", "huggingface-hub"}, f"{where}.model.kind" + ) + if kind == "local": + resolved = _exact_keys( + resolved, + {"kind", "logical_name", "content_digest", "file_count", "total_bytes"}, + f"{where}.model", + ) + normalized_model = { + "kind": kind, + "logical_name": _nonempty( + resolved["logical_name"], f"{where}.model.logical_name" + ), + "content_digest": _sha256( + resolved["content_digest"], f"{where}.model.content_digest" + ), + "file_count": _strict_int( + resolved["file_count"], f"{where}.model.file_count", minimum=1 + ), + "total_bytes": _strict_int( + resolved["total_bytes"], f"{where}.model.total_bytes", minimum=0 + ), + } + else: + resolved = _exact_keys( + resolved, + {"kind", "repo_id", "requested_revision", "resolved_commit"}, + f"{where}.model", + ) + resolved_commit = _nonempty( + resolved["resolved_commit"], f"{where}.model.resolved_commit" + ) + if not re.fullmatch(r"[0-9a-f]{40,64}", resolved_commit): + raise ImportSpecError( + f"{where}.model.resolved_commit must be a lowercase immutable commit." + ) + normalized_model = { + "kind": kind, + "repo_id": _nonempty(resolved["repo_id"], f"{where}.model.repo_id"), + "requested_revision": _optional_string( + resolved["requested_revision"], f"{where}.model.requested_revision" + ), + "resolved_commit": resolved_commit, + } + loader = _exact_keys( + raw["loader"], {"trust_remote_code", "torch_dtype"}, f"{where}.loader" + ) + return { + "backend": backend, + "model": normalized_model, + "device": _nonempty(raw["device"], f"{where}.device"), + "add_bos_token": _strict_bool(raw["add_bos_token"], f"{where}.add_bos_token"), + "generation_kwargs": normalized_generation, + "loader": { + "trust_remote_code": _strict_bool( + loader["trust_remote_code"], f"{where}.loader.trust_remote_code" + ), + "torch_dtype": _enum( + loader["torch_dtype"], + {"float16", "float32"}, + f"{where}.loader.torch_dtype", + ), + }, + } + + +_IMPLEMENTATION_KEYS = { + "qualified_name", "source_file", "source_digest", "distribution", + "distribution_version", "package_tree_digest", "package_file_count", + "callable_digest", +} + + +def _validate_implementation(value: Any, where: str) -> dict[str, Any]: + raw = _allowed_keys(value, _IMPLEMENTATION_KEYS, where) + if "qualified_name" not in raw: + raise ImportSpecError(f"Missing required field 'qualified_name' in {where}.") + result = _canonical_mapping(raw, where) + _nonempty(result["qualified_name"], f"{where}.qualified_name") + for key in ("source_file", "distribution", "distribution_version"): + if key in result: + _nonempty(result[key], f"{where}.{key}") + for key in ("source_digest", "package_tree_digest", "callable_digest"): + if key in result: + _sha256(result[key], f"{where}.{key}") + has_package_digest = "package_tree_digest" in result + has_package_count = "package_file_count" in result + if has_package_digest != has_package_count: + raise ImportSpecError( + f"{where}.package_tree_digest and {where}.package_file_count must appear together." + ) + if has_package_count: + _strict_int(result["package_file_count"], f"{where}.package_file_count", minimum=1) + return result + + +def _validate_method_payload(value: Any, where: str) -> dict[str, Any]: + raw = _exact_keys( + value, {"name", "effective_params", "preflight", "implementation"}, where + ) + preflight = raw["preflight"] + if preflight is not None: + preflight = _exact_keys( + preflight, {"source", "split", "n"}, f"{where}.preflight" + ) + preflight = { + "source": _nonempty(preflight["source"], f"{where}.preflight.source"), + "split": _nonempty(preflight["split"], f"{where}.preflight.split"), + "n": _strict_int(preflight["n"], f"{where}.preflight.n", minimum=1), + } + return { + "name": _nonempty(raw["name"], f"{where}.name"), + "effective_params": _canonical_mapping(raw["effective_params"], f"{where}.effective_params"), + "preflight": preflight, + "implementation": _validate_implementation(raw["implementation"], f"{where}.implementation"), + } + + +def _validate_prompt_payload(value: Any, where: str) -> dict[str, Any]: + raw = _exact_keys(value, {"version", "files"}, where) + files_raw = _mapping(raw["files"], f"{where}.files") + template_names = {"direct_mcq", "free_text", "option_matching"} + if set(files_raw) != template_names: + raise ImportSpecError( + f"{where}.files must contain exactly the current template names " + f"{sorted(template_names)}." + ) + files: dict[str, Any] = {} + for name in ("direct_mcq", "free_text", "option_matching"): + item = files_raw[name] + record = _exact_keys(item, {"sha256", "content"}, f"{where}.files.{name}") + content = record["content"] + if not isinstance(content, str): + raise ImportSpecError(f"{where}.files.{name}.content must be a string.") + digest = _sha256(record["sha256"], f"{where}.files.{name}.sha256") + if digest != integrity_digest(content): + raise ImportSpecError(f"{where}.files.{name}.sha256 does not match content.") + files[name] = {"sha256": digest, "content": content} + return {"version": _nonempty(raw["version"], f"{where}.version"), "files": files} + + +def _validate_prompt_native(value: Any, where: str) -> dict[str, Any]: + raw = _exact_keys(value, {"prompt_id", "version", "files"}, where) + payload = _validate_prompt_payload( + {"version": raw["version"], "files": raw["files"]}, where + ) + prompt_id = _nonempty(raw["prompt_id"], f"{where}.prompt_id") + expected_id = f"prompt_{integrity_digest(payload)[:16]}" + if prompt_id != expected_id: + raise ImportSpecError( + f"{where}.prompt_id does not match prompt_bundle_identity semantics." + ) + return {"prompt_id": prompt_id, **payload} + + +def validate_csv_dialect_identity(value: Any) -> dict[str, Any]: + """Validate and normalize an identity-bearing CSV dialect declaration.""" + prepared = dict(value) if isinstance(value, Mapping) else value + if isinstance(prepared, dict) and isinstance( + prepared.get("line_terminators"), tuple + ): + prepared["line_terminators"] = list(prepared["line_terminators"]) + return asdict(_build_dialect(prepared, "parsing_policy.dialect")) + + +def validate_numeric_columns_identity(value: Any) -> list[dict[str, Any]]: + """Validate identity-bearing numeric-column declarations.""" + if not isinstance(value, (list, tuple)): + raise ImportSpecError("parsing_policy.numeric_columns must be a list.") + records = [ + asdict(_build_numeric(item, f"parsing_policy.numeric_columns[{index}]")) + for index, item in enumerate(value) + ] + columns = [record["source_column"] for record in records] + if len(columns) != len(set(columns)): + raise ImportSpecError( + "parsing_policy.numeric_columns contains duplicate source columns." + ) + return records + + +def validate_option_mapping_identity(value: Any) -> dict[str, Any]: + """Validate and normalize an identity-bearing option mapping.""" + prepared = dict(value) if isinstance(value, Mapping) else value + if isinstance(prepared, dict) and isinstance( + prepared.get("ordered_columns"), tuple + ): + prepared["ordered_columns"] = list(prepared["ordered_columns"]) + return asdict(_build_option(prepared, "parsing_policy.option_mapping")) + + +def validate_native_model_payload(value: Any) -> dict[str, Any]: + """Validate a claimed current-native model identity payload.""" + return _validate_model_payload(value, "native_model.payload") + + +def validate_native_method_payload(value: Any) -> dict[str, Any]: + """Validate a claimed current-native method identity payload.""" + return _validate_method_payload(value, "native_method.payload") + + +def validate_implementation_identity_record(value: Any) -> dict[str, Any]: + """Validate the closed runtime implementation-identity record shape.""" + return _validate_implementation(value, "implementation") + + +def _validate_simple_native( + value: Any, where: str, prefix: str, payload_validator +) -> dict[str, Any]: + id_key = f"{prefix}_id" + raw = _exact_keys(value, {"payload", "digest", id_key}, where) + normalized = { + "payload": payload_validator(raw["payload"], f"{where}.payload"), + "digest": _sha256(raw["digest"], f"{where}.digest"), + id_key: _nonempty(raw[id_key], f"{where}.{id_key}"), + } + _validate_identity_claim( + normalized, payload_key="payload", digest_key="digest", id_key=id_key, + prefix=prefix, where=where, + ) + return normalized + + +def _build_dataset(value: Any, index: int) -> DatasetReferenceSpec: + where = f"datasets[{index}]" + raw = _exact_keys(value, _DATASET_KEYS, where) + native = raw["native_compatibility_identity"] + selection_seed = _optional_int(raw["selection_seed"], f"{where}.selection_seed") + selection_n_samples = _optional_int( + raw["selection_n_samples"], f"{where}.selection_n_samples", minimum=1 + ) + selection_unknown_reasons = _string_mapping( + raw["selection_unknown_reasons"], f"{where}.selection_unknown_reasons" + ) + _validate_unknown_reasons( + { + "selection_seed": selection_seed, + "selection_n_samples": selection_n_samples, + }, + selection_unknown_reasons, + f"{where}.selection_unknown_reasons", + ) + return DatasetReferenceSpec( + dataset_id=_nonempty(raw["dataset_id"], f"{where}.dataset_id"), + benchmark_name=_nonempty(raw["benchmark_name"], f"{where}.benchmark_name"), + split=_nonempty(raw["split"], f"{where}.split"), + reference_kind=_enum( + raw["reference_kind"], + {"independent_input_snapshot", "profile_derived_reference_snapshot"}, + f"{where}.reference_kind", + ), + trust_label=_nonempty(raw["trust_label"], f"{where}.trust_label"), + source_ids=_strings(raw["source_ids"], f"{where}.source_ids", nonempty=True, unique=True), + selection_source_id=_nonempty(raw["selection_source_id"], f"{where}.selection_source_id"), + expected_question_ids=_strings(raw["expected_question_ids"], f"{where}.expected_question_ids", nonempty=True, unique=True), + selection_seed=selection_seed, + selection_n_samples=selection_n_samples, + subject_filter=_strings(raw["subject_filter"], f"{where}.subject_filter", unique=True), + selection_unknown_reasons=selection_unknown_reasons, + columns=_string_mapping(raw["columns"], f"{where}.columns"), + revision=_optional_string(raw["revision"], f"{where}.revision"), + fingerprint=_optional_string(raw["fingerprint"], f"{where}.fingerprint"), + derivation=_canonical_mapping(raw["derivation"], f"{where}.derivation"), + limitations=_strings(raw["limitations"], f"{where}.limitations"), + native_compatibility_identity=( + None if native is None else _validate_dataset_native(native, f"{where}.native_compatibility_identity") + ), + ) + + +def _build_model(value: Any, index: int) -> ImportModelSpec: + where = f"models[{index}]" + raw = _exact_keys(value, _MODEL_KEYS, where) + native = raw["native_compatibility_identity"] + backend = _optional_string(raw["backend"], f"{where}.backend") + provider = _optional_string(raw["provider"], f"{where}.provider") + revision = _optional_string(raw["revision"], f"{where}.revision") + unknown_reasons = _string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons") + _validate_unknown_reasons( + {"backend": backend, "provider": provider, "revision": revision}, + unknown_reasons, + f"{where}.unknown_reasons", + ) + return ImportModelSpec( + model_key=_nonempty(raw["model_key"], f"{where}.model_key"), + display_name=_nonempty(raw["display_name"], f"{where}.display_name"), + backend=backend, + provider=provider, + revision=revision, + effective_parameters=_canonical_mapping(raw["effective_parameters"], f"{where}.effective_parameters"), + unknown_reasons=unknown_reasons, + native_compatibility_identity=( + None if native is None else _validate_simple_native( + native, f"{where}.native_compatibility_identity", "model", _validate_model_payload + ) + ), + ) + + +def _build_method(value: Any, index: int) -> ImportMethodSpec: + where = f"methods[{index}]" + raw = _exact_keys(value, _METHOD_KEYS, where) + implementation = raw["implementation"] + native = raw["native_compatibility_identity"] + implementation = ( + None + if implementation is None + else _validate_implementation(implementation, f"{where}.implementation") + ) + unknown_reasons = _string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons") + _validate_unknown_reasons( + {"implementation": implementation}, + unknown_reasons, + f"{where}.unknown_reasons", + ) + return ImportMethodSpec( + method_key=_nonempty(raw["method_key"], f"{where}.method_key"), + name=_nonempty(raw["name"], f"{where}.name"), + effective_parameters=_canonical_mapping(raw["effective_parameters"], f"{where}.effective_parameters"), + implementation=implementation, + unknown_reasons=unknown_reasons, + native_compatibility_identity=( + None if native is None else _validate_simple_native( + native, f"{where}.native_compatibility_identity", "method", _validate_method_payload + ) + ), + ) + + +def _build_prompt(value: Any, index: int) -> ImportPromptSpec: + where = f"prompts[{index}]" + raw = _exact_keys(value, _PROMPT_KEYS, where) + contents = raw["template_contents"] + native = raw["native_compatibility_identity"] + template_identity = _optional_string( + raw["template_identity"], f"{where}.template_identity" + ) + template_digest = _optional_sha256( + raw["template_digest"], f"{where}.template_digest" + ) + template_contents = ( + None + if contents is None + else _string_mapping(contents, f"{where}.template_contents") + ) + unknown_reason = _optional_string(raw["unknown_reason"], f"{where}.unknown_reason") + has_unknown = any( + item is None for item in (template_identity, template_digest, template_contents) + ) + if has_unknown and unknown_reason is None: + raise ImportSpecError( + f"{where}.unknown_reason must explain unrecoverable prompt fields." + ) + if not has_unknown and unknown_reason is not None: + raise ImportSpecError( + f"{where}.unknown_reason contradicts fully known prompt fields." + ) + return ImportPromptSpec( + prompt_key=_nonempty(raw["prompt_key"], f"{where}.prompt_key"), + template_identity=template_identity, + template_digest=template_digest, + template_contents=template_contents, + unknown_reason=unknown_reason, + native_compatibility_identity=( + None + if native is None + else _validate_prompt_native(native, f"{where}.native_compatibility_identity") + ), + ) + + +def _build_result_origin(value: Any, where: str) -> ResultOriginSpec: + raw = _exact_keys(value, _RESULT_ORIGIN_KEYS, where) + default = raw["default_prediction_origin"] + per_question_raw = _mapping( + raw["per_question_prediction_origins"], + f"{where}.per_question_prediction_origins", + ) + return ResultOriginSpec( + derivation_origin=_enum(raw["derivation_origin"], _DERIVATION_ORIGINS, f"{where}.derivation_origin"), + default_prediction_origin=(None if default is None else _enum(default, _PREDICTION_ORIGINS, f"{where}.default_prediction_origin")), + per_question_prediction_origins={ + _nonempty(key, f"{where}.per_question_prediction_origins key"): _enum( + item, _PREDICTION_ORIGINS, f"{where}.per_question_prediction_origins.{key}" + ) + for key, item in per_question_raw.items() + }, + ) + + +def _build_condition(value: Any, index: int) -> ImportConditionSpec: + where = f"conditions[{index}]" + raw = _exact_keys(value, _CONDITION_KEYS, where) + seed = _optional_int(raw["seed"], f"{where}.seed") + calibration_identity = _canonical_mapping_or_none( + raw["calibration_identity"], f"{where}.calibration_identity" + ) + preflight_identity = _canonical_mapping_or_none( + raw["preflight_identity"], f"{where}.preflight_identity" + ) + unknown_reasons = _string_mapping(raw["unknown_reasons"], f"{where}.unknown_reasons") + protocol_settings = _canonical_mapping( + raw["protocol_settings"], f"{where}.protocol_settings" + ) + unsupported_protocol = sorted( + set(protocol_settings) - _PROTOCOL_SETTING_KEYS + ) + if unsupported_protocol: + raise ImportSpecError( + f"{where}.protocol_settings contains unsupported field(s) " + f"{unsupported_protocol}." + ) + if "pride_modal_k_threshold" in protocol_settings: + _strict_number( + protocol_settings["pride_modal_k_threshold"], + f"{where}.protocol_settings.pride_modal_k_threshold", + ) + _validate_unknown_reasons( + { + "seed": seed, + "calibration_identity": calibration_identity, + "preflight_identity": preflight_identity, + }, + unknown_reasons, + f"{where}.unknown_reasons", + ) + return ImportConditionSpec( + condition_key=_nonempty(raw["condition_key"], f"{where}.condition_key"), + source_ids=_strings(raw["source_ids"], f"{where}.source_ids", nonempty=True, unique=True), + dataset_id=_nonempty(raw["dataset_id"], f"{where}.dataset_id"), + model_key=_nonempty(raw["model_key"], f"{where}.model_key"), + method_key=_nonempty(raw["method_key"], f"{where}.method_key"), + prompt_key=_nonempty(raw["prompt_key"], f"{where}.prompt_key"), + seed=seed, + calibration_identity=calibration_identity, + preflight_identity=preflight_identity, + protocol_settings=protocol_settings, + generation_parameters=_canonical_mapping(raw["generation_parameters"], f"{where}.generation_parameters"), + unknown_reasons=unknown_reasons, + expected_question_ids=_strings(raw["expected_question_ids"], f"{where}.expected_question_ids", nonempty=True, unique=True), + evidence_status=_enum(raw["evidence_status"], _EVIDENCE_STATUSES, f"{where}.evidence_status"), + scope_disposition=_enum(raw["scope_disposition"], _SCOPE_DISPOSITIONS, f"{where}.scope_disposition"), + executable=_optional_bool(raw["executable"], f"{where}.executable"), + qualifications=_mapping_sequence(raw["qualifications"], f"{where}.qualifications"), + limitations=_mapping_sequence(raw["limitations"], f"{where}.limitations"), + damaged_question_ids=_strings(raw["damaged_question_ids"], f"{where}.damaged_question_ids", unique=True), + recoverable_question_ids=_strings(raw["recoverable_question_ids"], f"{where}.recoverable_question_ids", unique=True), + result_origin=_build_result_origin(raw["result_origin"], f"{where}.result_origin"), + ) + + +def _build_authorization(value: Any, index: int) -> AuthorizationSpec: + where = f"authorizations[{index}]" + raw = _exact_keys(value, _AUTHORIZATION_KEYS, where) + reasons_raw = _mapping(raw["condition_question_reasons"], f"{where}.condition_question_reasons") + reasons = { + _nonempty(condition, f"{where}.condition_question_reasons key"): _string_mapping( + question_reasons, f"{where}.condition_question_reasons.{condition}" + ) + for condition, question_reasons in reasons_raw.items() + } + authorization_type = _enum( + raw["authorization_type"], + {"inference_repair", "offline_transformation"}, + f"{where}.authorization_type", + ) + executable = _strict_bool(raw["executable"], f"{where}.executable") + expected_executable = authorization_type == "inference_repair" + if executable is not expected_executable: + raise ImportSpecError( + f"{where}.authorization_type={authorization_type!r} requires " + f"executable={expected_executable!r}." + ) + return AuthorizationSpec( + authorization_id=_nonempty(raw["authorization_id"], f"{where}.authorization_id"), + authorization_type=authorization_type, + source_id=_nonempty(raw["source_id"], f"{where}.source_id"), + condition_question_reasons=reasons, + authority=_nonempty(raw["authority"], f"{where}.authority"), + purpose=_nonempty(raw["purpose"], f"{where}.purpose"), + executable=executable, + input_evidence_digests=_sha_mapping(raw["input_evidence_digests"], f"{where}.input_evidence_digests"), + expected_snapshot_digests=_sha_mapping(raw["expected_snapshot_digests"], f"{where}.expected_snapshot_digests"), + ) + + +def _build_overlay(value: Any, index: int) -> OverlaySpec: + where = f"overlays[{index}]" + raw = _exact_keys(value, _OVERLAY_KEYS, where) + return OverlaySpec( + overlay_id=_nonempty(raw["overlay_id"], f"{where}.overlay_id"), + base_run_path=Path(_nonempty(raw["base_run_path"], f"{where}.base_run_path")), + base_condition_digest=_sha256(raw["base_condition_digest"], f"{where}.base_condition_digest"), + base_realization_id=_nonempty(raw["base_realization_id"], f"{where}.base_realization_id"), + base_realization_digest=_sha256(raw["base_realization_digest"], f"{where}.base_realization_digest"), + base_evidence_digests=_sha_mapping(raw["base_evidence_digests"], f"{where}.base_evidence_digests"), + base_validation_artifact_sha256=_sha256(raw["base_validation_artifact_sha256"], f"{where}.base_validation_artifact_sha256"), + base_result_sha256=_optional_sha256(raw["base_result_sha256"], f"{where}.base_result_sha256"), + source_id=_nonempty(raw["source_id"], f"{where}.source_id"), + authorization_id=_nonempty(raw["authorization_id"], f"{where}.authorization_id"), + replacement_reasons=_string_mapping(raw["replacement_reasons"], f"{where}.replacement_reasons"), + result_origin=_build_result_origin(raw["result_origin"], f"{where}.result_origin"), + lineage_notes=_canonical_mapping(raw["lineage_notes"], f"{where}.lineage_notes"), + implementation=_canonical_mapping(raw["implementation"], f"{where}.implementation"), + input_digest=_sha256(raw["input_digest"], f"{where}.input_digest"), + preownership_output_digest=_sha256(raw["preownership_output_digest"], f"{where}.preownership_output_digest"), + expected_evidence_status=_enum(raw["expected_evidence_status"], _EVIDENCE_STATUSES, f"{where}.expected_evidence_status"), + ) + + +def _unique(items: tuple[Any, ...], attribute: str, where: str) -> set[str]: + values = [getattr(item, attribute) for item in items] + if len(values) != len(set(values)): + raise ImportSpecError(f"Duplicate {attribute} in {where}.") + return set(values) + + +def _require_references(spec: ImportSpec) -> None: + source_ids = _unique(spec.sources, "source_id", "sources") + dataset_ids = _unique(spec.datasets, "dataset_id", "datasets") + model_keys = _unique(spec.models, "model_key", "models") + method_keys = _unique(spec.methods, "method_key", "methods") + prompt_keys = _unique(spec.prompts, "prompt_key", "prompts") + condition_keys = _unique(spec.conditions, "condition_key", "conditions") + authorization_ids = _unique(spec.authorizations, "authorization_id", "authorizations") + _unique(spec.overlays, "overlay_id", "overlays") + + for dataset in spec.datasets: + missing = sorted(set(dataset.source_ids) - source_ids) + if missing: + raise ImportSpecError( + f"Dataset {dataset.dataset_id!r} has unknown source reference(s) {missing}." + ) + if dataset.selection_source_id not in dataset.source_ids: + raise ImportSpecError( + f"Dataset {dataset.dataset_id!r} selection_source_id must reference one of its source_ids." + ) + dataset_by_id = {item.dataset_id: item for item in spec.datasets} + conditions_by_key = {item.condition_key: item for item in spec.conditions} + for condition in spec.conditions: + references = ( + (condition.dataset_id, dataset_ids, "dataset"), + (condition.model_key, model_keys, "model"), + (condition.method_key, method_keys, "method"), + (condition.prompt_key, prompt_keys, "prompt"), + ) + for reference, known, label in references: + if reference not in known: + raise ImportSpecError( + f"Condition {condition.condition_key!r} has unknown {label} reference {reference!r}." + ) + missing_sources = sorted(set(condition.source_ids) - source_ids) + if missing_sources: + raise ImportSpecError( + f"Condition {condition.condition_key!r} has unknown source reference(s) {missing_sources}." + ) + dataset_questions = set(dataset_by_id[condition.dataset_id].expected_question_ids) + expected = set(condition.expected_question_ids) + if not expected <= dataset_questions: + raise ImportSpecError( + f"Condition {condition.condition_key!r} expects question IDs absent from its dataset." + ) + damaged = set(condition.damaged_question_ids) + recoverable = set(condition.recoverable_question_ids) + if not damaged <= expected or not recoverable <= expected: + raise ImportSpecError( + f"Condition {condition.condition_key!r} damage references unknown question IDs." + ) + if not recoverable <= damaged: + raise ImportSpecError( + f"Condition {condition.condition_key!r} recoverable questions must also be damaged." + ) + origin_questions = set(condition.result_origin.per_question_prediction_origins) + if not origin_questions <= expected: + raise ImportSpecError( + f"Condition {condition.condition_key!r} has origins for unknown question IDs." + ) + evaluable = ( + condition.evidence_status in {"complete", "qualified"} + and condition.scope_disposition == "included" + ) + if ( + evaluable + and condition.result_origin.default_prediction_origin is None + and origin_questions != expected + ): + missing_origins = sorted(expected - origin_questions) + raise ImportSpecError( + f"Condition {condition.condition_key!r} has evaluable rows without a " + f"prediction origin: {missing_origins}." + ) + + for authorization in spec.authorizations: + if authorization.source_id not in source_ids: + raise ImportSpecError( + f"Authorization {authorization.authorization_id!r} has an unknown source reference." + ) + for condition_key, reasons in authorization.condition_question_reasons.items(): + if condition_key not in condition_keys: + raise ImportSpecError( + f"Authorization {authorization.authorization_id!r} has an unknown condition reference." + ) + allowed_questions = set(conditions_by_key[condition_key].expected_question_ids) + if not set(reasons) <= allowed_questions: + raise ImportSpecError( + f"Authorization {authorization.authorization_id!r} references unknown question IDs." + ) + for overlay in spec.overlays: + if overlay.source_id not in source_ids: + raise ImportSpecError(f"Overlay {overlay.overlay_id!r} has an unknown source reference.") + if overlay.authorization_id not in authorization_ids: + raise ImportSpecError( + f"Overlay {overlay.overlay_id!r} has an unknown authorization reference." + ) + + +def _build_import_spec(raw_value: Any) -> ImportSpec: + raw = _exact_keys(raw_value, _TOP_LEVEL_KEYS, "top level") + credentials = sorted(_credential_keys_in(raw)) + if credentials: + raise ImportSpecError( + f"Import specification contains credential-named field(s) {credentials}; " + "credentials must never be stored in import declarations." + ) + sources = tuple( + _build_source(item, index) + for index, item in enumerate(_sequence(raw["sources"], "sources")) + ) + datasets = tuple( + _build_dataset(item, index) + for index, item in enumerate(_sequence(raw["datasets"], "datasets")) + ) + models = tuple( + _build_model(item, index) + for index, item in enumerate(_sequence(raw["models"], "models")) + ) + methods = tuple( + _build_method(item, index) + for index, item in enumerate(_sequence(raw["methods"], "methods")) + ) + prompts = tuple( + _build_prompt(item, index) + for index, item in enumerate(_sequence(raw["prompts"], "prompts")) + ) + conditions = tuple( + _build_condition(item, index) + for index, item in enumerate(_sequence(raw["conditions"], "conditions")) + ) + authorizations = tuple( + _build_authorization(item, index) + for index, item in enumerate(_sequence(raw["authorizations"], "authorizations")) + ) + overlays = tuple( + _build_overlay(item, index) + for index, item in enumerate(_sequence(raw["overlays"], "overlays")) + ) + for name, values in ( + ("sources", sources), ("datasets", datasets), ("models", models), + ("methods", methods), ("prompts", prompts), ("conditions", conditions), + ): + if not values: + raise ImportSpecError(f"{name} must be a non-empty list.") + metrics = _strings(raw["metrics"], "metrics", nonempty=True, unique=True) + invalid_metrics = sorted(set(metrics) - set(BUILTIN_METRICS)) + if invalid_metrics: + raise ImportSpecError( + f"Import metrics must use the closed built-in registry; unknown entries: {invalid_metrics}. " + f"Built-ins: {sorted(BUILTIN_METRICS)}." + ) + spec = ImportSpec( + schema_version=_enum( + raw["schema_version"], {"choicebench.import-spec.v1"}, "schema_version" + ), + import_name=_nonempty(raw["import_name"], "import_name"), + sources=sources, + datasets=datasets, + models=models, + methods=methods, + prompts=prompts, + conditions=conditions, + authorizations=authorizations, + overlays=overlays, + metrics=metrics, + provenance=_build_provenance(raw["provenance"]), + audit=_canonical_mapping(raw["audit"], "audit"), + ) + _require_references(spec) + return spec + + +def load_import_spec(path: Path) -> ImportSpec: + """Load a YAML import declaration using a closed, non-coercing schema.""" + source_path = Path(path) + try: + raw = yaml.safe_load(source_path.read_text(encoding="utf-8")) + except yaml.YAMLError as exc: + raise ImportSpecError(f"Could not parse {source_path} as safe YAML: {exc}") from exc + except (OSError, UnicodeError) as exc: + raise ImportSpecError(f"Could not read import specification {source_path}: {exc}") from exc + if not isinstance(raw, Mapping): + raise ImportSpecError( + f"{source_path} must contain a YAML mapping at the top level." + ) + return _build_import_spec(raw) + + +def stable_import_projection(spec: ImportSpec) -> dict[str, Any]: + """Return the identity-bearing declaration without machine-local locations.""" + if not isinstance(spec, ImportSpec): + raise TypeError(f"spec must be an ImportSpec; got {type(spec).__name__}.") + projection = asdict(spec) + projection.pop("audit", None) + for source in projection["sources"]: + source.pop("path", None) + for overlay in projection["overlays"]: + overlay.pop("base_run_path", None) + return canonicalize(projection) + + +def import_spec_digest(spec: ImportSpec) -> str: + """Return the stable full SHA-256 identity of an import declaration.""" + return integrity_digest(stable_import_projection(spec)) diff --git a/src/choicebench/importing/transaction.py b/src/choicebench/importing/transaction.py new file mode 100644 index 0000000..6a335c2 --- /dev/null +++ b/src/choicebench/importing/transaction.py @@ -0,0 +1,158 @@ +"""Atomic, no-replace staged publication for imported runs. + +Reduced scope: covers the concrete safety guarantees the repair-import +workflow needs -- a manifest lock, sibling staging on the same filesystem, +full-tree fsync before publication, and a genuine no-replace rename (never +an existence-check-then-os.replace race) -- without the fuller historical +crash-recovery/cleanup-scope richness of the original plan. +""" + +from __future__ import annotations + +import ctypes +import ctypes.util +import os +from pathlib import Path +from typing import Callable +import shutil +import uuid + +from choicebench.infra.artifacts import FileLock + +_RENAME_NOREPLACE = 0x1 +_AT_FDCWD = -100 + + +class TransactionError(ValueError): + """Raised when staged publication cannot proceed safely.""" + + +def fsync_tree(root: Path) -> None: + """Best-effort fsync of every file and directory under root, so a crash + immediately after this call cannot lose or half-write staged content.""" + root = Path(root) + for dirpath, dirnames, filenames in os.walk(root): + for name in filenames: + path = Path(dirpath) / name + try: + fd = os.open(path, os.O_RDONLY) + except OSError: + continue + try: + os.fsync(fd) + finally: + os.close(fd) + try: + dir_fd = os.open(dirpath, os.O_RDONLY) + except OSError: + continue + try: + os.fsync(dir_fd) + except OSError: + pass + finally: + os.close(dir_fd) + + +def _renameat2_no_replace(old: Path, new: Path) -> None: + """Thin wrapper around the libc renameat2(2) syscall with + RENAME_NOREPLACE, isolated so tests can mock its absence without + depending on kernel version. Never falls back to os.replace or an + existence-check race: if the primitive is unavailable, this fails + closed.""" + library_name = ctypes.util.find_library("c") + if not library_name: + raise TransactionError("libc is not available; refusing an unsafe rename fallback.") + libc = ctypes.CDLL(library_name, use_errno=True) + if not hasattr(libc, "renameat2"): + raise TransactionError( + "renameat2 is not available on this system; refusing an unsafe rename fallback." + ) + result = libc.renameat2( + ctypes.c_int(_AT_FDCWD), + os.fsencode(str(old)), + ctypes.c_int(_AT_FDCWD), + os.fsencode(str(new)), + ctypes.c_uint(_RENAME_NOREPLACE), + ) + if result != 0: + errno = ctypes.get_errno() + raise TransactionError( + f"renameat2 failed renaming {old} -> {new}: errno {errno} ({os.strerror(errno)})." + ) + + +def atomic_publish_directory_no_replace(source: Path, destination: Path) -> None: + """Atomically move a fully-staged directory into its final location. + Refuses (never silently overwrites or races an existence check) if the + destination already exists.""" + source = Path(source) + destination = Path(destination) + destination.parent.mkdir(parents=True, exist_ok=True) + _renameat2_no_replace(source, destination) + + +class ImportTransaction: + """Stage a new run under a same-filesystem sibling directory, run a + caller-supplied full-graph validator, and either publish it atomically + (new run) or verify an already-published run in place (existing run) -- + never partially write the final run directory.""" + + def __init__(self, *, runs_dir: Path, run_id: str) -> None: + self.runs_dir = Path(runs_dir) + self.run_id = run_id + self.final_run = self.runs_dir / run_id + self._staged_run: Path | None = None + self._lock: FileLock | None = None + self._published = False + + def __enter__(self) -> "ImportTransaction": + lock_path = self.runs_dir / ".locks" / f"{self.run_id}.manifest.lock" + self._lock = FileLock(lock_path, f"import run {self.run_id}") + self._lock.__enter__() + if not self.final_run.exists(): + staged = self.runs_dir / f".staging-{self.run_id}-{uuid.uuid4().hex}" + staged.mkdir(parents=True) + (staged / ".staging-owner").write_text(f"pid={os.getpid()}\n") + self._staged_run = staged + return self + + @property + def staged_run(self) -> Path: + if self._staged_run is None: + raise TransactionError( + f"Run {self.run_id!r} already exists; there is no staging directory to " + "write into. Only verify_import_run-style validation is available." + ) + return self._staged_run + + def publish(self, validator: Callable[[Path], None]) -> Path: + if self.final_run.exists(): + validator(self.final_run) + return self.final_run + staged = self.staged_run + if not (staged / ".staging-owner").is_file(): + raise TransactionError( + f"Refusing to publish {staged}: it is not marked as this transaction's " + "own staging directory." + ) + validator(staged) + fsync_tree(staged) + (staged / ".staging-owner").unlink() + atomic_publish_directory_no_replace(staged, self.final_run) + self._published = True + self._staged_run = None + return self.final_run + + def __exit__(self, exc_type, exc, traceback) -> None: + try: + if self._staged_run is not None and self._staged_run.exists(): + # Publication never happened (error, or caller chose not to + # publish); clean up only this transaction's own + # owner-marked staging directory. + if (self._staged_run / ".staging-owner").is_file(): + shutil.rmtree(self._staged_run, ignore_errors=True) + finally: + if self._lock is not None: + self._lock.__exit__(exc_type, exc, traceback) + self._lock = None diff --git a/src/choicebench/importing/validation.py b/src/choicebench/importing/validation.py new file mode 100644 index 0000000..3ecae30 --- /dev/null +++ b/src/choicebench/importing/validation.py @@ -0,0 +1,494 @@ +"""Validate imported source rows by question identity and derive per-realization +evidence status and evaluable content. + +Reduced scope: this covers the concrete guarantees the repair-import workflow +needs (exact question-identity join, content cross-check against the expected +dataset snapshot, coverage/defect-driven evidence status with a fail-closed +declaration check, and evaluable-row derivation) without the full historical +option-count/method-specific-extension richness of the original plan. +""" + +from __future__ import annotations + +from dataclasses import dataclass +from hashlib import sha256 +import json +from pathlib import Path +from typing import Any, Mapping, Sequence + +from choicebench.identity import canonicalize, integrity_digest +from choicebench.importing.csv_adapter import AdaptedTable, LogicalRecordSpan, SourceRow +from choicebench.importing.dataset_reference import ExpectedDataset +from choicebench.importing.schema import ImportConditionSpec +from choicebench.infra.artifacts import atomic_write_bytes + +_EVALUABLE_STATUSES = {"complete", "qualified"} +_COVERAGE_ONLY_FINDING_CODES = {"MISSING_QUESTION_ID"} +VALIDATION_ARTIFACT_SCHEMA_VERSION = "choicebench.realization-validation.v1" + + +class ImportValidationError(ValueError): + """Raised when imported row content fails fail-closed validation.""" + + +@dataclass(frozen=True) +class ValidationFinding: + code: str + source_id: str + field: str | None + question_id: str | None + record_span: LogicalRecordSpan | None + raw_row_sha256: str | None + message: str + + def to_json(self) -> dict[str, Any]: + return { + "code": self.code, + "source_id": self.source_id, + "field": self.field, + "question_id": self.question_id, + "record_span": ( + None + if self.record_span is None + else { + "record_index": self.record_span.index, + "start": self.record_span.start, + "end": self.record_span.end, + "terminator_hex": self.record_span.terminator.hex(), + } + ), + "raw_row_sha256": self.raw_row_sha256, + "message": self.message, + } + + +@dataclass(frozen=True) +class ValidatedSourceRows: + source_id: str + rows_by_question_id: Mapping[str, SourceRow] + expected_question_ids: tuple[str, ...] + findings: tuple[ValidationFinding, ...] + findings_digest: str + + +def _expected_row(expected: ExpectedDataset, question_id: str) -> Mapping[str, Any] | None: + matches = expected.frame[expected.frame["question_id"].astype(str) == question_id] + if matches.empty: + return None + return matches.iloc[0].to_dict() + + +def _expected_choice_letters(expected_row: Mapping[str, Any]) -> list[str]: + """The valid answer letters for a question, derived from its ordered + choices_json array (index 0 -> "A", 1 -> "B", ...) -- the same ordering + ChoiceBench's own option map already uses to define correct_option.""" + raw = expected_row.get("choices_json") + if not raw: + return [] + try: + choices = json.loads(raw) + except (TypeError, ValueError): + return [] + if not isinstance(choices, list): + return [] + return [chr(ord("A") + index) for index in range(len(choices))] + + +def validate_source_rows( + table: AdaptedTable, + *, + source_id: str, + mapping: Mapping[str, str], + condition: ImportConditionSpec, + expected: ExpectedDataset, +) -> ValidatedSourceRows: + """Join source rows to the expected dataset by exact string question_id. + + Every declared field is compared against the expected snapshot; source + values never override it. No row is dropped, padded, deduplicated, + reordered, or repaired -- every input row is either kept (with any + findings recorded alongside it) or excluded only when its question_id is + genuinely ambiguous (duplicated within this source). + """ + if "question_id" not in mapping: + raise ImportValidationError(f"{source_id}: mapping declares no question_id column.") + question_id_column = mapping["question_id"] + expected_ids = tuple(condition.expected_question_ids) + expected_id_set = set(expected_ids) + + findings: list[ValidationFinding] = [] + rows_by_question_id: dict[str, SourceRow] = {} + rows_by_raw_id: dict[str, list[SourceRow]] = {} + + for row in table.rows: + raw_qid = row.values.get(question_id_column) + if raw_qid is None or not str(raw_qid).strip(): + findings.append( + ValidationFinding( + code="NULL_QUESTION_ID", + source_id=source_id, + field=question_id_column, + question_id=None, + record_span=row.span, + raw_row_sha256=row.raw_sha256, + message="Row has a null or empty question_id.", + ) + ) + continue + rows_by_raw_id.setdefault(str(raw_qid), []).append(row) + + for qid, rows in rows_by_raw_id.items(): + if len(rows) > 1: + for row in rows: + findings.append( + ValidationFinding( + code="DUPLICATE_QUESTION_ID", + source_id=source_id, + field=question_id_column, + question_id=qid, + record_span=row.span, + raw_row_sha256=row.raw_sha256, + message=f"question_id {qid!r} occurs {len(rows)} times in this source.", + ) + ) + continue + row = rows[0] + rows_by_question_id[qid] = row + if qid not in expected_id_set: + findings.append( + ValidationFinding( + code="UNEXPECTED_QUESTION_ID", + source_id=source_id, + field=question_id_column, + question_id=qid, + record_span=row.span, + raw_row_sha256=row.raw_sha256, + message=f"question_id {qid!r} is not in the expected dataset selection.", + ) + ) + continue + expected_row = _expected_row(expected, qid) + if expected_row is None: + continue + for semantic_field in ("question_text", "correct_option"): + source_column = mapping.get(semantic_field) + if source_column is None: + continue + source_value = row.values.get(source_column) + expected_value = expected_row.get(semantic_field) + normalized_source = None if source_value is None else str(source_value).strip() + normalized_expected = None if expected_value is None else str(expected_value).strip() + if semantic_field == "correct_option": + normalized_source = normalized_source and normalized_source.upper() + normalized_expected = normalized_expected and normalized_expected.upper() + if normalized_source is None or normalized_source != normalized_expected: + findings.append( + ValidationFinding( + code=f"{semantic_field.upper()}_MISMATCH", + source_id=source_id, + field=source_column, + question_id=qid, + record_span=row.span, + raw_row_sha256=row.raw_sha256, + message=f"{semantic_field} does not match the expected dataset snapshot.", + ) + ) + expected_choice_letters = _expected_choice_letters(expected_row) + prediction_column = mapping.get("prediction") + if prediction_column is not None: + prediction_value = row.values.get(prediction_column) + if prediction_value is None or not str(prediction_value).strip(): + findings.append( + ValidationFinding( + code="MISSING_PREDICTION", + source_id=source_id, + field=prediction_column, + question_id=qid, + record_span=row.span, + raw_row_sha256=row.raw_sha256, + message="Row has no parsed prediction.", + ) + ) + else: + letter = str(prediction_value).strip().upper() + if letter not in expected_choice_letters: + findings.append( + ValidationFinding( + code="INVALID_PREDICTION", + source_id=source_id, + field=prediction_column, + question_id=qid, + record_span=row.span, + raw_row_sha256=row.raw_sha256, + message=( + f"Parsed prediction {prediction_value!r} is not one of " + f"this question's options {expected_choice_letters}." + ), + ) + ) + + for qid in sorted(expected_id_set - set(rows_by_question_id)): + findings.append( + ValidationFinding( + code="MISSING_QUESTION_ID", + source_id=source_id, + field=None, + question_id=qid, + record_span=None, + raw_row_sha256=None, + message=f"Expected question_id {qid!r} is absent from this source.", + ) + ) + + findings_tuple = tuple(findings) + findings_digest = integrity_digest([finding.to_json() for finding in findings_tuple]) + return ValidatedSourceRows( + source_id=source_id, + rows_by_question_id=rows_by_question_id, + expected_question_ids=expected_ids, + findings=findings_tuple, + findings_digest=findings_digest, + ) + + +@dataclass(frozen=True) +class RealizationValidation: + evaluable_rows: tuple[Mapping[str, Any], ...] + findings: tuple[ValidationFinding, ...] + validation_digest: str + computed_evidence_status: str + evaluable: bool + qualifications: tuple[Mapping[str, Any], ...] + limitations: tuple[Mapping[str, Any], ...] + defect_question_ids: tuple[str, ...] + # The full canonical row set (every expected question ID present exactly + # once), populated whenever the realization is *structurally complete* and + # in scope -- independent of whether its current content is evaluable. A + # malformed-but-structurally-complete cell (all IDs present, but some rows + # carry an unparseable prediction) is NOT evaluable yet still has a full, + # publishable, overlay-able row set. Empty for a genuinely incomplete + # (missing/duplicate/unexpected IDs) or out-of-scope realization. + published_rows: tuple[Mapping[str, Any], ...] = () + structurally_complete: bool = False + + +def compute_evidence_status( + validated: ValidatedSourceRows, *, condition: ImportConditionSpec +) -> tuple[str, tuple[str, ...]]: + """Pure evidence-status computation from exact row coverage/defects, + independent of what the condition itself declares. Returns + (computed_status, defect_question_ids). Exposed separately from + normalize_realization_rows so a caller assembling declarations (e.g. a + paper-freeze profile) can discover the correct status to declare before + constructing the condition, instead of guessing and retrying against the + declaration-mismatch check below. + """ + expected_ids = set(validated.expected_question_ids) + present_ids = set(validated.rows_by_question_id) + missing_ids = expected_ids - present_ids + defect_question_ids = tuple( + sorted( + { + finding.question_id + for finding in validated.findings + if finding.question_id is not None + and finding.code not in _COVERAGE_ONLY_FINDING_CODES + } + ) + ) + recoverable_ids = set(condition.recoverable_question_ids) + + if not present_ids and expected_ids: + computed = "failed" + elif defect_question_ids: + computed = "malformed" + elif missing_ids and missing_ids <= recoverable_ids and recoverable_ids: + computed = "recoverable" + elif missing_ids: + computed = "partial" + else: + computed = "qualified" if condition.qualifications else "complete" + return computed, defect_question_ids + + +def normalize_realization_rows( + validated: ValidatedSourceRows, + *, + condition: ImportConditionSpec, + expected: ExpectedDataset, + mapping: Mapping[str, str] | None = None, + experiment_id: str | None = None, + realization_id: str | None = None, +) -> RealizationValidation: + """Recompute evidence status from exact coverage/defects and refuse a + declaration mismatch. Fill evaluative content only from ExpectedDataset; + source values were already checked, never trusted, in validate_source_rows. + """ + computed, defect_question_ids = compute_evidence_status(validated, condition=condition) + + if computed != condition.evidence_status: + raise ImportValidationError( + f"Computed evidence status {computed!r} does not match the declared " + f"evidence status {condition.evidence_status!r}." + ) + + evaluable = computed in _EVALUABLE_STATUSES and condition.scope_disposition == "included" + + # Structural completeness (a data-identity property) is independent of + # evaluability (a scoring property): every expected question ID is present + # exactly once. Duplicates are dropped from rows_by_question_id (so a + # duplicated ID leaves its slot empty) and unexpected IDs add a member the + # expected set lacks -- either makes the present set differ from expected. + structurally_complete = set(validated.rows_by_question_id) == set( + validated.expected_question_ids + ) + # A structurally-complete, in-scope realization has a full canonical row + # set that can be published and later overlaid even when not evaluable. + publishable = structurally_complete and condition.scope_disposition == "included" + + prediction_column = None if mapping is None else mapping.get("prediction") + + def _normalized_rows() -> list[dict[str, Any]]: + per_question_origin = condition.result_origin.per_question_prediction_origins + default_origin = condition.result_origin.default_prediction_origin + rows: list[dict[str, Any]] = [] + for qid in validated.expected_question_ids: + if qid not in validated.rows_by_question_id: + continue + expected_row = _expected_row(expected, qid) + origin = per_question_origin.get(qid, default_origin) + if origin is None: + raise ImportValidationError( + f"No declared prediction_origin for question {qid!r}." + ) + predicted_option = None + if prediction_column is not None: + raw_prediction = validated.rows_by_question_id[qid].values.get(prediction_column) + predicted_option = None if raw_prediction is None else str(raw_prediction).strip().upper() + rows.append( + { + "question_id": qid, + "question_text": expected_row.get("question_text"), + "correct_option": expected_row.get("correct_option"), + "choices_json": expected_row.get("choices_json"), + "prediction_origin": origin, + "predicted_option": predicted_option, + } + ) + return rows + + published_rows = _normalized_rows() if publishable else [] + # An evaluable realization is always structurally complete and in scope, so + # its published row set is exactly its evaluable row set. Keeping + # evaluable_rows empty for a non-evaluable (e.g. malformed) realization + # preserves the existing scoring contract; published_rows carries the full + # set separately for the publish/overlay path. + evaluable_rows: list[dict[str, Any]] = list(published_rows) if evaluable else [] + + validation_payload = { + "schema_version": "choicebench.realization-validation-content.v1", + "computed_evidence_status": computed, + "evaluable": evaluable, + "findings_digest": validated.findings_digest, + "evaluable_rows": canonicalize(evaluable_rows), + } + validation_digest = integrity_digest(validation_payload) + + return RealizationValidation( + evaluable_rows=tuple(evaluable_rows), + findings=validated.findings, + validation_digest=validation_digest, + computed_evidence_status=computed, + evaluable=evaluable, + qualifications=tuple(condition.qualifications), + limitations=tuple(condition.limitations), + defect_question_ids=defect_question_ids, + published_rows=tuple(published_rows), + structurally_complete=structurally_complete, + ) + + +@dataclass(frozen=True) +class PreparedValidationArtifact: + relative_path: str + json_bytes: bytes + file_sha256: str + validation_digest: str + + +def prepare_realization_validation_artifact( + validation: RealizationValidation, + *, + realization: Mapping[str, Any], + evidence_records: Sequence[Mapping[str, Any]], +) -> PreparedValidationArtifact: + realization_id = realization["realization_id"] + payload = canonicalize( + { + "schema_version": VALIDATION_ARTIFACT_SCHEMA_VERSION, + "realization_id": realization_id, + "realization_digest": realization["realization_digest"], + "validation_digest": validation.validation_digest, + "computed_evidence_status": validation.computed_evidence_status, + "evaluable": validation.evaluable, + "qualifications": list(validation.qualifications), + "limitations": list(validation.limitations), + "defect_question_ids": list(validation.defect_question_ids), + "findings": [finding.to_json() for finding in validation.findings], + "evidence_records": list(evidence_records), + } + ) + json_bytes = ( + json.dumps(payload, indent=2, sort_keys=True, allow_nan=False) + "\n" + ).encode("utf-8") + return PreparedValidationArtifact( + relative_path=f"artifacts/imports/validation/{realization_id}.json", + json_bytes=json_bytes, + file_sha256=sha256(json_bytes).hexdigest(), + validation_digest=validation.validation_digest, + ) + + +def write_realization_validation_artifact( + staged_run: Path, prepared: PreparedValidationArtifact +) -> tuple[Path, bool]: + path = Path(staged_run) / prepared.relative_path + if path.exists(): + existing = path.read_bytes() + if existing == prepared.json_bytes: + return path, True + raise RuntimeError(f"Refusing to overwrite a divergent validation artifact: {path}") + atomic_write_bytes(path, prepared.json_bytes) + return path, False + + +def validate_realization_validation_artifact( + path: Path, + *, + manifest: Mapping[str, Any], + realization: Mapping[str, Any], + expected_sha256: str, +) -> Mapping[str, Any]: + path = Path(path) + try: + actual_bytes = path.read_bytes() + except OSError as exc: + raise RuntimeError(f"Validation artifact is unreadable: {path}: {exc}") from exc + actual_sha256 = sha256(actual_bytes).hexdigest() + if actual_sha256 != expected_sha256: + raise RuntimeError( + f"Validation artifact content integrity check failed: {path}; expected " + f"{expected_sha256}, found {actual_sha256}." + ) + try: + payload = json.loads(actual_bytes) + except json.JSONDecodeError as exc: + raise RuntimeError(f"Validation artifact is not valid JSON: {path}: {exc}") from exc + if payload.get("schema_version") != VALIDATION_ARTIFACT_SCHEMA_VERSION: + raise RuntimeError(f"Unsupported validation artifact schema: {path}") + if ( + payload.get("realization_id") != realization.get("realization_id") + or payload.get("realization_digest") != realization.get("realization_digest") + ): + raise RuntimeError(f"Validation artifact realization binding is inconsistent: {path}") + return payload diff --git a/src/choicebench/io/readers.py b/src/choicebench/io/readers.py index 0cf0781..3504fe2 100644 --- a/src/choicebench/io/readers.py +++ b/src/choicebench/io/readers.py @@ -1,10 +1,17 @@ # src/choicebench/io/readers.py +from dataclasses import dataclass from pathlib import Path +from typing import Any, Literal, Mapping, Sequence import json import pandas as pd -from choicebench.manifest import ManifestCompatibilityError, validate_manifest, validate_run_state +from choicebench.manifest import ( + RUN_STATE_FILENAME, + ManifestCompatibilityError, + validate_manifest, + validate_run_state, +) from choicebench.io.writers import result_artifact_path, validate_result_artifact from choicebench.identity import integrity_digest from choicebench.datasets import dataset_content_digest @@ -251,3 +258,129 @@ def read_all_run_results( return pd.DataFrame() return pd.concat(frames, ignore_index=True) + + +@dataclass(frozen=True) +class RealizationSelection: + policy: Literal["single_evaluable_per_condition", "explicit"] + realization_ids: tuple[str, ...] + + +@dataclass(frozen=True) +class ManifestResultSet: + rows: pd.DataFrame + manifest: Mapping[str, Any] + state: Mapping[str, Any] + selection: RealizationSelection + verified_realizations: Mapping[str, Any] + + +def read_manifest_result_set( + run_dir: Path, *, realization_ids: Sequence[str] | None = None +) -> ManifestResultSet: + """Read a v3 imported/repaired run's result rows through the shared + full-graph verifier (never a bare CSV read): unselected/ineligible + realizations remain visible via verified_realizations for accounting, + but never concatenated into rows. Auto-selection only applies when a + condition has at most one eligible realization; ambiguity is a refusal. + """ + from choicebench.importing.engine import verify_import_run + + run_dir = Path(run_dir) + verified = verify_import_run(run_dir) + manifest = verified.manifest + state = json.loads((run_dir / RUN_STATE_FILENAME).read_text()) + + by_condition: dict[str, list[str]] = {} + for realization_id, record in manifest["payload"]["realizations"].items(): + by_condition.setdefault(record["condition_id"], []).append(realization_id) + + def _eligible(realization_id: str) -> bool: + return verified.realizations[realization_id].result_sha256 is not None + + if realization_ids is None: + selected: list[str] = [] + for condition_id, ids in by_condition.items(): + eligible_ids = sorted(rid for rid in ids if _eligible(rid)) + if len(eligible_ids) > 1: + raise ResultSetError( + f"Condition {condition_id!r} has multiple eligible realizations " + f"{eligible_ids}; pass an explicit realization_ids selection." + ) + selected.extend(eligible_ids) + policy: Literal["single_evaluable_per_condition", "explicit"] = ( + "single_evaluable_per_condition" + ) + else: + unknown = sorted(set(realization_ids) - set(manifest["payload"]["realizations"])) + if unknown: + raise ResultSetError(f"Unknown realization ID(s) {unknown}.") + ineligible = sorted(rid for rid in realization_ids if not _eligible(rid)) + if ineligible: + raise ResultSetError(f"Realization(s) {ineligible} have no result to evaluate.") + selected = sorted(set(realization_ids)) + policy = "explicit" + + frames = [ + pd.read_csv(run_dir / f"results/{realization_id}.csv", dtype={"question_id": "string"}) + for realization_id in selected + ] + rows = pd.concat(frames, ignore_index=True) if frames else pd.DataFrame() + + return ManifestResultSet( + rows=rows, + manifest=manifest, + state=state, + selection=RealizationSelection(policy=policy, realization_ids=tuple(selected)), + verified_realizations=verified.realizations, + ) + + +def build_evaluation_report_for_result_set( + run_id: str, result_set: ManifestResultSet, *, reparse: bool +) -> dict[str, Any]: + """Reduced-scope evaluation-v2 report: per-condition realization lists, + per-realization selection/status/scope accounting, and accuracy for + selected evaluable realizations. Every realization -- selected or not -- + is accounted for; only selected ones contribute metrics or rows.""" + manifest = result_set.manifest + conditions: dict[str, dict[str, Any]] = {} + realizations: dict[str, dict[str, Any]] = {} + selected_ids = set(result_set.selection.realization_ids) + + for realization_id, realization in manifest["payload"]["realizations"].items(): + condition_id = realization["condition_id"] + conditions.setdefault(condition_id, {"realization_ids": []}) + conditions[condition_id]["realization_ids"].append(realization_id) + + evidence = realization["identity"]["realization"]["evidence"] + result_origin = realization["identity"]["realization"]["result_origin"] + selected = realization_id in selected_ids + entry: dict[str, Any] = { + "selected": selected, + "evidence_status": evidence["evidence_status"], + "scope_disposition": evidence["scope_disposition"], + "prediction_origins": result_origin["prediction_origins"], + "metrics": {}, + } + if selected: + rows = result_set.rows[result_set.rows["realization_id"] == realization_id] + if len(rows): + predicted = rows["predicted_option"].astype(str).str.upper() + correct = rows["correct_option"].astype(str).str.upper() + entry["metrics"] = { + "accuracy": float((predicted == correct).mean()), + "n": int(len(rows)), + } + realizations[realization_id] = entry + + for condition in conditions.values(): + condition["realization_ids"] = sorted(condition["realization_ids"]) + + return { + "schema_version": "choicebench.evaluation.v2", + "run_id": run_id, + "selection_policy": result_set.selection.policy, + "conditions": conditions, + "realizations": realizations, + } diff --git a/src/choicebench/io/writers.py b/src/choicebench/io/writers.py index ac90cde..4c8b5bc 100644 --- a/src/choicebench/io/writers.py +++ b/src/choicebench/io/writers.py @@ -1,38 +1,78 @@ # src/choicebench/io/writers.py +from dataclasses import dataclass +from hashlib import sha256 from pathlib import Path -from typing import Any +from typing import Any, Mapping, Sequence import pandas as pd import json -from choicebench.identity import file_digest, integrity_digest +from choicebench.identity import canonicalize, file_digest, integrity_digest, short_id from choicebench.infra.artifacts import atomic_write_json, atomic_write_text RESULT_ARTIFACT_SCHEMA_VERSION = "choicebench.result-artifact.v1" +RESULT_ARTIFACT_V2_SCHEMA_VERSION = "choicebench.result-artifact.v2" def result_artifact_path(result_path: Path) -> Path: return Path(result_path).with_suffix(".artifact.json") -def validate_result_artifact(result_path: Path) -> dict[str, Any]: +def validate_result_artifact( + result_path: Path, + *, + manifest: Mapping[str, Any] | None = None, + realization: Mapping[str, Any] | None = None, +) -> dict[str, Any]: metadata_path = result_artifact_path(result_path) try: metadata = json.loads(metadata_path.read_text()) except (OSError, json.JSONDecodeError) as exc: raise RuntimeError(f"Result metadata is missing or unreadable: {metadata_path}: {exc}") from exc - payload = {key: value for key, value in metadata.items() if key != "metadata_digest"} - if metadata.get("schema_version") != RESULT_ARTIFACT_SCHEMA_VERSION: + schema_version = metadata.get("schema_version") + if schema_version == RESULT_ARTIFACT_SCHEMA_VERSION: + payload = {key: value for key, value in metadata.items() if key != "metadata_digest"} + if metadata.get("metadata_digest") != integrity_digest(payload): + raise RuntimeError(f"Result metadata integrity check failed: {metadata_path}") + actual = file_digest(result_path) + if metadata.get("file_sha256") != actual: + raise RuntimeError( + f"Result content integrity check failed: {result_path}; expected " + f"{metadata.get('file_sha256')}, found {actual}." + ) + return metadata + if schema_version != RESULT_ARTIFACT_V2_SCHEMA_VERSION: raise RuntimeError(f"Unsupported result metadata schema: {metadata_path}") - if metadata.get("metadata_digest") != integrity_digest(payload): + payload = { + key: value + for key, value in metadata.items() + if key not in {"result_artifact_id", "result_artifact_digest"} + } + if metadata.get("result_artifact_digest") != integrity_digest(payload): raise RuntimeError(f"Result metadata integrity check failed: {metadata_path}") + if metadata.get("result_artifact_id") != short_id("result", payload): + raise RuntimeError(f"Result metadata ID does not match its own digest: {metadata_path}") actual = file_digest(result_path) if metadata.get("file_sha256") != actual: raise RuntimeError( f"Result content integrity check failed: {result_path}; expected " f"{metadata.get('file_sha256')}, found {actual}." ) + if manifest is not None: + realization_id = metadata.get("realization_id") + record = ( + realization + if realization is not None + else manifest.get("payload", {}).get("realizations", {}).get(realization_id) + ) + if record is None or record.get("realization_digest") != metadata.get("realization_digest"): + raise RuntimeError(f"Result metadata realization binding is inconsistent: {metadata_path}") + if ( + metadata.get("experiment_id") != manifest.get("experiment_id") + or metadata.get("experiment_digest") != manifest.get("experiment_digest") + ): + raise RuntimeError(f"Result metadata experiment binding is inconsistent: {metadata_path}") return metadata @@ -113,3 +153,172 @@ def write_run_results( metadata["metadata_digest"] = integrity_digest(metadata) atomic_write_json(result_artifact_path(output_path), metadata) return output_path + + +@dataclass(frozen=True) +class PreparedResultArtifact: + result_path: str + metadata_path: str + csv_bytes: bytes + metadata: Mapping[str, Any] + + +def prepare_manifest_result( + results: Sequence[Mapping[str, Any]], + *, + manifest: Mapping[str, Any], + realization_id: str, +) -> PreparedResultArtifact: + """Render one realization's rows and compute result identity after the + manifest is fixed. Never writes anything; identity here can never depend + on its own child (the manifest/realization are already immutable inputs). + """ + payload = manifest["payload"] + realization = payload["realizations"].get(realization_id) + if realization is None: + raise RuntimeError(f"Unknown realization {realization_id!r} in manifest.") + condition_id = realization["condition_id"] + condition = payload["semantic_conditions"].get(condition_id) + if condition is None or condition["condition_digest"] != realization["condition_digest"]: + raise RuntimeError(f"Realization {realization_id!r} condition binding is inconsistent.") + condition_identity = condition["identity"] + result_origin = realization["identity"]["realization"]["result_origin"] + + expected_assignments: dict[str, tuple[str, str]] = {} + for row in result_origin["row_assignments"]: + qid = str(row["question_id"]) + if qid in expected_assignments: + raise RuntimeError( + f"Realization {realization_id!r} result origin has duplicate question_id {qid!r}." + ) + expected_assignments[qid] = (row["prediction_origin"], row["prediction_lineage_id"]) + + rows_by_qid: dict[str, dict[str, Any]] = {} + for row in results: + qid = row.get("question_id") + qid = None if qid is None else str(qid) + if qid is None or qid not in expected_assignments: + raise RuntimeError( + f"Result row for realization {realization_id!r} references an " + f"unexpected question_id {row.get('question_id')!r}." + ) + if qid in rows_by_qid: + raise RuntimeError( + f"Result rows for realization {realization_id!r} declare " + f"question_id {qid!r} more than once." + ) + expected_origin, expected_lineage_id = expected_assignments[qid] + if ( + row.get("prediction_origin") != expected_origin + or row.get("prediction_lineage_id") != expected_lineage_id + ): + raise RuntimeError( + f"Result row {qid!r} prediction_origin/prediction_lineage_id does " + f"not match its declared realization lineage for {realization_id!r}." + ) + rows_by_qid[qid] = dict(row) + + missing = sorted(set(expected_assignments) - set(rows_by_qid)) + if missing: + raise RuntimeError( + f"Realization {realization_id!r} is missing result rows for {missing}." + ) + + ordered_question_ids = [str(row["question_id"]) for row in result_origin["row_assignments"]] + ordered_rows = [rows_by_qid[qid] for qid in ordered_question_ids] + + benchmark = condition_identity["benchmark"] + identity_columns = { + "condition_id": condition_id, + "realization_id": realization_id, + "experiment_id": manifest["experiment_id"], + "dataset_artifact_id": benchmark["artifact_id"], + "dataset_selection_id": benchmark["selection_id"], + "model_id": condition_identity["model_id"], + "method_id": condition_identity["method_id"], + "prompt_id": condition_identity["prompt_id"], + "benchmark_name": benchmark["name"], + "benchmark_split": benchmark["split"], + } + for row in ordered_rows: + for column, expected_value in identity_columns.items(): + if row.get(column) != expected_value: + raise RuntimeError( + f"Result row {row.get('question_id')!r} for realization " + f"{realization_id!r} has missing or incorrect {column!r}." + ) + + df = pd.DataFrame(ordered_rows) + csv_bytes = df.to_csv(index=False).encode("utf-8") + result_path = f"results/{realization_id}.csv" + metadata_path = f"results/{realization_id}.artifact.json" + + prediction_origin_counts: dict[str, int] = {} + for row in ordered_rows: + origin = row["prediction_origin"] + prediction_origin_counts[origin] = prediction_origin_counts.get(origin, 0) + 1 + + metadata_core = { + "schema_version": RESULT_ARTIFACT_V2_SCHEMA_VERSION, + "experiment_id": manifest["experiment_id"], + "experiment_digest": manifest["experiment_digest"], + "condition_id": condition_id, + "condition_digest": condition["condition_digest"], + "realization_id": realization_id, + "realization_digest": realization["realization_digest"], + "row_count": len(ordered_rows), + "columns": list(df.columns), + "question_ids": ordered_question_ids, + "rows_digest": integrity_digest(ordered_rows), + "file_sha256": sha256(csv_bytes).hexdigest(), + "prediction_origins": sorted(prediction_origin_counts), + "prediction_origin_counts": { + key: prediction_origin_counts[key] for key in sorted(prediction_origin_counts) + }, + "result_path": result_path, + "metadata_path": metadata_path, + } + result_artifact_digest = integrity_digest(metadata_core) + metadata = { + **metadata_core, + "result_artifact_id": short_id("result", metadata_core), + "result_artifact_digest": result_artifact_digest, + } + return PreparedResultArtifact( + result_path=result_path, + metadata_path=metadata_path, + csv_bytes=csv_bytes, + metadata=canonicalize(metadata), + ) + + +def publish_manifest_result( + prepared: PreparedResultArtifact, *, run_dir: Path +) -> tuple[Path, Mapping[str, Any], bool]: + """Write a prepared result artifact, or verify-only if it already exists. + + Returns (result_path, metadata, reused). `reused=True` means an identical + artifact already existed; any divergence is refused rather than overwritten. + """ + run_dir = Path(run_dir) + result_path = run_dir / prepared.result_path + metadata_path = run_dir / prepared.metadata_path + if result_path.exists() != metadata_path.exists(): + raise RuntimeError(f"Realization result/sidecar pair is incomplete: {result_path}") + if result_path.exists(): + existing_bytes = result_path.read_bytes() + try: + existing_metadata = json.loads(metadata_path.read_text()) + except (OSError, json.JSONDecodeError) as exc: + raise RuntimeError( + f"Existing result metadata is unreadable: {metadata_path}: {exc}" + ) from exc + if existing_bytes == prepared.csv_bytes and existing_metadata == dict(prepared.metadata): + return result_path, existing_metadata, True + raise RuntimeError( + f"Refusing to overwrite a divergent existing result artifact: {result_path}" + ) + result_path.parent.mkdir(parents=True, exist_ok=True) + atomic_write_text(result_path, prepared.csv_bytes.decode("utf-8")) + atomic_write_json(metadata_path, dict(prepared.metadata)) + return result_path, prepared.metadata, False diff --git a/src/choicebench/manifest.py b/src/choicebench/manifest.py index 46205f8..2996854 100644 --- a/src/choicebench/manifest.py +++ b/src/choicebench/manifest.py @@ -9,7 +9,7 @@ from datetime import datetime, timezone from importlib.metadata import PackageNotFoundError, version from pathlib import Path -from typing import Any +from typing import Any, Mapping from choicebench import __version__ from choicebench.identity import CANONICALIZATION_VERSION, canonicalize, integrity_digest, short_id @@ -21,6 +21,16 @@ MANIFEST_FILENAME = "manifest.json" RUN_STATE_FILENAME = "run_state.json" +PROTOCOL_V3_VERSION = "choicebench.protocol.v3" +MANIFEST_V3_SCHEMA_VERSION = "choicebench.manifest.v3" +RUN_STATE_V3_SCHEMA_VERSION = "choicebench.run-state.v3" +_MANIFEST_V3_PAYLOAD_KEYS = { + "protocol_version", + "canonicalization_version", + "semantic_conditions", + "realizations", +} + class ManifestCompatibilityError(RuntimeError): pass @@ -304,3 +314,145 @@ def validate_run_state(state: dict[str, Any], manifest: dict[str, Any]) -> None: def write_run_state(run_dir: Path, state: dict[str, Any]) -> None: path = Path(run_dir) / RUN_STATE_FILENAME atomic_write_json(path, canonicalize(state, redact_secrets=False)) + + +def make_manifest_v3( + payload: Mapping[str, Any], *, audit: Mapping[str, Any] | None = None +) -> dict[str, Any]: + """Build an immutable v3 manifest from semantic-condition/realization tables. + + Additive alongside make_manifest(); v2 manifests are unaffected. `payload` + holds only identity-bearing content (no audit fields), so the experiment + digest is simply the payload digest -- unlike v2 there is no "config" key + to exclude. + """ + if set(payload) != _MANIFEST_V3_PAYLOAD_KEYS: + raise ManifestCompatibilityError("Manifest v3 payload fields are invalid.") + payload = canonicalize(dict(payload)) + if payload["protocol_version"] != PROTOCOL_V3_VERSION: + raise ManifestCompatibilityError( + f"Unsupported protocol version {payload['protocol_version']!r}; " + f"expected {PROTOCOL_V3_VERSION!r}." + ) + if payload["canonicalization_version"] != CANONICALIZATION_VERSION: + raise ManifestCompatibilityError("Unsupported canonicalization version.") + experiment_digest = integrity_digest(payload) + return { + "schema_version": MANIFEST_V3_SCHEMA_VERSION, + "experiment_id": f"exp_{experiment_digest[:16]}", + "experiment_digest": experiment_digest, + "payload_digest": integrity_digest(payload), + "created_at": datetime.now(timezone.utc).isoformat(), + "payload": payload, + "audit": canonicalize(dict(audit or {})), + } + + +def validate_manifest_v3(manifest: Mapping[str, Any]) -> None: + """Validate a v3 manifest, recomputing every child ID/digest from its + own identity payload rather than trusting the stored value.""" + if manifest.get("schema_version") != MANIFEST_V3_SCHEMA_VERSION: + raise ManifestCompatibilityError("Unsupported or missing manifest schema version.") + payload = manifest.get("payload") + if not isinstance(payload, dict) or set(payload) != _MANIFEST_V3_PAYLOAD_KEYS: + raise ManifestCompatibilityError("Manifest v3 payload fields are invalid.") + if payload.get("protocol_version") != PROTOCOL_V3_VERSION: + raise ManifestCompatibilityError( + f"Unsupported protocol version {payload.get('protocol_version')!r}; " + f"expected {PROTOCOL_V3_VERSION!r}." + ) + if payload.get("canonicalization_version") != CANONICALIZATION_VERSION: + raise ManifestCompatibilityError("Unsupported canonicalization version.") + if manifest.get("payload_digest") != integrity_digest(payload): + raise ManifestCompatibilityError("Manifest payload integrity check failed.") + experiment_digest = integrity_digest(payload) + experiment_id = f"exp_{experiment_digest[:16]}" + if ( + manifest.get("experiment_digest") != experiment_digest + or manifest.get("experiment_id") != experiment_id + ): + raise ManifestCompatibilityError( + "Manifest integrity check failed; payload or identity was modified." + ) + + semantic_conditions = payload["semantic_conditions"] + if not isinstance(semantic_conditions, dict): + raise ManifestCompatibilityError("Manifest semantic_conditions must be a mapping.") + for condition_id, record in semantic_conditions.items(): + if not isinstance(record, dict): + raise ManifestCompatibilityError("Manifest condition record is invalid.") + recomputed_id = short_id("cond", record.get("identity")) + recomputed_digest = integrity_digest(record.get("identity")) + if ( + condition_id != record.get("condition_id") + or condition_id != recomputed_id + or record.get("condition_digest") != recomputed_digest + ): + raise ManifestCompatibilityError( + f"Manifest condition {condition_id!r} ID/digest do not match its " + "own identity payload." + ) + + realizations = payload["realizations"] + if not isinstance(realizations, dict): + raise ManifestCompatibilityError("Manifest realizations must be a mapping.") + for realization_id, record in realizations.items(): + if not isinstance(record, dict): + raise ManifestCompatibilityError("Manifest realization record is invalid.") + recomputed_id = short_id("real", record.get("identity")) + recomputed_digest = integrity_digest(record.get("identity")) + if ( + realization_id != record.get("realization_id") + or realization_id != recomputed_id + or record.get("realization_digest") != recomputed_digest + ): + raise ManifestCompatibilityError( + f"Manifest realization {realization_id!r} ID/digest do not match " + "its own identity payload." + ) + owning_condition = semantic_conditions.get(record.get("condition_id")) + if ( + owning_condition is None + or owning_condition.get("condition_digest") != record.get("condition_digest") + ): + raise ManifestCompatibilityError( + f"Manifest realization {realization_id!r} condition binding is " + "inconsistent with its declared semantic condition." + ) + + +def initial_run_state_v3(manifest: Mapping[str, Any]) -> dict[str, Any]: + realizations = manifest["payload"]["realizations"] + return { + "schema_version": RUN_STATE_V3_SCHEMA_VERSION, + "experiment_id": manifest["experiment_id"], + "realizations": {realization_id: {"status": "pending"} for realization_id in realizations}, + } + + +def validate_run_state_v3(state: Mapping[str, Any], manifest: Mapping[str, Any]) -> None: + if ( + state.get("schema_version") != RUN_STATE_V3_SCHEMA_VERSION + or state.get("experiment_id") != manifest["experiment_id"] + ): + raise ManifestCompatibilityError( + f"Run state does not belong to manifest {manifest['experiment_id']}." + ) + expected = set(manifest["payload"]["realizations"]) + if set(state.get("realizations", {})) != expected: + raise ManifestCompatibilityError( + "Run state realization grid differs from the immutable manifest." + ) + + +def load_run_state_v3(run_dir: Path, manifest: Mapping[str, Any]) -> dict[str, Any]: + """Load and validate existing v3 run state without creating mutable state.""" + path = Path(run_dir) / RUN_STATE_FILENAME + if not path.is_file(): + raise ManifestCompatibilityError(f"Missing {RUN_STATE_FILENAME}.") + try: + state = json.loads(path.read_text()) + except (OSError, json.JSONDecodeError) as exc: + raise ManifestCompatibilityError(f"Unreadable {RUN_STATE_FILENAME}: {exc}") from exc + validate_run_state_v3(state, manifest) + return state diff --git a/src/choicebench/provenance.py b/src/choicebench/provenance.py index 0bb1375..31ed925 100644 --- a/src/choicebench/provenance.py +++ b/src/choicebench/provenance.py @@ -20,6 +20,23 @@ class ProvenanceResolutionError(RuntimeError): pass +_packages_distributions_cache: dict[str, list[str]] | None = None + + +def _cached_packages_distributions() -> dict[str, list[str]]: + """packages_distributions() rescans every installed distribution's + metadata (a full parse of each package's METADATA file) on every call, + with no caching of its own. The set of installed distributions cannot + change within a process's lifetime, so cache it process-wide -- this + matters because implementation_identity() is called once per lineage + component (i.e. potentially once per row), and the uncached scan alone + made a real multi-hundred-row import run take tens of minutes.""" + global _packages_distributions_cache + if _packages_distributions_cache is None: + _packages_distributions_cache = packages_distributions() + return _packages_distributions_cache + + def directory_digest(root: Path) -> tuple[str, list[dict[str, Any]]]: """Hash every regular file in a local model directory by logical path/content.""" root = Path(root).expanduser().resolve(strict=True) @@ -120,7 +137,7 @@ def implementation_identity(target: Any) -> dict[str, Any]: record["source_file"] = Path(source_path).name record["source_digest"] = file_digest(Path(source_path)) top_level = module_name.split(".", 1)[0] if module_name else None - distributions = packages_distributions().get(top_level, []) if top_level else [] + distributions = _cached_packages_distributions().get(top_level, []) if top_level else [] if distributions: distribution = sorted(distributions)[0] try: diff --git a/tests/importing/__init__.py b/tests/importing/__init__.py new file mode 100644 index 0000000..e69de29 diff --git a/tests/importing/_cross_module_helpers.py b/tests/importing/_cross_module_helpers.py new file mode 100644 index 0000000..2db247b --- /dev/null +++ b/tests/importing/_cross_module_helpers.py @@ -0,0 +1,6 @@ +def helper_v1(value: str) -> str: + return value.strip() + + +def helper_v2(value: str) -> str: + return "tampered-" + value diff --git a/tests/importing/conftest.py b/tests/importing/conftest.py new file mode 100644 index 0000000..93025e8 --- /dev/null +++ b/tests/importing/conftest.py @@ -0,0 +1,212 @@ +from pathlib import Path + +import pandas as pd +import pytest + +from choicebench.cli import evaluate_run +from choicebench.datasets import dataset_content_digest +from choicebench.identity import integrity_digest +from choicebench.infra.artifacts import atomic_write_text +from choicebench.io.readers import read_manifest_results +from choicebench.io.writers import validate_result_artifact, write_run_results +from choicebench.manifest import ( + build_manifest_payload, + ensure_manifest, + initial_run_state, + make_manifest, + write_run_state, +) +from choicebench.metrics import BUILTIN_METRICS +from choicebench.provenance import implementation_identity + + +EXPECTED_EXPERIMENT_ID = "exp_cd53979b257867df" +EXPECTED_EVALUATION_ID = "eval_93b022690174ec36" + + +@pytest.fixture +def synthetic_v2_run(tmp_path, monkeypatch) -> tuple[Path, dict, str]: + questions = pd.DataFrame( + [ + { + "question_id": "q1", + "subject": "arithmetic", + "question_text": "What is 2 + 2?", + "choice_a": "3", + "choice_b": "4", + "choice_c": "5", + "choice_d": "6", + "correct_option": "B", + "correct_answer_text": "4", + }, + { + "question_id": "q2", + "subject": "geography", + "question_text": "What is the capital of France?", + "choice_a": "Berlin", + "choice_b": "Madrid", + "choice_c": "Paris", + "choice_d": "Rome", + "correct_option": "C", + "correct_answer_text": "Paris", + }, + ] + ) + prompt_content = "Question: {question_text}\nChoices:\n{choices}\nAnswer:" + selection_id = "sel_legacyfixture" + artifact_id = "dataset_legacyfixture" + prompt_id = "prompt_legacyfixture" + model_id = "model_legacyfixture" + method_id = "method_legacyfixture" + condition_id = "cond_legacyfixture" + + dataset_record = { + "selection_id": selection_id, + "artifact_id": artifact_id, + "benchmark": "toy", + "split": "test", + "prepared_path": "toy/test.csv", + "prepared_content_digest": "prepared-legacy-fixture", + "prepared_metadata": {"schema_version": "choicebench.dataset-artifact.v1"}, + "selected_content_digest": dataset_content_digest(questions), + "selected_question_ids": ["q1", "q2"], + "selected_sample_identities": ["sample_q1", "sample_q2"], + "run_snapshot_path": f"artifacts/datasets/{selection_id}.csv", + "row_count": 2, + } + prompt_record = { + "prompt_id": prompt_id, + "version": "legacy-v1", + "files": { + "direct_mcq": { + "content": prompt_content, + "sha256": integrity_digest(prompt_content), + } + }, + "run_snapshot_path": f"artifacts/prompts/{prompt_id}", + } + model_record = { + "model_id": model_id, + "config": {"backend": "dummy", "model_name_or_path": "dummy-model"}, + "resolved_model": None, + } + method_record = { + "method_id": method_id, + "config": {"name": "direct_mcq", "params": {}, "preflight": None}, + "implementation": {"qualified_name": "legacy.fixture:DirectMCQ"}, + } + condition_record = { + "condition_id": condition_id, + "benchmark_name": "toy", + "split": "test", + "selection_id": selection_id, + "artifact_id": artifact_id, + "model_id": model_id, + "method_id": method_id, + "prompt_id": prompt_id, + "prompt_snapshot_path": prompt_record["run_snapshot_path"], + "result_path": f"results/{condition_id}.csv", + "checkpoint_path": f"checkpoints/{condition_id}.json", + "result_metadata_path": f"results/{condition_id}.artifact.json", + "model_display_name": "dummy-model", + "resolved_model": None, + "identity": {"fixture": "native-v2"}, + } + source = { + "git_commit": "0123456789abcdef0123456789abcdef01234567", + "dirty": False, + "dirty_tracked_digest": None, + "source_tree_digest": "source-tree-legacy-fixture", + } + environment = { + "choicebench_version": "0.2.0", + "python": "3.12.0", + "implementation": "CPython", + "dependencies": {"pandas": "2.2.0", "pytest": "8.3.0"}, + } + payload = build_manifest_payload( + config={"run": {"seed": 7}, "metrics": ["accuracy"]}, + datasets=[dataset_record], + prompts=prompt_record, + models=[model_record], + methods=[method_record], + conditions=[condition_record], + source=source, + environment=environment, + ) + payload["evaluation"] = [ + { + "name": "accuracy", + "implementation": implementation_identity(BUILTIN_METRICS["accuracy"]), + } + ] + manifest = make_manifest(payload) + assert manifest["experiment_id"] == EXPECTED_EXPERIMENT_ID + + run_dir = tmp_path / "native-v2-fixture" + ensure_manifest(run_dir, manifest) + atomic_write_text( + run_dir / dataset_record["run_snapshot_path"], + questions.to_csv(index=False), + ) + atomic_write_text( + run_dir + / prompt_record["run_snapshot_path"] + / prompt_record["version"] + / "direct_mcq.txt", + prompt_content, + ) + + rows = [] + for question, parsed_choice in zip( + questions.to_dict("records"), ("B", "A"), strict=True + ): + rows.append( + { + **question, + "raw_text": f"The answer is {parsed_choice}", + "parsed_choice": parsed_choice, + "parse_status": "parse_ok", + "normalized_text": parsed_choice, + "parse_reason": "Answer successfully parsed", + "score_status": "scored", + "is_correct": parsed_choice == question["correct_option"], + "transport_status": "ok", + "experiment_id": manifest["experiment_id"], + "condition_id": condition_id, + "dataset_artifact_id": artifact_id, + "dataset_selection_id": selection_id, + "model_id": model_id, + "method_id": method_id, + "prompt_id": prompt_id, + "benchmark_split": "test", + "model_name": "dummy-model", + "method_name": "direct_mcq", + } + ) + result_path = write_run_results( + rows, + run_dir, + run_dir.name, + "direct_mcq", + "dummy-model", + "toy", + condition_id, + ) + artifact = validate_result_artifact(result_path) + state = initial_run_state(manifest) + state["conditions"][condition_id] = { + "status": "completed", + "result_path": condition_record["result_path"], + "result_sha256": artifact["file_sha256"], + } + write_run_state(run_dir, state) + + monkeypatch.setattr(evaluate_run, "RUNS_DIR", run_dir.parent) + frame = pd.read_csv(result_path) + report = evaluate_run.build_evaluation_report( + run_dir.name, frame, manifest, reparse=False + ) + assert report["evaluation_id"] == EXPECTED_EVALUATION_ID + + return run_dir, manifest, EXPECTED_EVALUATION_ID diff --git a/tests/importing/test_authorization.py b/tests/importing/test_authorization.py new file mode 100644 index 0000000..680e336 --- /dev/null +++ b/tests/importing/test_authorization.py @@ -0,0 +1,255 @@ +"""Tests for typed authorization bundle validation (Task 10, reduced scope).""" + +from __future__ import annotations + +from hashlib import sha256 +from pathlib import Path + +import pytest + +from choicebench.identity import integrity_digest, short_id +from choicebench.importing.authorization import ( + AuthorizationError, + authorization_for_condition, + validate_authorization_bundle, +) +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.schema import AuthorizationSpec +from tests.importing.test_identity import _dataset + + +def _source(data: bytes = b"authorization,payload\n1,2\n") -> OpenedSource: + return OpenedSource( + source_id="auth-source", + audit_path=Path("/machine-a/authorization.csv"), + logical_path="inputs/authorization.csv", + data=data, + sha256=sha256(data).hexdigest(), + ) + + +def _bundle_payload(*, authorization_type, executable, condition_digest, grants, source, purpose="repair"): + return { + "schema_version": "choicebench.authorization-bundle.v1", + "authorization_type": authorization_type, + "grants": {condition_digest: dict(grants)}, + "authority": "principal-investigator", + "purpose": purpose, + "executable": executable, + "source_sha256": source.sha256, + "input_evidence_digests": {"ev1": "1" * 64}, + "expected_snapshot_digests": {condition_digest: "2" * 64}, + } + + +def _declaration(*, authorization_type, executable, condition_digest, grants, source, authorization_id=None): + payload = _bundle_payload( + authorization_type=authorization_type, executable=executable, + condition_digest=condition_digest, grants=grants, source=source, + ) + computed_id = short_id("auth", payload) + return AuthorizationSpec( + authorization_id=authorization_id or computed_id, + authorization_type=authorization_type, + source_id="auth-source", + condition_question_reasons={condition_digest: dict(grants)}, + authority="principal-investigator", + purpose="repair", + executable=executable, + input_evidence_digests={"ev1": "1" * 64}, + expected_snapshot_digests={condition_digest: "2" * 64}, + ) + + +def test_validate_authorization_bundle_accepts_a_well_formed_inference_repair_grant(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q1": "malformed prediction"}, source=source, + ) + bundle = validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + assert bundle.authorization_type == "inference_repair" + assert bundle.grants[condition_digest] == {"q1": "malformed prediction"} + + +def test_validate_authorization_bundle_rejects_source_id_mismatch(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + wrong_id_source = OpenedSource( + source_id="a-different-source-id", + audit_path=Path("/machine-a/authorization.csv"), + logical_path="inputs/authorization.csv", + data=source.data, + sha256=source.sha256, + ) + with pytest.raises(AuthorizationError, match="source"): + validate_authorization_bundle( + declaration, opened_source=wrong_id_source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_tampered_source_bytes(): + """Bundle identity is bound to the exact opened source bytes: swapping the + source content for a different byte-identical-source_id file changes the + recomputed authorization_id, so a declared ID from the original bytes is + refused rather than silently accepted against different content.""" + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + tampered_source = _source(data=b"different,bytes\n9,9\n") + with pytest.raises(AuthorizationError, match="recomputed digest"): + validate_authorization_bundle( + declaration, opened_source=tampered_source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_inference_repair_marked_nonexecutable(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=False, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + with pytest.raises(AuthorizationError, match="executable"): + validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_offline_transformation_marked_executable(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="offline_transformation", executable=True, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + with pytest.raises(AuthorizationError, match="executable"): + validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_unknown_condition(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + with pytest.raises(AuthorizationError, match="known conditions"): + validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": "d" * 64}, # different digest + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_question_outside_selection(): + expected = _dataset() # selected_question_ids are q2, q1 + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q9": "reason"}, source=source, + ) + with pytest.raises(AuthorizationError, match="expected selection"): + validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_empty_reason(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q1": " "}, source=source, + ) + with pytest.raises(AuthorizationError, match="explicit reason"): + validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_validate_authorization_bundle_rejects_self_assigned_id(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="inference_repair", executable=True, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + authorization_id="auth_" + "0" * 16, + ) + with pytest.raises(AuthorizationError, match="recomputed digest"): + validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + + +def test_authorization_for_condition_slices_one_condition(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="offline_transformation", executable=False, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + bundle = validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + sliced = authorization_for_condition(bundle, condition_digest=condition_digest) + assert sliced.bundle_id == bundle.authorization_id + assert sliced.bundle_digest == bundle.authorization_digest + assert sliced.question_reasons == {"q1": "reason"} + + +def test_authorization_for_condition_refuses_unauthorized_condition(): + expected = _dataset() + condition_digest = "c" * 64 + source = _source() + declaration = _declaration( + authorization_type="offline_transformation", executable=False, + condition_digest=condition_digest, grants={"q1": "reason"}, source=source, + ) + bundle = validate_authorization_bundle( + declaration, opened_source=source, + condition_digests={"cond_key": condition_digest}, + expected={condition_digest: expected}, + ) + with pytest.raises(AuthorizationError, match="self-authorize"): + authorization_for_condition(bundle, condition_digest="d" * 64) diff --git a/tests/importing/test_csv_adapter.py b/tests/importing/test_csv_adapter.py new file mode 100644 index 0000000..d238096 --- /dev/null +++ b/tests/importing/test_csv_adapter.py @@ -0,0 +1,529 @@ +from __future__ import annotations + +from dataclasses import replace +from hashlib import sha256 +from pathlib import Path + +import pytest + +from choicebench.importing.csv_adapter import ( + CsvAdapterError, + OpenedSource, + effective_csv_adapter_projection, + parse_csv_source, + scan_csv_logical_records, +) +from choicebench.importing.schema import ( + CsvDialectSpec, + NumericColumnSpec, + OptionMappingSpec, + SourceArtifactSpec, +) + + +def _opened(tmp_path: Path, data: bytes) -> OpenedSource: + path = tmp_path / "source.csv" + path.write_bytes(data) + return OpenedSource( + source_id="source", + audit_path=path, + logical_path="freeze/source.csv", + data=data, + sha256=sha256(data).hexdigest(), + ) + + +def _source_spec( + data: bytes, + *, + expected_columns: tuple[str, ...] = ("id", "value"), + columns: dict[str, str] | None = None, + ignored_columns: dict[str, str] | None = None, + null_values: tuple[str, ...] = (), + numeric_columns: tuple[NumericColumnSpec, ...] = (), + option_mapping: OptionMappingSpec | None = None, + extra_field_policy: str = "preserve_unmapped", + dialect: CsvDialectSpec | None = None, +) -> SourceArtifactSpec: + return SourceArtifactSpec( + source_id="source", + path=Path("/explicit/read-only/source.csv"), + logical_path="freeze/source.csv", + expected_sha256=sha256(data).hexdigest(), + format="csv", + format_version="producer-v1", + classification="raw", + dialect=dialect or CsvDialectSpec(), + columns=columns or {"question_id": "id", "prediction": "value"}, + expected_columns=expected_columns, + ignored_columns=ignored_columns or {}, + null_values=null_values, + numeric_columns=numeric_columns, + option_mapping=option_mapping + or OptionMappingSpec("ordered_columns", ("value",), None, None, None), + extra_field_policy=extra_field_policy, # type: ignore[arg-type] + preserve_namespace="producer", + source_run_id=None, + source_repository=None, + source_commit=None, + notes={}, + ) + + +def _parse( + tmp_path: Path, + data: bytes, + *, + strict: bool = False, + **spec_kwargs: object, +): + return parse_csv_source( + _opened(tmp_path, data), _source_spec(data, **spec_kwargs), strict=strict + ) + + +@pytest.mark.parametrize( + ("data", "terminators"), + [ + (b"id,value\nq1,x\n", (b"\n", b"\n")), + (b"id,value\r\nq1,x\r\n", (b"\r\n", b"\r\n")), + (b"id,value\rq1,x\r", (b"\r", b"\r")), + (b"id,value\r\nq1,x\nq2,y\r", (b"\r\n", b"\n", b"\r")), + ], +) +def test_scanner_recognizes_declared_terminators(data: bytes, terminators: tuple[bytes, ...]): + spans = scan_csv_logical_records(data, CsvDialectSpec()) + assert tuple(span.terminator for span in spans) == terminators + assert b"".join(data[span.start : span.end] for span in spans) == data + + +def test_embedded_newline_span_is_exact(): + data = b'id,text\r\nq1,"line one\r\nline two"\r\nq2,end\n' + spans = scan_csv_logical_records(data, CsvDialectSpec()) + assert data[spans[1].start : spans[1].end] == b'q1,"line one\r\nline two"\r\n' + assert spans[1].terminator == b"\r\n" + + +@pytest.mark.parametrize("embedded", [b"\r", b"\n", b"\r\n"]) +def test_each_embedded_terminator_remains_inside_quoted_record(embedded: bytes): + data = b'id,text\nq1,"a' + embedded + b'b"\n' + spans = scan_csv_logical_records(data, CsvDialectSpec()) + assert len(spans) == 2 + assert data[spans[1].start : spans[1].end] == b'q1,"a' + embedded + b'b"\n' + + +def test_doubled_quote_does_not_end_quoted_field(): + data = b'id,text\nq1,"a""b\nc"\n' + spans = scan_csv_logical_records(data, CsvDialectSpec(double_quote=True)) + assert len(spans) == 2 + assert data[spans[1].start : spans[1].end] == b'q1,"a""b\nc"\n' + + +def test_explicit_escape_protects_quote_and_newline(): + data = b'id,text\nq1,"a\\"b\nc"\n' + dialect = CsvDialectSpec(escape_character="\\", double_quote=False) + spans = scan_csv_logical_records(data, dialect) + assert len(spans) == 2 + assert data[spans[1].start : spans[1].end] == b'q1,"a\\"b\nc"\n' + + +def test_explicit_escape_outside_quotes_protects_one_lf_byte(): + data = b"id,text\nq1,a\\\nb\n" + dialect = CsvDialectSpec(escape_character="\\") + spans = scan_csv_logical_records(data, dialect) + assert len(spans) == 2 + assert data[spans[1].start : spans[1].end] == b"q1,a\\\nb\n" + + +def test_unclosed_quote_is_rejected_with_exact_record_start(): + data = b'id,text\nq1,"unterminated\n' + with pytest.raises(CsvAdapterError, match=r"unclosed quote.*byte 8"): + scan_csv_logical_records(data, CsvDialectSpec()) + + +def test_final_record_without_terminator_is_preserved(): + data = b"id,value\nq1,x" + spans = scan_csv_logical_records(data, CsvDialectSpec()) + assert spans[-1].terminator == b"" + assert data[spans[-1].start : spans[-1].end] == b"q1,x" + + +def test_forbidden_final_record_without_terminator_is_rejected(): + data = b"id,value\nq1,x" + dialect = CsvDialectSpec(final_record_without_terminator="forbid") + with pytest.raises(CsvAdapterError, match="final logical record"): + scan_csv_logical_records(data, dialect) + + +def test_forbidden_mixed_terminators_are_rejected(): + data = b"id,value\r\nq1,x\n" + dialect = CsvDialectSpec(mixed_line_terminators="forbid") + with pytest.raises(CsvAdapterError, match="mixed line terminators"): + scan_csv_logical_records(data, dialect) + + +@pytest.mark.parametrize( + ("declared", "actual"), + [ + (("lf",), b"\r\n"), + (("lf",), b"\r"), + (("crlf",), b"\n"), + (("crlf",), b"\r"), + (("cr",), b"\r\n"), + (("cr",), b"\n"), + ], +) +def test_actual_undeclared_terminator_is_rejected( + declared: tuple[str, ...], actual: bytes +): + data = b"id,value" + actual + b"q1,x" + actual + dialect = CsvDialectSpec(line_terminators=declared) + with pytest.raises(CsvAdapterError, match="undeclared line terminator"): + scan_csv_logical_records(data, dialect) + + +def test_mixed_policy_compares_actual_crlf_and_lf_terminators(): + data = b"id,value\r\nq1,x\n" + dialect = CsvDialectSpec( + line_terminators=("crlf", "lf"), mixed_line_terminators="forbid" + ) + with pytest.raises(CsvAdapterError, match="mixed line terminators"): + scan_csv_logical_records(data, dialect) + + +@pytest.mark.parametrize("blank", [b"\n", b"\r", b"\r\n"]) +def test_blank_logical_records_are_rejected(blank: bytes): + data = b"id,value\n" + blank + b"q1,x\n" + with pytest.raises(CsvAdapterError, match="blank logical record"): + scan_csv_logical_records(data, CsvDialectSpec()) + + +@pytest.mark.parametrize("terminator", [b"\n", b"\r", b"\r\n"]) +def test_bom_strip_rejects_blank_first_logical_record_directly(terminator: bytes): + data = b"\xef\xbb\xbf" + terminator + b"q1,x" + terminator + dialect = CsvDialectSpec(bom_policy="strip_utf8_bom") + with pytest.raises(CsvAdapterError, match="blank logical record"): + scan_csv_logical_records(data, dialect) + + +def test_bom_is_forbidden_by_default(tmp_path: Path): + data = b"\xef\xbb\xbfid,value\nq1,x\n" + with pytest.raises(CsvAdapterError, match="UTF-8 BOM"): + _parse(tmp_path, data) + + +def test_bom_strip_preserves_original_byte_offsets(tmp_path: Path): + data = b"\xef\xbb\xbfid,value\nq1,x\n" + dialect = CsvDialectSpec(bom_policy="strip_utf8_bom") + table = _parse(tmp_path, data, dialect=dialect) + assert table.columns == ("id", "value") + assert table.rows[0].span.start == len(b"\xef\xbb\xbfid,value\n") + assert table.rows[0].values == {"id": "q1", "value": "x"} + + +def test_bom_strip_keeps_a_quoted_first_header_field_in_quote_state(): + data = b'\xef\xbb\xbf"id\ncontinued",value\nq1,x\n' + dialect = CsvDialectSpec(bom_policy="strip_utf8_bom") + spans = scan_csv_logical_records(data, dialect) + assert len(spans) == 2 + assert data[spans[0].start : spans[0].end] == b'\xef\xbb\xbf"id\ncontinued",value\n' + + +def test_invalid_utf8_is_rejected_without_replacement(tmp_path: Path): + data = b"id,value\nq1,\xff\n" + with pytest.raises(CsvAdapterError, match="UTF-8.*logical record 1"): + _parse(tmp_path, data) + + +def test_invalid_utf8_offset_in_bom_stripped_header_uses_original_bytes(tmp_path: Path): + data = b"\xef\xbb\xbfid,\xff\nq1,x\n" + dialect = CsvDialectSpec(bom_policy="strip_utf8_bom") + with pytest.raises(CsvAdapterError, match=r"logical record 0 at source byte 6"): + _parse(tmp_path, data, dialect=dialect) + + +def test_header_order_and_raw_row_hash_are_preserved(tmp_path: Path): + data = b"value,id\r\nx,q1\r\n" + table = _parse( + tmp_path, + data, + expected_columns=("value", "id"), + columns={"question_id": "id", "prediction": "value"}, + ) + assert table.columns == ("value", "id") + assert table.rows[0].raw_sha256 == sha256(b"x,q1\r\n").hexdigest() + + +def test_duplicate_header_is_rejected(tmp_path: Path): + data = b"id,id\nq1,x\n" + with pytest.raises(CsvAdapterError, match="duplicate header"): + _parse(tmp_path, data, expected_columns=("id",)) + + +@pytest.mark.parametrize("credential_column", ["api_key", "nested_access_token", "password"]) +def test_credential_shaped_header_is_rejected_in_every_mode( + tmp_path: Path, credential_column: str +): + data = f"id,value,{credential_column}\nq1,x,secret\n".encode() + for strict in (False, True): + with pytest.raises(CsvAdapterError, match="credential-named"): + _parse(tmp_path, data, strict=strict) + + +def test_normal_mode_preserves_unmapped_extra_columns(tmp_path: Path): + data = b"id,value,diagnostic\nq1,x,kept\n" + table = _parse(tmp_path, data, strict=False) + assert table.rows[0].values["diagnostic"] == "kept" + + +def test_cli_strict_mode_tightens_preserve_unmapped_policy(tmp_path: Path): + data = b"id,value,diagnostic\nq1,x,kept\n" + with pytest.raises(CsvAdapterError, match="unexpected source columns.*diagnostic"): + _parse(tmp_path, data, strict=True) + + +def test_cli_normal_mode_cannot_loosen_declared_reject_policy(tmp_path: Path): + data = b"id,value,diagnostic\nq1,x,kept\n" + with pytest.raises(CsvAdapterError, match="unexpected source columns.*diagnostic"): + _parse(tmp_path, data, strict=False, extra_field_policy="reject_unmapped") + + +def test_effective_strictness_is_identity_bearing(): + data = b"id,value\nq1,x\n" + declaration = _source_spec(data) + assert effective_csv_adapter_projection( + declaration, strict=False + ) != effective_csv_adapter_projection(declaration, strict=True) + assert effective_csv_adapter_projection( + replace(declaration, dialect=replace(declaration.dialect, delimiter=";")), + strict=False, + ) != effective_csv_adapter_projection(declaration, strict=False) + + +def test_declared_ignored_column_requires_reason_and_is_retained(tmp_path: Path): + data = b"id,value,score\nq1,x,0.5\n" + table = _parse( + tmp_path, + data, + expected_columns=("id", "value", "score"), + ignored_columns={"score": "producer aggregate only"}, + ) + assert table.rows[0].values["score"] == "0.5" + + +def test_missing_declared_column_is_rejected(tmp_path: Path): + data = b"id,value\nq1,x\n" + with pytest.raises(CsvAdapterError, match="missing declared source columns.*score"): + _parse(tmp_path, data, expected_columns=("id", "value", "score")) + + +def test_empty_string_is_distinct_from_null_unless_declared(tmp_path: Path): + data = b"id,value\nq1,\n" + assert _parse(tmp_path, data).rows[0].values["value"] == "" + assert _parse(tmp_path, data, null_values=("",)).rows[0].values["value"] is None + + +def test_literal_nan_is_not_implicit_null(tmp_path: Path): + data = b"id,value\nq1,NaN\n" + table = _parse(tmp_path, data, strict=True) + assert table.rows[0].values["value"] == "NaN" + + +@pytest.mark.parametrize("value", [b"1", b"-2", b"+3"]) +def test_integer_numeric_policy_accepts_exact_integer_strings(tmp_path: Path, value: bytes): + data = b"id,value\nq1," + value + b"\n" + numeric = (NumericColumnSpec("value", "integer", False),) + assert _parse(tmp_path, data, numeric_columns=numeric).rows[0].values["value"] == value.decode() + + +@pytest.mark.parametrize("value", [b"1.0", b"1e2", b"x", b" 1"]) +def test_integer_numeric_policy_rejects_malformed_strings(tmp_path: Path, value: bytes): + data = b"id,value\nq1," + value + b"\n" + numeric = (NumericColumnSpec("value", "integer", False),) + with pytest.raises(CsvAdapterError, match="integer.*value"): + _parse(tmp_path, data, numeric_columns=numeric) + + +@pytest.mark.parametrize("value", [b"1", b"-1.25", b"1e2"]) +def test_float_numeric_policy_accepts_finite_numbers(tmp_path: Path, value: bytes): + data = b"id,value\nq1," + value + b"\n" + numeric = (NumericColumnSpec("value", "float", False),) + _parse(tmp_path, data, numeric_columns=numeric) + + +@pytest.mark.parametrize("value", [b"NaN", b"inf", b"-Infinity", b"x", b" 1"]) +def test_finite_float_policy_rejects_nonfinite_or_malformed_values(tmp_path: Path, value: bytes): + data = b"id,value\nq1," + value + b"\n" + numeric = (NumericColumnSpec("value", "float", False, finite_only=True),) + with pytest.raises(CsvAdapterError, match="finite float.*value"): + _parse(tmp_path, data, numeric_columns=numeric) + + +def test_nonfinite_float_is_allowed_only_when_explicit(tmp_path: Path): + data = b"id,value\nq1,NaN\n" + numeric = (NumericColumnSpec("value", "float", False, finite_only=False),) + assert _parse(tmp_path, data, numeric_columns=numeric).rows[0].values["value"] == "NaN" + + +def test_numeric_null_policy_is_explicit(tmp_path: Path): + data = b"id,value\nq1,NA\n" + with pytest.raises(CsvAdapterError, match="null.*value"): + _parse( + tmp_path, + data, + null_values=("NA",), + numeric_columns=(NumericColumnSpec("value", "float", False),), + ) + table = _parse( + tmp_path, + data, + null_values=("NA",), + numeric_columns=(NumericColumnSpec("value", "float", True),), + ) + assert table.rows[0].values["value"] is None + + +@pytest.mark.parametrize("option_count", [3, 6]) +def test_ordered_option_mapping_supports_variable_option_counts( + tmp_path: Path, option_count: int +): + option_columns = tuple(f"choice_{index}" for index in range(option_count)) + header = ("id", *option_columns) + data = ( + ",".join(header) + + "\nq1," + + ",".join(f"v{index}" for index in range(option_count)) + + "\n" + ).encode() + table = _parse( + tmp_path, + data, + expected_columns=header, + columns={"question_id": "id"}, + option_mapping=OptionMappingSpec("ordered_columns", option_columns, None, None, None), + ) + assert tuple(table.rows[0].values[name] for name in option_columns) == tuple( + f"v{index}" for index in range(option_count) + ) + + +def test_empty_trailing_ordered_option_stays_null_not_phantom_choice(tmp_path: Path): + data = b"id,a,b,c,d\nq1,A,B,C,\n" + table = _parse( + tmp_path, + data, + expected_columns=("id", "a", "b", "c", "d"), + columns={"question_id": "id"}, + null_values=("",), + option_mapping=OptionMappingSpec("ordered_columns", ("a", "b", "c", "d"), None, None, None), + ) + assert [table.rows[0].values[name] for name in ("a", "b", "c", "d")] == ["A", "B", "C", None] + + +def test_structured_json_choices_accept_variable_option_list(tmp_path: Path): + data = ( + b'id,choices\nq1,"[{""label"":""A"",""text"":""one""},' + b'{""label"":""B"",""text"":""two""},' + b'{""label"":""C"",""text"":""three""}]"\n' + ) + mapping = OptionMappingSpec("structured_json", (), "choices", "label", "text") + table = _parse( + tmp_path, + data, + expected_columns=("id", "choices"), + columns={"question_id": "id"}, + option_mapping=mapping, + ) + assert table.rows[0].values["choices"].startswith('[{"label":"A"') + + +@pytest.mark.parametrize( + "payload", + [b"not-json", b"{}", b'[{"label":"A"}]', b'[{"label":"A","text":1}]'], +) +def test_structured_json_choices_reject_malformed_payloads(tmp_path: Path, payload: bytes): + escaped = payload.replace(b'"', b'""') + data = b'id,choices\nq1,"' + escaped + b'"\n' + mapping = OptionMappingSpec("structured_json", (), "choices", "label", "text") + with pytest.raises(CsvAdapterError, match="structured choices"): + _parse( + tmp_path, + data, + expected_columns=("id", "choices"), + columns={"question_id": "id"}, + option_mapping=mapping, + ) + + +def test_structured_json_choices_accept_positional_source_index_labels(tmp_path: Path): + """Regression: real fourth-cell CSVs use {"text": ..., "source_index": N} + -- an integer positional index, no explicit string label -- unlike the + explicit {"label": "A", "text": ...} shape tested above. This must be + letterized (0 -> A, 1 -> B, ...), not rejected.""" + data = ( + b'id,choices\nq1,"[{""text"":""slower"",""source_index"":0},' + b'{""text"":""faster"",""source_index"":1},' + b'{""text"":""at the same speed"",""source_index"":2}]"\n' + ) + mapping = OptionMappingSpec("structured_json", (), "choices", "source_index", "text") + table = _parse( + tmp_path, + data, + expected_columns=("id", "choices"), + columns={"question_id": "id"}, + option_mapping=mapping, + ) + assert table.rows[0].values["choices"].startswith('[{"text":"slower"') + + +def test_structured_json_choices_reject_source_index_out_of_letter_range(tmp_path: Path): + data = b'id,choices\nq1,"[{""text"":""x"",""source_index"":99}]"\n' + mapping = OptionMappingSpec("structured_json", (), "choices", "source_index", "text") + with pytest.raises(CsvAdapterError, match="structured choices"): + _parse( + tmp_path, + data, + expected_columns=("id", "choices"), + columns={"question_id": "id"}, + option_mapping=mapping, + ) + + +def test_structured_json_choices_reject_duplicate_source_index(tmp_path: Path): + data = ( + b'id,choices\nq1,"[{""text"":""x"",""source_index"":0},' + b'{""text"":""y"",""source_index"":0}]"\n' + ) + mapping = OptionMappingSpec("structured_json", (), "choices", "source_index", "text") + with pytest.raises(CsvAdapterError, match="duplicate labels"): + _parse( + tmp_path, + data, + expected_columns=("id", "choices"), + columns={"question_id": "id"}, + option_mapping=mapping, + ) + + +def test_row_with_more_fields_than_header_is_rejected(tmp_path: Path): + data = b"id,value\nq1,x,unexpected\n" + with pytest.raises(CsvAdapterError, match="field count"): + _parse(tmp_path, data) + + +def test_source_checksum_and_declaration_source_id_are_verified(tmp_path: Path): + data = b"id,value\nq1,x\n" + source = _opened(tmp_path, data) + bad_digest = replace(_source_spec(data), expected_sha256="0" * 64) + with pytest.raises(CsvAdapterError, match="checksum"): + parse_csv_source(source, bad_digest, strict=False) + bad_source_id = replace(_source_spec(data), source_id="different") + with pytest.raises(CsvAdapterError, match="source_id"): + parse_csv_source(source, bad_source_id, strict=False) + + +def test_source_logical_path_must_match_declaration(tmp_path: Path): + data = b"id,value\nq1,x\n" + source = replace(_opened(tmp_path, data), logical_path="other/source.csv") + with pytest.raises(CsvAdapterError, match="logical_path"): + parse_csv_source(source, _source_spec(data), strict=False) diff --git a/tests/importing/test_dataset_reference.py b/tests/importing/test_dataset_reference.py new file mode 100644 index 0000000..0dc2efc --- /dev/null +++ b/tests/importing/test_dataset_reference.py @@ -0,0 +1,690 @@ +from __future__ import annotations + +from dataclasses import replace +from hashlib import sha256 +import json +from pathlib import Path + +import pandas as pd +import pytest + +from choicebench.datasets import ( + NORMALIZATION_VERSION, + dataset_content_digest, + dataset_sample_identities, +) +from choicebench.identity import integrity_digest, short_id +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.dataset_reference import ( + DatasetReferenceError, + build_expected_dataset, + validate_expected_snapshot, + write_expected_snapshot, +) +from choicebench.importing.schema import DatasetReferenceSpec +from choicebench.pipeline.options import build_option_map + + +def _csv_bytes(rows: list[dict[str, str]]) -> bytes: + frame = pd.DataFrame(rows) + return frame.to_csv(index=False, lineterminator="\n").encode("utf-8") + + +def _source( + source_id: str, + rows: list[dict[str, str]], + *, + logical_path: str | None = None, + audit_path: Path | None = None, +) -> OpenedSource: + data = _csv_bytes(rows) + return OpenedSource( + source_id=source_id, + audit_path=audit_path or Path(f"/machine-a/{source_id}.csv"), + logical_path=logical_path or f"publisher/{source_id}.csv", + data=data, + sha256=sha256(data).hexdigest(), + ) + + +def _rows(*, six_options: bool = False) -> list[dict[str, str]]: + option_count = 6 if six_options else 3 + questions = ( + ("q1", "First?", "B", ("one", "two", "three", "four", "five", "six")), + ("q2", "Second?", "A", ("red", "green", "blue", "cyan", "magenta", "yellow")), + ("q3", "Third?", "C", ("cat", "dog", "owl", "fox", "yak", "eel")), + ) + result: list[dict[str, str]] = [] + for question_id, question, gold, options in questions: + row = {"qid": question_id, "stem": question, "gold": gold} + row.update( + {f"option_{index + 1}": option for index, option in enumerate(options[:option_count])} + ) + result.append(row) + return result + + +def _columns(option_count: int = 3) -> dict[str, str]: + columns = { + "question_id": "qid", + "question_text": "stem", + "correct_option": "gold", + } + columns.update( + {f"choice_{chr(97 + index)}": f"option_{index + 1}" for index in range(option_count)} + ) + return columns + + +def _declaration( + *, + source_ids: tuple[str, ...] = ("reference",), + selection_source_id: str = "reference", + expected_question_ids: tuple[str, ...] = ("q2", "q1"), + reference_kind: str = "independent_input_snapshot", + trust_label: str = "publisher-input", + columns: dict[str, str] | None = None, + revision: str | None = "publisher-r1", + derivation: dict | None = None, + limitations: tuple[str, ...] = (), +) -> DatasetReferenceSpec: + return DatasetReferenceSpec( + dataset_id="historical-test", + benchmark_name="synthetic", + split="test", + reference_kind=reference_kind, # type: ignore[arg-type] + trust_label=trust_label, + source_ids=source_ids, + selection_source_id=selection_source_id, + expected_question_ids=expected_question_ids, + selection_seed=None, + selection_n_samples=None, + subject_filter=(), + selection_unknown_reasons={ + "selection_seed": "not recorded by producer", + "selection_n_samples": "not recorded by producer", + }, + columns=columns or _columns(), + revision=revision, + fingerprint="publisher-fingerprint", + derivation=derivation or {}, + limitations=limitations, + native_compatibility_identity=None, + ) + + +def _derived_declaration(**changes: object) -> DatasetReferenceSpec: + base = _declaration( + source_ids=("results-a", "results-b"), + selection_source_id="results-a", + reference_kind="profile_derived_reference_snapshot", + trust_label="cross-source-internal-consistency", + derivation={ + "method": "exact selected-row agreement", + "independent_source_groups": [["results-a"], ["results-b"]], + }, + ) + return replace(base, **changes) + + +def _semantic_payloads(dataset) -> tuple[dict, dict]: + artifact_payload = { + "schema_version": "choicebench.semantic-dataset.v1", + "benchmark": dataset.benchmark_name, + "split": dataset.split, + "content_digest": dataset_content_digest(dataset.artifact_frame), + } + selection_payload = { + "artifact_id": short_id("ds", artifact_payload), + "content_digest": dataset_content_digest(dataset.frame), + "sample_identities": dataset_sample_identities(dataset.frame), + "seed": None, + "n_samples": None, + "subject_filter": [], + } + return artifact_payload, selection_payload + + +def test_independent_reference_builds_full_semantic_identities_and_stable_order(): + dataset = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + + artifact_payload, selection_payload = _semantic_payloads(dataset) + assert dataset.selected_question_ids == ("q2", "q1") + assert dataset.frame["question_id"].tolist() == ["q2", "q1"] + assert dataset.artifact_digest == integrity_digest(artifact_payload) + assert dataset.artifact_id == short_id("ds", artifact_payload) + assert dataset.artifact_payload == artifact_payload + assert dataset.selection_digest == integrity_digest(selection_payload) + assert dataset.selection_id == short_id("sel", selection_payload) + assert dataset.selection_payload == selection_payload + assert dataset.identity_mode == "imported_semantic_fallback" + assert dataset.selection_unknown_reasons == { + "selection_seed": "not recorded by producer", + "selection_n_samples": "not recorded by producer", + } + assert dataset.question_set_digest == integrity_digest(["q2", "q1"]) + assert dataset.reference_kind == "independent_input_snapshot" + + +def test_reference_kinds_are_not_equivalent(): + rows = _rows() + independent = build_expected_dataset( + _declaration(), {"reference": _source("reference", rows)} + ) + derived = build_expected_dataset( + _derived_declaration(), + { + "results-a": _source("results-a", rows), + "results-b": _source("results-b", rows), + }, + ) + + assert independent.artifact_id == derived.artifact_id + assert independent.selection_id == derived.selection_id + assert independent.reference_kind == "independent_input_snapshot" + assert derived.reference_kind == "profile_derived_reference_snapshot" + assert independent.derivation_digest != derived.derivation_digest + assert independent.snapshot_digest != derived.snapshot_digest + assert "not independently authenticated" in derived.limitations + + +def test_profile_derived_reference_requires_two_independent_source_groups(): + declaration = _derived_declaration( + derivation={ + "method": "exact agreement", + "independent_source_groups": [["results-a", "results-b"]], + } + ) + with pytest.raises(DatasetReferenceError, match="two independent source groups"): + build_expected_dataset( + declaration, + { + "results-a": _source("results-a", _rows()), + "results-b": _source("results-b", _rows()), + }, + ) + + +@pytest.mark.parametrize( + ("field", "replacement", "message"), + [ + ("stem", "Disagrees?", "question_text"), + ("gold", "C", "correct_option"), + ("option_2", "different option", "ordered options"), + ], +) +def test_profile_derived_reference_refuses_cross_source_semantic_disagreement( + field: str, replacement: str, message: str +): + changed = _rows() + changed[0][field] = replacement + with pytest.raises(DatasetReferenceError, match=message): + build_expected_dataset( + _derived_declaration(expected_question_ids=("q1", "q2")), + { + "results-a": _source("results-a", _rows()), + "results-b": _source("results-b", changed), + }, + ) + + +def test_derivation_digest_binds_complete_stable_reference_chain_not_audit_path(): + declaration = _declaration() + first = build_expected_dataset( + declaration, + { + "reference": _source( + "reference", _rows(), audit_path=Path("/machine-a/input.csv") + ) + }, + ) + relocated = build_expected_dataset( + declaration, + { + "reference": _source( + "reference", _rows(), audit_path=Path("/machine-b/input.csv") + ) + }, + ) + changed_chain = build_expected_dataset( + declaration, + { + "reference": _source( + "reference", _rows(), logical_path="publisher/mirror/input.csv" + ) + }, + ) + + assert first.derivation_digest == integrity_digest(first.derivation) + assert relocated.derivation_digest == first.derivation_digest + assert changed_chain.derivation_digest != first.derivation_digest + + +def test_unknown_publisher_revision_is_an_explicit_limitation(): + dataset = build_expected_dataset( + _declaration(revision=None), {"reference": _source("reference", _rows())} + ) + assert "publisher revision was not recorded" in dataset.limitations + + +def test_duplicate_selected_question_id_is_refused(): + with pytest.raises(DatasetReferenceError, match="duplicate selected question ID.*q1"): + build_expected_dataset( + _declaration(expected_question_ids=("q1", "q1")), + {"reference": _source("reference", _rows())}, + ) + + +def test_missing_selected_question_id_is_refused(): + with pytest.raises(DatasetReferenceError, match="missing selected question ID.*absent"): + build_expected_dataset( + _declaration(expected_question_ids=("q1", "absent")), + {"reference": _source("reference", _rows())}, + ) + + +def test_duplicate_source_question_id_is_refused(): + rows = _rows() + rows.append(dict(rows[0])) + with pytest.raises(DatasetReferenceError, match="duplicate question_id.*q1"): + build_expected_dataset( + _declaration(expected_question_ids=("q1",)), + {"reference": _source("reference", rows)}, + ) + + +def test_six_options_are_preserved_in_order(): + dataset = build_expected_dataset( + _declaration(columns=_columns(6), expected_question_ids=("q1",)), + {"reference": _source("reference", _rows(six_options=True))}, + ) + assert build_option_map(dataset.frame.iloc[0].to_dict()) == { + "A": "one", + "B": "two", + "C": "three", + "D": "four", + "E": "five", + "F": "six", + } + + +def test_arc_style_three_options_have_no_phantom_fourth_option(): + dataset = build_expected_dataset( + _declaration(expected_question_ids=("q1",)), + {"reference": _source("reference", _rows())}, + ) + assert build_option_map(dataset.frame.iloc[0].to_dict()) == { + "A": "one", + "B": "two", + "C": "three", + } + assert "choice_d" not in dataset.frame.columns + + +def test_source_mapping_trust_and_logical_provenance_are_not_semantic_identity(): + ordinary = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + renamed_rows = [ + { + "id": row["qid"], + "prompt": row["stem"], + "answer": row["gold"], + "a": row["option_1"], + "b": row["option_2"], + "c": row["option_3"], + } + for row in _rows() + ] + remapped = build_expected_dataset( + _declaration( + trust_label="archive-copy", + columns={ + "question_id": "id", + "question_text": "prompt", + "correct_option": "answer", + "choice_a": "a", + "choice_b": "b", + "choice_c": "c", + }, + ), + { + "reference": _source( + "reference", renamed_rows, logical_path="archive/remapped.csv" + ) + }, + ) + + assert remapped.artifact_id == ordinary.artifact_id + assert remapped.artifact_digest == ordinary.artifact_digest + assert remapped.selection_id == ordinary.selection_id + assert remapped.selection_digest == ordinary.selection_digest + assert remapped.derivation_digest != ordinary.derivation_digest + assert remapped.snapshot_digest != ordinary.snapshot_digest + + +def test_semantic_content_membership_and_order_change_applicable_identities(): + base_rows = _rows() + base = build_expected_dataset( + _declaration(), {"reference": _source("reference", base_rows)} + ) + changed_gold_rows = [dict(row) for row in base_rows] + changed_gold_rows[1]["gold"] = "B" + changed_gold = build_expected_dataset( + _declaration(), {"reference": _source("reference", changed_gold_rows)} + ) + changed_option_rows = [dict(row) for row in base_rows] + changed_option_rows[1]["option_1"] = "scarlet" + changed_option = build_expected_dataset( + _declaration(), {"reference": _source("reference", changed_option_rows)} + ) + changed_membership = build_expected_dataset( + _declaration(expected_question_ids=("q2", "q3")), + {"reference": _source("reference", base_rows)}, + ) + changed_order = build_expected_dataset( + _declaration(expected_question_ids=("q1", "q2")), + {"reference": _source("reference", base_rows)}, + ) + + for changed in (changed_gold, changed_option): + assert changed.artifact_id != base.artifact_id + assert changed.selection_id != base.selection_id + for changed in (changed_membership, changed_order): + assert changed.artifact_id == base.artifact_id + assert changed.selection_id != base.selection_id + + +def test_snapshot_is_written_atomically_and_self_validates(tmp_path: Path): + dataset = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + record = write_expected_snapshot(tmp_path, dataset) + + assert record["artifact_digest"] == dataset.artifact_digest + assert record["selection_digest"] == dataset.selection_digest + assert record["derivation_digest"] == dataset.derivation_digest + assert record["reference_kind"] == dataset.reference_kind + assert record["selected_question_ids"] == ["q2", "q1"] + assert (tmp_path / record["run_snapshot_path"]).is_file() + assert (tmp_path / record["reference_metadata_path"]).is_file() + validate_expected_snapshot(tmp_path, record) + + snapshot = tmp_path / record["run_snapshot_path"] + snapshot.write_text(snapshot.read_text(encoding="utf-8") + "tampered", encoding="utf-8") + with pytest.raises(DatasetReferenceError, match="snapshot.*integrity"): + validate_expected_snapshot(tmp_path, record) + + +def test_snapshot_self_validation_preserves_known_selection_semantics(tmp_path: Path): + declaration = replace( + _declaration(), + selection_seed=17, + selection_n_samples=2, + subject_filter=("science",), + selection_unknown_reasons={}, + ) + dataset = build_expected_dataset( + declaration, {"reference": _source("reference", _rows())} + ) + record = write_expected_snapshot(tmp_path, dataset) + + assert record["selection_semantics"] == { + "seed": 17, + "n_samples": 2, + "subject_filter": ["science"], + } + validate_expected_snapshot(tmp_path, record) + + +def test_known_selection_count_must_equal_selected_question_coverage(): + declaration = replace( + _declaration(), + selection_n_samples=99, + selection_unknown_reasons={"selection_seed": "not recorded"}, + ) + with pytest.raises(DatasetReferenceError, match="n_samples|selected|coverage"): + build_expected_dataset( + declaration, + {"reference": _source("reference", _rows())}, + ) + + +def test_validated_native_compatibility_identity_is_preserved_exactly(): + full_artifact = build_expected_dataset( + _declaration(expected_question_ids=("q1", "q2", "q3")), + {"reference": _source("reference", _rows())}, + ) + native_artifact_payload = { + "spec": { + "benchmark": "synthetic", + "split": "test", + "hf_path": "publisher/synthetic", + "hf_subset": None, + "source_revision": "immutable-r1", + "normalization_version": NORMALIZATION_VERSION, + "transforms": [], + "output_name": "synthetic", + }, + "content_digest": dataset_content_digest(full_artifact.frame), + "source": {"revision": "immutable-r1"}, + } + artifact_digest = integrity_digest(native_artifact_payload) + artifact_id = short_id("ds", native_artifact_payload) + native_selection_payload = { + "artifact_id": artifact_id, + "content_digest": dataset_content_digest(full_artifact.frame.iloc[[1, 0]]), + "sample_identities": dataset_sample_identities(full_artifact.frame.iloc[[1, 0]]), + "seed": 23, + "n_samples": 2, + "subject_filter": ["science"], + } + native = { + "artifact_payload": native_artifact_payload, + "artifact_digest": artifact_digest, + "artifact_id": artifact_id, + "selection_payload": native_selection_payload, + "selection_digest": integrity_digest(native_selection_payload), + "selection_id": short_id("sel", native_selection_payload), + } + declaration = replace( + _declaration(), + selection_seed=23, + selection_n_samples=2, + subject_filter=("science",), + selection_unknown_reasons={}, + native_compatibility_identity=native, + ) + + dataset = build_expected_dataset( + declaration, {"reference": _source("reference", _rows())} + ) + + assert dataset.identity_mode == "native_compatibility" + assert dataset.artifact_payload == native_artifact_payload + assert dataset.selection_payload == native_selection_payload + assert dataset.artifact_id == artifact_id + assert dataset.selection_id == native["selection_id"] + assert len(dataset.artifact_frame) == 3 + assert len(dataset.frame) == 2 + + +def test_native_compatibility_identity_must_own_the_selected_semantic_rows(): + fallback = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + artifact_payload = { + "spec": { + "benchmark": "synthetic", + "split": "test", + "hf_path": None, + "hf_subset": None, + "source_revision": None, + "normalization_version": NORMALIZATION_VERSION, + "transforms": [], + "output_name": "synthetic", + }, + "content_digest": "0" * 64, + "source": {}, + } + artifact_id = short_id("ds", artifact_payload) + selection_payload = { + "artifact_id": artifact_id, + "content_digest": dataset_content_digest(fallback.frame), + "sample_identities": dataset_sample_identities(fallback.frame), + "seed": 1, + "n_samples": 2, + "subject_filter": [], + } + native = { + "artifact_payload": artifact_payload, + "artifact_digest": integrity_digest(artifact_payload), + "artifact_id": artifact_id, + "selection_payload": selection_payload, + "selection_digest": integrity_digest(selection_payload), + "selection_id": short_id("sel", selection_payload), + } + declaration = replace( + _declaration(), + selection_seed=1, + selection_n_samples=2, + selection_unknown_reasons={}, + native_compatibility_identity=native, + ) + + with pytest.raises(DatasetReferenceError, match="native compatibility.*content"): + build_expected_dataset( + declaration, {"reference": _source("reference", _rows())} + ) + + +def test_unselected_semantic_content_changes_artifact_and_selection_identity(): + rows = _rows() + original = build_expected_dataset( + _declaration(), {"reference": _source("reference", rows)} + ) + changed_rows = [dict(row) for row in rows] + changed_rows[2]["stem"] = "Changed unselected question?" + changed = build_expected_dataset( + _declaration(), {"reference": _source("reference", changed_rows)} + ) + + assert dataset_content_digest(original.frame) == dataset_content_digest(changed.frame) + assert original.artifact_id != changed.artifact_id + assert original.selection_id != changed.selection_id + + +@pytest.mark.parametrize("source_index", ["not-an-integer", -1, 1]) +def test_structured_choice_source_indices_fail_closed(source_index): + rows = [ + { + "qid": "q1", + "stem": "First?", + "gold": "A", + "choices": json.dumps( + [ + {"text": "one", "source_index": source_index}, + {"text": "two", "source_index": 1}, + ] + ), + } + ] + declaration = _declaration( + expected_question_ids=("q1",), + columns={ + "question_id": "qid", + "question_text": "stem", + "correct_option": "gold", + "choices_json": "choices", + }, + ) + + with pytest.raises(DatasetReferenceError, match="source_index"): + build_expected_dataset( + declaration, {"reference": _source("reference", rows)} + ) + + +def test_ordered_options_reject_an_empty_middle_value(): + rows = _rows() + rows[0]["option_2"] = "" + + with pytest.raises(DatasetReferenceError, match="empty option.*followed"): + build_expected_dataset( + _declaration(expected_question_ids=("q1",)), + {"reference": _source("reference", rows)}, + ) + + +def test_ordered_option_mapping_requires_contiguous_semantic_keys(): + declaration = _declaration( + expected_question_ids=("q1",), + columns={ + "question_id": "qid", + "question_text": "stem", + "correct_option": "gold", + "choice_a": "option_1", + "choice_c": "option_3", + }, + ) + + with pytest.raises(DatasetReferenceError, match="contiguous"): + build_expected_dataset( + declaration, {"reference": _source("reference", _rows())} + ) + + +def test_current_normalized_integer_fields_preserve_native_semantic_types(): + rows = _rows() + for row in rows: + row["correct_index"] = str(ord(row["gold"]) - ord("A")) + row["n_choices"] = "3" + declaration = _declaration( + columns={ + **_columns(), + "correct_index": "correct_index", + "n_choices": "n_choices", + } + ) + + dataset = build_expected_dataset( + declaration, {"reference": _source("reference", rows)} + ) + + assert dataset.artifact_frame["correct_index"].tolist() == [1, 0, 2] + assert dataset.artifact_frame["n_choices"].tolist() == [3, 3, 3] + + +@pytest.mark.parametrize("mutation", ["schema", "extra_field"]) +def test_snapshot_record_schema_is_exact_and_fail_closed(tmp_path: Path, mutation: str): + dataset = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + record = write_expected_snapshot(tmp_path, dataset) + if mutation == "schema": + record["schema_version"] = "choicebench.expected-dataset.v999" + else: + record["unexpected"] = "not allowed" + unsigned = dict(record) + unsigned.pop("record_digest") + record["record_digest"] = integrity_digest(unsigned) + + with pytest.raises(DatasetReferenceError, match="schema|fields"): + validate_expected_snapshot(tmp_path, record) + + +def test_snapshot_validation_rejects_symlinked_artifact_paths(tmp_path: Path): + dataset = build_expected_dataset( + _declaration(), {"reference": _source("reference", _rows())} + ) + record = write_expected_snapshot(tmp_path, dataset) + snapshot = tmp_path / record["run_snapshot_path"] + outside = tmp_path.parent / f"{tmp_path.name}-outside.csv" + outside.write_bytes(snapshot.read_bytes()) + snapshot.unlink() + snapshot.symlink_to(outside) + + with pytest.raises(DatasetReferenceError, match="symlink|unsafe"): + validate_expected_snapshot(tmp_path, record) diff --git a/tests/importing/test_engine.py b/tests/importing/test_engine.py new file mode 100644 index 0000000..5bc69e2 --- /dev/null +++ b/tests/importing/test_engine.py @@ -0,0 +1,296 @@ +"""End-to-end tests for the reduced-scope import engine: base (non-overlay) +plan/dry-run/real-import/idempotence/verify. Overlay merge-into-new-run +orchestration (execute_overlay_import) is a separate function covered by +test_engine_overlay.py; this file covers only the base import path. +""" + +from __future__ import annotations + +from hashlib import sha256 +from pathlib import Path + +import pytest +import yaml + +from choicebench.importing.engine import ( + ImportEngineError, + ImportRequest, + build_import_plan, + execute_import, + serialize_import_report, + verify_import_run, +) +from choicebench.importing.schema import load_import_spec + +_RESULTS_CSV = "qid,question,choice_a,choice_b,answer,gold,score\nq1,One?,x,y,a,a,0.9\nq2,Two?,m,n,b,b,0.7\n" + + +def _raw_spec( + tmp_path: Path, results_csv: str = _RESULTS_CSV, question_ids: tuple[str, ...] = ("q1", "q2") +) -> dict: + results_path = tmp_path / "results.csv" + results_path.write_bytes(results_csv.encode("utf-8")) + expected_sha256 = sha256(results_csv.encode("utf-8")).hexdigest() + return { + "schema_version": "choicebench.import-spec.v1", + "import_name": "engine test import", + "sources": [ + { + "source_id": "results", + "path": str(results_path), + "logical_path": "freeze/results.csv", + "expected_sha256": expected_sha256, + "format": "csv", + "format_version": "producer-v1", + "classification": "raw", + "dialect": { + "encoding": "utf-8", "bom_policy": "forbid", "decoding_errors": "strict", + "delimiter": ",", "quote_character": '"', "escape_character": None, + "double_quote": True, "line_terminators": ["crlf", "lf", "cr"], + "mixed_line_terminators": "allow", "final_record_without_terminator": "allow", + "blank_record_policy": "reject", "skip_initial_space": False, + "header": "first_logical_record", "strict_syntax": True, + }, + "columns": { + "question_id": "qid", "question_text": "question", + "correct_option": "gold", "prediction": "answer", + }, + "expected_columns": ["qid", "question", "choice_a", "choice_b", "answer", "gold", "score"], + "ignored_columns": {"score": "producer aggregate only"}, + "null_values": ["", "NA"], + "numeric_columns": [ + {"source_column": "score", "value_type": "float", "null_allowed": True, "finite_only": True} + ], + "option_mapping": { + "mode": "ordered_columns", "ordered_columns": ["choice_a", "choice_b"], + "structured_column": None, "structured_label_key": None, "structured_text_key": None, + }, + "extra_field_policy": "preserve_unmapped", + "preserve_namespace": "producer", + "source_run_id": None, "source_repository": None, "source_commit": None, + "notes": {}, + } + ], + "datasets": [ + { + "dataset_id": "dataset", + "benchmark_name": "historical-benchmark", + "split": "test", + "reference_kind": "independent_input_snapshot", + "trust_label": "producer-supplied", + "source_ids": ["results"], + "selection_source_id": "results", + "expected_question_ids": list(question_ids), + "selection_seed": None, + "selection_n_samples": None, + "subject_filter": [], + "selection_unknown_reasons": { + "selection_seed": "not recorded by producer", + "selection_n_samples": "not recorded by producer", + }, + "columns": { + "question_id": "qid", "question_text": "question", "correct_option": "gold", + "choice_a": "choice_a", "choice_b": "choice_b", + }, + "revision": None, "fingerprint": None, "derivation": {}, + "limitations": ["publisher revision was not recorded"], + "native_compatibility_identity": None, + } + ], + "models": [ + { + "model_key": "model", "display_name": "historical-model", "backend": None, + "provider": None, "revision": None, "effective_parameters": {}, + "unknown_reasons": { + "backend": "not recorded by producer", "provider": "not recorded by producer", + "revision": "not recorded by producer", + }, + "native_compatibility_identity": None, + } + ], + "methods": [ + { + "method_key": "method", "name": "historical-direct", "effective_parameters": {}, + "implementation": None, + "unknown_reasons": {"implementation": "not recorded by producer"}, + "native_compatibility_identity": None, + } + ], + "prompts": [ + { + "prompt_key": "prompt", "template_identity": None, "template_digest": None, + "template_contents": None, "unknown_reason": "not recorded by producer", + "native_compatibility_identity": None, + } + ], + "conditions": [ + { + "condition_key": "condition", "source_ids": ["results"], "dataset_id": "dataset", + "model_key": "model", "method_key": "method", "prompt_key": "prompt", "seed": None, + "calibration_identity": None, "preflight_identity": None, "protocol_settings": {}, + "generation_parameters": {}, + "unknown_reasons": { + "seed": "not recorded by producer", + "calibration_identity": "not recorded by producer", + "preflight_identity": "not recorded by producer", + }, + "expected_question_ids": list(question_ids), + "evidence_status": "complete", "scope_disposition": "included", "executable": None, + "qualifications": [], "limitations": [], "damaged_question_ids": [], + "recoverable_question_ids": [], + "result_origin": { + "derivation_origin": "external_import", + "default_prediction_origin": "external_historical_inference", + "per_question_prediction_origins": {}, + }, + } + ], + "authorizations": [], + "overlays": [], + "metrics": ["accuracy"], + "provenance": {"producer_request_id": {"value": None, "reason": "not recorded by producer"}}, + "audit": {"source_path": str(tmp_path), "imported_at": "2026-07-18T00:00:00Z"}, + } + + +def _spec(tmp_path: Path, **kwargs): + raw = _raw_spec(tmp_path, **kwargs) + spec_path = tmp_path / "import.yaml" + spec_path.write_text(yaml.safe_dump(raw, sort_keys=False), encoding="utf-8") + return load_import_spec(spec_path) + + +def test_build_import_plan_produces_a_valid_v3_manifest(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + plan = build_import_plan(request) + assert plan.manifest["schema_version"] == "choicebench.manifest.v3" + assert len(plan.manifest["payload"]["semantic_conditions"]) == 1 + assert len(plan.manifest["payload"]["realizations"]) == 1 + (result,) = plan.realizations.values() + assert result.computed_evidence_status == "complete" + assert result.evaluable is True + assert len(result.evaluable_rows) == 2 + + +def test_build_import_plan_uses_a_supplied_expected_dataset_override(tmp_path): + spec = _spec(tmp_path) + (dataset,) = spec.datasets + (source,) = spec.sources + from choicebench.importing.dataset_reference import build_expected_dataset + from choicebench.importing.evidence import open_verified_source + + opened = open_verified_source(source, containment_root=tmp_path) + override = build_expected_dataset(dataset, {"results": opened}) + request = ImportRequest( + spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True, + expected_datasets={"dataset": override}, + ) + plan = build_import_plan(request) + assert plan.expected_datasets["dataset"] is override + + +def test_build_import_plan_rejects_an_incomplete_expected_dataset_override(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest( + spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True, + expected_datasets={}, + ) + with pytest.raises(ImportEngineError, match="missing dataset"): + build_import_plan(request) + + +def test_build_import_plan_rejects_a_source_outside_the_workspace_by_default(tmp_path): + external_root = tmp_path / "external" + external_root.mkdir() + workspace_root = tmp_path / "workspace" + workspace_root.mkdir() + spec = _spec(external_root) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=workspace_root, strict=True) + with pytest.raises(Exception, match="escapes containment root"): + build_import_plan(request) + + +def test_build_import_plan_accepts_a_source_outside_the_workspace_with_an_explicit_containment_root( + tmp_path, +): + external_root = tmp_path / "external" + external_root.mkdir() + workspace_root = tmp_path / "workspace" + workspace_root.mkdir() + spec = _spec(external_root) + request = ImportRequest( + spec=spec, run_id="run-1", workspace_root=workspace_root, strict=True, + source_containment_root=external_root, + ) + plan = build_import_plan(request) + assert len(plan.manifest["payload"]["realizations"]) == 1 + + +def test_dry_run_writes_nothing(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + report = execute_import(request, dry_run=True) + assert report.import_state == "validated" + assert report.wrote_artifacts is False + assert not (tmp_path / "runs").exists() + assert report.counts.evidence_status == {"complete": 1} + assert report.counts.scope_disposition == {"included": 1} + + +def test_real_import_writes_manifest_state_and_result(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + report = execute_import(request, dry_run=False) + assert report.import_state == "imported" + assert report.wrote_artifacts is True + assert report.idempotent_noop is False + run_dir = tmp_path / "runs" / "run-1" + assert (run_dir / "manifest.json").is_file() + assert (run_dir / "run_state.json").is_file() + (realization_id,) = report.realization_digests.keys() + assert (run_dir / f"results/{realization_id}.csv").is_file() + assert (run_dir / f"artifacts/imports/validation/{realization_id}.json").is_file() + assert (run_dir / "artifacts/imports/evidence/index.json").is_file() + + +def test_real_import_is_idempotent_noop_on_repeat(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + execute_import(request, dry_run=False) + second = execute_import(request, dry_run=False) + assert second.idempotent_noop is True + assert second.wrote_artifacts is False + + +def test_verify_import_run_validates_the_published_graph(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + execute_import(request, dry_run=False) + verified = verify_import_run(tmp_path / "runs" / "run-1") + (realization,) = verified.realizations.values() + assert realization.result_sha256 is not None + assert set(realization.rows_by_question_id) == {"q1", "q2"} + assert realization.prediction_origins == { + "q1": "external_historical_inference", "q2": "external_historical_inference", + } + + +def test_execute_import_reports_failure_without_writing_on_invalid_declaration(tmp_path): + spec = _spec(tmp_path, results_csv="qid,question,choice_a,choice_b,answer,gold,score\nq1,One?,x,y,a,a,0.9\n") + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + report = execute_import(request, dry_run=False) + assert report.import_state == "failed" + assert report.wrote_artifacts is False + assert not (tmp_path / "runs").exists() + assert report.failures + + +def test_serialize_import_report_round_trips_to_plain_dict(tmp_path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + report = execute_import(request, dry_run=True) + serialized = serialize_import_report(report) + assert serialized["import_state"] == "validated" + assert serialized["counts"]["evidence_status"] == {"complete": 1} + assert serialized["failures"] == [] diff --git a/tests/importing/test_engine_overlay.py b/tests/importing/test_engine_overlay.py new file mode 100644 index 0000000..fd23bc9 --- /dev/null +++ b/tests/importing/test_engine_overlay.py @@ -0,0 +1,705 @@ +"""Tests for the overlay merge-into-new-run orchestration +(execute_overlay_import / _merge_overlay_realization). Builds a real base run +via test_engine.py's fixtures, then applies an authorized overlay on top. +""" + +from __future__ import annotations + +from hashlib import sha256 +import json +from pathlib import Path + +import pytest + +from choicebench.identity import short_id +from choicebench.importing.dataset_reference import build_expected_dataset +from choicebench.importing.engine import ( + ImportEngineError, + ImportRequest, + OverlayImportRequest, + execute_import, + execute_overlay_import, + verify_import_run, +) +from choicebench.importing.evidence import open_verified_source +from choicebench.importing.schema import ( + AuthorizationSpec, + CsvDialectSpec, + OptionMappingSpec, + OverlaySpec, + ResultOriginSpec, + SourceArtifactSpec, +) +import yaml + +from choicebench.importing.schema import load_import_spec +from tests.importing.test_engine import _raw_spec, _spec + +_AUTH_CSV = b"authorization,payload\n1,2\n" +_OVERLAY_CSV = "qid,predicted_letter\nq1,B\n" + + +def _make_base_run(tmp_path: Path): + spec = _spec(tmp_path) + request = ImportRequest(spec=spec, run_id="base-run", workspace_root=tmp_path, strict=True) + execute_import(request, dry_run=False) + base_run_dir = tmp_path / "runs" / "base-run" + verified = verify_import_run(base_run_dir) + (base,) = verified.realizations.values() + return base_run_dir, base + + +def _base_expected_dataset(tmp_path: Path, **spec_kwargs): + spec = _spec(tmp_path, **spec_kwargs) + (dataset_ref,) = spec.datasets + (results_source,) = spec.sources + opened_results = open_verified_source(results_source, containment_root=tmp_path) + return build_expected_dataset(dataset_ref, opened_sources={"results": opened_results}) + + +def _authorization_source(tmp_path: Path) -> SourceArtifactSpec: + auth_path = tmp_path / "authorization.csv" + auth_path.write_bytes(_AUTH_CSV) + return SourceArtifactSpec( + source_id="auth-source", + path=auth_path, + logical_path="inputs/authorization.csv", + expected_sha256=sha256(_AUTH_CSV).hexdigest(), + format="csv", + format_version="producer-v1", + classification="raw", + dialect=CsvDialectSpec(), + columns={}, + expected_columns=("authorization", "payload"), + ignored_columns={}, + null_values=(), + numeric_columns=(), + option_mapping=OptionMappingSpec( + mode="ordered_columns", ordered_columns=(), structured_column=None, + structured_label_key=None, structured_text_key=None, + ), + extra_field_policy="preserve_unmapped", + preserve_namespace="producer", + source_run_id=None, source_repository=None, source_commit=None, + notes={}, + ) + + +def _overlay_source(tmp_path: Path, csv_text: str = _OVERLAY_CSV, filename: str = "overlay.csv") -> SourceArtifactSpec: + overlay_path = tmp_path / filename + overlay_path.write_bytes(csv_text.encode("utf-8")) + return SourceArtifactSpec( + source_id="overlay-source", + path=overlay_path, + logical_path="inputs/overlay.csv", + expected_sha256=sha256(csv_text.encode("utf-8")).hexdigest(), + format="csv", + format_version="producer-v1", + classification="repaired", + dialect=CsvDialectSpec(), + columns={"question_id": "qid", "prediction": "predicted_letter"}, + expected_columns=("qid", "predicted_letter"), + ignored_columns={}, + null_values=(), + numeric_columns=(), + option_mapping=OptionMappingSpec( + mode="ordered_columns", ordered_columns=(), structured_column=None, + structured_label_key=None, structured_text_key=None, + ), + extra_field_policy="preserve_unmapped", + preserve_namespace="producer", + source_run_id=None, source_repository=None, source_commit=None, + notes={}, + ) + + +def _authorization( + *, base, authorization_type: str, executable: bool, + grants: dict[str, str] | None = None, +) -> AuthorizationSpec: + grants = grants or {"q1": "malformed prediction"} + payload = { + "schema_version": "choicebench.authorization-bundle.v1", + "authorization_type": authorization_type, + "grants": {base.condition_digest: grants}, + "authority": "principal-investigator", + "purpose": "repair", + "executable": executable, + "source_sha256": sha256(_AUTH_CSV).hexdigest(), + "input_evidence_digests": dict(base.evidence_source_digests), + "expected_snapshot_digests": {base.condition_digest: "2" * 64}, + } + authorization_id = short_id("auth", payload) + return AuthorizationSpec( + authorization_id=authorization_id, + authorization_type=authorization_type, + source_id="auth-source", + condition_question_reasons={base.condition_digest: grants}, + authority="principal-investigator", + purpose="repair", + executable=executable, + input_evidence_digests=dict(base.evidence_source_digests), + expected_snapshot_digests={base.condition_digest: "2" * 64}, + ) + + +def _overlay_spec(*, base, authorization: AuthorizationSpec) -> OverlaySpec: + return OverlaySpec( + overlay_id="ovl_1", + base_run_path=Path("/tmp/base-run"), + base_condition_digest=base.condition_digest, + base_realization_id=base.realization_id, + base_realization_digest=base.realization_digest, + base_evidence_digests=dict(base.evidence_source_digests), + base_validation_artifact_sha256=base.validation_artifact_sha256, + base_result_sha256=base.result_sha256, + source_id="overlay-source", + authorization_id=authorization.authorization_id, + replacement_reasons={"q1": "repair malformed answer"}, + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", + default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + lineage_notes={}, + implementation={ + "qualified_name": "external:repair_tool", + "source_digest": sha256(_OVERLAY_CSV.encode("utf-8")).hexdigest(), + }, + input_digest="9" * 64, + preownership_output_digest="a" * 64, + expected_evidence_status="complete", + ) + + +def _make_overlay_request(tmp_path: Path, base_run_dir: Path, base, *, run_id: str = "overlay-run"): + auth_source = _authorization_source(tmp_path) + authorization = _authorization(base=base, authorization_type="inference_repair", executable=True) + overlay_source = _overlay_source(tmp_path) + overlay = _overlay_spec(base=base, authorization=authorization) + expected_dataset = _base_expected_dataset(tmp_path) + + return OverlayImportRequest( + base_run_dir=base_run_dir, + base_realization_id=base.realization_id, + authorization=authorization, + authorization_source=auth_source, + overlay=overlay, + overlay_source=overlay_source, + overlay_mapping={"question_id": "qid", "prediction": "predicted_letter"}, + condition_digests={"cond_key": base.condition_digest}, + expected_dataset=expected_dataset, + run_id=run_id, + workspace_root=tmp_path, + ) + + +def test_execute_overlay_import_merges_repair_into_a_new_run(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + request = _make_overlay_request(tmp_path, base_run_dir, base) + + report = execute_overlay_import(request, dry_run=False) + assert report.import_state == "imported" + assert report.wrote_artifacts is True + assert report.idempotent_noop is False + + overlay_run_dir = tmp_path / "runs" / "overlay-run" + verified = verify_import_run(overlay_run_dir) + (derived,) = verified.realizations.values() + assert set(derived.rows_by_question_id) == {"q1", "q2"} + assert derived.rows_by_question_id["q1"]["predicted_option"] == "B" + assert derived.prediction_origins["q1"] == "native_inference" + assert derived.prediction_origins["q2"] == "external_historical_inference" + assert (overlay_run_dir / "artifacts/imports/evidence/index.json").is_file() + + # the immutable base run is untouched + reverified_base = verify_import_run(base_run_dir) + (base_again,) = reverified_base.realizations.values() + assert base_again.realization_digest == base.realization_digest + + +_THREE_QUESTION_RESULTS_CSV = ( + "qid,question,choice_a,choice_b,answer,gold,score\n" + "q1,One?,x,y,a,a,0.9\n" + "q2,Two?,m,n,b,b,0.7\n" + "q3,Three?,p,q,a,a,0.8\n" +) + + +def test_execute_overlay_import_lineage_components_are_in_deterministic_dataset_order(tmp_path): + """Regression for a real bug found by review: with >=2 retained rows, + lineage_components was assembled by iterating a `set` of lineage_ids, + an order that depends on Python's per-process string hash + randomization -- and canonicalize() hashes list order, so the merged + realization/experiment digest was not reproducible across processes. + Uses a 3-question base (q1 replaced, q2/q3 retained -- 2 retained rows, + the minimum needed to exercise set-iteration nondeterminism) and + asserts lineage_components is in the same deterministic order as + row_assignments (expected_dataset question order), not "retained then + replaced" or any other order that could vary by run. + """ + spec = _spec(tmp_path, results_csv=_THREE_QUESTION_RESULTS_CSV, question_ids=("q1", "q2", "q3")) + request = ImportRequest(spec=spec, run_id="base-run", workspace_root=tmp_path, strict=True) + execute_import(request, dry_run=False) + base_run_dir = tmp_path / "runs" / "base-run" + verified = verify_import_run(base_run_dir) + (base,) = verified.realizations.values() + + auth_source = _authorization_source(tmp_path) + authorization = _authorization(base=base, authorization_type="inference_repair", executable=True) + overlay_source = _overlay_source(tmp_path) + overlay = _overlay_spec(base=base, authorization=authorization) + expected_dataset = _base_expected_dataset( + tmp_path, results_csv=_THREE_QUESTION_RESULTS_CSV, question_ids=("q1", "q2", "q3"), + ) + request = OverlayImportRequest( + base_run_dir=base_run_dir, + base_realization_id=base.realization_id, + authorization=authorization, + authorization_source=auth_source, + overlay=overlay, + overlay_source=overlay_source, + overlay_mapping={"question_id": "qid", "prediction": "predicted_letter"}, + condition_digests={"cond_key": base.condition_digest}, + expected_dataset=expected_dataset, + run_id="overlay-run", + workspace_root=tmp_path, + ) + report = execute_overlay_import(request, dry_run=False) + assert report.import_state == "imported", report.failures + + overlay_run_dir = tmp_path / "runs" / "overlay-run" + manifest = json.loads((overlay_run_dir / "manifest.json").read_text()) + (realization,) = manifest["payload"]["realizations"].values() + realization_identity = realization["identity"]["realization"] + + row_assignment_order = [ + assignment["question_id"] + for assignment in realization_identity["result_origin"]["row_assignments"] + ] + lineage_component_order = [ + component["identity"]["question_id"] + for component in realization_identity["lineage_components"] + ] + assert row_assignment_order == list(expected_dataset.selected_question_ids), ( + "row_assignments must follow expected_dataset's own question order, " + "not 'retained then replaced' concatenation order" + ) + assert lineage_component_order == row_assignment_order, ( + "lineage_components must be in the same deterministic order as " + "row_assignments (previously built from set(retained_lineage_ids), " + "which is nondeterministic across processes)" + ) + + +def test_execute_overlay_import_merges_offline_transformation_retaining_origin(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + + auth_source = _authorization_source(tmp_path) + authorization = _authorization( + base=base, authorization_type="offline_transformation", executable=False, + grants={"q2": "offline semantic rematch"}, + ) + overlay_csv = "qid,predicted_letter\nq2,A\n" + overlay_source = _overlay_source(tmp_path, csv_text=overlay_csv, filename="offline-overlay.csv") + overlay = OverlaySpec( + overlay_id="ovl_offline", + base_run_path=Path("/tmp/base-run"), + base_condition_digest=base.condition_digest, + base_realization_id=base.realization_id, + base_realization_digest=base.realization_digest, + base_evidence_digests=dict(base.evidence_source_digests), + base_validation_artifact_sha256=base.validation_artifact_sha256, + base_result_sha256=base.result_sha256, + source_id="overlay-source", + authorization_id=authorization.authorization_id, + replacement_reasons={"q2": "offline semantic rematch"}, + result_origin=ResultOriginSpec( + derivation_origin="offline_transformation", + default_prediction_origin=None, + per_question_prediction_origins={}, + ), + lineage_notes={}, + implementation={ + "qualified_name": "external:offline_rematcher", + "source_digest": sha256(overlay_csv.encode("utf-8")).hexdigest(), + }, + input_digest="9" * 64, + preownership_output_digest="a" * 64, + expected_evidence_status="complete", + ) + request = OverlayImportRequest( + base_run_dir=base_run_dir, + base_realization_id=base.realization_id, + authorization=authorization, + authorization_source=auth_source, + overlay=overlay, + overlay_source=overlay_source, + overlay_mapping={"question_id": "qid", "prediction": "predicted_letter"}, + condition_digests={"cond_key": base.condition_digest}, + expected_dataset=_base_expected_dataset(tmp_path), + run_id="offline-overlay-run", + workspace_root=tmp_path, + ) + + report = execute_overlay_import(request, dry_run=False) + assert report.import_state == "imported" + + overlay_run_dir = tmp_path / "runs" / "offline-overlay-run" + verified = verify_import_run(overlay_run_dir) + (derived,) = verified.realizations.values() + assert derived.rows_by_question_id["q2"]["predicted_option"] == "A" + # offline_transformation retains the underlying response's own origin + # rather than reassigning it (base.prediction_origins["q2"] is + # external_historical_inference) -- unlike inference_repair. + assert derived.prediction_origins["q2"] == "external_historical_inference" + assert derived.prediction_origins["q1"] == "external_historical_inference" + + +_REVERSE_ORDER_RESULTS_CSV = ( + "qid,question,choice_a,choice_b,answer,gold,score\n" + "qb,Qb?,x,y,a,a,0.9\n" + "qa,Qa?,m,n,b,b,0.7\n" + "qc,Qc?,p,q,a,a,0.8\n" +) + + +def test_execute_overlay_import_offline_transformation_digests_use_dataset_order(tmp_path): + """Regression for Fix B: the offline_transformation overlay's + transformation_input/preownership_output digests were computed from + derive_overlay's sorted(replacement_ids) order, but identity.py recomputes + them from the realization's own dataset-ordered lineage_components. When two + replaced IDs' dataset order (qb, qa) differs from their sorted order + (qa, qb), the two integrity_digests diverged and identity validation raised + a false-positive 'offline_transformation overlay ... digest conflicts with + its lineage components.' This exercises exactly that >=2-replaced, + non-sorted-dataset-order case with a real lineage graph (no mock digests). + """ + spec = _spec(tmp_path, results_csv=_REVERSE_ORDER_RESULTS_CSV, question_ids=("qb", "qa", "qc")) + request = ImportRequest(spec=spec, run_id="base-run", workspace_root=tmp_path, strict=True) + execute_import(request, dry_run=False) + base_run_dir = tmp_path / "runs" / "base-run" + verified = verify_import_run(base_run_dir) + (base,) = verified.realizations.values() + + auth_source = _authorization_source(tmp_path) + authorization = _authorization( + base=base, authorization_type="offline_transformation", executable=False, + grants={"qb": "offline rematch", "qa": "offline rematch"}, + ) + overlay_csv = "qid,predicted_letter\nqb,B\nqa,A\n" + overlay_source = _overlay_source(tmp_path, csv_text=overlay_csv, filename="offline-reverse.csv") + overlay = OverlaySpec( + overlay_id="ovl_offline_reverse", + base_run_path=Path("/tmp/base-run"), + base_condition_digest=base.condition_digest, + base_realization_id=base.realization_id, + base_realization_digest=base.realization_digest, + base_evidence_digests=dict(base.evidence_source_digests), + base_validation_artifact_sha256=base.validation_artifact_sha256, + base_result_sha256=base.result_sha256, + source_id="overlay-source", + authorization_id=authorization.authorization_id, + replacement_reasons={"qb": "offline rematch", "qa": "offline rematch"}, + result_origin=ResultOriginSpec( + derivation_origin="offline_transformation", + default_prediction_origin=None, + per_question_prediction_origins={}, + ), + lineage_notes={}, + implementation={ + "qualified_name": "external:offline_rematcher", + "source_digest": sha256(overlay_csv.encode("utf-8")).hexdigest(), + }, + input_digest="9" * 64, + preownership_output_digest="a" * 64, + expected_evidence_status="complete", + ) + request = OverlayImportRequest( + base_run_dir=base_run_dir, + base_realization_id=base.realization_id, + authorization=authorization, + authorization_source=auth_source, + overlay=overlay, + overlay_source=overlay_source, + overlay_mapping={"question_id": "qid", "prediction": "predicted_letter"}, + condition_digests={"cond_key": base.condition_digest}, + expected_dataset=_base_expected_dataset( + tmp_path, results_csv=_REVERSE_ORDER_RESULTS_CSV, question_ids=("qb", "qa", "qc"), + ), + run_id="offline-reverse-run", + workspace_root=tmp_path, + ) + report = execute_overlay_import(request, dry_run=False) + assert report.import_state == "imported", report.failures + + overlay_run_dir = tmp_path / "runs" / "offline-reverse-run" + verified = verify_import_run(overlay_run_dir) + (derived,) = verified.realizations.values() + assert derived.rows_by_question_id["qb"]["predicted_option"] == "B" + assert derived.rows_by_question_id["qa"]["predicted_option"] == "A" + # deterministic: a second import against a fresh workspace yields the same + # realization digest (identity re-verifies the digests each run). + second_ws = tmp_path / "second" + second_ws.mkdir() + spec2 = _spec(second_ws, results_csv=_REVERSE_ORDER_RESULTS_CSV, question_ids=("qb", "qa", "qc")) + execute_import( + ImportRequest(spec=spec2, run_id="base-run", workspace_root=second_ws, strict=True), + dry_run=False, + ) + assert derived.realization_digest # non-empty, and verify_import_run above already re-checked it + + +def test_execute_overlay_import_dry_run_writes_nothing(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + request = _make_overlay_request(tmp_path, base_run_dir, base) + + report = execute_overlay_import(request, dry_run=True) + assert report.import_state == "validated" + assert report.wrote_artifacts is False + assert not (tmp_path / "runs" / "overlay-run").exists() + + +def test_execute_overlay_import_is_idempotent_noop_on_repeat(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + request = _make_overlay_request(tmp_path, base_run_dir, base) + + execute_overlay_import(request, dry_run=False) + second = execute_overlay_import(request, dry_run=False) + assert second.idempotent_noop is True + assert second.wrote_artifacts is False + + +def test_execute_overlay_import_rejects_an_unknown_base_realization_id(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + request = _make_overlay_request(tmp_path, base_run_dir, base) + unknown = OverlayImportRequest( + base_run_dir=request.base_run_dir, + base_realization_id="not-the-real-id", + authorization=request.authorization, + authorization_source=request.authorization_source, + overlay=request.overlay, + overlay_source=request.overlay_source, + overlay_mapping=request.overlay_mapping, + condition_digests=request.condition_digests, + expected_dataset=request.expected_dataset, + run_id=request.run_id, + workspace_root=request.workspace_root, + ) + report = execute_overlay_import(unknown, dry_run=False) + assert report.import_state == "failed" + assert report.wrote_artifacts is False + assert not (tmp_path / "runs" / "overlay-run").exists() + + +def test_execute_overlay_import_refuses_to_overwrite_a_divergent_existing_run(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + request = _make_overlay_request(tmp_path, base_run_dir, base) + execute_overlay_import(request, dry_run=False) + + # a second, different overlay (different replaced answer) targeting the + # same run_id must be refused, not silently treated as a no-op + divergent_csv = "qid,predicted_letter\nq1,C\n" + divergent_overlay_source = _overlay_source(tmp_path, csv_text=divergent_csv, filename="divergent-overlay.csv") + divergent_overlay = OverlaySpec( + overlay_id="ovl_2", + base_run_path=request.overlay.base_run_path, + base_condition_digest=request.overlay.base_condition_digest, + base_realization_id=request.overlay.base_realization_id, + base_realization_digest=request.overlay.base_realization_digest, + base_evidence_digests=request.overlay.base_evidence_digests, + base_validation_artifact_sha256=request.overlay.base_validation_artifact_sha256, + base_result_sha256=request.overlay.base_result_sha256, + source_id="overlay-source", + authorization_id=request.overlay.authorization_id, + replacement_reasons={"q1": "repair malformed answer"}, + result_origin=request.overlay.result_origin, + lineage_notes={}, + implementation={ + "qualified_name": "external:repair_tool", + "source_digest": sha256(divergent_csv.encode("utf-8")).hexdigest(), + }, + input_digest="9" * 64, + preownership_output_digest="a" * 64, + expected_evidence_status="complete", + ) + divergent_request = OverlayImportRequest( + base_run_dir=request.base_run_dir, + base_realization_id=request.base_realization_id, + authorization=request.authorization, + authorization_source=request.authorization_source, + overlay=divergent_overlay, + overlay_source=divergent_overlay_source, + overlay_mapping=request.overlay_mapping, + condition_digests=request.condition_digests, + expected_dataset=request.expected_dataset, + run_id=request.run_id, + workspace_root=request.workspace_root, + ) + # matches execute_import's existing base-path contract (engine.py's + # verify_import_run(expected=...) check): a divergent overwrite is a + # loud failure that propagates out of the transaction, not a returned + # "failed" report -- publish() has already fully released its lock/ + # staging state by the time this raises. + with pytest.raises(ImportEngineError, match="different experiment"): + execute_overlay_import(divergent_request, dry_run=False) + + +def test_execute_overlay_import_refuses_a_forged_authorization_source(tmp_path): + base_run_dir, base = _make_base_run(tmp_path) + request = _make_overlay_request(tmp_path, base_run_dir, base) + request.authorization_source.path.write_bytes(b"tampered,bytes\n9,9\n") + + # the forged bytes no longer match expected_sha256; EvidenceError (a + # ValueError) is caught by execute_overlay_import's own error handling + # and surfaces as a failed report rather than a silent success. + report = execute_overlay_import(request, dry_run=False) + assert report.import_state == "failed" + assert report.wrote_artifacts is False + + +# --- Overlayable-vs-evaluable regression tests ------------------------------- +# +# Regression coverage for the fix distinguishing "overlayable" (a base with a +# full published row set of exactly the expected question-ID identity) from +# "evaluable" (that content currently scores cleanly). A malformed-but- +# structurally-complete base -- e.g. one unrelated row with an unparseable +# prediction, having nothing to do with the overlay's 3 authorized IDs -- must +# still accept an authorized overlay; a structurally incomplete base (missing +# question IDs) must still be refused. + +_MALFORMED_BUT_COMPLETE_RESULTS_CSV = ( + "qid,question,choice_a,choice_b,answer,gold,score\n" + "q1,One?,x,y,a,a,0.9\n" + "q2,Two?,m,n,b,b,0.7\n" + "q3,Three?,p,q,,a,0.5\n" # blank answer -> parse-missing, unrelated to the q1 overlay target +) + + +def _spec_with_declared_evidence_status(tmp_path: Path, *, evidence_status: str, **kwargs): + """Like test_engine.py's `_spec`, but overrides the condition's declared + `evidence_status` -- needed because strict=True requires the declared + status to match the COMPUTED one (`build_import_plan` fails closed on a + mismatch), and the default fixture always declares 'complete'.""" + raw = _raw_spec(tmp_path, **kwargs) + raw["conditions"][0]["evidence_status"] = evidence_status + spec_path = tmp_path / "import.yaml" + spec_path.write_text(yaml.safe_dump(raw, sort_keys=False), encoding="utf-8") + return load_import_spec(spec_path) + + +def _malformed_but_complete_base_run(tmp_path: Path): + spec = _spec_with_declared_evidence_status( + tmp_path, evidence_status="malformed", + results_csv=_MALFORMED_BUT_COMPLETE_RESULTS_CSV, question_ids=("q1", "q2", "q3"), + ) + request = ImportRequest(spec=spec, run_id="base-run", workspace_root=tmp_path, strict=True) + report = execute_import(request, dry_run=False) + assert report.import_state == "imported", report.failures + base_run_dir = tmp_path / "runs" / "base-run" + verified = verify_import_run(base_run_dir) + (base,) = verified.realizations.values() + return base_run_dir, base, verified + + +def test_malformed_but_structurally_complete_base_is_not_evaluable_but_is_overlayable(tmp_path): + base_run_dir, base, verified = _malformed_but_complete_base_run(tmp_path) + # structurally complete: all 3 expected IDs present exactly once, source + # hash verified (verify_import_run already re-verified this to build `base`). + assert set(base.rows_by_question_id) == {"q1", "q2", "q3"} + assert base.result_sha256 is not None + realization_id = base.realization_id + evidence = verified.manifest["payload"]["realizations"][realization_id]["identity"]["realization"]["evidence"] + # not evaluable: q3's blank answer makes this base malformed, not scoreable. + assert evidence["evidence_status"] == "malformed" + + +def test_execute_overlay_import_accepts_a_malformed_but_structurally_complete_base(tmp_path): + base_run_dir, base, _ = _malformed_but_complete_base_run(tmp_path) + auth_source = _authorization_source(tmp_path) + authorization = _authorization(base=base, authorization_type="inference_repair", executable=True) + overlay_source = _overlay_source(tmp_path) # replaces q1 -> "B" + overlay = _overlay_spec(base=base, authorization=authorization) + expected_dataset = _base_expected_dataset( + tmp_path, results_csv=_MALFORMED_BUT_COMPLETE_RESULTS_CSV, question_ids=("q1", "q2", "q3"), + ) + request = OverlayImportRequest( + base_run_dir=base_run_dir, + base_realization_id=base.realization_id, + authorization=authorization, + authorization_source=auth_source, + overlay=overlay, + overlay_source=overlay_source, + overlay_mapping={"question_id": "qid", "prediction": "predicted_letter"}, + condition_digests={"cond_key": base.condition_digest}, + expected_dataset=expected_dataset, + run_id="overlay-run", + workspace_root=tmp_path, + ) + + # previously this raised "has no published result to retain rows from; + # only an evaluable base realization can be overlaid" -- now it succeeds, + # because the base is overlayable (structurally complete) even though it + # is not evaluable (malformed). + report = execute_overlay_import(request, dry_run=False) + assert report.import_state == "imported", report.failures + + overlay_run_dir = tmp_path / "runs" / "overlay-run" + verified = verify_import_run(overlay_run_dir) + (derived,) = verified.realizations.values() + + # authorized row replaced. + assert derived.rows_by_question_id["q1"]["predicted_option"] == "B" + # unrelated rows retained byte/value-identical -- q2's original prediction + # unchanged, and q3's unparseable (blank) prediction is NOT repaired, + # filtered, or reinterpreted by the overlay. + assert derived.rows_by_question_id["q2"]["predicted_option"] == base.rows_by_question_id["q2"]["predicted_option"] + q3_predicted = derived.rows_by_question_id["q3"]["predicted_option"] + assert q3_predicted is None or str(q3_predicted).strip() == "" or str(q3_predicted).lower() == "nan" + + # evidence_status is RECOMPUTED from the actual merged content, not + # forced to the overlay's declared "complete" -- q3 is still unparseable, + # so the derived realization is still malformed. + realization_id = derived.realization_id + evidence = verified.manifest["payload"]["realizations"][realization_id]["identity"]["realization"]["evidence"] + assert evidence["evidence_status"] == "malformed" + # the overlay's own declared/intended historical status is preserved + # separately, never conflated with the recomputed observed status. + state = json_load_run_state(overlay_run_dir) + assert state["realizations"][realization_id]["declared_evidence_status"] == "complete" + assert state["realizations"][realization_id]["status"] == "published_unscored" + + # idempotent on repeat, same as the evaluable-base overlay path. + second = execute_overlay_import(request, dry_run=False) + assert second.idempotent_noop is True + assert second.wrote_artifacts is False + + +def json_load_run_state(run_dir: Path) -> dict: + return json.loads((run_dir / "run_state.json").read_text()) + + +def test_execute_overlay_import_rejects_a_structurally_incomplete_base(tmp_path): + # only 2 of the 3 expected question IDs actually appear in the results + # file -- a genuinely broken base, distinct from "malformed but complete". + incomplete_csv = ( + "qid,question,choice_a,choice_b,answer,gold,score\n" + "q1,One?,x,y,a,a,0.9\n" + "q2,Two?,m,n,b,b,0.7\n" + ) + spec = _spec_with_declared_evidence_status( + tmp_path, evidence_status="malformed", results_csv=incomplete_csv, + question_ids=("q1", "q2", "q3"), + ) + request = ImportRequest(spec=spec, run_id="base-run", workspace_root=tmp_path, strict=True) + # the generic engine already refuses an incomplete base at base-import + # time under strict=True (missing expected rows -> declared-vs-computed + # evidence-status mismatch, or an outright validation failure) -- meaning + # such a base can never even reach the overlay path in the first place, + # which is itself the "structurally incomplete base is rejected" + # guarantee the overlayable gate depends on. + report = execute_import(request, dry_run=False) + assert report.import_state == "failed" + assert report.wrote_artifacts is False + assert not (tmp_path / "runs" / "base-run").exists() + assert not (tmp_path / "runs" / "overlay-run").exists() diff --git a/tests/importing/test_evaluation.py b/tests/importing/test_evaluation.py new file mode 100644 index 0000000..5de630b --- /dev/null +++ b/tests/importing/test_evaluation.py @@ -0,0 +1,71 @@ +"""Tests for the reduced-scope v3 reader/evaluator (Task 13). Reuses +tests/importing/test_engine.py's ImportSpec fixture builder rather than +re-deriving a full spec here.""" + +from __future__ import annotations + +import pytest + +from choicebench.io.readers import ( + ResultSetError, + build_evaluation_report_for_result_set, + read_manifest_result_set, +) +from choicebench.importing.engine import ImportRequest, execute_import +from tests.importing.test_engine import _spec + + +def _imported_run(tmp_path, **spec_kwargs): + spec = _spec(tmp_path, **spec_kwargs) + request = ImportRequest(spec=spec, run_id="run-1", workspace_root=tmp_path, strict=True) + execute_import(request, dry_run=False) + return tmp_path / "runs" / "run-1" + + +def test_read_manifest_result_set_auto_selects_the_single_eligible_realization(tmp_path): + run_dir = _imported_run(tmp_path) + result_set = read_manifest_result_set(run_dir) + assert result_set.selection.policy == "single_evaluable_per_condition" + assert len(result_set.selection.realization_ids) == 1 + assert set(result_set.rows["question_id"].astype(str)) == {"q1", "q2"} + + +def test_read_manifest_result_set_rejects_unknown_realization_id(tmp_path): + run_dir = _imported_run(tmp_path) + with pytest.raises(ResultSetError, match="Unknown realization"): + read_manifest_result_set(run_dir, realization_ids=("real_" + "0" * 16,)) + + +def test_read_manifest_result_set_explicit_selection(tmp_path): + run_dir = _imported_run(tmp_path) + auto = read_manifest_result_set(run_dir) + (realization_id,) = auto.selection.realization_ids + explicit = read_manifest_result_set(run_dir, realization_ids=(realization_id,)) + assert explicit.selection.policy == "explicit" + assert explicit.selection.realization_ids == (realization_id,) + + +def test_build_evaluation_report_accounts_for_every_realization(tmp_path): + run_dir = _imported_run(tmp_path) + result_set = read_manifest_result_set(run_dir) + report = build_evaluation_report_for_result_set("run-1", result_set, reparse=False) + assert report["schema_version"] == "choicebench.evaluation.v2" + (condition_id,) = report["conditions"] + (realization_id,) = report["conditions"][condition_id]["realization_ids"] + entry = report["realizations"][realization_id] + assert entry["selected"] is True + assert entry["evidence_status"] == "complete" + assert entry["scope_disposition"] == "included" + assert entry["metrics"]["n"] == 2 + assert entry["metrics"]["accuracy"] == 1.0 # both rows' answer matches gold + + +def test_build_evaluation_report_scores_incorrect_predictions(tmp_path): + run_dir = _imported_run( + tmp_path, + results_csv="qid,question,choice_a,choice_b,answer,gold,score\nq1,One?,x,y,b,a,0.9\nq2,Two?,m,n,b,b,0.7\n", + ) + result_set = read_manifest_result_set(run_dir) + report = build_evaluation_report_for_result_set("run-1", result_set, reparse=False) + (realization_id,) = result_set.selection.realization_ids + assert report["realizations"][realization_id]["metrics"]["accuracy"] == 0.5 diff --git a/tests/importing/test_evidence.py b/tests/importing/test_evidence.py new file mode 100644 index 0000000..2ec08b1 --- /dev/null +++ b/tests/importing/test_evidence.py @@ -0,0 +1,168 @@ +"""Tests for verified source opening and content-addressed evidence storage +(Task 11, reduced scope).""" + +from __future__ import annotations + +from dataclasses import replace +from hashlib import sha256 + +import pytest + +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.evidence import ( + EvidenceError, + open_verified_source, + validate_evidence_blob, + validate_evidence_index, + write_evidence_blob, + write_evidence_index, +) +from choicebench.importing.schema import CsvDialectSpec, OptionMappingSpec, SourceArtifactSpec + + +def _declaration(tmp_path, data: bytes, source_id="results"): + path = tmp_path / "results.csv" + path.write_bytes(data) + return SourceArtifactSpec( + source_id=source_id, + path=path, + logical_path="inputs/results.csv", + expected_sha256=sha256(data).hexdigest(), + format="csv", + format_version="1", + classification="raw", + dialect=CsvDialectSpec(), + columns={"question_id": "qid"}, + expected_columns=("qid",), + ignored_columns={}, + null_values=(), + numeric_columns=(), + option_mapping=OptionMappingSpec( + mode="ordered_columns", ordered_columns=("a", "b"), + structured_column=None, structured_label_key=None, structured_text_key=None, + ), + extra_field_policy="reject_unmapped", + preserve_namespace="ext", + source_run_id=None, + source_repository=None, + source_commit=None, + notes={}, + ) + + +def test_open_verified_source_accepts_matching_checksum(tmp_path): + declaration = _declaration(tmp_path, b"qid,a,b\nq1,x,y\n") + opened = open_verified_source(declaration) + assert opened.source_id == "results" + assert opened.sha256 == declaration.expected_sha256 + assert opened.data == b"qid,a,b\nq1,x,y\n" + + +def test_open_verified_source_rejects_checksum_mismatch(tmp_path): + declaration = _declaration(tmp_path, b"qid,a,b\nq1,x,y\n") + declaration.path.write_bytes(b"qid,a,b\nq1,TAMPERED,y\n") + with pytest.raises(EvidenceError, match="checksum"): + open_verified_source(declaration) + + +def test_open_verified_source_rejects_missing_file(tmp_path): + declaration = _declaration(tmp_path, b"qid,a,b\nq1,x,y\n") + declaration.path.unlink() + with pytest.raises(EvidenceError, match="regular file"): + open_verified_source(declaration) + + +def test_open_verified_source_rejects_symlink(tmp_path): + declaration = _declaration(tmp_path, b"qid,a,b\nq1,x,y\n") + real_path = declaration.path + link_path = tmp_path / "link.csv" + link_path.symlink_to(real_path) + linked_declaration = replace(declaration, path=link_path) + with pytest.raises(EvidenceError, match="symlink"): + open_verified_source(linked_declaration) + + +def test_open_verified_source_rejects_path_outside_containment_root(tmp_path): + outside_dir = tmp_path / "outside" + outside_dir.mkdir() + declaration = _declaration(outside_dir, b"qid,a,b\nq1,x,y\n") + containment_root = tmp_path / "inside" + containment_root.mkdir() + with pytest.raises(EvidenceError, match="containment"): + open_verified_source(declaration, containment_root=containment_root) + + +def test_write_and_validate_evidence_blob_round_trip(tmp_path): + data = b"qid,a,b\nq1,x,y\n" + source = OpenedSource( + source_id="results", audit_path=tmp_path / "results.csv", + logical_path="inputs/results.csv", data=data, sha256=sha256(data).hexdigest(), + ) + record = write_evidence_blob(tmp_path, source, references=("q1",)) + validate_evidence_blob(tmp_path, record) # must not raise + assert record["references"] == ["q1"] + + +def test_write_evidence_blob_is_idempotent_for_identical_bytes(tmp_path): + data = b"qid,a,b\nq1,x,y\n" + source = OpenedSource( + source_id="results", audit_path=tmp_path / "results.csv", + logical_path="inputs/results.csv", data=data, sha256=sha256(data).hexdigest(), + ) + first = write_evidence_blob(tmp_path, source, references=("q1",)) + second = write_evidence_blob(tmp_path, source, references=("q1",)) + assert first == second + + +def test_validate_evidence_blob_rejects_tampered_bytes(tmp_path): + data = b"qid,a,b\nq1,x,y\n" + source = OpenedSource( + source_id="results", audit_path=tmp_path / "results.csv", + logical_path="inputs/results.csv", data=data, sha256=sha256(data).hexdigest(), + ) + record = write_evidence_blob(tmp_path, source, references=("q1",)) + from choicebench.importing.evidence import evidence_blob_path + + blob_path = evidence_blob_path(tmp_path, source.sha256, "csv") + blob_path.write_bytes(data + b"tampered") + with pytest.raises(EvidenceError, match="content integrity"): + validate_evidence_blob(tmp_path, record) + + +def test_evidence_index_round_trips(tmp_path): + data = b"qid,a,b\nq1,x,y\n" + source = OpenedSource( + source_id="results", audit_path=tmp_path / "results.csv", + logical_path="inputs/results.csv", data=data, sha256=sha256(data).hexdigest(), + ) + record = write_evidence_blob(tmp_path, source, references=("q1",)) + index = write_evidence_index(tmp_path, [record]) + validated = validate_evidence_index(tmp_path, index["evidence_index_digest"]) + assert validated["records"] == [record] + + +def test_evidence_index_rejects_digest_mismatch(tmp_path): + data = b"qid,a,b\nq1,x,y\n" + source = OpenedSource( + source_id="results", audit_path=tmp_path / "results.csv", + logical_path="inputs/results.csv", data=data, sha256=sha256(data).hexdigest(), + ) + record = write_evidence_blob(tmp_path, source, references=("q1",)) + write_evidence_index(tmp_path, [record]) + with pytest.raises(EvidenceError, match="digest"): + validate_evidence_index(tmp_path, "0" * 64) + + +def test_evidence_index_rejects_missing_blob(tmp_path): + data = b"qid,a,b\nq1,x,y\n" + source = OpenedSource( + source_id="results", audit_path=tmp_path / "results.csv", + logical_path="inputs/results.csv", data=data, sha256=sha256(data).hexdigest(), + ) + record = write_evidence_blob(tmp_path, source, references=("q1",)) + index = write_evidence_index(tmp_path, [record]) + from choicebench.importing.evidence import evidence_blob_path + + evidence_blob_path(tmp_path, source.sha256, "csv").unlink() + with pytest.raises(EvidenceError, match="Missing evidence blob"): + validate_evidence_index(tmp_path, index["evidence_index_digest"]) diff --git a/tests/importing/test_identity.py b/tests/importing/test_identity.py new file mode 100644 index 0000000..e2e9288 --- /dev/null +++ b/tests/importing/test_identity.py @@ -0,0 +1,3065 @@ +from __future__ import annotations + +from dataclasses import asdict, replace +from hashlib import sha256 +import importlib +from pathlib import Path + +import pandas as pd +import pytest + +from choicebench.config.schema import ( + BenchmarkConfig, + ExperimentConfig, + MethodConfig, + ModelConfig, + RunConfig, +) +from choicebench.datasets import dataset_content_digest, dataset_sample_identities +from choicebench.identity import integrity_digest, short_id +from choicebench.importing.csv_adapter import OpenedSource +from choicebench.importing.dataset_reference import build_expected_dataset +from choicebench.importing.identity import ( + ImportIdentityError, + build_import_semantic_identity, + importer_implementation_identity, + make_lineage_component, + make_realization as _make_realization, + make_result_origin, + make_semantic_condition, +) +from choicebench.importing.schema import ( + DatasetReferenceSpec, + CsvDialectSpec, + ImportConditionSpec, + ImportMethodSpec, + ImportModelSpec, + ImportPromptSpec, + NumericColumnSpec, + OptionMappingSpec, + ResultOriginSpec, +) +from tests.importing import _cross_module_helpers + + +def _adapter(value: str) -> str: + return value + + +def _validator(value: str) -> bool: + return bool(value) + + +_GLOBAL_VALIDATOR_RESULT = True + + +def _global_validator(value: str) -> bool: + return bool(value) and _GLOBAL_VALIDATOR_RESULT + + +_cross_module_helper = _cross_module_helpers.helper_v1 + + +def _cross_module_adapter(value: str) -> str: + return _cross_module_helper(value) + + +def _transform(value: str) -> str: + return value.strip() + + +def _transform_required(value: str, *, threshold: float) -> str: + return value if threshold >= 0 else "" + + +def _redefined_adapter(value: str, prefix: str = "first") -> str: + return prefix + value + + +_first_redefined_adapter = _redefined_adapter + + +def _redefined_adapter(value: str, prefix: str = "second") -> str: + return prefix + value + + +_second_redefined_adapter = _redefined_adapter + + +class _StatefulCallable: + def __init__(self, value: str) -> None: + self.value = value + + def __call__(self, text: str) -> str: + return self.value + text + + def transform(self, text: str) -> str: + return self.value + text + + +def make_realization(**kwargs): + identity = kwargs["identity"] + runtime_lineages = { + component["lineage_id"]: _transform + for component in identity.get("lineage_components", ()) + if component.get("identity", {}) + .get("implementation", {}) + .get("identity_mode") + == "runtime_callable" + } + return _make_realization( + adapter=_adapter, + validator=_validator, + expected_dataset=kwargs.pop("expected_dataset", _dataset()), + semantic_condition=kwargs.pop("semantic_condition", _records().condition), + lineage_runtime_callables=runtime_lineages, + **kwargs, + ) + + +def _dataset(): + frame = pd.DataFrame( + [ + {"qid": "q1", "stem": "One?", "gold": "A", "a": "x", "b": "y"}, + {"qid": "q2", "stem": "Two?", "gold": "B", "a": "m", "b": "n"}, + {"qid": "q3", "stem": "Three?", "gold": "A", "a": "i", "b": "j"}, + ] + ) + data = frame.to_csv(index=False, lineterminator="\n").encode() + source = OpenedSource( + source_id="benchmark", + audit_path=Path("/machine-a/input.csv"), + logical_path="inputs/benchmark.csv", + data=data, + sha256=sha256(data).hexdigest(), + ) + declaration = DatasetReferenceSpec( + dataset_id="dataset", + benchmark_name="benchmark", + split="test", + reference_kind="independent_input_snapshot", + trust_label="declared-independent-input", + source_ids=("benchmark",), + selection_source_id="benchmark", + expected_question_ids=("q2", "q1"), + selection_seed=None, + selection_n_samples=2, + subject_filter=(), + selection_unknown_reasons={"selection_seed": "not recorded"}, + columns={ + "question_id": "qid", + "question_text": "stem", + "correct_option": "gold", + "choice_a": "a", + "choice_b": "b", + }, + revision=None, + fingerprint=None, + derivation={"source_role": "benchmark_input"}, + limitations=("publisher revision unknown",), + native_compatibility_identity=None, + ) + return build_expected_dataset(declaration, {"benchmark": source}) + + +def _model() -> ImportModelSpec: + return ImportModelSpec( + model_key="model", + display_name="historical-model", + backend=None, + provider=None, + revision=None, + effective_parameters={"temperature": 0}, + unknown_reasons={ + "backend": "not recorded", + "provider": "not recorded", + "revision": "not recorded", + }, + native_compatibility_identity=None, + ) + + +def _method() -> ImportMethodSpec: + return ImportMethodSpec( + method_key="method", + name="semantic_matching_v1", + effective_parameters={"matching": "semantic"}, + implementation=None, + unknown_reasons={"implementation": "historical code unavailable"}, + native_compatibility_identity=None, + ) + + +def _prompt() -> ImportPromptSpec: + return ImportPromptSpec( + prompt_key="prompt", + template_identity=None, + template_digest=None, + template_contents=None, + unknown_reason="raw prompt was not preserved", + native_compatibility_identity=None, + ) + + +def _condition() -> ImportConditionSpec: + return ImportConditionSpec( + condition_key="condition", + source_ids=("results",), + dataset_id="dataset", + model_key="model", + method_key="method", + prompt_key="prompt", + seed=None, + calibration_identity=None, + preflight_identity=None, + protocol_settings={}, + generation_parameters={"max_tokens": 64}, + unknown_reasons={ + "seed": "not recorded", + "calibration_identity": "not applicable", + "preflight_identity": "not recorded", + }, + expected_question_ids=("q2", "q1"), + evidence_status="complete", + scope_disposition="included", + executable=None, + qualifications=(), + limitations=(), + damaged_question_ids=(), + recoverable_question_ids=(), + result_origin=ResultOriginSpec( + derivation_origin="external_import", + default_prediction_origin="external_historical_inference", + per_question_prediction_origins={}, + ), + ) + + +def _records(**condition_changes): + if condition_changes.get("seed") is not None and "unknown_reasons" not in condition_changes: + condition_changes["unknown_reasons"] = { + key: value + for key, value in _condition().unknown_reasons.items() + if key != "seed" + } + condition = replace(_condition(), **condition_changes) + return build_import_semantic_identity( + condition=condition, + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +def test_fallback_child_payloads_are_exact_and_unknowns_are_explicit(): + records = _records() + + assert records.dataset_artifact["identity"] == { + "benchmark": "benchmark", + "content_digest": records.dataset_artifact["identity"]["content_digest"], + "schema_version": "choicebench.semantic-dataset.v1", + "split": "test", + } + assert records.model["identity"] == { + "backend": {"reason": "not recorded", "value": None}, + "effective_parameters": {"max_tokens": 64, "temperature": 0}, + "model": "historical-model", + "provider": {"reason": "not recorded", "value": None}, + "revision": {"reason": "not recorded", "value": None}, + "schema_version": "choicebench.semantic-model.v1", + } + assert records.method["identity"] == { + "effective_params": {"matching": "semantic"}, + "implementation": { + "reason": "historical code unavailable", + "value": None, + }, + "name": "semantic_matching_v1", + "preflight": {"reason": "not recorded", "value": None}, + "schema_version": "choicebench.semantic-method.v1", + } + unknown_prompt = {"reason": "raw prompt was not preserved", "value": None} + assert records.prompt["identity"] == { + "schema_version": "choicebench.semantic-prompt.v1", + "template_contents": unknown_prompt, + "template_digest": unknown_prompt, + "template_identity": unknown_prompt, + } + assert records.selection["identity"] == { + "artifact_id": records.dataset_artifact["artifact_id"], + "content_digest": records.selection["identity"]["content_digest"], + "sample_identities": records.selection["identity"]["sample_identities"], + "seed": None, + "n_samples": 2, + "subject_filter": [], + } + + +def test_child_records_have_full_digests_and_existing_short_prefixes(): + records = _records() + for record, prefix, id_key, digest_key in ( + (records.dataset_artifact, "ds", "artifact_id", "artifact_digest"), + (records.selection, "sel", "selection_id", "selection_digest"), + (records.model, "model", "model_id", "model_digest"), + (records.method, "method", "method_id", "method_digest"), + (records.prompt, "prompt", "prompt_id", "prompt_digest"), + ): + assert record[digest_key] == integrity_digest(record["identity"]) + assert record[id_key] == short_id(prefix, record["identity"]) + + +def test_condition_payload_matches_current_choicebench_keys_and_frozen_prompt_path(): + records = _records() + identity = records.condition["identity"] + + assert set(identity) == { + "benchmark", + "preflight", + "model_id", + "method_id", + "prompt_id", + "prompt_snapshot_path", + "seed", + } + assert identity["benchmark"] == { + "name": "benchmark", + "split": "test", + "artifact_id": records.dataset_artifact["artifact_id"], + "selection_id": records.selection["selection_id"], + } + assert identity["prompt_snapshot_path"] == ( + f"artifacts/prompts/{records.prompt['prompt_id']}" + ) + assert records.condition["condition_digest"] == integrity_digest(identity) + assert records.condition["condition_id"] == short_id("cond", identity) + + +@pytest.mark.parametrize( + ("field", "value"), + [ + ("seed", 7), + ("protocol_settings", {"pride_modal_k_threshold": 3}), + ("generation_parameters", {"max_tokens": 65}), + ], +) +def test_scientific_condition_changes_change_condition_identity(field, value): + assert _records(**{field: value}).condition["condition_id"] != _records().condition[ + "condition_id" + ] + + +@pytest.mark.parametrize( + "field", + ["selection_id", "artifact_id", "model_id", "method_id", "prompt_id"], +) +def test_semantic_child_change_changes_condition_identity(field): + original = _records().condition + identity = dict(original["identity"]) + if field in {"selection_id", "artifact_id"}: + identity["benchmark"] = dict(identity["benchmark"]) + prefix = "sel" if field == "selection_id" else "ds" + identity["benchmark"][field] = f"{prefix}_{'f' * 16}" + else: + prefix = field.removesuffix("_id") + identity[field] = f"{prefix}_{'f' * 16}" + if field == "prompt_id": + identity["prompt_snapshot_path"] = ( + f"artifacts/prompts/{identity['prompt_id']}" + ) + changed = make_semantic_condition(identity=identity, fields={}) + assert changed["condition_id"] != original["condition_id"] + + +def test_make_semantic_condition_rejects_arbitrary_prompt_path_injection(): + identity = dict(_records().condition["identity"]) + identity["prompt_snapshot_path"] = "/tmp/attacker/prompt" + + with pytest.raises(ImportIdentityError, match="prompt.*path"): + make_semantic_condition(identity=identity, fields={"condition_key": "x"}) + + +def test_semantic_condition_fields_are_closed_and_cannot_contradict_identity(): + identity = _records().condition["identity"] + for fields in ( + {"model_id": "model_conflict"}, + {"source_sha256": "1" * 64}, + {"realization_id": "real_child"}, + ): + with pytest.raises(ImportIdentityError, match="condition fields"): + make_semantic_condition(identity=identity, fields=fields) + + +@pytest.mark.parametrize( + ("field", "replacement"), + [ + ("dataset_id", "wrong-dataset"), + ("model_key", "wrong-model"), + ("method_key", "wrong-method"), + ("prompt_key", "wrong-prompt"), + ], +) +def test_semantic_builder_refuses_cross_wired_child_registry_keys(field, replacement): + with pytest.raises(ImportIdentityError, match=field): + _records(**{field: replacement}) + + +def test_unregistered_method_cannot_claim_native_compatibility(): + model_payload = {"backend": "dummy", "model_name_or_path": "native"} + method_payload = { + "name": "direct", + "effective_params": {}, + "preflight": None, + "implementation": {"qualified_name": "choicebench.methods:Direct"}, + } + native_model = { + "payload": model_payload, + "digest": integrity_digest(model_payload), + "model_id": short_id("model", model_payload), + } + native_method = { + "payload": method_payload, + "digest": integrity_digest(method_payload), + "method_id": short_id("method", method_payload), + } + prompt_payload = { + "version": "v1", + "files": { + name: {"sha256": integrity_digest(name), "content": name} + for name in ("direct_mcq", "free_text", "option_matching") + }, + } + native_prompt = { + "prompt_id": short_id("prompt", prompt_payload), + **prompt_payload, + } + with pytest.raises(ImportIdentityError, match="registered current runner"): + build_import_semantic_identity( + condition=replace( + _condition(), + generation_parameters={}, + unknown_reasons={ + **_condition().unknown_reasons, + "preflight_identity": "not applicable", + }, + ), + dataset=_dataset(), + model=replace( + _model(), + display_name="native", + backend="dummy", + effective_parameters={}, + unknown_reasons={ + "provider": "not applicable", + "revision": "not applicable", + }, + native_compatibility_identity=native_model, + ), + method=replace( + _method(), + name="direct", + effective_parameters={}, + implementation=method_payload["implementation"], + unknown_reasons={}, + native_compatibility_identity=native_method, + ), + prompt=replace( + _prompt(), + template_identity="v1", + template_digest=integrity_digest(prompt_payload["files"]), + template_contents={ + name: item["content"] + for name, item in prompt_payload["files"].items() + }, + unknown_reason=None, + native_compatibility_identity=native_prompt, + ), + ) + + +def test_native_model_identity_refuses_unbound_generation_parameters(): + payload = {"backend": "dummy", "model_name_or_path": "historical-model"} + native = { + "payload": payload, + "digest": integrity_digest(payload), + "model_id": short_id("model", payload), + } + with pytest.raises(ImportIdentityError, match="generation"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=replace( + _model(), + backend="dummy", + unknown_reasons={ + "provider": "not applicable", + "revision": "not applicable", + }, + native_compatibility_identity=native, + ), + method=_method(), + prompt=_prompt(), + ) + + +def test_native_child_claims_must_match_their_semantic_declarations(): + model_payload = {"backend": "dummy", "model_name_or_path": "historical-model"} + native_model = { + "payload": model_payload, + "digest": integrity_digest(model_payload), + "model_id": short_id("model", model_payload), + } + with pytest.raises(ImportIdentityError, match="backend"): + build_import_semantic_identity( + condition=replace(_condition(), generation_parameters={}), + dataset=_dataset(), + model=replace( + _model(), + backend="api", + provider="provider", + unknown_reasons={"revision": "not applicable"}, + native_compatibility_identity=native_model, + ), + method=_method(), + prompt=_prompt(), + ) + + method_payload = { + "name": "direct_mcq", + "effective_params": {}, + "preflight": None, + "implementation": {"qualified_name": "choicebench.methods:Direct"}, + } + native_method = { + "payload": method_payload, + "digest": integrity_digest(method_payload), + "method_id": short_id("method", method_payload), + } + with pytest.raises( + ImportIdentityError, match="method.*(?:declaration|registered|runner)" + ): + build_import_semantic_identity( + condition=replace( + _condition(), + preflight_identity=None, + unknown_reasons={ + **_condition().unknown_reasons, + "preflight_identity": "not applicable", + }, + ), + dataset=_dataset(), + model=_model(), + method=replace( + _method(), + implementation={"qualified_name": "historical:SemanticMatching"}, + unknown_reasons={}, + native_compatibility_identity=native_method, + ), + prompt=_prompt(), + ) + + +def test_native_prompt_contents_must_match_the_semantic_declaration(): + payload = { + "version": "v1", + "files": { + name: {"sha256": integrity_digest(name), "content": name} + for name in ("direct_mcq", "free_text", "option_matching") + }, + } + native = {"prompt_id": short_id("prompt", payload), **payload} + with pytest.raises(ImportIdentityError, match="prompt.*contents"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=replace( + _prompt(), + template_identity="v1", + template_digest=integrity_digest(payload["files"]), + template_contents={"direct_mcq": "different"}, + unknown_reason=None, + native_compatibility_identity=native, + ), + ) + + bad_payload = { + **payload, + "files": { + **payload["files"], + "direct_mcq": {"sha256": "0" * 64, "content": "direct_mcq"}, + }, + } + bad_native = {"prompt_id": short_id("prompt", bad_payload), **bad_payload} + with pytest.raises(ImportIdentityError, match="prompt.*hash|prompt.*digest"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=replace( + _prompt(), + template_identity="v1", + template_digest=integrity_digest(bad_payload["files"]), + template_contents={ + name: item["content"] for name, item in bad_payload["files"].items() + }, + unknown_reason=None, + native_compatibility_identity=bad_native, + ), + ) + + +def test_native_api_payload_is_validated_by_the_authoritative_schema(): + payload = { + "backend": "api", + "provider": 7, + "model_name_or_path": "historical-model", + "base_url": None, + "generation_kwargs": ["not", "a", "mapping"], + } + native = { + "payload": payload, + "digest": integrity_digest(payload), + "model_id": short_id("model", payload), + } + with pytest.raises(ImportIdentityError, match="native model|provider|generation"): + build_import_semantic_identity( + condition=replace(_condition(), generation_parameters={}), + dataset=_dataset(), + model=replace( + _model(), + backend="api", + provider="provider", + unknown_reasons={"revision": "not applicable"}, + native_compatibility_identity=native, + ), + method=_method(), + prompt=_prompt(), + ) + + +def test_native_huggingface_payload_binds_name_revision_and_parameters(): + payload = { + "backend": "huggingface", + "model": { + "kind": "huggingface-hub", + "repo_id": "right/model", + "requested_revision": "right-revision", + "resolved_commit": "a" * 40, + }, + "device": "cpu", + "add_bos_token": False, + "generation_kwargs": { + "max_new_tokens": 64, + "temperature": 0, + "do_sample": False, + }, + "loader": {"trust_remote_code": True, "torch_dtype": "float32"}, + } + native = { + "payload": payload, + "digest": integrity_digest(payload), + "model_id": short_id("model", payload), + } + with pytest.raises(ImportIdentityError, match="name|revision|parameter|declaration"): + build_import_semantic_identity( + condition=replace(_condition(), generation_parameters={}), + dataset=_dataset(), + model=replace( + _model(), + display_name="wrong/model", + backend="huggingface", + revision="wrong-revision", + effective_parameters={"temperature": 1}, + unknown_reasons={"provider": "not applicable"}, + native_compatibility_identity=native, + ), + method=_method(), + prompt=_prompt(), + ) + + +def test_native_method_payload_requires_authoritative_implementation_shape(): + payload = { + "name": "semantic_matching_v1", + "effective_params": {"matching": "semantic"}, + "preflight": None, + "implementation": {"attacker": "not-code-identity"}, + } + native = { + "payload": payload, + "digest": integrity_digest(payload), + "method_id": short_id("method", payload), + } + with pytest.raises(ImportIdentityError, match="native method|implementation"): + build_import_semantic_identity( + condition=replace( + _condition(), + unknown_reasons={ + **_condition().unknown_reasons, + "preflight_identity": "not applicable", + }, + ), + dataset=_dataset(), + model=_model(), + method=replace( + _method(), + implementation={"qualified_name": "historical:Matcher"}, + unknown_reasons={}, + native_compatibility_identity=native, + ), + prompt=_prompt(), + ) + + +def test_unknown_declarations_cannot_be_upgraded_by_native_claims(): + method_payload = { + "name": "semantic_matching_v1", + "effective_params": {"matching": "semantic"}, + "preflight": None, + "implementation": {"qualified_name": "fabricated:Current"}, + } + native_method = { + "payload": method_payload, + "digest": integrity_digest(method_payload), + "method_id": short_id("method", method_payload), + } + with pytest.raises( + ImportIdentityError, + match="unknown.*native|implementation|registered current runner", + ): + build_import_semantic_identity( + condition=replace( + _condition(), + unknown_reasons={ + **_condition().unknown_reasons, + "preflight_identity": "not applicable", + }, + ), + dataset=_dataset(), + model=_model(), + method=replace(_method(), native_compatibility_identity=native_method), + prompt=_prompt(), + ) + + prompt_payload = { + "version": "v1", + "files": { + name: {"sha256": integrity_digest(name), "content": name} + for name in ("direct_mcq", "free_text", "option_matching") + }, + } + native_prompt = {"prompt_id": short_id("prompt", prompt_payload), **prompt_payload} + with pytest.raises(ImportIdentityError, match="unknown.*native|prompt"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=replace(_prompt(), native_compatibility_identity=native_prompt), + ) + + +def test_fully_validated_native_children_reproduce_build_execution_plan_condition(): + import choicebench.cli.run_experiment as run_experiment + + run_experiment = importlib.reload(run_experiment) + config = ExperimentConfig( + "native-compat", + [ModelConfig("dummy", "same")], + [BenchmarkConfig("toy", n_samples=2)], + [MethodConfig("direct_mcq")], + ["accuracy"], + RunConfig(seed=42), + ) + plan = run_experiment.build_execution_plan(config) + native_condition = next(iter(plan.conditions.values())) + native_selection = plan.selections[0] + artifact = native_selection.artifact + artifact_payload = { + "spec": artifact.metadata["spec"], + "content_digest": artifact.content_digest, + "source": artifact.metadata["source"], + } + selection_payload = { + "artifact_id": artifact.artifact_id, + "content_digest": dataset_content_digest(native_selection.questions), + "sample_identities": list(native_selection.sample_identities), + "seed": 42, + "n_samples": 2, + "subject_filter": [], + } + native_dataset = { + "artifact_payload": artifact_payload, + "artifact_digest": integrity_digest(artifact_payload), + "artifact_id": artifact.artifact_id, + "selection_payload": selection_payload, + "selection_digest": integrity_digest(selection_payload), + "selection_id": native_selection.selection_id, + } + source_bytes = artifact.dataframe.to_csv(index=False, lineterminator="\n").encode() + source = OpenedSource( + source_id="native-input", + audit_path=artifact.path, + logical_path="prepared/toy.csv", + data=source_bytes, + sha256=sha256(source_bytes).hexdigest(), + ) + dataset = build_expected_dataset( + DatasetReferenceSpec( + dataset_id="toy", + benchmark_name="toy", + split="test", + reference_kind="independent_input_snapshot", + trust_label="choicebench-verified", + source_ids=("native-input",), + selection_source_id="native-input", + expected_question_ids=tuple( + native_selection.questions["question_id"].astype(str) + ), + selection_seed=42, + selection_n_samples=2, + subject_filter=(), + selection_unknown_reasons={}, + columns={column: column for column in artifact.dataframe.columns}, + revision=None, + fingerprint=None, + derivation={"kind": "native compatibility fixture"}, + limitations=(), + native_compatibility_identity=native_dataset, + ), + {"native-input": source}, + ) + native_model_record = plan.manifest["payload"]["models"][0] + native_model_payload = { + "backend": "dummy", + "model_name_or_path": "same", + } + native_method_record = plan.manifest["payload"]["methods"][0] + native_method_payload = { + "name": native_method_record["config"]["name"], + "effective_params": native_method_record["config"]["params"], + "preflight": native_method_record["config"]["preflight"], + "implementation": native_method_record["implementation"], + } + native_prompt_record = plan.manifest["payload"]["prompts"] + condition = replace( + _condition(), + dataset_id="toy", + seed=42, + generation_parameters={}, + unknown_reasons={ + "calibration_identity": "not applicable", + "preflight_identity": "not applicable", + }, + expected_question_ids=tuple( + native_selection.questions["question_id"].astype(str) + ), + ) + records = build_import_semantic_identity( + condition=condition, + dataset=dataset, + model=replace( + _model(), + display_name="same", + backend="dummy", + provider=None, + revision=None, + unknown_reasons={"provider": "not applicable", "revision": "not applicable"}, + effective_parameters={}, + native_compatibility_identity={ + "payload": native_model_payload, + "digest": integrity_digest(native_model_payload), + "model_id": native_model_record["model_id"], + }, + ), + method=replace( + _method(), + name="direct_mcq", + effective_parameters=native_method_record["config"]["params"], + implementation=native_method_record["implementation"], + unknown_reasons={}, + native_compatibility_identity={ + "payload": native_method_payload, + "digest": integrity_digest(native_method_payload), + "method_id": native_method_record["method_id"], + }, + ), + prompt=replace( + _prompt(), + template_identity=native_prompt_record["version"], + template_digest=integrity_digest(native_prompt_record["files"]), + template_contents={ + name: item["content"] + for name, item in native_prompt_record["files"].items() + }, + unknown_reason=None, + native_compatibility_identity={ + "prompt_id": native_prompt_record["prompt_id"], + "version": native_prompt_record["version"], + "files": native_prompt_record["files"], + }, + ), + ) + + assert records.condition["identity"] == native_condition["identity"] + assert records.condition["condition_id"] == native_condition["condition_id"] + + +def _attach_result_origin( + identity: dict, + *, + derivation_origin: str, + assignments: tuple[tuple[str, str], ...], + authorization_digest: str | None = None, + source_digest: str = "2" * 64, + overlay_source_digest: str = "e" * 64, +) -> None: + derived = derivation_origin in {"repair_overlay", "offline_transformation"} + components = [ + make_lineage_component( + operation_type=derivation_origin, + question_id=question_id, + parent_digests=(("d" * 64,) if derived else ()), + source_digests=( + tuple(sorted({source_digest, overlay_source_digest})) + if derived + else (source_digest,) + ), + authorization_digest=(authorization_digest if derived else None), + implementation=_transform, + parameters={}, + input_digest=integrity_digest( + {"question_id": question_id, "stage": "input"} + ), + preownership_output_digest=integrity_digest( + { + "question_id": question_id, + "prediction_origin": prediction_origin, + "stage": "preownership-output", + } + ), + prediction_origin=prediction_origin, + ) + for question_id, prediction_origin in assignments + ] + identity["lineage_components"] = components + identity["result_origin"] = make_result_origin( + derivation_origin=derivation_origin, + row_assignments=tuple( + ( + question_id, + prediction_origin, + component["lineage_id"], + ) + for (question_id, prediction_origin), component in zip( + assignments, components, strict=True + ) + ), + ) + + +def _overlay_identity(identity: dict, derivation_origin: str) -> dict: + derived_components = [ + component + for component in identity["lineage_components"] + if component["identity"]["operation_type"] == derivation_origin + ] + offline = derivation_origin == "offline_transformation" + return { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": ( + integrity_digest( + [ + component["identity"]["input_digest"] + for component in derived_components + ] + ) + if offline + else None + ), + "preownership_output_digest": ( + integrity_digest( + [ + component["identity"]["preownership_output_digest"] + for component in derived_components + ] + ) + if offline + else None + ), + "implementation_digest": ( + integrity_digest(derived_components[0]["identity"]["implementation"]) + if offline + else None + ), + } + + +def _realization_identity() -> dict: + expected_dataset = _dataset() + identity = { + "import_spec_digest": "1" * 64, + "sources": [ + { + "source_id": "results", + "logical_path": "publisher/results.csv", + "classification": "canonical", + "format": "csv", + "format_version": "1", + "sha256": "2" * 64, + "provenance": { + "source_run_id": None, + "source_repository": None, + "source_commit": None, + "notes_digest": None, + "evidence_digest": None, + "unknown_reasons": { + "source_run_id": "not recorded", + "source_repository": "not recorded", + "source_commit": "not recorded", + }, + }, + } + ], + "expected_dataset": { + "snapshot_digest": expected_dataset.snapshot_digest, + "question_set_digest": expected_dataset.question_set_digest, + "derivation_digest": expected_dataset.derivation_digest, + "question_ids": list(expected_dataset.selected_question_ids), + }, + "importer_implementation": importer_implementation_identity( + adapter=_adapter, validator=_validator + ), + "parsing_policy": { + "dialect": asdict(CsvDialectSpec()), + "mapping": {"question_id": "qid"}, + "null_values": [], + "numeric_columns": [ + asdict( + NumericColumnSpec( + source_column="score", + value_type="float", + null_allowed=True, + ) + ) + ], + "option_mapping": asdict( + OptionMappingSpec( + mode="ordered_columns", + ordered_columns=("a", "b"), + structured_column=None, + structured_label_key=None, + structured_text_key=None, + ) + ), + "extra_field_policy": "preserve_unmapped", + }, + "validation": {"findings_digest": "6" * 64}, + "evidence": { + "evidence_status": "complete", + "qualification_digest": "7" * 64, + "limitation_digest": "8" * 64, + "defect_digest": "9" * 64, + "scope_disposition": "included", + }, + "lineage_components": [], + "result_origin": {}, + "parent_digests": { + "realization_digests": [], + "evidence_digests": [], + "result_digests": [], + }, + "authorization_digest": None, + "overlay": None, + } + _attach_result_origin( + identity, + derivation_origin="external_import", + assignments=( + ("q2", "external_historical_inference"), + ("q1", "external_historical_inference"), + ), + ) + return identity + + +def _mutate_realization(identity: dict, field: str, value) -> None: + if field == "source_sha256": + identity["sources"][0]["sha256"] = value + _attach_result_origin( + identity, + derivation_origin="external_import", + assignments=( + ("q2", "external_historical_inference"), + ("q1", "external_historical_inference"), + ), + source_digest=value, + ) + elif field == "mapping": + identity["parsing_policy"]["mapping"] = value + elif field == "dialect": + identity["parsing_policy"]["dialect"].update(value) + elif field == "evidence_status": + identity["evidence"]["evidence_status"] = value + elif field == "scope": + identity["evidence"]["scope_disposition"] = value + elif field == "authorization_digest": + identity["authorization_digest"] = value + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": None, + "preownership_output_digest": None, + "implementation_digest": None, + } + _attach_result_origin( + identity, + derivation_origin="repair_overlay", + assignments=( + ("q2", "external_historical_inference"), + ("q1", "external_repair_inference"), + ), + authorization_digest=value, + ) + elif field == "overlay_sha256": + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = { + "source_sha256": value, + "replacement_digest": "a" * 64, + "transformation_input_digest": None, + "preownership_output_digest": None, + "implementation_digest": None, + } + _attach_result_origin( + identity, + derivation_origin="repair_overlay", + assignments=( + ("q2", "external_historical_inference"), + ("q1", "external_repair_inference"), + ), + authorization_digest="c" * 64, + overlay_source_digest=value, + ) + elif field == "prediction_origins": + _attach_result_origin( + identity, + derivation_origin="external_import", + assignments=value, + ) + else: + raise AssertionError(field) + + +@pytest.mark.parametrize( + ("field", "value"), + [ + ("source_sha256", "3" * 64), + ("mapping", {"question_id": "id"}), + ("dialect", {"delimiter": ";"}), + ("evidence_status", "qualified"), + ("scope", "superseded"), + ("authorization_digest", "4" * 64), + ("overlay_sha256", "5" * 64), + ( + "prediction_origins", + ( + ("q2", "native_inference"), + ("q1", "external_historical_inference"), + ), + ), + ], +) +def test_realization_changes_leave_condition_fixed(field, value): + condition = _records().condition + original = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=_realization_identity(), + fields={"audit": {"source_path": "/machine-a/results.csv"}}, + ) + changed_identity = _realization_identity() + _mutate_realization(changed_identity, field, value) + changed = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=changed_identity, + fields={"audit": {"source_path": "/machine-b/results.csv"}}, + ) + + assert changed["condition_id"] == original["condition_id"] + assert changed["realization_id"] != original["realization_id"] + + +def test_audit_only_realization_fields_do_not_change_identity(): + condition = _records().condition + first = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=_realization_identity(), + fields={"audit": {"source_path": "/machine-a/a.csv", "timestamp": "now"}}, + ) + moved = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=_realization_identity(), + fields={"audit": {"source_path": "/machine-b/a.csv", "timestamp": "later"}}, + ) + assert first["realization_id"] == moved["realization_id"] + assert first["realization_digest"] == moved["realization_digest"] + + +@pytest.mark.parametrize( + ("field", "value"), + [ + ("bom_policy", "strip_utf8_bom"), + ("delimiter", ";"), + ("quote_character", "'"), + ("escape_character", "\\"), + ("double_quote", False), + ("line_terminators", ["lf"]), + ("mixed_line_terminators", "forbid"), + ("final_record_without_terminator", "forbid"), + ("skip_initial_space", True), + ("strict_syntax", False), + ], +) +def test_every_csv_decoding_and_dialect_field_changes_realization_not_condition( + field, value +): + condition = _records().condition + base_dialect = asdict(CsvDialectSpec()) + first_identity = _realization_identity() + first_identity["parsing_policy"]["dialect"] = base_dialect + changed_identity = _realization_identity() + changed_identity["parsing_policy"]["dialect"] = { + **base_dialect, + field: value, + } + first = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=first_identity, + fields={}, + ) + changed = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=changed_identity, + fields={}, + ) + assert first["condition_id"] == changed["condition_id"] + assert first["realization_id"] != changed["realization_id"] + + +@pytest.mark.parametrize( + ("field", "value"), + [ + ("encoding", "utf-8-sig"), + ("decoding_errors", "replace"), + ("delimiter", 7), + ("blank_record_policy", "allow"), + ("header", "explicit"), + ("strict_syntax", "yes"), + ], +) +def test_realization_refuses_invalid_csv_dialect_values(field, value): + condition = _records().condition + identity = _realization_identity() + identity["parsing_policy"]["dialect"][field] = value + with pytest.raises(ImportIdentityError, match="dialect|parsing"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_refuses_invalid_nested_parsing_policy_records(): + condition = _records().condition + invalid_mutations = ( + lambda policy: policy["mapping"].update({"prediction": 7}), + lambda policy: policy["numeric_columns"].append( + {"source_column": "n", "value_type": "decimal", "null_allowed": False} + ), + lambda policy: policy["option_mapping"].update({"audit": "host-a"}), + ) + for mutate in invalid_mutations: + identity = _realization_identity() + mutate(identity["parsing_policy"]) + with pytest.raises(ImportIdentityError, match="parsing|mapping|numeric|option"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_base_repair_and_transformation_are_distinct_realizations_of_one_condition(): + condition = _records().condition + identities = [] + for derivation in ("external_import", "repair_overlay", "offline_transformation"): + identity = _realization_identity() + _attach_result_origin( + identity, + derivation_origin=derivation, + assignments=( + ("q2", "external_historical_inference"), + ( + "q1", + ( + "external_repair_inference" + if derivation == "repair_overlay" + else "external_historical_inference" + ), + ), + ), + authorization_digest=("c" * 64 if derivation != "external_import" else None), + ) + if derivation != "external_import": + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = _overlay_identity(identity, derivation) + identities.append( + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + ) + assert {item["condition_id"] for item in identities} == { + condition["condition_id"] + } + assert len({item["realization_id"] for item in identities}) == 3 + + +@pytest.mark.parametrize("derivation", ["repair_overlay", "offline_transformation"]) +def test_derived_realization_requires_authorization_parent_and_overlay(derivation): + condition = _records().condition + identity = _realization_identity() + _attach_result_origin( + identity, + derivation_origin=derivation, + assignments=( + ("q2", "external_historical_inference"), + ( + "q1", + ( + "external_repair_inference" + if derivation == "repair_overlay" + else "external_historical_inference" + ), + ), + ), + authorization_digest="c" * 64, + ) + for field in ("authorization_digest", "overlay", "parent", "evidence"): + candidate = _realization_identity() + candidate["result_origin"] = identity["result_origin"] + candidate["lineage_components"] = identity["lineage_components"] + candidate["authorization_digest"] = "c" * 64 + candidate["parent_digests"]["realization_digests"] = ["d" * 64] + candidate["parent_digests"]["evidence_digests"] = ["0" * 64] + candidate["overlay"] = _overlay_identity(candidate, derivation) + if field == "parent": + candidate["parent_digests"]["realization_digests"] = [] + elif field == "evidence": + candidate["parent_digests"]["evidence_digests"] = [] + else: + candidate[field] = None + with pytest.raises( + ImportIdentityError, match="authorization|overlay|parent|source" + ): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=candidate, + fields={}, + ) + + +def test_realization_refuses_post_publication_or_machine_local_identity_fields(): + condition = _records().condition + for forbidden in ( + "result_artifact_id", + "experiment_id", + "output_path", + "source_path", + "timestamp", + ): + identity = _realization_identity() + identity[forbidden] = "forbidden" + with pytest.raises(ImportIdentityError, match=forbidden): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_schema_is_closed_complete_and_origin_consistent(): + condition = _records().condition + cases = [] + missing = _realization_identity() + missing.pop("validation") + cases.append(missing) + extra = _realization_identity() + extra["import_state"] = "validated" + cases.append(extra) + incomplete_dialect = _realization_identity() + incomplete_dialect["parsing_policy"]["dialect"] = {"delimiter": ","} + cases.append(incomplete_dialect) + bare_mixed = _realization_identity() + bare_mixed["result_origin"] = { + "derivation_origin": "repair_overlay", + "prediction_origins": ["mixed"], + } + cases.append(bare_mixed) + for identity in cases: + with pytest.raises(ImportIdentityError): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_condition_short_id_must_match_full_digest(): + condition = _records().condition + with pytest.raises(ImportIdentityError, match="condition_id.*digest"): + make_realization( + condition_id="cond_0000000000000000", + condition_digest=condition["condition_digest"], + identity=_realization_identity(), + fields={}, + ) + + +def test_realization_fields_are_audit_only_and_cannot_repeat_identity_claims(): + condition = _records().condition + for fields in ( + {"evidence_status": "qualified"}, + {"scope_disposition": "held"}, + {"result_origin": {}}, + ): + with pytest.raises(ImportIdentityError, match="Realization fields"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=_realization_identity(), + fields=fields, + ) + + +@pytest.mark.parametrize( + "provenance", + [ + {"machine_id": "host-a"}, + {"cwd": "/home/alice/project"}, + {"host": "machine-a"}, + {"import_date": "2026-07-19"}, + ], +) +def test_realization_source_provenance_is_closed_and_machine_independent(provenance): + condition = _records().condition + identity = _realization_identity() + identity["sources"][0]["provenance"].update(provenance) + with pytest.raises(ImportIdentityError, match="provenance|forbidden"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_requires_runtime_derived_importer_implementation_identity(): + condition = _records().condition + identity = _realization_identity() + identity["importer_implementation"] = { + "schema_version": "choicebench.importer-implementation.v1", + "package": {"name": "choicebench", "version": "0.2.0"}, + "adapter": {"qualified_name": "fake:adapter", "source_digest": "a" * 64}, + "validator": {"qualified_name": "fake:validator", "source_digest": "b" * 64}, + } + with pytest.raises(ImportIdentityError, match="runtime|callable|implementation"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_recomputes_implementation_identity_from_supplied_callables(): + condition = _records().condition + identity = _realization_identity() + with pytest.raises(ImportIdentityError, match="runtime|callable|implementation"): + _make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + adapter=_validator, + validator=_adapter, + expected_dataset=_dataset(), + semantic_condition=condition, + lineage_runtime_callables={ + component["lineage_id"]: _transform + for component in identity["lineage_components"] + }, + ) + + +def test_realization_refuses_stateful_callable_instances_with_shared_source(): + first = _StatefulCallable("first") + with pytest.raises(ImportIdentityError, match="stateless|callable|implementation"): + importer_implementation_identity( + adapter=first, + validator=_validator, + ) + + +def test_semantic_condition_refuses_nested_audit_metadata_and_invalid_children(): + original = _records().condition + nested_audit = { + **original["identity"], + "benchmark": {**original["identity"]["benchmark"], "audit": {"machine_id": "x"}}, + } + invalid_child = {**original["identity"], "model_id": "model_not-a-digest"} + for identity in (nested_audit, invalid_child): + with pytest.raises(ImportIdentityError, match="condition|benchmark|model"): + make_semantic_condition(identity=identity, fields={}) + + +def test_expected_dataset_identity_mode_is_closed(): + with pytest.raises(ImportIdentityError, match="identity mode"): + build_import_semantic_identity( + condition=_condition(), + dataset=replace(_dataset(), identity_mode="invented"), + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +@pytest.mark.parametrize( + "mutation", ["unexpected_key", "false_content_digest", "forged_selected_row"] +) +def test_expected_dataset_payloads_are_bound_to_exact_schema_and_frames(mutation): + dataset = _dataset() + selected_frame = dataset.frame.copy() + artifact_payload = dict(dataset.artifact_payload) + if mutation == "unexpected_key": + artifact_payload["unexpected"] = "smuggled" + elif mutation == "false_content_digest": + artifact_payload["content_digest"] = "f" * 64 + else: + selected_frame.loc[selected_frame.index[0], "question_text"] = "forged" + artifact_digest = integrity_digest(artifact_payload) + artifact_id = short_id("ds", artifact_payload) + selection_payload = { + **dataset.selection_payload, + "artifact_id": artifact_id, + "content_digest": dataset_content_digest(selected_frame), + "sample_identities": dataset_sample_identities(selected_frame), + } + forged = replace( + dataset, + artifact_payload=artifact_payload, + artifact_digest=artifact_digest, + artifact_id=artifact_id, + selection_payload=selection_payload, + selection_digest=integrity_digest(selection_payload), + selection_id=short_id("sel", selection_payload), + frame=selected_frame, + ) + with pytest.raises(ImportIdentityError, match="dataset|artifact|content"): + build_import_semantic_identity( + condition=_condition(), + dataset=forged, + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +def test_expected_dataset_unknown_reasons_must_be_nonempty(): + dataset = replace(_dataset(), selection_unknown_reasons={"selection_seed": ""}) + with pytest.raises(ImportIdentityError, match="unknown reason"): + build_import_semantic_identity( + condition=_condition(), + dataset=dataset, + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +@pytest.mark.parametrize( + ("target", "field", "value"), + [ + ("model", "source_sha256", "a" * 64), + ("model", "cache_path", "/machine-a/cache"), + ("method", "recorded_at", "2026-07-19T12:00:00Z"), + ("condition", "result_digest", "b" * 64), + ], +) +def test_fallback_parameters_refuse_realization_and_audit_metadata( + target, field, value +): + model = _model() + method = _method() + condition = _condition() + if target == "model": + model = replace(model, effective_parameters={field: value}) + elif target == "method": + method = replace(method, effective_parameters={field: value}) + else: + condition = replace(condition, generation_parameters={field: value}) + with pytest.raises(ImportIdentityError, match="parameter|metadata|path|forbidden"): + build_import_semantic_identity( + condition=condition, + dataset=_dataset(), + model=model, + method=method, + prompt=_prompt(), + ) + + +def test_fallback_model_logical_identity_refuses_machine_local_path(): + with pytest.raises(ImportIdentityError, match="model|path|machine"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=replace(_model(), display_name="/machine-a/models/model"), + method=_method(), + prompt=_prompt(), + ) + + +def test_fallback_method_implementation_refuses_machine_local_source_file(): + method = replace( + _method(), + implementation={ + "qualified_name": "historical.matcher:match", + "source_file": "/machine-a/matcher.py", + "source_digest": "a" * 64, + }, + unknown_reasons={}, + ) + with pytest.raises(ImportIdentityError, match="path|implementation|machine"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=method, + prompt=_prompt(), + ) + + +def test_fallback_preflight_refuses_source_provenance_metadata(): + condition = replace( + _condition(), + preflight_identity={"source_sha256": "a" * 64}, + unknown_reasons={ + key: value + for key, value in _condition().unknown_reasons.items() + if key != "preflight_identity" + }, + ) + with pytest.raises(ImportIdentityError, match="preflight|metadata|forbidden"): + build_import_semantic_identity( + condition=condition, + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +def test_fallback_prompt_digest_owns_exact_unredacted_contents(): + def build(content: str): + contents = {"direct_mcq": content} + return build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=ImportPromptSpec( + prompt_key="prompt", + template_identity="historical-v1", + template_digest=integrity_digest(contents), + template_contents=contents, + unknown_reason=None, + native_compatibility_identity=None, + ), + ) + + first = build("literal password=alpha") + second = build("literal password=beta") + assert first.prompt["prompt_id"] != second.prompt["prompt_id"] + assert first.condition["condition_id"] != second.condition["condition_id"] + assert first.prompt["identity"]["template_contents"] == { + "direct_mcq": "literal password=alpha" + } + + +def test_fallback_prompt_refuses_digest_that_does_not_own_contents(): + with pytest.raises(ImportIdentityError, match="prompt.*digest|contents"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=ImportPromptSpec( + prompt_key="prompt", + template_identity="historical-v1", + template_digest="a" * 64, + template_contents={"direct_mcq": "literal prompt"}, + unknown_reason=None, + native_compatibility_identity=None, + ), + ) + + +def test_direct_method_implementation_must_have_the_closed_code_identity_shape(): + method = replace( + _method(), + implementation={"attacker": "not an implementation identity"}, + unknown_reasons={}, + ) + with pytest.raises(ImportIdentityError, match="implementation|field"): + build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=method, + prompt=_prompt(), + ) + + +@pytest.mark.parametrize( + "mutation", + [ + {"reference_kind": "profile_derived_reference_snapshot"}, + {"trust_label": "forged-trust"}, + {"derivation": {"forged": True}}, + {"derivation_digest": "a" * 64}, + {"snapshot_digest": "b" * 64}, + ], +) +def test_expected_dataset_trust_chain_is_owned_and_recomputed(mutation): + with pytest.raises(ImportIdentityError, match="dataset|derivation|snapshot|trust|reference"): + build_import_semantic_identity( + condition=_condition(), + dataset=replace(_dataset(), **mutation), + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +def test_expected_dataset_derivation_revision_refuses_nested_audit_metadata(): + dataset = _dataset() + derivation = { + **dataset.derivation, + "revision": {"audit_path": "/home/alice/source", "timestamp": "now"}, + } + derivation_digest = integrity_digest(derivation) + forged = replace( + dataset, + derivation=derivation, + derivation_digest=derivation_digest, + snapshot_digest=integrity_digest( + { + "artifact_digest": dataset.artifact_digest, + "selection_digest": dataset.selection_digest, + "question_set_digest": dataset.question_set_digest, + "derivation_digest": derivation_digest, + "reference_kind": dataset.reference_kind, + "trust_label": dataset.trust_label, + } + ), + ) + with pytest.raises(ImportIdentityError, match="revision|audit|path|timestamp"): + build_import_semantic_identity( + condition=_condition(), + dataset=forged, + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +def test_profile_derived_dataset_requires_independent_source_groups_at_identity_boundary(): + dataset = _dataset() + derivation = { + **dataset.derivation, + "reference_kind": "profile_derived_reference_snapshot", + "trust_label": "cross-source-internal-consistency", + "declared_derivation": {"independent_source_groups": [["benchmark"]]}, + } + derivation_digest = integrity_digest(derivation) + snapshot_digest = integrity_digest( + { + "artifact_digest": dataset.artifact_digest, + "selection_digest": dataset.selection_digest, + "question_set_digest": dataset.question_set_digest, + "derivation_digest": derivation_digest, + "reference_kind": "profile_derived_reference_snapshot", + "trust_label": "cross-source-internal-consistency", + } + ) + forged = replace( + dataset, + reference_kind="profile_derived_reference_snapshot", + trust_label="cross-source-internal-consistency", + derivation=derivation, + derivation_digest=derivation_digest, + snapshot_digest=snapshot_digest, + ) + + with pytest.raises(ImportIdentityError, match="profile|independent|group|source"): + build_import_semantic_identity( + condition=_condition(), + dataset=forged, + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +def test_complete_realization_question_set_is_owned_by_validated_dataset(): + condition = _records().condition + identity = _realization_identity() + identity["expected_dataset"] = { + **identity["expected_dataset"], + "question_ids": ["q2"], + "question_set_digest": integrity_digest(["q2"]), + } + identity["result_origin"] = make_result_origin( + derivation_origin="external_import", + row_assignments=( + ("q2", "external_historical_inference", f"lin_{'a' * 16}"), + ), + ) + with pytest.raises(ImportIdentityError, match="dataset|question set|question IDs"): + _make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + adapter=_adapter, + validator=_validator, + expected_dataset=_dataset(), + semantic_condition=condition, + lineage_runtime_callables={ + component["lineage_id"]: _transform + for component in identity["lineage_components"] + }, + ) + + +@pytest.mark.parametrize( + "protocol_settings", + [ + {"audit": {"operator": "alice"}}, + {"source_sha256": "a" * 64}, + {"evidence_status": "complete"}, + {"authorization_digest": "b" * 64}, + {"machine": "host-a"}, + {"recorded_at": "2026-07-19"}, + {"working_directory": "work"}, + {"artifact_location": "results/final.csv"}, + {"node_name": "node-a"}, + {"executed_at": "now"}, + {"run_date": "2026-07-19"}, + {"platform": "linux"}, + {"operator": "alice"}, + ], +) +def test_semantic_protocol_settings_refuse_realization_and_audit_metadata( + protocol_settings, +): + identity = dict(_records().condition["identity"]) + identity["protocol_settings"] = protocol_settings + with pytest.raises(ImportIdentityError, match="protocol|forbidden|metadata"): + make_semantic_condition(identity=identity, fields={}) + + +def test_native_dummy_cannot_discard_a_known_revision(): + payload = {"backend": "dummy", "model_name_or_path": "historical-model"} + native = { + "payload": payload, + "digest": integrity_digest(payload), + "model_id": short_id("model", payload), + } + with pytest.raises(ImportIdentityError, match="revision|fallback"): + build_import_semantic_identity( + condition=replace(_condition(), generation_parameters={}), + dataset=_dataset(), + model=replace( + _model(), + backend="dummy", + revision="publisher-r1", + effective_parameters={}, + unknown_reasons={"provider": "not applicable"}, + native_compatibility_identity=native, + ), + method=_method(), + prompt=_prompt(), + ) + + +def test_native_method_claim_must_match_registered_current_runner_identity(): + fake_implementation = { + "qualified_name": "historical.fake:Runner", + "source_file": "fake.py", + "source_digest": "a" * 64, + "callable_digest": "b" * 64, + } + payload = { + "name": "direct_mcq", + "effective_params": {}, + "preflight": None, + "implementation": fake_implementation, + } + method = replace( + _method(), + name="direct_mcq", + effective_parameters={}, + implementation=fake_implementation, + unknown_reasons={}, + native_compatibility_identity={ + "payload": payload, + "digest": integrity_digest(payload), + "method_id": short_id("method", payload), + }, + ) + condition = replace( + _condition(), + preflight_identity=None, + unknown_reasons={ + **_condition().unknown_reasons, + "preflight_identity": "not applicable", + }, + ) + with pytest.raises(ImportIdentityError, match="native method|registered|runner"): + build_import_semantic_identity( + condition=condition, + dataset=_dataset(), + model=_model(), + method=method, + prompt=_prompt(), + ) + + +def test_native_prompt_identity_preserves_credential_shaped_literal_text(): + contents = { + "direct_mcq": "Answer the question; literal example password=alpha", + "free_text": "Respond freely", + "option_matching": "Match the answer", + } + files = { + name: {"sha256": integrity_digest(content), "content": content} + for name, content in contents.items() + } + prompt_payload = {"version": "test-v1", "files": files} + prompt = ImportPromptSpec( + prompt_key="prompt", + template_identity="test-v1", + template_digest=integrity_digest(files), + template_contents=contents, + unknown_reason=None, + native_compatibility_identity={ + "prompt_id": f"prompt_{integrity_digest(prompt_payload)[:16]}", + **prompt_payload, + }, + ) + records = build_import_semantic_identity( + condition=_condition(), + dataset=_dataset(), + model=_model(), + method=_method(), + prompt=prompt, + ) + assert records.prompt["identity"]["files"]["direct_mcq"]["content"] == ( + contents["direct_mcq"] + ) + + +@pytest.mark.parametrize( + "source", + [ + {"audit": {"operator": "alice"}, "node": "machine-a"}, + {"hf_path": "/home/alice/machine-only/dataset"}, + {"generator": "/home/alice/generator.py"}, + { + "hf_dataset_info": { + "builder_name": "/home/alice/private-builder", + "config_name": "cfg", + "version": "1", + } + }, + ], +) +def test_native_dataset_source_refuses_audit_and_machine_metadata(source): + dataset = _dataset() + artifact_payload = { + "spec": { + "benchmark": dataset.benchmark_name, + "split": dataset.split, + "hf_path": None, + "hf_subset": None, + "source_revision": None, + "normalization_version": "v1", + "transforms": [], + "output_name": None, + }, + "content_digest": dataset_content_digest(dataset.artifact_frame), + "source": source, + } + artifact_id = short_id("ds", artifact_payload) + selection_payload = {**dataset.selection_payload, "artifact_id": artifact_id} + derivation = {**dataset.derivation, "identity_mode": "native_compatibility"} + derivation_digest = integrity_digest(derivation) + forged = replace( + dataset, + identity_mode="native_compatibility", + artifact_payload=artifact_payload, + artifact_digest=integrity_digest(artifact_payload), + artifact_id=artifact_id, + selection_payload=selection_payload, + selection_digest=integrity_digest(selection_payload), + selection_id=short_id("sel", selection_payload), + derivation=derivation, + derivation_digest=derivation_digest, + snapshot_digest=integrity_digest( + { + "artifact_digest": integrity_digest(artifact_payload), + "selection_digest": integrity_digest(selection_payload), + "question_set_digest": dataset.question_set_digest, + "derivation_digest": derivation_digest, + "reference_kind": dataset.reference_kind, + "trust_label": dataset.trust_label, + } + ), + ) + with pytest.raises(ImportIdentityError, match="source|audit|path|timestamp"): + build_import_semantic_identity( + condition=_condition(), + dataset=forged, + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +def test_native_dataset_spec_refuses_machine_local_hf_path(): + dataset = _dataset() + artifact_payload = { + "spec": { + "benchmark": dataset.benchmark_name, + "split": dataset.split, + "hf_path": "/home/alice/machine-only/dataset", + "hf_subset": None, + "source_revision": None, + "normalization_version": "v1", + "transforms": [], + "output_name": None, + }, + "content_digest": dataset_content_digest(dataset.artifact_frame), + "source": {}, + } + artifact_id = short_id("ds", artifact_payload) + selection_payload = {**dataset.selection_payload, "artifact_id": artifact_id} + forged = replace( + dataset, + identity_mode="native_compatibility", + artifact_payload=artifact_payload, + artifact_digest=integrity_digest(artifact_payload), + artifact_id=artifact_id, + selection_payload=selection_payload, + selection_digest=integrity_digest(selection_payload), + selection_id=short_id("sel", selection_payload), + ) + with pytest.raises(ImportIdentityError, match="spec|path|machine"): + build_import_semantic_identity( + condition=_condition(), + dataset=forged, + model=_model(), + method=_method(), + prompt=_prompt(), + ) + + +def test_partial_realization_origin_may_be_an_ordered_expected_subset(): + condition = _records().condition + identity = _realization_identity() + identity["evidence"]["evidence_status"] = "partial" + _attach_result_origin( + identity, + derivation_origin="external_import", + assignments=( + ("q2", "external_historical_inference"), + ), + ) + realization = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + assignments = realization["identity"]["realization"]["result_origin"][ + "row_assignments" + ] + assert assignments[0]["question_id"] == "q2" + + +def test_malformed_realization_origin_may_have_no_evaluable_rows(): + condition = _records().condition + identity = _realization_identity() + identity["evidence"]["evidence_status"] = "malformed" + _attach_result_origin( + identity, + derivation_origin="external_import", + assignments=(), + ) + realization = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + assert realization["identity"]["realization"]["result_origin"]["row_assignments"] == [] + + +@pytest.mark.parametrize("kind", ["model", "prompt", "condition"]) +def test_direct_dataclasses_reject_contradictory_unknown_reasons(kind): + condition = _condition() + model = _model() + prompt = _prompt() + if kind == "model": + model = replace( + model, + backend="dummy", + unknown_reasons={**model.unknown_reasons, "backend": "not recorded"}, + ) + elif kind == "prompt": + prompt = replace( + prompt, + template_identity="known", + template_digest="a" * 64, + template_contents={"direct_mcq": "known"}, + unknown_reason="not recorded", + ) + else: + condition = replace( + condition, + seed=7, + unknown_reasons={**condition.unknown_reasons, "seed": "not recorded"}, + ) + with pytest.raises(ImportIdentityError, match="unknown reason|contradict"): + build_import_semantic_identity( + condition=condition, + dataset=_dataset(), + model=model, + method=_method(), + prompt=prompt, + ) + + +def test_lineage_component_is_parent_derived_and_has_no_child_edge(): + component = make_lineage_component( + operation_type="offline_transformation", + question_id="q1", + parent_digests=("2" * 64, "1" * 64), + source_digests=("4" * 64, "3" * 64), + authorization_digest="5" * 64, + implementation=_transform, + parameters={}, + input_digest="6" * 64, + preownership_output_digest="7" * 64, + prediction_origin="external_historical_inference", + ) + + assert component["lineage_digest"] == integrity_digest(component["identity"]) + assert component["lineage_id"] == short_id("lin", component["identity"]) + assert component["identity"]["parent_digests"] == ["1" * 64, "2" * 64] + assert not ({"realization_id", "result_artifact_id", "experiment_id"} & set(component["identity"])) + + +def test_lineage_refuses_child_paths_timestamps_and_final_csv_hashes(): + for forbidden in ( + "realization_id", + "child_realization_id", + "output_path", + "result_path", + "timestamp", + "created_at", + "final_csv_sha256", + "machine_id", + "child_lineage_id", + "final_result_digest", + "lineage_id", + "next_lineage_id", + "child_component_id", + "machine", + "recorded_at", + "working_directory", + "artifact_location", + "node_name", + "executed_at", + "run_date", + "platform", + "operator", + ): + with pytest.raises(ImportIdentityError, match=forbidden): + make_lineage_component( + operation_type="offline_transformation", + question_id="q1", + parent_digests=(), + source_digests=("1" * 64,), + authorization_digest=None, + implementation=_transform, + parameters={forbidden: "bad"}, + input_digest="2" * 64, + preownership_output_digest="3" * 64, + prediction_origin="external_historical_inference", + ) + + +def test_lineage_implementation_requires_a_valid_code_identity(): + with pytest.raises(ImportIdentityError, match="implementation|source"): + make_lineage_component( + operation_type="offline_transformation", + question_id="q1", + parent_digests=("1" * 64,), + source_digests=("2" * 64,), + authorization_digest="3" * 64, + implementation={"attacker": "not-code-identity"}, + parameters={}, + input_digest="4" * 64, + preownership_output_digest="5" * 64, + prediction_origin="external_historical_inference", + ) + + +def test_lineage_refuses_bound_methods_with_unbound_instance_state(): + with pytest.raises(ImportIdentityError, match="stateless|bound|callable"): + make_lineage_component( + operation_type="offline_transformation", + question_id="q1", + parent_digests=("1" * 64,), + source_digests=("2" * 64,), + authorization_digest="3" * 64, + implementation=_StatefulCallable("state").transform, + parameters={}, + input_digest="4" * 64, + preownership_output_digest="5" * 64, + prediction_origin="external_historical_inference", + ) + + +def test_lineage_requires_mandatory_runtime_keyword_parameters(): + with pytest.raises(ImportIdentityError, match="required|threshold"): + make_lineage_component( + operation_type="offline_transformation", + question_id="q1", + parent_digests=("1" * 64,), + source_digests=("2" * 64,), + authorization_digest="3" * 64, + implementation=_transform_required, + parameters={}, + input_digest="4" * 64, + preownership_output_digest="5" * 64, + prediction_origin="external_historical_inference", + ) + + +def test_runtime_identity_refuses_ambiguous_module_lambdas(): + first = lambda value: value # noqa: E731 + second = lambda value: not value # noqa: E731 + for target in (first, second): + with pytest.raises(ImportIdentityError, match="named|lambda|addressable"): + importer_implementation_identity(adapter=target, validator=_validator) + + +def test_runtime_identity_distinguishes_same_named_top_level_function_objects(): + first = importer_implementation_identity( + adapter=_first_redefined_adapter, + validator=_validator, + ) + second = importer_implementation_identity( + adapter=_second_redefined_adapter, + validator=_validator, + ) + assert first != second + + +def test_runtime_identity_binds_behavior_affecting_resolved_globals(monkeypatch): + first = importer_implementation_identity( + adapter=_adapter, + validator=_global_validator, + ) + monkeypatch.setitem(_global_validator.__globals__, "_GLOBAL_VALIDATOR_RESULT", False) + second = importer_implementation_identity( + adapter=_adapter, + validator=_global_validator, + ) + assert first != second + + +def test_runtime_identity_binds_cross_module_helper_globals(monkeypatch): + monkeypatch.setitem( + _cross_module_adapter.__globals__, + "_cross_module_helper", + _cross_module_helpers.helper_v1, + ) + first = importer_implementation_identity( + adapter=_cross_module_adapter, + validator=_validator, + ) + monkeypatch.setitem( + _cross_module_adapter.__globals__, + "_cross_module_helper", + _cross_module_helpers.helper_v2, + ) + second = importer_implementation_identity( + adapter=_cross_module_adapter, + validator=_validator, + ) + assert first != second + + +def test_runtime_identity_refuses_uninspectable_resolved_global(monkeypatch): + monkeypatch.setitem(_global_validator.__globals__, "_GLOBAL_VALIDATOR_RESULT", object()) + with pytest.raises(ImportIdentityError, match="global|unsupported"): + importer_implementation_identity( + adapter=_adapter, + validator=_global_validator, + ) + + +def test_importer_implementation_identity_binds_importer_core_code(): + implementation = importer_implementation_identity( + adapter=_adapter, + validator=_validator, + ) + assert implementation["importer"]["source_digest"] + assert implementation["importer"]["qualified_name"].startswith( + "choicebench.importing.identity:" + ) + + +def test_lineage_supports_nonexecuted_declared_external_implementation_identity(): + source_digest = "2" * 64 + component = make_lineage_component( + operation_type="precomputed_external_transformation", + question_id="q1", + parent_digests=("1" * 64,), + source_digests=(source_digest,), + authorization_digest="3" * 64, + implementation={ + "qualified_name": "historical.matcher:rematch", + "source_file": "matcher.py", + "source_digest": source_digest, + }, + implementation_mode="declared_external", + parameters={}, + input_digest="4" * 64, + preownership_output_digest="5" * 64, + prediction_origin="external_historical_inference", + ) + assert component["identity"]["implementation"]["identity_mode"] == ( + "declared_external" + ) + assert component["identity"]["implementation"]["identity"]["source_digest"] == ( + source_digest + ) + + +def test_lineage_refuses_duplicate_parent_or_source_edges(): + for parents, sources in ( + (("1" * 64, "1" * 64), ("2" * 64,)), + (("1" * 64,), ("2" * 64, "2" * 64)), + ): + with pytest.raises(ImportIdentityError, match="duplicate"): + make_lineage_component( + operation_type="external_import", + question_id="q1", + parent_digests=parents, + source_digests=sources, + authorization_digest=None, + implementation=_transform, + parameters={}, + input_digest="4" * 64, + preownership_output_digest="5" * 64, + prediction_origin="external_historical_inference", + ) + + +def test_result_origin_records_constituents_counts_and_ordered_per_row_mapping(): + origin = make_result_origin( + derivation_origin="repair_overlay", + row_assignments=( + ("q1", "external_historical_inference", f"lin_{'a' * 16}"), + ("q2", "native_inference", f"lin_{'b' * 16}"), + ("q3", "native_inference", f"lin_{'c' * 16}"), + ), + ) + + assert origin["prediction_origins"] == [ + "external_historical_inference", + "native_inference", + ] + assert origin["prediction_origin_counts"] == { + "external_historical_inference": 1, + "native_inference": 2, + } + assert [item["question_id"] for item in origin["row_assignments"]] == [ + "q1", + "q2", + "q3", + ] + assert origin["ordered_row_origin_digest"] == integrity_digest( + origin["row_assignments"] + ) + identity = { + key: value + for key, value in origin.items() + if key not in {"origin_id", "origin_digest"} + } + assert origin["origin_digest"] == integrity_digest(identity) + assert origin["origin_id"] == short_id("origin", identity) + + +def test_realization_result_origin_question_ids_match_expected_question_set(): + condition = _records().condition + identity = _realization_identity() + identity["result_origin"] = make_result_origin( + derivation_origin="external_import", + row_assignments=( + ( + "unrelated-question", + "external_historical_inference", + f"lin_{'a' * 16}", + ), + ), + ) + with pytest.raises(ImportIdentityError, match="question set|question IDs"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_refuses_unowned_lineage_component_id(): + condition = _records().condition + identity = _realization_identity() + identity["result_origin"] = make_result_origin( + derivation_origin="external_import", + row_assignments=( + ("q2", "external_historical_inference", f"lin_{'c' * 16}"), + ("q1", "external_historical_inference", f"lin_{'d' * 16}"), + ), + ) + with pytest.raises(ImportIdentityError, match="lineage|owned|component"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_refuses_partial_coverage_for_qualified_and_included(): + condition = _records().condition + identity = _realization_identity() + identity["evidence"]["evidence_status"] = "qualified" + identity["evidence"]["scope_disposition"] = "included" + _attach_result_origin( + identity, derivation_origin="external_import", + assignments=(("q1", "external_historical_inference"),), # missing q2 + ) + with pytest.raises(ImportIdentityError, match="match the full expected"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_allows_partial_coverage_for_qualified_excluded_from_paper_matrix(): + """A realization can be individually row-complete/qualified yet still + excluded from the paper's evaluation scope for an unrelated scientific + reason (e.g. a method whose scoring mechanism isn't comparable to the + others) -- normalize_realization_rows's own evaluable computation + already treats scope_disposition != "included" as non-evaluable + regardless of evidence_status, so the identity layer must not demand + full row_origin coverage in that case either. Discovered via a real + Stage 1 freeze cell (excluded_from_paper_matrix + qualified) that this + check previously refused unconditionally.""" + condition = _records().condition + identity = _realization_identity() + identity["evidence"]["evidence_status"] = "qualified" + identity["evidence"]["scope_disposition"] = "excluded_from_paper_matrix" + _attach_result_origin( + identity, derivation_origin="external_import", + assignments=(("q1", "external_historical_inference"),), # missing q2 + ) + realization = make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + assert realization["condition_id"] == condition["condition_id"] + + +@pytest.mark.parametrize( + "logical_path", + [r"C:\machine\results.csv", r"\\server\share\results.csv"], +) +def test_realization_refuses_windows_machine_local_logical_paths(logical_path): + condition = _records().condition + identity = _realization_identity() + identity["sources"][0]["logical_path"] = logical_path + with pytest.raises(ImportIdentityError, match="path|unsafe|machine"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_runtime_lineage_must_match_registered_callable(): + condition = _records().condition + identity = _realization_identity() + original = identity["lineage_components"][0] + forged_payload = { + **original["identity"], + "implementation": { + "identity_mode": "runtime_callable", + "identity": importer_implementation_identity( + adapter=_adapter, validator=_validator + )["adapter"], + }, + } + forged = { + "lineage_id": short_id("lin", forged_payload), + "lineage_digest": integrity_digest(forged_payload), + "identity": forged_payload, + } + identity["lineage_components"][0] = forged + assignments = identity["result_origin"]["row_assignments"] + identity["result_origin"] = make_result_origin( + derivation_origin="external_import", + row_assignments=( + ( + assignments[0]["question_id"], + assignments[0]["prediction_origin"], + forged["lineage_id"], + ), + ( + assignments[1]["question_id"], + assignments[1]["prediction_origin"], + assignments[1]["prediction_lineage_id"], + ), + ), + ) + with pytest.raises(ImportIdentityError, match="runtime|callable|implementation"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +@pytest.mark.parametrize("edge", ["authorization", "parent"]) +def test_derived_realization_lineage_edges_are_owned_by_realization(edge): + condition = _records().condition + identity = _realization_identity() + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": None, + "preownership_output_digest": None, + "implementation_digest": None, + } + _attach_result_origin( + identity, + derivation_origin="repair_overlay", + assignments=( + ("q2", "external_historical_inference"), + ("q1", "external_repair_inference"), + ), + authorization_digest="c" * 64, + ) + if edge == "authorization": + identity["authorization_digest"] = "a" * 64 + else: + identity["parent_digests"]["realization_digests"] = ["b" * 64] + with pytest.raises(ImportIdentityError, match="lineage|authorization|parent"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_repair_row_itself_must_own_authorization_and_parent_lineage_edges(): + condition = _records().condition + identity = _realization_identity() + base_component = make_lineage_component( + operation_type="external_import", + question_id="q2", + parent_digests=("d" * 64,), + source_digests=("2" * 64,), + authorization_digest="c" * 64, + implementation=_transform, + parameters={}, + input_digest=integrity_digest("base-input"), + preownership_output_digest=integrity_digest("base-output"), + prediction_origin="external_historical_inference", + ) + repair_component = make_lineage_component( + operation_type="repair_overlay", + question_id="q1", + parent_digests=(), + source_digests=("2" * 64,), + authorization_digest=None, + implementation=_transform, + parameters={}, + input_digest=integrity_digest("repair-input"), + preownership_output_digest=integrity_digest("repair-output"), + prediction_origin="external_repair_inference", + ) + identity["lineage_components"] = [base_component, repair_component] + identity["result_origin"] = make_result_origin( + derivation_origin="repair_overlay", + row_assignments=( + ( + "q2", + "external_historical_inference", + base_component["lineage_id"], + ), + ("q1", "external_repair_inference", repair_component["lineage_id"]), + ), + ) + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": None, + "preownership_output_digest": None, + "implementation_digest": None, + } + with pytest.raises(ImportIdentityError, match="repair|authorization|parent|lineage"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_derived_lineage_must_reference_declared_base_realization_digest(): + condition = _records().condition + identity = _realization_identity() + _attach_result_origin( + identity, + derivation_origin="repair_overlay", + assignments=( + ("q2", "external_historical_inference"), + ("q1", "external_repair_inference"), + ), + authorization_digest="c" * 64, + ) + for component in identity["lineage_components"]: + payload = { + **component["identity"], + "parent_digests": ["0" * 64], + } + component.update( + { + "lineage_id": short_id("lin", payload), + "lineage_digest": integrity_digest(payload), + "identity": payload, + } + ) + identity["result_origin"] = make_result_origin( + derivation_origin="repair_overlay", + row_assignments=tuple( + ( + component["identity"]["question_id"], + component["identity"]["prediction_origin"], + component["lineage_id"], + ) + for component in identity["lineage_components"] + ), + ) + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": None, + "preownership_output_digest": None, + "implementation_digest": None, + } + with pytest.raises(ImportIdentityError, match="base|realization|parent|lineage"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +@pytest.mark.parametrize("derivation", ["repair_overlay", "offline_transformation"]) +def test_derived_lineage_must_bind_overlay_operational_digests(derivation): + condition = _records().condition + identity = _realization_identity() + _attach_result_origin( + identity, + derivation_origin=derivation, + assignments=( + ("q2", "external_historical_inference"), + ( + "q1", + ( + "external_repair_inference" + if derivation == "repair_overlay" + else "external_historical_inference" + ), + ), + ), + authorization_digest="c" * 64, + ) + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = { + "source_sha256": "a" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": ( + "a" * 64 if derivation == "offline_transformation" else None + ), + "preownership_output_digest": ( + "b" * 64 if derivation == "offline_transformation" else None + ), + "implementation_digest": ( + "c" * 64 if derivation == "offline_transformation" else None + ), + } + with pytest.raises(ImportIdentityError, match="overlay|source|input|output|implementation"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_offline_transformation_rows_must_reference_offline_lineage_component(): + condition = _records().condition + identity = _realization_identity() + assigned_components = [ + make_lineage_component( + operation_type="external_import", + question_id=question_id, + parent_digests=("d" * 64,), + source_digests=("2" * 64,), + authorization_digest="c" * 64, + implementation=_transform, + parameters={}, + input_digest=integrity_digest({"question_id": question_id, "input": "base"}), + preownership_output_digest=integrity_digest( + {"question_id": question_id, "output": "base"} + ), + prediction_origin="external_historical_inference", + ) + for question_id in ("q2", "q1") + ] + orphan_component = make_lineage_component( + operation_type="offline_transformation", + question_id="q1", + parent_digests=("d" * 64,), + source_digests=("2" * 64, "e" * 64), + authorization_digest="c" * 64, + implementation=_transform, + parameters={}, + input_digest=integrity_digest({"question_id": "q1", "input": "transform"}), + preownership_output_digest=integrity_digest( + {"question_id": "q1", "output": "transform"} + ), + prediction_origin="external_historical_inference", + ) + identity["lineage_components"] = [*assigned_components, orphan_component] + identity["result_origin"] = make_result_origin( + derivation_origin="offline_transformation", + row_assignments=tuple( + ( + component["identity"]["question_id"], + component["identity"]["prediction_origin"], + component["lineage_id"], + ) + for component in assigned_components + ), + ) + identity["authorization_digest"] = "c" * 64 + identity["parent_digests"]["realization_digests"] = ["d" * 64] + identity["parent_digests"]["evidence_digests"] = ["0" * 64] + identity["overlay"] = { + "source_sha256": "e" * 64, + "replacement_digest": "f" * 64, + "transformation_input_digest": integrity_digest( + [orphan_component["identity"]["input_digest"]] + ), + "preownership_output_digest": integrity_digest( + [orphan_component["identity"]["preownership_output_digest"]] + ), + "implementation_digest": integrity_digest( + orphan_component["identity"]["implementation"] + ), + } + + with pytest.raises(ImportIdentityError, match="offline|row|lineage|assignment"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_orphan_lineage_component_is_unreachable_from_any_row(): + condition = _records().condition + identity = _realization_identity() + orphan_component = make_lineage_component( + operation_type="external_import", + question_id="q3", + parent_digests=(), + source_digests=("2" * 64,), + authorization_digest=None, + implementation=_transform, + parameters={}, + input_digest=integrity_digest({"question_id": "q3", "stage": "input"}), + preownership_output_digest=integrity_digest( + { + "question_id": "q3", + "prediction_origin": "external_historical_inference", + "stage": "preownership-output", + } + ), + prediction_origin="external_historical_inference", + ) + identity["lineage_components"] = [*identity["lineage_components"], orphan_component] + with pytest.raises(ImportIdentityError, match="reachable|row"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_declared_source_must_be_reachable_from_lineage(): + condition = _records().condition + identity = _realization_identity() + identity["sources"] = [ + *identity["sources"], + { + "source_id": "unused", + "logical_path": "publisher/unused.csv", + "classification": "canonical", + "format": "csv", + "format_version": "1", + "sha256": "9" * 64, + "provenance": { + "source_run_id": None, + "source_repository": None, + "source_commit": None, + "notes_digest": None, + "evidence_digest": None, + "unknown_reasons": { + "source_run_id": "not recorded", + "source_repository": "not recorded", + "source_commit": "not recorded", + }, + }, + }, + ] + with pytest.raises(ImportIdentityError, match="reachable|source"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_lineage_operation_type_must_match_non_derived_derivation_origin(): + condition = _records().condition + identity = _realization_identity() + mismatched_component = make_lineage_component( + operation_type="repair_overlay", + question_id="q1", + parent_digests=(), + source_digests=("2" * 64,), + authorization_digest=None, + implementation=_transform, + parameters={}, + input_digest=integrity_digest({"question_id": "q1", "stage": "input"}), + preownership_output_digest=integrity_digest( + { + "question_id": "q1", + "prediction_origin": "external_historical_inference", + "stage": "preownership-output", + } + ), + prediction_origin="external_historical_inference", + ) + remaining = [ + component + for component in identity["lineage_components"] + if component["identity"]["question_id"] != "q1" + ] + identity["lineage_components"] = [*remaining, mismatched_component] + identity["result_origin"] = make_result_origin( + derivation_origin="external_import", + row_assignments=tuple( + ( + component["identity"]["question_id"], + component["identity"]["prediction_origin"], + component["lineage_id"], + ) + for component in identity["lineage_components"] + ), + ) + with pytest.raises(ImportIdentityError, match="operation type|incompatible"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_source_provenance_requires_explicit_unknown_reason(): + condition = _records().condition + identity = _realization_identity() + identity["sources"][0]["provenance"]["unknown_reasons"] = { + "source_run_id": "not recorded", + "source_repository": "not recorded", + } + with pytest.raises(ImportIdentityError, match="unknown reasons|provenance"): + make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + + +def test_realization_condition_is_bound_to_expected_dataset_selection(): + condition = _records().condition + foreign_identity = { + **condition["identity"], + "benchmark": { + **condition["identity"]["benchmark"], + "artifact_id": f"ds_{'a' * 16}", + "selection_id": f"sel_{'b' * 16}", + }, + } + foreign_condition = make_semantic_condition( + identity=foreign_identity, + fields={ + "condition_key": "foreign", + "expected_question_ids": ["q2", "q1"], + }, + ) + with pytest.raises(ImportIdentityError, match="condition|dataset|selection|artifact"): + _make_realization( + condition_id=foreign_condition["condition_id"], + condition_digest=foreign_condition["condition_digest"], + identity=_realization_identity(), + fields={}, + adapter=_adapter, + validator=_validator, + expected_dataset=_dataset(), + semantic_condition=foreign_condition, + lineage_runtime_callables={ + component["lineage_id"]: _transform + for component in _realization_identity()["lineage_components"] + }, + ) + + +def test_result_origin_rejects_bare_mixed_duplicate_rows_and_invalid_origins(): + for assignments in ( + (("q1", "mixed", f"lin_{'a' * 16}"),), + ( + ("q1", "external_historical_inference", f"lin_{'a' * 16}"), + ("q1", "native_inference", f"lin_{'b' * 16}"), + ), + ): + with pytest.raises(ImportIdentityError): + make_result_origin( + derivation_origin="repair_overlay", row_assignments=assignments + ) + + +def test_offline_transformation_keeps_underlying_prediction_origin(): + origin = make_result_origin( + derivation_origin="offline_transformation", + row_assignments=( + ("q1", "external_historical_inference", f"lin_{'a' * 16}"), + ), + ) + assert origin["derivation_origin"] == "offline_transformation" + assert origin["prediction_origins"] == ["external_historical_inference"] + + +def test_importer_implementation_identity_binds_package_adapter_and_validator(): + identity = importer_implementation_identity(adapter=_adapter, validator=_validator) + assert identity["schema_version"] == "choicebench.importer-implementation.v1" + assert identity["package"]["name"] == "choicebench" + assert identity["package"]["version"] + assert identity["adapter"]["qualified_name"].endswith(":_adapter") + assert identity["adapter"]["source_digest"] + assert identity["validator"]["qualified_name"].endswith(":_validator") + changed = importer_implementation_identity(adapter=_validator, validator=_adapter) + assert integrity_digest(identity) != integrity_digest(changed) + + +def test_importer_implementation_identity_refuses_uninspectable_callables(): + with pytest.raises(ImportIdentityError, match="source identity|stateless|callable"): + importer_implementation_identity(adapter=len, validator=_validator) diff --git a/tests/importing/test_manifest_v2_compat.py b/tests/importing/test_manifest_v2_compat.py new file mode 100644 index 0000000..9b061eb --- /dev/null +++ b/tests/importing/test_manifest_v2_compat.py @@ -0,0 +1,65 @@ +import pytest + +from choicebench.cli import evaluate_run +from choicebench.io.readers import read_manifest_results +from choicebench.manifest import validate_manifest +from tests.importing import conftest as importing_conftest + + +@pytest.fixture +def publication_reader_calls(monkeypatch): + calls = [] + original = importing_conftest.read_manifest_results + + def tracked_read_manifest_results(run_dir): + calls.append(run_dir) + return original(run_dir) + + monkeypatch.setattr( + importing_conftest, "read_manifest_results", tracked_read_manifest_results + ) + return calls + + +def test_v2_fixture_validates_without_rewrite( + publication_reader_calls, synthetic_v2_run +): + run_dir, manifest, _ = synthetic_v2_run + assert publication_reader_calls == [] + before = { + p.relative_to(run_dir): p.read_bytes() + for p in run_dir.rglob("*") + if p.is_file() + } + validate_manifest(manifest) + frame, loaded = importing_conftest.read_manifest_results(run_dir) + after = { + p.relative_to(run_dir): p.read_bytes() + for p in run_dir.rglob("*") + if p.is_file() + } + assert loaded["schema_version"] == "choicebench.manifest.v2" + assert frame["question_id"].astype(str).tolist() == ["q1", "q2"] + assert publication_reader_calls == [run_dir] + assert after == before + + +def test_v2_evaluation_identity_and_shape_are_frozen( + synthetic_v2_run, monkeypatch +): + run_dir, manifest, expected_evaluation_id = synthetic_v2_run + monkeypatch.setattr(evaluate_run, "RUNS_DIR", run_dir.parent) + frame, _ = read_manifest_results(run_dir) + report = evaluate_run.build_evaluation_report( + run_dir.name, frame, manifest, reparse=False + ) + assert report["schema_version"] == "choicebench.evaluation.v1" + assert report["evaluation_id"] == expected_evaluation_id + assert set(report["conditions"]) == {"cond_legacyfixture"} + + +def test_v2_missing_result_origin_implies_native_only_in_compat_view( + synthetic_v2_run, +): + _, manifest, _ = synthetic_v2_run + assert "result_origin" not in manifest["payload"]["conditions"][0] diff --git a/tests/importing/test_manifest_v3.py b/tests/importing/test_manifest_v3.py new file mode 100644 index 0000000..b7a67f1 --- /dev/null +++ b/tests/importing/test_manifest_v3.py @@ -0,0 +1,203 @@ +"""Additive manifest-v3 tests: new schema alongside the untouched v2 path. + +Scope is deliberately reduced from the original plan: no ManifestView +cross-version normalization and no legacy-v2-in-v3 synthesis. Every manifest +this module builds is a fresh v3 import; existing v2 manifests keep using the +unmodified v2 functions exercised by tests/importing/test_manifest_v2_compat.py. +""" + +from __future__ import annotations + +import pytest + +from choicebench.identity import CANONICALIZATION_VERSION, integrity_digest, short_id +from choicebench.manifest import ( + MANIFEST_SCHEMA_VERSION, + MANIFEST_V3_SCHEMA_VERSION, + PROTOCOL_V3_VERSION, + RUN_STATE_V3_SCHEMA_VERSION, + ManifestCompatibilityError, + initial_run_state_v3, + make_manifest_v3, + validate_manifest, + validate_manifest_v3, + validate_run_state_v3, +) + + +def _condition_identity_triple(): + """A condition_id/digest pair that genuinely matches its own identity payload, + so validate_manifest_v3's recomputation check passes on the happy path.""" + identity = {"seed": 1} + digest = integrity_digest(identity) + condition_id = short_id("cond", identity) + return condition_id, digest, identity + + +def _realization_record(condition_id, condition_digest, realization_identity): + payload = { + "schema_version": "choicebench.realization.v1", + "condition_id": condition_id, + "condition_digest": condition_digest, + "realization": realization_identity, + } + return { + "realization_id": short_id("real", payload), + "realization_digest": integrity_digest(payload), + "condition_id": condition_id, + "condition_digest": condition_digest, + "identity": payload, + } + + +@pytest.fixture +def v3_payload(): + condition_id, condition_digest, identity = _condition_identity_triple() + condition = { + "condition_id": condition_id, + "condition_digest": condition_digest, + "identity": identity, + } + realization = _realization_record(condition_id, condition_digest, {"source": "x"}) + return { + "protocol_version": PROTOCOL_V3_VERSION, + "canonicalization_version": CANONICALIZATION_VERSION, + "semantic_conditions": {condition_id: condition}, + "realizations": {realization["realization_id"]: realization}, + }, condition_id, realization["realization_id"] + + +def test_make_manifest_v3_computes_schema_experiment_and_payload_digests(v3_payload): + payload, condition_id, realization_id = v3_payload + manifest = make_manifest_v3(payload) + assert manifest["schema_version"] == MANIFEST_V3_SCHEMA_VERSION + assert manifest["experiment_id"].startswith("exp_") + assert manifest["experiment_digest"] == integrity_digest(manifest["payload"]) + assert manifest["payload_digest"] == integrity_digest(manifest["payload"]) + assert manifest["payload"]["semantic_conditions"][condition_id]["condition_id"] == condition_id + assert realization_id in manifest["payload"]["realizations"] + + +def test_make_manifest_v3_excludes_audit_from_identity(v3_payload): + payload, _, _ = v3_payload + first = make_manifest_v3(payload, audit={"source_location": "/a/one"}) + second = make_manifest_v3(payload, audit={"source_location": "/b/two"}) + assert first["experiment_id"] == second["experiment_id"] + assert first["experiment_digest"] == second["experiment_digest"] + assert first["audit"] != second["audit"] + + +def test_validate_manifest_v3_accepts_a_well_formed_manifest(v3_payload): + payload, _, _ = v3_payload + manifest = make_manifest_v3(payload) + validate_manifest_v3(manifest) # must not raise + + +def test_validate_manifest_v3_rejects_wrong_schema_version(v3_payload): + payload, _, _ = v3_payload + manifest = make_manifest_v3(payload) + manifest["schema_version"] = MANIFEST_SCHEMA_VERSION + with pytest.raises(ManifestCompatibilityError): + validate_manifest_v3(manifest) + + +def test_validate_manifest_v3_rejects_tampered_payload_digest(v3_payload): + payload, _, _ = v3_payload + manifest = make_manifest_v3(payload) + manifest["payload"]["semantic_conditions"] = {} + with pytest.raises(ManifestCompatibilityError): + validate_manifest_v3(manifest) + + +def test_validate_manifest_v3_recomputes_condition_id_and_digest(v3_payload): + payload, condition_id, _ = v3_payload + manifest = make_manifest_v3(payload) + forged = dict(manifest["payload"]["semantic_conditions"][condition_id]) + forged["condition_digest"] = "f" * 64 + manifest["payload"]["semantic_conditions"][condition_id] = forged + manifest["payload_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_id"] = f"exp_{manifest['experiment_digest'][:16]}" + with pytest.raises(ManifestCompatibilityError, match="condition"): + validate_manifest_v3(manifest) + + +def test_validate_manifest_v3_recomputes_realization_id_and_digest(v3_payload): + payload, _, realization_id = v3_payload + manifest = make_manifest_v3(payload) + forged = dict(manifest["payload"]["realizations"][realization_id]) + forged["realization_digest"] = "f" * 64 + manifest["payload"]["realizations"][realization_id] = forged + manifest["payload_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_id"] = f"exp_{manifest['experiment_digest'][:16]}" + with pytest.raises(ManifestCompatibilityError, match="realization"): + validate_manifest_v3(manifest) + + +def test_validate_manifest_v3_requires_realization_condition_binding(v3_payload): + payload, condition_id, realization_id = v3_payload + manifest = make_manifest_v3(payload) + forged_realization = dict(manifest["payload"]["realizations"][realization_id]) + forged_realization["condition_digest"] = "e" * 64 + manifest["payload"]["realizations"][realization_id] = forged_realization + manifest["payload_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_id"] = f"exp_{manifest['experiment_digest'][:16]}" + with pytest.raises(ManifestCompatibilityError, match="condition"): + validate_manifest_v3(manifest) + + +def test_validate_manifest_v3_rejects_realization_with_unknown_condition(v3_payload): + payload, condition_id, realization_id = v3_payload + manifest = make_manifest_v3(payload) + del manifest["payload"]["semantic_conditions"][condition_id] + manifest["payload_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_digest"] = integrity_digest(manifest["payload"]) + manifest["experiment_id"] = f"exp_{manifest['experiment_digest'][:16]}" + with pytest.raises(ManifestCompatibilityError, match="condition"): + validate_manifest_v3(manifest) + + +def test_run_state_v3_is_keyed_by_realization(v3_payload): + payload, _, realization_id = v3_payload + manifest = make_manifest_v3(payload) + state = initial_run_state_v3(manifest) + assert state["schema_version"] == RUN_STATE_V3_SCHEMA_VERSION + assert state["experiment_id"] == manifest["experiment_id"] + assert state["realizations"] == {realization_id: {"status": "pending"}} + validate_run_state_v3(state, manifest) # must not raise + + +def test_run_state_v3_rejects_state_for_a_different_manifest(v3_payload): + payload, _, _ = v3_payload + manifest = make_manifest_v3(payload) + other_manifest = make_manifest_v3( + {**payload, "semantic_conditions": {}, "realizations": {}} + ) + state = initial_run_state_v3(manifest) + with pytest.raises(ManifestCompatibilityError): + validate_run_state_v3(state, other_manifest) + + +def test_run_state_v3_rejects_realization_grid_drift(v3_payload): + payload, _, realization_id = v3_payload + manifest = make_manifest_v3(payload) + state = initial_run_state_v3(manifest) + del state["realizations"][realization_id] + with pytest.raises(ManifestCompatibilityError): + validate_run_state_v3(state, manifest) + + +def test_v2_manifest_functions_are_unaffected_by_v3_additions(): + """Sanity check that importing the v3 symbols doesn't change v2 behavior.""" + v2_manifest = { + "schema_version": MANIFEST_SCHEMA_VERSION, + "experiment_id": "exp_0000000000000000", + "experiment_digest": "0" * 64, + "payload_digest": "0" * 64, + "created_at": "now", + "payload": {"protocol_version": "choicebench.protocol.v2", "config": {}}, + } + with pytest.raises(ManifestCompatibilityError): + validate_manifest(v2_manifest) # digest mismatch, but still v2-only code path diff --git a/tests/importing/test_overlays.py b/tests/importing/test_overlays.py new file mode 100644 index 0000000..6e27a43 --- /dev/null +++ b/tests/importing/test_overlays.py @@ -0,0 +1,305 @@ +"""Tests for pure overlay derivation (Task 10, reduced scope).""" + +from __future__ import annotations + +from pathlib import Path + +import pytest + +from choicebench.importing.authorization import ValidatedAuthorization +from choicebench.importing.csv_adapter import AdaptedTable, LogicalRecordSpan, SourceRow +from choicebench.importing.overlays import ( + OverlayError, + VerifiedBaseRealization, + derive_overlay, +) +from choicebench.importing.schema import OverlaySpec, ResultOriginSpec +from tests.importing.test_identity import _dataset + +_BASE_CONDITION_DIGEST = "c" * 64 +_BASE_REALIZATION_ID = "real_" + "a" * 16 +_BASE_REALIZATION_DIGEST = "d" * 64 +_EVIDENCE_DIGESTS = {"ev1": "1" * 64} +_VALIDATION_SHA = "3" * 64 +_RESULT_SHA = "4" * 64 +_AUTH_BUNDLE_ID = "auth_" + "b" * 16 +_AUTH_BUNDLE_DIGEST = "5" * 64 + + +def _base(**overrides): + defaults = dict( + condition_id="cond_a", + condition_digest=_BASE_CONDITION_DIGEST, + realization_id=_BASE_REALIZATION_ID, + realization_digest=_BASE_REALIZATION_DIGEST, + evidence_index_digest="2" * 64, + evidence_source_digests=_EVIDENCE_DIGESTS, + validation_artifact_sha256=_VALIDATION_SHA, + result_sha256=_RESULT_SHA, + rows_by_question_id={"q1": {"question_id": "q1"}, "q2": {"question_id": "q2"}}, + prediction_origins={"q1": "external_historical_inference", "q2": "external_historical_inference"}, + ) + defaults.update(overrides) + return VerifiedBaseRealization(**defaults) + + +def _authorization(*, authorization_type, executable, question_reasons): + return ValidatedAuthorization( + bundle_id=_AUTH_BUNDLE_ID, + bundle_digest=_AUTH_BUNDLE_DIGEST, + authorization_type=authorization_type, + condition_digest=_BASE_CONDITION_DIGEST, + question_reasons=dict(question_reasons), + authority="principal-investigator", + purpose="repair", + executable=executable, + source_sha256="6" * 64, + input_evidence_digests=_EVIDENCE_DIGESTS, + expected_snapshot_digest="7" * 64, + ) + + +def _overlay(*, result_origin, replacement_reasons, authorization_id=_AUTH_BUNDLE_ID, **overrides): + defaults = dict( + overlay_id="ovl_1", + base_run_path=Path("/tmp/base-run"), + base_condition_digest=_BASE_CONDITION_DIGEST, + base_realization_id=_BASE_REALIZATION_ID, + base_realization_digest=_BASE_REALIZATION_DIGEST, + base_evidence_digests=_EVIDENCE_DIGESTS, + base_validation_artifact_sha256=_VALIDATION_SHA, + base_result_sha256=_RESULT_SHA, + source_id="overlay-source", + authorization_id=authorization_id, + replacement_reasons=dict(replacement_reasons), + result_origin=result_origin, + lineage_notes={}, + # source_digest must match the overlay source's own checksum ("b" * 64, + # see _overlay_table) -- make_lineage_component requires a declared + # external implementation's code to be one of the declared sources. + implementation={"qualified_name": "external:repair_tool", "source_digest": "b" * 64}, + input_digest="9" * 64, + preownership_output_digest="a" * 64, + expected_evidence_status="complete", + ) + defaults.update(overrides) + return OverlaySpec(**defaults) + + +def _overlay_table(rows): + return AdaptedTable(columns=("qid", "predicted_letter"), rows=rows, source_sha256="b" * 64) + + +def _overlay_row(qid, predicted): + return SourceRow( + values={"qid": qid, "predicted_letter": predicted}, + span=LogicalRecordSpan(index=0, start=0, end=9, terminator=b"\n"), + raw_sha256="0" * 64, + ) + + +_MAPPING = {"question_id": "qid", "prediction": "predicted_letter"} + + +def test_derive_overlay_valid_inference_repair(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", + default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + replacement_reasons={"q1": "repair malformed answer"}, + ) + derived = derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + assert derived.replacement_question_ids == ("q1",) + assert derived.preownership_rows[0]["prediction_origin"] == "native_inference" + assert derived.preownership_rows[0]["predicted_option"] == "B" + assert derived.result_origin["derivation_origin"] == "repair_overlay" + assert derived.lineage_components[0]["identity"]["operation_type"] == "repair_overlay" + assert derived.lineage_components[0]["identity"]["authorization_digest"] == _AUTH_BUNDLE_DIGEST + + +def test_derive_overlay_valid_offline_transformation_retains_underlying_origin(): + authorization = _authorization( + authorization_type="offline_transformation", executable=False, question_reasons={"q2": "rematch"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="offline_transformation", + default_prediction_origin=None, + per_question_prediction_origins={}, + ), + replacement_reasons={"q2": "offline semantic rematch"}, + ) + derived = derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q2", "A"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + # base.prediction_origins["q2"] == "external_historical_inference" -- retained, not reassigned + assert derived.preownership_rows[0]["prediction_origin"] == "external_historical_inference" + assert derived.result_origin["derivation_origin"] == "offline_transformation" + + +def test_derive_overlay_refuses_cross_condition_authorization(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + other_base = _base(condition_digest="e" * 64) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + replacement_reasons={"q1": "reason"}, + base_condition_digest="e" * 64, + ) + with pytest.raises(OverlayError, match="condition"): + derive_overlay( + base=other_base, overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_base_realization_mismatch(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + replacement_reasons={"q1": "reason"}, + base_realization_digest="f" * 64, # forged + ) + with pytest.raises(OverlayError, match="base"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_evidence_digest_mismatch(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + replacement_reasons={"q1": "reason"}, + base_evidence_digests={"ev1": "9" * 64}, # diverges from authorization/base + ) + with pytest.raises(OverlayError, match="evidence"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_unauthorized_replacement(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q2": "native_inference"}, + ), + replacement_reasons={"q2": "not actually authorized"}, + ) + with pytest.raises(OverlayError, match="not granted"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q2", "A"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_missing_repair_origin_assignment(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={}, # no assignment for q1 + ), + replacement_reasons={"q1": "reason"}, + ) + with pytest.raises(OverlayError, match="origin assignment"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_invalid_repair_origin_value(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q1": "mixed"}, + ), + replacement_reasons={"q1": "reason"}, + ) + with pytest.raises(OverlayError, match="valid repair prediction origin"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_duplicate_replacement_row(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + replacement_reasons={"q1": "reason"}, + ) + with pytest.raises(OverlayError, match="exactly one"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"), _overlay_row("q1", "A"))), + overlay_mapping=_MAPPING, expected=_dataset(), + ) + + +def test_derive_overlay_refuses_authorization_bundle_mismatch(): + authorization = _authorization( + authorization_type="inference_repair", executable=True, question_reasons={"q1": "malformed"} + ) + overlay = _overlay( + result_origin=ResultOriginSpec( + derivation_origin="repair_overlay", default_prediction_origin=None, + per_question_prediction_origins={"q1": "native_inference"}, + ), + replacement_reasons={"q1": "reason"}, + authorization_id="auth_" + "9" * 16, # different bundle than the one supplied + ) + with pytest.raises(OverlayError, match="authorization bundle"): + derive_overlay( + base=_base(), overlay=overlay, authorization=authorization, + overlay_table=_overlay_table((_overlay_row("q1", "B"),)), + overlay_mapping=_MAPPING, expected=_dataset(), + ) diff --git a/tests/importing/test_result_artifact.py b/tests/importing/test_result_artifact.py new file mode 100644 index 0000000..88fd0f6 --- /dev/null +++ b/tests/importing/test_result_artifact.py @@ -0,0 +1,260 @@ +"""Result-artifact-v2 tests: non-circular identity for imported realizations. + +Reuses tests/importing/test_identity.py's condition/realization builders +instead of re-deriving a full semantic identity fixture here. +""" + +from __future__ import annotations + +import hashlib +import json + +import pandas as pd +import pytest + +from choicebench.identity import CANONICALIZATION_VERSION, integrity_digest +from choicebench.manifest import PROTOCOL_V3_VERSION, make_manifest_v3 +from choicebench.io.writers import ( + RESULT_ARTIFACT_V2_SCHEMA_VERSION, + prepare_manifest_result, + publish_manifest_result, + validate_result_artifact, +) +from tests.importing.test_identity import ( + _realization_identity, + _records, + make_realization as _make_realization, +) + + +def _build_manifest_and_realization(): + condition = _records().condition + identity = _realization_identity() + realization = _make_realization( + condition_id=condition["condition_id"], + condition_digest=condition["condition_digest"], + identity=identity, + fields={}, + ) + payload = { + "protocol_version": PROTOCOL_V3_VERSION, + "canonicalization_version": CANONICALIZATION_VERSION, + "semantic_conditions": {condition["condition_id"]: condition}, + "realizations": {realization["realization_id"]: realization}, + } + manifest = make_manifest_v3(payload) + return manifest, realization + + +def _rows_for(manifest, realization): + """Rows shaped as Task 9's normalization step would deliver them: already + carrying the fixed ownership/identity columns prepare_manifest_result + validates (never fabricates) per row.""" + realization_id = realization["realization_id"] + condition_id = realization["condition_id"] + condition = manifest["payload"]["semantic_conditions"][condition_id] + benchmark = condition["identity"]["benchmark"] + identity_columns = { + "condition_id": condition_id, + "realization_id": realization_id, + "experiment_id": manifest["experiment_id"], + "dataset_artifact_id": benchmark["artifact_id"], + "dataset_selection_id": benchmark["selection_id"], + "model_id": condition["identity"]["model_id"], + "method_id": condition["identity"]["method_id"], + "prompt_id": condition["identity"]["prompt_id"], + "benchmark_name": benchmark["name"], + "benchmark_split": benchmark["split"], + } + row_assignments = realization["identity"]["realization"]["result_origin"]["row_assignments"] + return [ + { + **identity_columns, + "question_id": row["question_id"], + "prediction_origin": row["prediction_origin"], + "prediction_lineage_id": row["prediction_lineage_id"], + "parsed_choice": "A", + "is_correct": True, + } + for row in row_assignments + ] + + +def test_prepare_manifest_result_computes_identity_after_manifest_is_fixed(): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + assert "result_artifact_id" not in manifest + assert prepared.metadata["experiment_id"] == manifest["experiment_id"] + assert prepared.metadata["realization_id"] == realization_id + assert prepared.metadata["file_sha256"] == hashlib.sha256(prepared.csv_bytes).hexdigest() + assert prepared.metadata["result_artifact_id"].startswith("result_") + assert prepared.result_path == f"results/{realization_id}.csv" + assert prepared.metadata_path == f"results/{realization_id}.artifact.json" + + +def test_prepare_manifest_result_injects_identity_columns_per_row(): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + df = pd.read_csv(pd.io.common.BytesIO(prepared.csv_bytes)) + for column in ( + "condition_id", + "realization_id", + "experiment_id", + "dataset_artifact_id", + "dataset_selection_id", + "model_id", + "method_id", + "prompt_id", + "benchmark_name", + "benchmark_split", + "prediction_origin", + "prediction_lineage_id", + ): + assert column in df.columns + assert df[column].nunique(dropna=False) == 1 or column in ( + "prediction_origin", + "prediction_lineage_id", + ) + assert set(df["realization_id"]) == {realization_id} + + +def test_prepare_manifest_result_orders_rows_by_declared_row_assignments(): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + rows = _rows_for(manifest, realization) + prepared = prepare_manifest_result( + list(reversed(rows)), manifest=manifest, realization_id=realization_id + ) + expected_order = [row["question_id"] for row in rows] + df = pd.read_csv(pd.io.common.BytesIO(prepared.csv_bytes), dtype={"question_id": "string"}) + assert df["question_id"].tolist() == expected_order + + +def test_prepare_manifest_result_rejects_unknown_question_id(): + manifest, realization = _build_manifest_and_realization() + rows = _rows_for(manifest, realization) + rows[0]["question_id"] = "not-a-declared-question" + with pytest.raises(RuntimeError, match="unexpected question_id"): + prepare_manifest_result( + rows, manifest=manifest, realization_id=realization["realization_id"] + ) + + +def test_prepare_manifest_result_rejects_missing_row(): + manifest, realization = _build_manifest_and_realization() + rows = _rows_for(manifest, realization)[:-1] + with pytest.raises(RuntimeError, match="missing result rows"): + prepare_manifest_result( + rows, manifest=manifest, realization_id=realization["realization_id"] + ) + + +def test_prepare_manifest_result_rejects_prediction_origin_mismatch(): + manifest, realization = _build_manifest_and_realization() + rows = _rows_for(manifest, realization) + rows[0]["prediction_origin"] = "native_inference" + with pytest.raises(RuntimeError, match="prediction_origin"): + prepare_manifest_result( + rows, manifest=manifest, realization_id=realization["realization_id"] + ) + + +def test_prepare_manifest_result_rejects_unknown_realization(): + manifest, _realization = _build_manifest_and_realization() + with pytest.raises(RuntimeError, match="[Uu]nknown realization"): + prepare_manifest_result( + [], manifest=manifest, realization_id="real_" + "0" * 16 + ) + + +def test_publish_manifest_result_writes_csv_and_sidecar(tmp_path): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + path, metadata, reused = publish_manifest_result(prepared, run_dir=tmp_path) + assert reused is False + assert path == tmp_path / "results" / f"{realization_id}.csv" + assert path.read_bytes() == prepared.csv_bytes + sidecar_path = tmp_path / "results" / f"{realization_id}.artifact.json" + assert json.loads(sidecar_path.read_text()) == dict(prepared.metadata) + assert metadata["result_artifact_id"] == prepared.metadata["result_artifact_id"] + + +def test_publish_manifest_result_is_idempotent_noop_for_identical_bytes(tmp_path): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + publish_manifest_result(prepared, run_dir=tmp_path) + _, _, second_reused = publish_manifest_result(prepared, run_dir=tmp_path) + assert second_reused is True + + +def test_publish_manifest_result_refuses_divergent_overwrite(tmp_path): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + rows = _rows_for(manifest, realization) + prepared = prepare_manifest_result(rows, manifest=manifest, realization_id=realization_id) + publish_manifest_result(prepared, run_dir=tmp_path) + rows[0]["parsed_choice"] = "B" + divergent = prepare_manifest_result(rows, manifest=manifest, realization_id=realization_id) + with pytest.raises(RuntimeError, match="divergent"): + publish_manifest_result(divergent, run_dir=tmp_path) + + +def test_validate_result_artifact_v2_accepts_a_well_formed_pair(tmp_path): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + path, _, _ = publish_manifest_result(prepared, run_dir=tmp_path) + metadata = validate_result_artifact(path, manifest=manifest, realization=realization) + assert metadata["schema_version"] == RESULT_ARTIFACT_V2_SCHEMA_VERSION + + +def test_validate_result_artifact_v2_rejects_tampered_csv_bytes(tmp_path): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + path, _, _ = publish_manifest_result(prepared, run_dir=tmp_path) + path.write_bytes(prepared.csv_bytes + b"\n# tampered") + with pytest.raises(RuntimeError, match="content integrity"): + validate_result_artifact(path, manifest=manifest, realization=realization) + + +def test_validate_result_artifact_v2_rejects_realization_binding_mismatch(tmp_path): + manifest, realization = _build_manifest_and_realization() + realization_id = realization["realization_id"] + prepared = prepare_manifest_result( + _rows_for(manifest, realization), manifest=manifest, realization_id=realization_id + ) + path, _, _ = publish_manifest_result(prepared, run_dir=tmp_path) + empty_manifest = make_manifest_v3( + { + "protocol_version": PROTOCOL_V3_VERSION, + "canonicalization_version": CANONICALIZATION_VERSION, + "semantic_conditions": {}, + "realizations": {}, + } + ) + with pytest.raises(RuntimeError, match="realization"): + validate_result_artifact(path, manifest=empty_manifest) + + +def test_validate_result_artifact_v1_path_is_unaffected(): + """Calling with no manifest/realization keeps the exact v1 behavior.""" + from choicebench.io.writers import RESULT_ARTIFACT_SCHEMA_VERSION + + assert RESULT_ARTIFACT_SCHEMA_VERSION == "choicebench.result-artifact.v1" diff --git a/tests/importing/test_schema.py b/tests/importing/test_schema.py new file mode 100644 index 0000000..ce109bc --- /dev/null +++ b/tests/importing/test_schema.py @@ -0,0 +1,1268 @@ +from __future__ import annotations + +from copy import deepcopy +from dataclasses import replace +from pathlib import Path + +import pytest +import yaml + +from choicebench.identity import integrity_digest, short_id +from choicebench.importing.schema import ( + ImportSpecError, + import_spec_digest, + load_import_spec, + stable_import_projection, +) +from choicebench.metrics import BUILTIN_METRICS +from choicebench.pipeline.prompt_builder import prompt_bundle_identity +from choicebench.provenance import implementation_identity + + +SHA_A = "a" * 64 +SHA_B = "b" * 64 +SHA_C = "c" * 64 + + +def _native_identity(prefix: str, payload: dict) -> dict: + return { + "payload": payload, + "digest": integrity_digest(payload), + f"{prefix}_id": short_id(prefix, payload), + } + + +@pytest.fixture +def minimal_raw(tmp_path: Path) -> dict: + return { + "schema_version": "choicebench.import-spec.v1", + "import_name": "minimal historical results", + "sources": [ + { + "source_id": "results", + "path": str(tmp_path / "results.csv"), + "logical_path": "freeze/results.csv", + "expected_sha256": SHA_A, + "format": "csv", + "format_version": "producer-v1", + "classification": "raw", + "dialect": { + "encoding": "utf-8", + "bom_policy": "forbid", + "decoding_errors": "strict", + "delimiter": ",", + "quote_character": '"', + "escape_character": None, + "double_quote": True, + "line_terminators": ["crlf", "lf", "cr"], + "mixed_line_terminators": "allow", + "final_record_without_terminator": "allow", + "blank_record_policy": "reject", + "skip_initial_space": False, + "header": "first_logical_record", + "strict_syntax": True, + }, + "columns": { + "question_id": "qid", + "prediction": "answer", + "gold": "gold", + }, + "expected_columns": [ + "qid", + "question", + "choice_a", + "choice_b", + "answer", + "gold", + "score", + ], + "ignored_columns": {"score": "producer aggregate only"}, + "null_values": ["", "NA"], + "numeric_columns": [ + { + "source_column": "score", + "value_type": "float", + "null_allowed": True, + "finite_only": True, + } + ], + "option_mapping": { + "mode": "ordered_columns", + "ordered_columns": ["choice_a", "choice_b"], + "structured_column": None, + "structured_label_key": None, + "structured_text_key": None, + }, + "extra_field_policy": "preserve_unmapped", + "preserve_namespace": "producer", + "source_run_id": None, + "source_repository": None, + "source_commit": None, + "notes": {}, + } + ], + "datasets": [ + { + "dataset_id": "dataset", + "benchmark_name": "historical-benchmark", + "split": "test", + "reference_kind": "independent_input_snapshot", + "trust_label": "producer-supplied", + "source_ids": ["results"], + "selection_source_id": "results", + "expected_question_ids": ["q1", "q2"], + "selection_seed": None, + "selection_n_samples": None, + "subject_filter": [], + "selection_unknown_reasons": { + "selection_seed": "not recorded by producer", + "selection_n_samples": "not recorded by producer", + }, + "columns": { + "question_id": "qid", + "question_text": "question", + "correct_option": "gold", + }, + "revision": None, + "fingerprint": None, + "derivation": {}, + "limitations": ["publisher revision was not recorded"], + "native_compatibility_identity": None, + } + ], + "models": [ + { + "model_key": "model", + "display_name": "historical-model", + "backend": None, + "provider": None, + "revision": None, + "effective_parameters": {}, + "unknown_reasons": { + "backend": "not recorded by producer", + "provider": "not recorded by producer", + "revision": "not recorded by producer", + }, + "native_compatibility_identity": None, + } + ], + "methods": [ + { + "method_key": "method", + "name": "historical-direct", + "effective_parameters": {}, + "implementation": None, + "unknown_reasons": { + "implementation": "not recorded by producer" + }, + "native_compatibility_identity": None, + } + ], + "prompts": [ + { + "prompt_key": "prompt", + "template_identity": None, + "template_digest": None, + "template_contents": None, + "unknown_reason": "not recorded by producer", + "native_compatibility_identity": None, + } + ], + "conditions": [ + { + "condition_key": "condition", + "source_ids": ["results"], + "dataset_id": "dataset", + "model_key": "model", + "method_key": "method", + "prompt_key": "prompt", + "seed": None, + "calibration_identity": None, + "preflight_identity": None, + "protocol_settings": {}, + "generation_parameters": {}, + "unknown_reasons": { + "seed": "not recorded by producer", + "calibration_identity": "not recorded by producer", + "preflight_identity": "not recorded by producer", + }, + "expected_question_ids": ["q1", "q2"], + "evidence_status": "complete", + "scope_disposition": "included", + "executable": None, + "qualifications": [], + "limitations": [], + "damaged_question_ids": [], + "recoverable_question_ids": [], + "result_origin": { + "derivation_origin": "external_import", + "default_prediction_origin": "external_historical_inference", + "per_question_prediction_origins": {}, + }, + } + ], + "authorizations": [], + "overlays": [], + "metrics": ["accuracy"], + "provenance": { + "producer_request_id": { + "value": None, + "reason": "not recorded by producer", + } + }, + "audit": { + "source_path": str(tmp_path), + "imported_at": "2026-07-18T00:00:00Z", + }, + } + + +def _load(tmp_path: Path, raw: object): + path = tmp_path / "import.yaml" + path.write_text(yaml.safe_dump(raw, sort_keys=False), encoding="utf-8") + return load_import_spec(path) + + +def test_protocol_settings_use_the_closed_current_choicebench_schema( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + raw["conditions"][0]["protocol_settings"] = {"operator": "alice"} + with pytest.raises(ImportSpecError, match="protocol_settings|unsupported"): + _load(tmp_path, raw) + + +@pytest.fixture +def minimal_spec(tmp_path: Path, minimal_raw: dict): + return _load(tmp_path, minimal_raw) + + +@pytest.fixture +def minimal_overlay_spec(tmp_path: Path, minimal_raw: dict): + raw = deepcopy(minimal_raw) + raw["authorizations"] = [ + { + "authorization_id": "auth", + "authorization_type": "inference_repair", + "source_id": "results", + "condition_question_reasons": { + "condition": {"q2": "repair explicitly approved"} + }, + "authority": "benchmark owner", + "purpose": "repair damaged prediction", + "executable": True, + "input_evidence_digests": {"results": SHA_A}, + "expected_snapshot_digests": {"q2": SHA_B}, + } + ] + raw["overlays"] = [ + { + "overlay_id": "overlay", + "base_run_path": str(tmp_path / "base-run"), + "base_condition_digest": SHA_A, + "base_realization_id": "realization_base", + "base_realization_digest": SHA_B, + "base_evidence_digests": {"results": SHA_A}, + "base_validation_artifact_sha256": SHA_C, + "base_result_sha256": None, + "source_id": "results", + "authorization_id": "auth", + "replacement_reasons": {"q2": "malformed producer row"}, + "result_origin": { + "derivation_origin": "repair_overlay", + "default_prediction_origin": None, + "per_question_prediction_origins": { + "q2": "external_repair_inference" + }, + }, + "lineage_notes": {}, + "implementation": {"name": "approved repair"}, + "input_digest": SHA_A, + "preownership_output_digest": SHA_B, + "expected_evidence_status": "qualified", + } + ] + return _load(tmp_path, raw) + + +def test_loads_valid_minimal_yaml(minimal_spec): + assert minimal_spec.schema_version == "choicebench.import-spec.v1" + assert minimal_spec.sources[0].path.is_absolute() + assert minimal_spec.sources[0].numeric_columns[0].value_type == "float" + assert minimal_spec.sources[0].option_mapping.mode == "ordered_columns" + assert minimal_spec.metrics == ("accuracy",) + + +def test_safe_loader_rejects_non_mapping_and_python_object_tag(tmp_path: Path): + path = tmp_path / "bad.yaml" + path.write_text("- not\n- a\n- mapping\n", encoding="utf-8") + with pytest.raises(ImportSpecError, match="top level"): + load_import_spec(path) + + path.write_text("!!python/object/apply:os.system [['echo unsafe']]\n", encoding="utf-8") + with pytest.raises(ImportSpecError, match="YAML"): + load_import_spec(path) + + +@pytest.mark.parametrize( + ("section", "unknown_key"), + [ + (None, "output_root"), + ("sources", "output_path"), + ("dialect", "sniff"), + ("numeric_columns", "coerce"), + ("option_mapping", "labels"), + ("datasets", "paper_name"), + ("models", "endpoint"), + ("methods", "entry_point"), + ("prompts", "prompt_path"), + ("conditions", "import_state"), + ("result_origin", "producer"), + ("authorizations", "approved_at"), + ("overlays", "output_path"), + ], +) +def test_rejects_unknown_keys_at_every_schema_layer( + tmp_path: Path, minimal_raw: dict, section: str | None, unknown_key: str +): + raw = deepcopy(minimal_raw) + if section is None: + raw[unknown_key] = "unsafe" + elif section == "dialect": + raw["sources"][0]["dialect"][unknown_key] = True + elif section == "numeric_columns": + raw["sources"][0]["numeric_columns"][0][unknown_key] = True + elif section == "option_mapping": + raw["sources"][0]["option_mapping"][unknown_key] = [] + elif section == "result_origin": + raw["conditions"][0]["result_origin"][unknown_key] = "historical" + elif section == "authorizations": + raw["authorizations"] = deepcopy( + minimal_overlay_raw(raw)["authorizations"] + ) + raw["authorizations"][0][unknown_key] = "later" + elif section == "overlays": + overlay_raw = minimal_overlay_raw(raw) + raw["authorizations"] = overlay_raw["authorizations"] + raw["overlays"] = overlay_raw["overlays"] + raw["overlays"][0][unknown_key] = "unsafe" + else: + raw[section][0][unknown_key] = "unsafe" + with pytest.raises(ImportSpecError, match="Unknown field"): + _load(tmp_path, raw) + + +def minimal_overlay_raw(raw: dict) -> dict: + result = deepcopy(raw) + result["authorizations"] = [ + { + "authorization_id": "auth", + "authorization_type": "inference_repair", + "source_id": "results", + "condition_question_reasons": {"condition": {"q2": "approved"}}, + "authority": "owner", + "purpose": "repair", + "executable": True, + "input_evidence_digests": {"results": SHA_A}, + "expected_snapshot_digests": {"q2": SHA_B}, + } + ] + result["overlays"] = [ + { + "overlay_id": "overlay", + "base_run_path": "/audit/base", + "base_condition_digest": SHA_A, + "base_realization_id": "realization_base", + "base_realization_digest": SHA_B, + "base_evidence_digests": {"results": SHA_A}, + "base_validation_artifact_sha256": SHA_C, + "base_result_sha256": None, + "source_id": "results", + "authorization_id": "auth", + "replacement_reasons": {"q2": "damaged"}, + "result_origin": { + "derivation_origin": "repair_overlay", + "default_prediction_origin": None, + "per_question_prediction_origins": { + "q2": "external_repair_inference" + }, + }, + "lineage_notes": {}, + "implementation": {"name": "repair"}, + "input_digest": SHA_A, + "preownership_output_digest": SHA_B, + "expected_evidence_status": "qualified", + } + ] + return result + + +@pytest.mark.parametrize( + ("path", "value"), + [ + (("sources", 0, "dialect", "double_quote"), 1), + (("sources", 0, "numeric_columns", 0, "null_allowed"), "false"), + (("conditions", 0, "seed"), True), + (("conditions", 0, "executable"), 0), + (("sources", 0, "notes", "temperature"), float("nan")), + (("models", 0, "effective_parameters", "temperature"), float("inf")), + ], +) +def test_rejects_coerced_booleans_integers_and_nonfinite_numbers( + tmp_path: Path, minimal_raw: dict, path: tuple, value: object +): + raw = deepcopy(minimal_raw) + target = raw + for part in path[:-1]: + target = target[part] + target[path[-1]] = value + with pytest.raises(ImportSpecError, match="must be|finite"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + ("path", "value"), + [ + (("sources", 0, "expected_sha256"), "abc"), + (("conditions", 0, "evidence_status"), "verified"), + (("conditions", 0, "scope_disposition"), "paper"), + (("conditions", 0, "result_origin", "derivation_origin"), "native"), + (("sources", 0, "classification"), "trusted"), + ], +) +def test_rejects_invalid_sha256_and_enums( + tmp_path: Path, minimal_raw: dict, path: tuple, value: str +): + raw = deepcopy(minimal_raw) + target = raw + for part in path[:-1]: + target = target[part] + target[path[-1]] = value + with pytest.raises(ImportSpecError): + _load(tmp_path, raw) + + +@pytest.mark.parametrize("credential_key", ["api_key", "nested_access_token", "password"]) +def test_rejects_recursive_credential_named_keys( + tmp_path: Path, minimal_raw: dict, credential_key: str +): + raw = deepcopy(minimal_raw) + raw["sources"][0]["notes"] = {"outer": [{"inner": {credential_key: "secret"}}]} + with pytest.raises(ImportSpecError, match="credential-named"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + "metric", + [ + "package.metrics:UnsafeMetric", + "package.metrics.UnsafeMetric", + "entry-point://unsafe", + "/tmp/metric.py:Metric", + "./metric.py", + ], +) +def test_import_metrics_are_closed_to_builtin_registry( + tmp_path: Path, minimal_raw: dict, metric: str +): + raw = deepcopy(minimal_raw) + raw["metrics"] = [metric] + with pytest.raises(ImportSpecError, match="built-in"): + _load(tmp_path, raw) + + assert set(BUILTIN_METRICS) + + +@pytest.mark.parametrize( + ("field", "value"), + [ + ("delimiter", "::"), + ("delimiter", "é"), + ("quote_character", ","), + ("escape_character", '"'), + ("encoding", "utf-8-sig"), + ("decoding_errors", "replace"), + ], +) +def test_csv_dialect_is_fixed_utf8_strict_and_uses_distinct_ascii_bytes( + tmp_path: Path, minimal_raw: dict, field: str, value: object +): + raw = deepcopy(minimal_raw) + raw["sources"][0]["dialect"][field] = value + with pytest.raises(ImportSpecError): + _load(tmp_path, raw) + + +def test_csv_dialect_accepts_distinct_ascii_escape(tmp_path: Path, minimal_raw: dict): + raw = deepcopy(minimal_raw) + raw["sources"][0]["dialect"].update( + {"delimiter": ";", "quote_character": "'", "escape_character": "\\"} + ) + spec = _load(tmp_path, raw) + assert spec.sources[0].dialect.escape_character == "\\" + + +def test_structured_json_options_and_reject_unmapped_policy( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + source = raw["sources"][0] + source["expected_columns"].append("options_json") + source["option_mapping"] = { + "mode": "structured_json", + "ordered_columns": [], + "structured_column": "options_json", + "structured_label_key": "label", + "structured_text_key": "text", + } + source["extra_field_policy"] = "reject_unmapped" + source["ignored_columns"].update( + { + "question": "reference snapshot owns the question text", + "choice_a": "superseded by structured choices", + "choice_b": "superseded by structured choices", + } + ) + spec = _load(tmp_path, raw) + assert spec.sources[0].option_mapping.structured_column == "options_json" + assert spec.sources[0].extra_field_policy == "reject_unmapped" + + +@pytest.mark.parametrize( + "conflict", + [ + "duplicate_mapping", + "mapped_option", + "mapped_ignored", + "option_ignored", + "undeclared_strict_column", + ], +) +def test_each_source_column_has_exactly_one_disposition( + tmp_path: Path, minimal_raw: dict, conflict: str +): + raw = deepcopy(minimal_raw) + source = raw["sources"][0] + if conflict == "duplicate_mapping": + source["columns"]["raw_prediction"] = "answer" + elif conflict == "mapped_option": + source["columns"]["prediction"] = "choice_a" + elif conflict == "mapped_ignored": + source["columns"]["prediction"] = "score" + elif conflict == "option_ignored": + source["ignored_columns"]["choice_a"] = "cannot also be an option" + else: + source["extra_field_policy"] = "reject_unmapped" + with pytest.raises(ImportSpecError, match="disposition|source column"): + _load(tmp_path, raw) + + +def test_reject_unmapped_accepts_exactly_partitioned_columns( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + source = raw["sources"][0] + source["extra_field_policy"] = "reject_unmapped" + source["ignored_columns"]["question"] = "question comes from reference snapshot" + spec = _load(tmp_path, raw) + assert spec.sources[0].extra_field_policy == "reject_unmapped" + + +@pytest.mark.parametrize( + "option_mapping", + [ + { + "mode": "ordered_columns", + "ordered_columns": [], + "structured_column": None, + "structured_label_key": None, + "structured_text_key": None, + }, + { + "mode": "ordered_columns", + "ordered_columns": ["choice_a"], + "structured_column": "options_json", + "structured_label_key": None, + "structured_text_key": None, + }, + { + "mode": "structured_json", + "ordered_columns": [], + "structured_column": "options_json", + "structured_label_key": None, + "structured_text_key": "text", + }, + { + "mode": "structured_json", + "ordered_columns": ["choice_a"], + "structured_column": "options_json", + "structured_label_key": "label", + "structured_text_key": "text", + }, + ], +) +def test_rejects_incomplete_or_contradictory_option_declarations( + tmp_path: Path, minimal_raw: dict, option_mapping: dict +): + raw = deepcopy(minimal_raw) + raw["sources"][0]["option_mapping"] = option_mapping + with pytest.raises(ImportSpecError, match="option_mapping"): + _load(tmp_path, raw) + + +def test_rejects_duplicate_numeric_column_rules(tmp_path: Path, minimal_raw: dict): + raw = deepcopy(minimal_raw) + raw["sources"][0]["numeric_columns"].append( + deepcopy(raw["sources"][0]["numeric_columns"][0]) + ) + with pytest.raises(ImportSpecError, match="Duplicate numeric"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + ("section", "id_field"), + [ + ("sources", "source_id"), + ("datasets", "dataset_id"), + ("models", "model_key"), + ("methods", "method_key"), + ("prompts", "prompt_key"), + ("conditions", "condition_key"), + ], +) +def test_rejects_duplicate_declared_ids( + tmp_path: Path, minimal_raw: dict, section: str, id_field: str +): + raw = deepcopy(minimal_raw) + raw[section].append(deepcopy(raw[section][0])) + assert raw[section][0][id_field] == raw[section][1][id_field] + with pytest.raises(ImportSpecError, match="Duplicate"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + ("path", "value"), + [ + (("datasets", 0, "source_ids"), ["missing"]), + (("datasets", 0, "selection_source_id"), "missing"), + (("conditions", 0, "dataset_id"), "missing"), + (("conditions", 0, "model_key"), "missing"), + (("conditions", 0, "method_key"), "missing"), + (("conditions", 0, "prompt_key"), "missing"), + ], +) +def test_rejects_missing_references( + tmp_path: Path, minimal_raw: dict, path: tuple, value: object +): + raw = deepcopy(minimal_raw) + target = raw + for part in path[:-1]: + target = target[part] + target[path[-1]] = value + with pytest.raises(ImportSpecError, match="reference|unknown"): + _load(tmp_path, raw) + + +def test_unknown_provenance_is_explicit_null_with_reason( + tmp_path: Path, minimal_raw: dict +): + spec = _load(tmp_path, minimal_raw) + assert spec.provenance["producer_request_id"] == { + "value": None, + "reason": "not recorded by producer", + } + + for bad in ( + None, + {"value": None}, + {"value": None, "reason": ""}, + {"value": None, "reason": "unknown", "guess": "request-1"}, + ): + raw = deepcopy(minimal_raw) + raw["provenance"]["producer_request_id"] = bad + with pytest.raises(ImportSpecError, match="provenance"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + "case", + [ + "dataset_missing", + "dataset_contradictory", + "dataset_unknown_key", + "model_missing", + "model_contradictory", + "model_unknown_key", + "method_missing", + "method_contradictory", + "prompt_missing", + "prompt_contradictory", + "condition_missing", + "condition_contradictory", + "condition_unknown_key", + ], +) +def test_nullable_scientific_fields_require_exact_unknown_reasons( + tmp_path: Path, minimal_raw: dict, case: str +): + raw = deepcopy(minimal_raw) + if case == "dataset_missing": + raw["datasets"][0]["selection_unknown_reasons"].pop("selection_seed") + elif case == "dataset_contradictory": + raw["datasets"][0]["selection_seed"] = 7 + elif case == "dataset_unknown_key": + raw["datasets"][0]["selection_unknown_reasons"]["revision"] = "unknown" + elif case == "model_missing": + raw["models"][0]["unknown_reasons"].pop("backend") + elif case == "model_contradictory": + raw["models"][0]["backend"] = "dummy" + elif case == "model_unknown_key": + raw["models"][0]["unknown_reasons"]["temperature"] = "unknown" + elif case == "method_missing": + raw["methods"][0]["unknown_reasons"].pop("implementation") + elif case == "method_contradictory": + raw["methods"][0]["implementation"] = { + "qualified_name": "historical:Method" + } + elif case == "prompt_missing": + raw["prompts"][0]["unknown_reason"] = None + elif case == "prompt_contradictory": + raw["prompts"][0].update( + { + "template_identity": "historical-template", + "template_digest": SHA_A, + "template_contents": {"prompt": "contents"}, + } + ) + elif case == "condition_missing": + raw["conditions"][0]["unknown_reasons"].pop("calibration_identity") + elif case == "condition_contradictory": + raw["conditions"][0]["seed"] = 7 + else: + raw["conditions"][0]["unknown_reasons"]["model"] = "unknown" + with pytest.raises(ImportSpecError, match="unknown.reason|unknown_reasons|unknown_reason"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + ("authorization_type", "executable"), + [("inference_repair", True), ("offline_transformation", False)], +) +def test_authorization_type_binds_executability( + tmp_path: Path, + minimal_raw: dict, + authorization_type: str, + executable: bool, +): + raw = deepcopy(minimal_raw) + raw["authorizations"] = [ + { + "authorization_id": "auth", + "authorization_type": authorization_type, + "source_id": "results", + "condition_question_reasons": {"condition": {"q1": "approved"}}, + "authority": "benchmark owner", + "purpose": "bounded operation", + "executable": executable, + "input_evidence_digests": {"results": SHA_A}, + "expected_snapshot_digests": {"q1": SHA_B}, + } + ] + spec = _load(tmp_path, raw) + assert spec.authorizations[0].executable is executable + + +@pytest.mark.parametrize( + ("authorization_type", "executable"), + [("inference_repair", False), ("offline_transformation", True)], +) +def test_rejects_authorization_executability_mismatch( + tmp_path: Path, + minimal_raw: dict, + authorization_type: str, + executable: bool, +): + raw = deepcopy(minimal_raw) + raw["authorizations"] = [ + { + "authorization_id": "auth", + "authorization_type": authorization_type, + "source_id": "results", + "condition_question_reasons": {"condition": {"q1": "approved"}}, + "authority": "benchmark owner", + "purpose": "bounded operation", + "executable": executable, + "input_evidence_digests": {"results": SHA_A}, + "expected_snapshot_digests": {"q1": SHA_B}, + } + ] + with pytest.raises(ImportSpecError, match="authorization_type|executable"): + _load(tmp_path, raw) + + +def test_included_evaluable_condition_requires_prediction_origin_coverage( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + origin = raw["conditions"][0]["result_origin"] + origin["default_prediction_origin"] = None + origin["per_question_prediction_origins"] = { + "q1": "external_historical_inference" + } + with pytest.raises(ImportSpecError, match="prediction origin"): + _load(tmp_path, raw) + + +def test_exact_per_question_origins_cover_evaluable_condition( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + origin = raw["conditions"][0]["result_origin"] + origin["default_prediction_origin"] = None + origin["per_question_prediction_origins"] = { + "q1": "external_historical_inference", + "q2": "external_repair_inference", + } + spec = _load(tmp_path, raw) + assert spec.conditions[0].result_origin.default_prediction_origin is None + + +@pytest.mark.parametrize( + ("status", "scope"), + [ + ("partial", "included"), + ("complete", "excluded_from_paper_matrix"), + ], +) +def test_non_evaluable_evidence_may_omit_prediction_origins( + tmp_path: Path, minimal_raw: dict, status: str, scope: str +): + raw = deepcopy(minimal_raw) + condition = raw["conditions"][0] + condition["evidence_status"] = status + condition["scope_disposition"] = scope + condition["result_origin"]["default_prediction_origin"] = None + condition["result_origin"]["per_question_prediction_origins"] = {} + spec = _load(tmp_path, raw) + assert spec.conditions[0].evidence_status == status + + +def test_rejects_paths_in_stable_provenance(tmp_path: Path, minimal_raw: dict): + raw = deepcopy(minimal_raw) + raw["provenance"]["producer_path"] = { + "value": "/home/producer/results.csv", + "reason": None, + } + with pytest.raises(ImportSpecError, match="path"): + _load(tmp_path, raw) + + +def _native_compatibility_payloads() -> dict[str, dict]: + artifact_payload = { + "spec": { + "benchmark": "toy", + "split": "test", + "hf_path": None, + "hf_subset": None, + "source_revision": None, + "normalization_version": "2", + "transforms": [], + "output_name": "toy", + }, + "content_digest": SHA_A, + "source": {}, + } + selection_payload = { + "artifact_id": short_id("ds", artifact_payload), + "content_digest": SHA_B, + "sample_identities": [SHA_A, SHA_B], + "seed": 7, + "n_samples": None, + "subject_filter": [], + } + dataset_identity = { + "artifact_payload": artifact_payload, + "artifact_digest": integrity_digest(artifact_payload), + "artifact_id": short_id("ds", artifact_payload), + "selection_payload": selection_payload, + "selection_digest": integrity_digest(selection_payload), + "selection_id": short_id("sel", selection_payload), + } + model_payload = {"backend": "dummy", "model_name_or_path": "dummy-model"} + method_payload = { + "name": "direct_mcq", + "effective_params": {}, + "preflight": None, + "implementation": { + "qualified_name": "choicebench.methods.direct_mcq:DirectMCQRunner", + }, + } + return { + "datasets": dataset_identity, + "models": _native_identity("model", model_payload), + "methods": _native_identity("method", method_payload), + "prompts": prompt_bundle_identity("v1"), + } + + +def test_validates_exact_native_compatibility_identities( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + identities = _native_compatibility_payloads() + for section, identity in identities.items(): + raw[section][0]["native_compatibility_identity"] = identity + spec = _load(tmp_path, raw) + assert spec.models[0].native_compatibility_identity["model_id"].startswith( + "model_" + ) + assert ( + spec.prompts[0].native_compatibility_identity + == prompt_bundle_identity("v1") + ) + + +def test_accepts_external_package_implementation_identity_in_native_method( + tmp_path: Path, minimal_raw: dict +): + implementation = implementation_identity(yaml.YAMLObject) + assert implementation["package_file_count"] > 0 + payload = deepcopy(_native_compatibility_payloads()["methods"]["payload"]) + payload["implementation"] = implementation + identity = _native_identity("method", payload) + raw = deepcopy(minimal_raw) + raw["methods"][0]["native_compatibility_identity"] = identity + + spec = _load(tmp_path, raw) + + assert spec.methods[0].native_compatibility_identity == identity + + +@pytest.mark.parametrize("package_file_count", [0, True]) +def test_rejects_non_positive_or_non_strict_package_file_count( + tmp_path: Path, minimal_raw: dict, package_file_count: object +): + implementation = implementation_identity(yaml.YAMLObject) + implementation["package_file_count"] = package_file_count + payload = deepcopy(_native_compatibility_payloads()["methods"]["payload"]) + payload["implementation"] = implementation + raw = deepcopy(minimal_raw) + raw["methods"][0]["native_compatibility_identity"] = _native_identity( + "method", payload + ) + + with pytest.raises( + ImportSpecError, + match=r"package_file_count.*(?:must be an integer|must be >= 1)", + ): + _load(tmp_path, raw) + + +@pytest.mark.parametrize("missing_key", ["package_tree_digest", "package_file_count"]) +def test_requires_package_tree_digest_and_file_count_together( + tmp_path: Path, minimal_raw: dict, missing_key: str +): + implementation = implementation_identity(yaml.YAMLObject) + implementation.pop(missing_key) + payload = deepcopy(_native_compatibility_payloads()["methods"]["payload"]) + payload["implementation"] = implementation + raw = deepcopy(minimal_raw) + raw["methods"][0]["native_compatibility_identity"] = _native_identity( + "method", payload + ) + + with pytest.raises(ImportSpecError, match="package_tree_digest.*package_file_count.*together"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + "field", ["source_file", "distribution", "distribution_version"] +) +def test_rejects_empty_optional_implementation_string( + tmp_path: Path, minimal_raw: dict, field: str +): + implementation = {"qualified_name": "external:Target", field: ""} + payload = deepcopy(_native_compatibility_payloads()["methods"]["payload"]) + payload["implementation"] = implementation + raw = deepcopy(minimal_raw) + raw["methods"][0]["native_compatibility_identity"] = _native_identity( + "method", payload + ) + + with pytest.raises(ImportSpecError, match=field): + _load(tmp_path, raw) + + +def test_native_prompt_identity_uses_exact_contents_without_redaction( + tmp_path: Path, minimal_raw: dict +): + prompt_root = tmp_path / "prompts" + version_dir = prompt_root / "credential-shaped-science" + version_dir.mkdir(parents=True) + contents = { + "direct_mcq": "Question: {question}\napi_key=scientific-label\nAnswer:", + "free_text": "Question: {question}\npassword=ordinary-text\nAnswer:", + "option_matching": "Question: {question}\n{options}\nAnswer:", + } + for name, content in contents.items(): + (version_dir / f"{name}.txt").write_text(content, encoding="utf-8") + identity = prompt_bundle_identity("credential-shaped-science", prompt_root) + raw = deepcopy(minimal_raw) + raw["prompts"][0]["native_compatibility_identity"] = identity + spec = _load(tmp_path, raw) + assert spec.prompts[0].native_compatibility_identity == identity + + +@pytest.mark.parametrize("mutation", ["missing_template", "extra_template", "short_id"]) +def test_rejects_non_native_prompt_bundle_shapes( + tmp_path: Path, minimal_raw: dict, mutation: str +): + identity = deepcopy(prompt_bundle_identity("v1")) + if mutation == "missing_template": + identity["files"].pop("free_text") + payload = {"version": identity["version"], "files": identity["files"]} + identity["prompt_id"] = f"prompt_{integrity_digest(payload)[:16]}" + elif mutation == "extra_template": + identity["files"]["unexpected"] = { + "content": "extra", + "sha256": integrity_digest("extra"), + } + payload = {"version": identity["version"], "files": identity["files"]} + identity["prompt_id"] = f"prompt_{integrity_digest(payload)[:16]}" + else: + content = identity["files"]["direct_mcq"]["content"] + content += "\napi_key=scientific-label" + identity["files"]["direct_mcq"] = { + "content": content, + "sha256": integrity_digest(content), + } + payload = {"version": identity["version"], "files": identity["files"]} + identity["prompt_id"] = short_id("prompt", payload) + raw = deepcopy(minimal_raw) + raw["prompts"][0]["native_compatibility_identity"] = identity + with pytest.raises(ImportSpecError, match="native_compatibility_identity"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize( + "payload", + [ + { + "backend": "api", + "provider": "openai", + "model_name_or_path": "gpt-test", + "base_url": None, + "generation_kwargs": { + "max_new_tokens": 32, + "temperature": 0.0, + "untrusted": True, + }, + }, + { + "backend": "huggingface", + "model": { + "kind": "huggingface-hub", + "repo_id": "org/model", + "requested_revision": None, + "resolved_commit": "a" * 40, + "untrusted": True, + }, + "device": "cuda", + "add_bos_token": True, + "generation_kwargs": { + "max_new_tokens": 32, + "temperature": 0.0, + "do_sample": False, + }, + "loader": {"trust_remote_code": True, "torch_dtype": "float16"}, + }, + { + "backend": "huggingface", + "model": { + "kind": "local", + "logical_name": "model", + "content_digest": SHA_A, + "file_count": 1, + "total_bytes": 10, + }, + "device": "cuda", + "add_bos_token": True, + "generation_kwargs": { + "max_new_tokens": 32, + "temperature": 0.0, + "do_sample": False, + }, + "loader": {"trust_remote_code": 1, "torch_dtype": "float16"}, + }, + ], +) +def test_rejects_unknown_or_coerced_nested_native_model_fields( + tmp_path: Path, minimal_raw: dict, payload: dict +): + raw = deepcopy(minimal_raw) + raw["models"][0]["native_compatibility_identity"] = _native_identity( + "model", payload + ) + with pytest.raises(ImportSpecError, match="native_compatibility_identity"): + _load(tmp_path, raw) + + +def test_rejects_unknown_native_method_preflight_field( + tmp_path: Path, minimal_raw: dict +): + raw = deepcopy(minimal_raw) + payload = deepcopy(_native_compatibility_payloads()["methods"]["payload"]) + payload["preflight"] = { + "source": "benchmark", + "split": "validation", + "n": 100, + "untrusted": True, + } + raw["methods"][0]["native_compatibility_identity"] = _native_identity( + "method", payload + ) + with pytest.raises(ImportSpecError, match="native_compatibility_identity"): + _load(tmp_path, raw) + + +@pytest.mark.parametrize("section", ["datasets", "models", "methods", "prompts"]) +@pytest.mark.parametrize("mutation", ["unknown", "partial", "mismatched_digest", "mismatched_id"]) +def test_rejects_unverified_native_compatibility_claims( + tmp_path: Path, minimal_raw: dict, section: str, mutation: str +): + raw = deepcopy(minimal_raw) + identity = deepcopy(_native_compatibility_payloads()[section]) + if mutation == "unknown": + identity["claimed_native_id"] = "native" + elif mutation == "partial": + identity.pop(next(iter(identity))) + elif mutation == "mismatched_digest": + if section == "prompts": + identity["files"]["direct_mcq"]["sha256"] = SHA_C + else: + digest_key = "artifact_digest" if section == "datasets" else "digest" + identity[digest_key] = SHA_C + else: + id_key = { + "datasets": "selection_id", + "models": "model_id", + "methods": "method_id", + "prompts": "prompt_id", + }[section] + identity[id_key] = "claimed_arbitrarily" + raw[section][0]["native_compatibility_identity"] = identity + with pytest.raises(ImportSpecError, match="native_compatibility_identity"): + _load(tmp_path, raw) + + +def test_audit_locations_do_not_change_import_spec_digest( + minimal_spec, minimal_overlay_spec +): + moved = replace( + minimal_spec, + audit={"source_path": "/different/host", "imported_at": "later"}, + ) + assert import_spec_digest(moved) == import_spec_digest(minimal_spec) + + moved_source = replace( + minimal_spec, + sources=( + replace( + minimal_spec.sources[0], path=Path("/other/freeze/source.csv") + ), + ), + ) + assert import_spec_digest(moved_source) == import_spec_digest(minimal_spec) + + moved_base = replace( + minimal_overlay_spec, + overlays=( + replace( + minimal_overlay_spec.overlays[0], + base_run_path=Path("/other/home/runs/base"), + ), + ), + ) + assert import_spec_digest(moved_base) == import_spec_digest( + minimal_overlay_spec + ) + + +def mutate_identity_field(spec, change: str): + source = spec.sources[0] + condition = spec.conditions[0] + if change == "mapping": + return replace(spec, sources=(replace(source, columns={**source.columns, "x": "y"}),)) + if change == "dialect": + return replace( + spec, + sources=(replace(source, dialect=replace(source.dialect, delimiter=";")),), + ) + if change == "numeric_policy": + numeric = source.numeric_columns[0] + return replace( + spec, + sources=( + replace( + source, + numeric_columns=(replace(numeric, null_allowed=False),), + ), + ), + ) + if change == "option_policy": + option = source.option_mapping + return replace( + spec, + sources=( + replace( + source, + option_mapping=replace( + option, ordered_columns=tuple(reversed(option.ordered_columns)) + ), + ), + ), + ) + if change == "extra_field_policy": + return replace( + spec, + sources=(replace(source, extra_field_policy="reject_unmapped"),), + ) + if change == "status": + return replace( + spec, + conditions=(replace(condition, evidence_status="qualified"),), + ) + if change == "scope": + return replace( + spec, + conditions=( + replace(condition, scope_disposition="excluded_from_paper_matrix"), + ), + ) + raise AssertionError(change) + + +@pytest.mark.parametrize( + "change", + [ + "mapping", + "dialect", + "numeric_policy", + "option_policy", + "extra_field_policy", + "status", + "scope", + ], +) +def test_stable_mapping_fields_change_import_spec_digest(minimal_spec, change): + changed = mutate_identity_field(minimal_spec, change) + assert import_spec_digest(changed) != import_spec_digest(minimal_spec) + + +def test_stable_projection_is_json_native_and_has_no_audit_locations( + minimal_overlay_spec, +): + projection = stable_import_projection(minimal_overlay_spec) + assert "audit" not in projection + assert "path" not in projection["sources"][0] + assert "base_run_path" not in projection["overlays"][0] diff --git a/tests/importing/test_stage1_import_spec.py b/tests/importing/test_stage1_import_spec.py new file mode 100644 index 0000000..ace7510 --- /dev/null +++ b/tests/importing/test_stage1_import_spec.py @@ -0,0 +1,329 @@ +"""Tests for full Stage 1 ImportSpec/AuthorizationSpec assembly (Task 18). + +Uses a small synthetic freeze fixture combining test_stage1_profile.py's +cell-manifest/queue fixtures with its ARC/MMLU dataset fixtures, plus real +per-cell canonical result CSVs -- never real historical data. Column names, +the 6-column common schema (question_id/correct_option/parsed_choice/ +choice_a..d), and the evidence-status reconciliation behavior below were all +verified directly against the real, read-only Stage 1 freeze at +/home/cotenthusiast/Projects/model-generalization/paper_data_freeze during +development (never committed): of 102 real cells, every one reconciles to +either "malformed" (any MISSING_PREDICTION/INVALID_PREDICTION/ +CORRECT_OPTION_MISMATCH finding) or "complete"/"qualified" -- the freeze's +own STATUS_MAP-mapped status disagrees with the independently-computed +status for a majority of cells (e.g. 20 of 28 "malformed_requires_inference" +cells have zero row-level defects and reconcile to "complete"; 28 of 57 +"canonical_complete" cells have some MISSING_PREDICTION rows the freeze +tolerates but ChoiceBench's generic validator does not, and reconcile to +"malformed"), confirming the reconciliation-by-probe design is load-bearing, +not a formality. +""" + +from __future__ import annotations + +import csv +from hashlib import sha256 +import json +from pathlib import Path + +import pytest + +from choicebench.importing.authorization import validate_authorization_bundle +from choicebench.importing.evidence import open_verified_source +from choicebench.importing.profiles.stage1_paper_freeze import ( + Stage1ProfileError, + build_stage1_cell_authorization, + build_stage1_expected_datasets, + build_stage1_import_spec, + translate_stage1_paper_freeze, +) +from tests.importing.test_stage1_profile import ( + _build_dataset_freeze, + _cell, + _default_arc_normalized_rows, + _default_mmlu_normalized_rows, + _write, +) + +_STAGE1_RESULT_COLUMNS = [ + "question_id", "correct_option", "choice_a", "choice_b", "choice_c", "choice_d", + "parsed_choice", +] + + +def _result_row(question_id, correct_option, choices, parsed_choice): + row = {"question_id": question_id, "correct_option": correct_option, "parsed_choice": parsed_choice} + for letter, text in zip("abcd", choices): + row[f"choice_{letter}"] = text + return row + + +def _write_result_csv(path: Path, rows: list[dict]) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + with path.open("w", newline="", encoding="utf-8") as handle: + writer = csv.DictWriter(handle, fieldnames=_STAGE1_RESULT_COLUMNS) + writer.writeheader() + writer.writerows(rows) + + +def _arc_result_rows(*, all_correct: bool = True): + normalized = _default_arc_normalized_rows() + rows = [] + for row in normalized: + choices = [row["choice_a"], row["choice_b"], row["choice_c"], row["choice_d"]] + parsed = row["correct_option"] if all_correct else "A" + rows.append(_result_row(row["question_id"], row["correct_option"], choices, parsed)) + return rows + + +def _build_full_freeze( + tmp_path: Path, + *, + cells, + cell_result_rows: dict[str, list[dict]], + approved_rows=(), + held_rows=(), + excluded_rows=(), + rerun_candidate_rows=(), +): + freeze_root = _build_dataset_freeze(tmp_path) + + manifest_json_path = freeze_root / "manifests" / "canonical_results_manifest.json" + _write(manifest_json_path, json.dumps(cells)) + _write(freeze_root / "manifests" / "canonical_results_manifest.csv", "cell_id\n") + _write(freeze_root / "manifests" / "cell_status_matrix.csv", "cell_id\n") + _write(freeze_root / "manifests" / "expected_matrix.csv", "cell_id\n") + + def _write_queue_csv(path, rows): + path.parent.mkdir(parents=True, exist_ok=True) + if not rows: + path.write_text("cell_id\n", encoding="utf-8") + return + with path.open("w", newline="", encoding="utf-8") as handle: + writer = csv.DictWriter(handle, fieldnames=list(rows[0])) + writer.writeheader() + writer.writerows(rows) + + _write_queue_csv(freeze_root / "manifests" / "approved_rerun_queue.csv", list(approved_rows)) + _write_queue_csv(freeze_root / "manifests" / "held_or_declined_reruns.csv", list(held_rows)) + _write_queue_csv(freeze_root / "manifests" / "paper_scope_excluded_reruns.csv", list(excluded_rows)) + _write_queue_csv(freeze_root / "manifests" / "rerun_queue.csv", list(rerun_candidate_rows)) + _write(freeze_root / "reports" / "canonical_freeze_report.md", "# report\n") + + for cell in cells: + rows = cell_result_rows[cell["cell_id"]] + _write_result_csv(freeze_root / cell["canonical_path"], rows) + digest = sha256((freeze_root / cell["canonical_path"]).read_bytes()).hexdigest() + cell["canonical_sha256"] = digest + _write(manifest_json_path, json.dumps(cells)) + + relative_paths = [ + "manifests/canonical_results_manifest.json", + "manifests/canonical_results_manifest.csv", + "manifests/cell_status_matrix.csv", + "manifests/expected_matrix.csv", + "manifests/approved_rerun_queue.csv", + "manifests/held_or_declined_reruns.csv", + "manifests/paper_scope_excluded_reruns.csv", + "manifests/rerun_queue.csv", + "reports/canonical_freeze_report.md", + *[cell["canonical_path"] for cell in cells], + ] + ledger_path = freeze_root / "checksums" / "checksums.sha256" + existing_lines = ledger_path.read_text(encoding="utf-8").splitlines() + all_lines = list(existing_lines) + for relative_path in relative_paths: + data = (freeze_root / relative_path).read_bytes() + digest = sha256(data).hexdigest() + all_lines.append(f"{digest} {relative_path}") + _write(ledger_path, "\n".join(all_lines) + "\n") + return freeze_root + + +def _default_cells_and_rows(): + arc_ids = [row["question_id"] for row in _default_arc_normalized_rows()] + clean_cell_id = "cbp__model-a__arc_challenge__baseline" + recoverable_cell_id = "cbp__model-a__arc_challenge__semantic_matching_v1" + malformed_cell_id = "cbp__model-a__arc_challenge__two_stage_v1" + + cells = [ + _cell(clean_cell_id, "baseline", "canonical_complete"), + _cell( + recoverable_cell_id, "semantic_matching_v1", "recoverable_from_existing_artifacts", + recoverable=(arc_ids[0],), + ), + _cell(malformed_cell_id, "two_stage_v1", "malformed_requires_inference"), + ] + for cell in cells: + cell["benchmark"] = "arc_challenge" + cell["expected_question_count"] = len(arc_ids) + + rows = { + clean_cell_id: _arc_result_rows(all_correct=True), + # recoverable: rows are all present/parseable (just semantically + # "wrong" per the freeze's own domain knowledge) -- reconciles to + # "complete", matching the real freeze's 6 recoverable cells. + recoverable_cell_id: _arc_result_rows(all_correct=False), + malformed_cell_id: [ + *_arc_result_rows(all_correct=True)[:-1], + _result_row(arc_ids[-1], "B", ["a4", "b4", "c4", "d4"], ""), # missing prediction + ], + } + return cells, rows, arc_ids + + +def test_build_stage1_import_spec_assembles_sources_conditions_and_children(tmp_path): + cells, rows, arc_ids = _default_cells_and_rows() + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + translation = translate_stage1_paper_freeze(freeze_root) + expected_datasets = build_stage1_expected_datasets(freeze_root) + + spec = build_stage1_import_spec(freeze_root, translation, expected_datasets) + assert len(spec.conditions) == 3 + assert len(spec.sources) == 3 + 2 # 3 cells + arc/mmlu normalized reference sources + assert {model.model_key for model in spec.models} == {"model-a"} + assert {method.method_key for method in spec.methods} == { + "baseline", "semantic_matching_v1", "two_stage_v1", + } + assert len(spec.prompts) == 1 + + +def test_build_stage1_import_spec_reconciles_recoverable_cell_to_complete(tmp_path): + """The freeze declares this cell 'recoverable_from_existing_artifacts', but + every row is present with a valid (if semantically wrong) parsed_choice -- + ChoiceBench's generic validator independently computes 'complete', and + build_stage1_import_spec must declare that, not the freeze's own label, + or the real import would fail its own declaration-mismatch check.""" + cells, rows, arc_ids = _default_cells_and_rows() + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + translation = translate_stage1_paper_freeze(freeze_root) + expected_datasets = build_stage1_expected_datasets(freeze_root) + + spec = build_stage1_import_spec(freeze_root, translation, expected_datasets) + by_key = {c.condition_key: c for c in spec.conditions} + recoverable = by_key["cbp__model-a__arc_challenge__semantic_matching_v1"] + assert recoverable.evidence_status == "complete" + assert recoverable.recoverable_question_ids == (arc_ids[0],) + + +def test_build_stage1_import_spec_reconciles_malformed_cell_with_missing_prediction(tmp_path): + cells, rows, arc_ids = _default_cells_and_rows() + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + translation = translate_stage1_paper_freeze(freeze_root) + expected_datasets = build_stage1_expected_datasets(freeze_root) + + spec = build_stage1_import_spec(freeze_root, translation, expected_datasets) + by_key = {c.condition_key: c for c in spec.conditions} + malformed = by_key["cbp__model-a__arc_challenge__two_stage_v1"] + assert malformed.evidence_status == "malformed" + + +def test_build_stage1_import_spec_keeps_a_clean_complete_cell_complete(tmp_path): + cells, rows, arc_ids = _default_cells_and_rows() + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + translation = translate_stage1_paper_freeze(freeze_root) + expected_datasets = build_stage1_expected_datasets(freeze_root) + + spec = build_stage1_import_spec(freeze_root, translation, expected_datasets) + by_key = {c.condition_key: c for c in spec.conditions} + clean = by_key["cbp__model-a__arc_challenge__baseline"] + assert clean.evidence_status == "complete" + assert clean.scope_disposition == "included" + + +def test_build_stage1_import_spec_marks_excluded_from_paper_matrix(tmp_path): + cells, rows, arc_ids = _default_cells_and_rows() + excluded_cell_id = "cbp__model-a__arc_challenge__pride" + excluded_cell = _cell(excluded_cell_id, "pride", "excluded_from_paper_matrix") + excluded_cell["benchmark"] = "arc_challenge" + excluded_cell["expected_question_count"] = len(arc_ids) + cells = [*cells, excluded_cell] + rows[excluded_cell_id] = _arc_result_rows(all_correct=True) + + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + translation = translate_stage1_paper_freeze(freeze_root) + expected_datasets = build_stage1_expected_datasets(freeze_root) + + spec = build_stage1_import_spec(freeze_root, translation, expected_datasets) + by_key = {c.condition_key: c for c in spec.conditions} + excluded = by_key[excluded_cell_id] + assert excluded.scope_disposition == "excluded_from_paper_matrix" + assert excluded.evidence_status == "complete" + + +def test_build_stage1_import_spec_rejects_canonical_sha256_ledger_disagreement(tmp_path): + cells, rows, arc_ids = _default_cells_and_rows() + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + expected_datasets = build_stage1_expected_datasets(freeze_root) + + # Tamper the manifest's own recorded canonical_sha256 for one cell (a + # forged claim about that cell's own result file), while re-checksumming + # the manifest.json file itself so translate_stage1_paper_freeze's own + # whole-file checksum check still passes -- isolating the per-cell + # cross-check build_stage1_import_spec performs against the untouched + # ledger entry for that cell's actual canonical CSV. + manifest_path = freeze_root / "manifests" / "canonical_results_manifest.json" + manifest = json.loads(manifest_path.read_text()) + manifest[0]["canonical_sha256"] = "f" * 64 + manifest_path.write_text(json.dumps(manifest), encoding="utf-8") + ledger_path = freeze_root / "checksums" / "checksums.sha256" + lines = ledger_path.read_text(encoding="utf-8").splitlines() + new_manifest_digest = sha256(manifest_path.read_bytes()).hexdigest() + updated_lines = [ + f"{new_manifest_digest} manifests/canonical_results_manifest.json" + if line.endswith("manifests/canonical_results_manifest.json") else line + for line in lines + ] + ledger_path.write_text("\n".join(updated_lines) + "\n", encoding="utf-8") + + translation = translate_stage1_paper_freeze(freeze_root) + with pytest.raises(Stage1ProfileError, match="disagrees with the checksum ledger"): + build_stage1_import_spec(freeze_root, translation, expected_datasets) + + +def test_build_stage1_cell_authorization_produces_a_bundle_the_generic_validator_accepts(tmp_path): + cells, rows, arc_ids = _default_cells_and_rows() + freeze_root = _build_full_freeze(tmp_path, cells=cells, cell_result_rows=rows) + translation = translate_stage1_paper_freeze(freeze_root) + expected_datasets = build_stage1_expected_datasets(freeze_root) + from choicebench.importing.profiles.stage1_paper_freeze import _load_checksum_ledger + + ledger = _load_checksum_ledger(freeze_root) + spec = build_stage1_import_spec(freeze_root, translation, expected_datasets) + by_key = {c.condition_key: c for c in spec.conditions} + cell_id = "cbp__model-a__arc_challenge__semantic_matching_v1" + condition = by_key[cell_id] + dataset = expected_datasets["arc_challenge"] + + from choicebench.importing.identity import build_import_semantic_identity + + models_by_key = {m.model_key: m for m in spec.models} + methods_by_key = {m.method_key: m for m in spec.methods} + prompts_by_key = {p.prompt_key: p for p in spec.prompts} + semantic = build_import_semantic_identity( + condition=condition, dataset=dataset, + model=models_by_key[condition.model_key], method=methods_by_key[condition.method_key], + prompt=prompts_by_key[condition.prompt_key], + ) + condition_digest = semantic.condition["condition_digest"] + + sources_by_id = {s.source_id: s for s in spec.sources} + cell_source = sources_by_id[cell_id] + + authorization, auth_source = build_stage1_cell_authorization( + freeze_root, ledger, + cell_id=cell_id, condition_digest=condition_digest, + authorization_type="offline_transformation", + question_reasons={arc_ids[0]: "option-hidden semantic rematch"}, + dataset_snapshot_digest=dataset.snapshot_digest, + cell_evidence_sha256=cell_source.expected_sha256, + ) + opened = open_verified_source(auth_source, containment_root=freeze_root) + bundle = validate_authorization_bundle( + authorization, opened_source=opened, + condition_digests={cell_id: condition_digest}, + expected={condition_digest: dataset}, + ) + assert bundle.authorization_type == "offline_transformation" + assert bundle.grants[condition_digest] == {arc_ids[0]: "option-hidden semantic rematch"} + assert bundle.executable is False diff --git a/tests/importing/test_stage1_profile.py b/tests/importing/test_stage1_profile.py new file mode 100644 index 0000000..38f3421 --- /dev/null +++ b/tests/importing/test_stage1_profile.py @@ -0,0 +1,538 @@ +"""Tests for Stage 1 profile translation (Task 15, reduced scope). + +Uses a small synthetic freeze fixture (never real historical data) whose +structure was verified against the real, read-only Stage 1 freeze at +/home/cotenthusiast/Projects/model-generalization/paper_data_freeze during +development: file names, column headers, and the exact real queue/matrix +counts (138 approved / 687 held / 3 excluded / 828 classified / 776 +rerun-queue-candidate pairs; 100/84/4/12/2 matrix breakdown; 6-cell/18-pair +offline-recoverable authority) were all cross-checked against that freeze +directly, not merely assumed from the plan. +""" + +from __future__ import annotations + +import csv +from hashlib import sha256 +import io +import json +from pathlib import Path + +import pytest + +from choicebench.importing.profiles.stage1_paper_freeze import ( + STATUS_MAP, + Stage1ProfileError, + build_stage1_expected_datasets, + translate_stage1_paper_freeze, +) + + +def _write(path: Path, content: str) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(content, encoding="utf-8") + + +def _cell(cell_id, method, status, *, damaged=(), recoverable=()): + return { + "cell_id": cell_id, "model": "model-a", "provider_backend": "provider", + "benchmark": "arc_challenge", "benchmark_split": "robustness", "method": method, + "expected_question_count": 10, "evidence_for_expected_existence": "x", + "evidence_for_expected_row_count": "y", "discovered_candidate_count": 1, + "final_status": status, "status": status, "canonical_artifact_id": f"artifact_{cell_id}", + "canonical_raw_path": f"raw/{cell_id}.csv", "canonical_path": f"canonical/{cell_id}.csv", + "canonical_sha256": "0" * 64, "observed_unique_count": 10, "duplicate_question_count": 0, + "exact_missing_question_ids": [], "exact_unexpected_question_ids": [], + "damaged_question_ids": list(damaged), "recoverable_question_ids": list(recoverable), + "qualification": "", + } + + +def _queue_row(cell_id, question_ids, reasons, **overrides): + row = { + "cell_id": cell_id, "model": "model-a", "provider_backend": "provider", + "model_revision": "not recorded", "benchmark": "arc_challenge", + "benchmark_snapshot_revision": "0" * 64, "method": "baseline", + "exact_question_ids": json.dumps(question_ids), + "number_of_questions": len(question_ids), "seed": 42, + "work_type": "complete_partial_or_replace_damaged_rows", + "configuration_identity": "0" * 64, + "question_reasons": json.dumps(reasons), + "source_configuration": "[]", "prompt_template_identity": "[]", + "generation_parameters": "{}", "expected_output_schema": "[]", + "supporting_artifact_paths": "[]", + } + row.update(overrides) + return row + + +def _write_csv(path: Path, rows: list[dict]) -> None: + import csv + + path.parent.mkdir(parents=True, exist_ok=True) + if not rows: + path.write_text("cell_id\n", encoding="utf-8") + return + with path.open("w", newline="", encoding="utf-8") as handle: + writer = csv.DictWriter(handle, fieldnames=list(rows[0])) + writer.writeheader() + writer.writerows(rows) + + +def _build_freeze( + tmp_path: Path, + *, + cells, + approved_rows=(), + held_rows=(), + excluded_rows=(), + rerun_candidate_rows=(), + tamper_checksum_for=None, +): + freeze_root = tmp_path / "freeze" + manifest_json_path = freeze_root / "manifests" / "canonical_results_manifest.json" + _write(manifest_json_path, json.dumps(cells)) + _write(freeze_root / "manifests" / "canonical_results_manifest.csv", "cell_id\n") + _write(freeze_root / "manifests" / "cell_status_matrix.csv", "cell_id\n") + _write(freeze_root / "manifests" / "expected_matrix.csv", "cell_id\n") + _write_csv(freeze_root / "manifests" / "approved_rerun_queue.csv", list(approved_rows)) + _write_csv(freeze_root / "manifests" / "held_or_declined_reruns.csv", list(held_rows)) + _write_csv(freeze_root / "manifests" / "paper_scope_excluded_reruns.csv", list(excluded_rows)) + _write_csv(freeze_root / "manifests" / "rerun_queue.csv", list(rerun_candidate_rows)) + _write(freeze_root / "reports" / "canonical_freeze_report.md", "# report\n") + + relative_paths = ( + "manifests/canonical_results_manifest.json", + "manifests/canonical_results_manifest.csv", + "manifests/cell_status_matrix.csv", + "manifests/expected_matrix.csv", + "manifests/approved_rerun_queue.csv", + "manifests/held_or_declined_reruns.csv", + "manifests/paper_scope_excluded_reruns.csv", + "manifests/rerun_queue.csv", + "reports/canonical_freeze_report.md", + ) + ledger_lines = [] + for relative_path in relative_paths: + data = (freeze_root / relative_path).read_bytes() + digest = sha256(data).hexdigest() + if tamper_checksum_for == relative_path: + digest = "f" * 64 + ledger_lines.append(f"{digest} {relative_path}") + _write(freeze_root / "checksums" / "checksums.sha256", "\n".join(ledger_lines) + "\n") + return freeze_root + + +def _default_cells(): + return [ + _cell("cbp__model-a__arc_challenge__baseline", "baseline", "canonical_complete"), + _cell("cbp__model-a__arc_challenge__two_stage_v1", "two_stage_v1", "canonical_qualified"), + _cell( + "cbp__model-a__arc_challenge__semantic_matching_v1", "semantic_matching_v1", + "recoverable_from_existing_artifacts", recoverable=("q1", "q2"), + ), + _cell( + "cbp__model-b__arc_challenge__semantic_matching_v1", "semantic_matching_v1", + "recoverable_from_existing_artifacts", recoverable=("q1", "q2"), + ), + _cell( + "cbp__model-a__arc_challenge__independent_hypothesis", "independent_hypothesis", + "malformed_requires_inference", damaged=("q3",), + ), + _cell("cbp__model-a__arc_challenge__pride", "pride", "canonical_complete"), + _cell("cbp__model-a__mmlu__pride", "pride", "excluded_from_paper_matrix"), + ] + + +def test_translate_stage1_paper_freeze_reproduces_matrix_and_offline_authority(tmp_path): + freeze_root = _build_freeze(tmp_path, cells=_default_cells()) + result = translate_stage1_paper_freeze(freeze_root) + assert result.matrix.intended_cells == 6 # 7 cells minus the one excluded_from_paper_matrix + # baseline + two_stage_v1 + both semantic_matching_v1 cells (only + # independent_hypothesis/pride are excluded from the "core method" bucket) + assert result.matrix.core_method_cells == 4 + assert result.matrix.ihs_cells == 1 + assert result.matrix.local_pride_cells == 1 # 2 pride cells minus 1 excluded + assert result.matrix.excluded_preserved_cells == 1 + assert result.offline_authority.condition_count == 2 + assert result.offline_authority.question_cell_count == 4 + assert result.offline_authority.recoverable_question_ids == ("q1", "q2") + assert result.offline_authority.inference_executable is False + + +def test_translate_stage1_paper_freeze_verifies_checksums(tmp_path): + freeze_root = _build_freeze( + tmp_path, cells=_default_cells(), + tamper_checksum_for="manifests/canonical_results_manifest.json", + ) + with pytest.raises(Stage1ProfileError, match="checksum mismatch"): + translate_stage1_paper_freeze(freeze_root) + + +def test_translate_stage1_paper_freeze_rejects_duplicate_cell_id(tmp_path): + cells = _default_cells() + cells.append(dict(cells[0])) + freeze_root = _build_freeze(tmp_path, cells=cells) + with pytest.raises(Stage1ProfileError, match="duplicate cell_id"): + translate_stage1_paper_freeze(freeze_root) + + +def test_translate_stage1_paper_freeze_requires_matching_recoverable_ids_across_cells(tmp_path): + cells = _default_cells() + # give the second recoverable cell a DIFFERENT question set than the first + for cell in cells: + if cell["cell_id"] == "cbp__model-b__arc_challenge__semantic_matching_v1": + cell["recoverable_question_ids"] = ["q1", "q9"] + freeze_root = _build_freeze(tmp_path, cells=cells) + with pytest.raises(Stage1ProfileError, match="do not all authorize the exact same question"): + translate_stage1_paper_freeze(freeze_root) + + +def test_translate_stage1_paper_freeze_computes_queue_pairs_and_disjointness(tmp_path): + approved = [ + _queue_row( + "cbp__model-a__arc_challenge__baseline", ["q1", "q2"], + {"q1": "reason1", "q2": "reason2"}, + queue_disposition="approved", execution_authority="authoritative", executable="true", + ) + ] + held = [ + _queue_row( + "cbp__model-a__arc_challenge__two_stage_v1", ["q3"], {"q3": "reason3"}, + queue_disposition="held", execution_authority="none", executable="false", + ) + ] + excluded = [ + _queue_row( + "cbp__model-a__arc_challenge__pride", ["q4"], {"q4": "reason4"}, + queue_disposition="excluded", execution_authority="none", executable="false", + ) + ] + rerun_candidates = [ + _queue_row("cbp__model-a__arc_challenge__baseline", ["q1"], {}), + ] + freeze_root = _build_freeze( + tmp_path, cells=_default_cells(), approved_rows=approved, held_rows=held, + excluded_rows=excluded, rerun_candidate_rows=rerun_candidates, + ) + result = translate_stage1_paper_freeze(freeze_root) + assert result.queue_counts.approved_executable_question_cells == 2 + assert result.queue_counts.held_nonexecutable_question_cells == 1 + assert result.queue_counts.excluded_nonexecutable_question_cells == 1 + assert result.queue_counts.classified_question_cells == 4 + assert result.queue_counts.rerun_candidate_question_cells == 1 + assert result.queue_counts.rerun_candidates_are_classified_subset is True + assert result.queue_pairs.approved[("cbp__model-a__arc_challenge__baseline", "q1")] == "reason1" + + +def test_translate_stage1_paper_freeze_rejects_overlapping_approved_and_held(tmp_path): + approved = [ + _queue_row( + "cbp__model-a__arc_challenge__baseline", ["q1"], {"q1": "r"}, + queue_disposition="approved", execution_authority="authoritative", executable="true", + ) + ] + held = [ + _queue_row( + "cbp__model-a__arc_challenge__baseline", ["q1"], {"q1": "r"}, + queue_disposition="held", execution_authority="none", executable="false", + ) + ] + freeze_root = _build_freeze(tmp_path, cells=_default_cells(), approved_rows=approved, held_rows=held) + with pytest.raises(Stage1ProfileError, match="not pairwise disjoint"): + translate_stage1_paper_freeze(freeze_root) + + +def test_translate_stage1_paper_freeze_rejects_unclassified_rerun_candidate(tmp_path): + rerun_candidates = [ + _queue_row("cbp__model-a__arc_challenge__baseline", ["q_never_classified"], {}), + ] + freeze_root = _build_freeze(tmp_path, cells=_default_cells(), rerun_candidate_rows=rerun_candidates) + with pytest.raises(Stage1ProfileError, match="never classified"): + translate_stage1_paper_freeze(freeze_root) + + +def test_translate_stage1_paper_freeze_rejects_approved_row_claiming_wrong_disposition(tmp_path): + approved = [ + _queue_row( + "cbp__model-a__arc_challenge__baseline", ["q1"], {"q1": "r"}, + queue_disposition="held", # wrong -- approved file must say "approved" + execution_authority="authoritative", executable="true", + ) + ] + freeze_root = _build_freeze(tmp_path, cells=_default_cells(), approved_rows=approved) + with pytest.raises(Stage1ProfileError, match="queue_disposition"): + translate_stage1_paper_freeze(freeze_root) + + +def test_translate_stage1_paper_freeze_rejects_held_row_claiming_executable(tmp_path): + held = [ + _queue_row( + "cbp__model-a__arc_challenge__baseline", ["q1"], {"q1": "r"}, + queue_disposition="held", execution_authority="none", executable="true", # wrong + ) + ] + freeze_root = _build_freeze(tmp_path, cells=_default_cells(), held_rows=held) + with pytest.raises(Stage1ProfileError, match="executable"): + translate_stage1_paper_freeze(freeze_root) + + +def test_status_map_matches_canonical_values(): + assert STATUS_MAP["canonical_complete"] == "complete" + assert STATUS_MAP["canonical_qualified"] == "qualified" + assert STATUS_MAP["recoverable_from_existing_artifacts"] == "recoverable" + assert STATUS_MAP["incomplete_requires_inference"] == "partial" + assert STATUS_MAP["malformed_requires_inference"] == "malformed" + assert "excluded_from_paper_matrix" not in STATUS_MAP # not an evidence-status input + + +# --- ARC/MMLU expected-dataset trust chain (Task 16) ------------------------ +# +# Synthetic fixture shaped exactly like the real freeze's raw/normalized/ +# split file formats (verified against the real freeze during development): +# ARC raw choices are {"text": [...], "label": [...]}; MMLU raw choices are a +# Python-repr list with an integer answer index; both normalized files share +# one schema (question_id/subject/question_text/choice_a..d/correct_option/ +# correct_answer_text). Includes ARC's real edge cases: a 3-option row +# (variable option count) and a 5-option row (the archived normalizer +# silently truncates to 4, verified in the real freeze to never drop the +# correct option) -- plus an MMLU field-identical duplicate question_id. + +_DATASET_RELATIVE_PATHS = ( + "raw/local_model_generalization/data/raw/arc_challenge_raw.csv", + "raw/local_model_generalization/data/processed/arc_challenge_normalized.csv", + "raw/local_model_generalization/data/splits/arc_challenge/robustness_ids.json", + "raw/local_model_generalization/data/splits/arc_challenge/robustness_metadata.json", + "raw/local_model_generalization/data/raw/mmlu_raw.csv", + "raw/local_model_generalization/data/processed/mmlu_normalized.csv", + "raw/local_model_generalization/data/splits/benchmark/robustness_ids.json", + "raw/local_model_generalization/data/splits/benchmark/robustness_metadata.json", +) + + +def _arc_raw_csv(rows): + buffer = io.StringIO() + writer = csv.DictWriter(buffer, fieldnames=["id", "question", "choices", "answerKey"]) + writer.writeheader() + for row_id, question, texts, labels, answer_key in rows: + writer.writerow({ + "id": row_id, "question": question, + "choices": json.dumps({"text": texts, "label": labels}), + "answerKey": answer_key, + }) + return buffer.getvalue() + + +def _normalized_csv(rows): + fieldnames = [ + "question_id", "subject", "question_text", "choice_a", "choice_b", + "choice_c", "choice_d", "correct_option", "correct_answer_text", + ] + buffer = io.StringIO() + writer = csv.DictWriter(buffer, fieldnames=fieldnames) + writer.writeheader() + for row in rows: + writer.writerow(row) + return buffer.getvalue() + + +def _default_arc_raw_rows(): + return [ + ("r1", "Q1?", ["a1", "b1", "c1", "d1"], ["A", "B", "C", "D"], "A"), + ("r2", "Q2?", ["a2", "b2", "c2", "d2"], ["A", "B", "C", "D"], "B"), + ("r3", "Q3?", ["a3", "b3", "c3"], ["A", "B", "C"], "A"), # 3-option + ("r4", "Q4?", ["a4", "b4", "c4", "d4", "e4"], ["A", "B", "C", "D", "E"], "B"), # 5-option + ] + + +def _default_arc_normalized_rows(): + return [ + {"question_id": "arc_q1", "subject": "arc_challenge", "question_text": "Q1?", + "choice_a": "a1", "choice_b": "b1", "choice_c": "c1", "choice_d": "d1", + "correct_option": "A", "correct_answer_text": "a1"}, + {"question_id": "arc_q2", "subject": "arc_challenge", "question_text": "Q2?", + "choice_a": "a2", "choice_b": "b2", "choice_c": "c2", "choice_d": "d2", + "correct_option": "B", "correct_answer_text": "b2"}, + {"question_id": "arc_q3", "subject": "arc_challenge", "question_text": "Q3?", + "choice_a": "a3", "choice_b": "b3", "choice_c": "c3", "choice_d": "", + "correct_option": "A", "correct_answer_text": "a3"}, + {"question_id": "arc_q4", "subject": "arc_challenge", "question_text": "Q4?", + "choice_a": "a4", "choice_b": "b4", "choice_c": "c4", "choice_d": "d4", + "correct_option": "B", "correct_answer_text": "b4"}, + ] + + +def _default_mmlu_raw_rows(): + # (question, subject, choices_repr, answer_index) + return [ + ("MQ1?", "history", "['x1', 'y1', 'z1', 'w1']", 0), + ("MQ2?", "history", "['x2', 'y2', 'z2', 'w2']", 2), + ("MQ2?", "history", "['x2', 'y2', 'z2', 'w2']", 2), # duplicate row -> same question_id + ] + + +def _default_mmlu_normalized_rows(): + return [ + {"question_id": "mmlu_q1", "subject": "history", "question_text": "MQ1?", + "choice_a": "x1", "choice_b": "y1", "choice_c": "z1", "choice_d": "w1", + "correct_option": "A", "correct_answer_text": "x1"}, + {"question_id": "mmlu_q2", "subject": "history", "question_text": "MQ2?", + "choice_a": "x2", "choice_b": "y2", "choice_c": "z2", "choice_d": "w2", + "correct_option": "C", "correct_answer_text": "z2"}, + {"question_id": "mmlu_q2", "subject": "history", "question_text": "MQ2?", + "choice_a": "x2", "choice_b": "y2", "choice_c": "z2", "choice_d": "w2", + "correct_option": "C", "correct_answer_text": "z2"}, + ] + + +def _build_dataset_freeze( + tmp_path, + *, + arc_raw_rows=None, + arc_normalized_rows=None, + mmlu_raw_rows=None, + mmlu_normalized_rows=None, + arc_selected_ids=None, + mmlu_selected_ids=None, + tamper_checksum_for=None, +): + freeze_root = tmp_path / "freeze" + arc_raw_rows = _default_arc_raw_rows() if arc_raw_rows is None else arc_raw_rows + arc_normalized_rows = ( + _default_arc_normalized_rows() if arc_normalized_rows is None else arc_normalized_rows + ) + mmlu_raw_rows = _default_mmlu_raw_rows() if mmlu_raw_rows is None else mmlu_raw_rows + mmlu_normalized_rows = ( + _default_mmlu_normalized_rows() if mmlu_normalized_rows is None else mmlu_normalized_rows + ) + arc_selected_ids = ( + [row["question_id"] for row in arc_normalized_rows] + if arc_selected_ids is None else arc_selected_ids + ) + mmlu_selected_ids = ( + sorted({row["question_id"] for row in mmlu_normalized_rows}) + if mmlu_selected_ids is None else mmlu_selected_ids + ) + + _write( + freeze_root / "raw/local_model_generalization/data/raw/arc_challenge_raw.csv", + _arc_raw_csv(arc_raw_rows), + ) + _write( + freeze_root / "raw/local_model_generalization/data/processed/arc_challenge_normalized.csv", + _normalized_csv(arc_normalized_rows), + ) + _write( + freeze_root / "raw/local_model_generalization/data/splits/arc_challenge/robustness_ids.json", + json.dumps(arc_selected_ids), + ) + _write( + freeze_root / "raw/local_model_generalization/data/splits/arc_challenge/robustness_metadata.json", + json.dumps({"seed": 42, "actual_size": len(arc_selected_ids)}), + ) + + mmlu_buffer = io.StringIO() + writer = csv.DictWriter(mmlu_buffer, fieldnames=["question", "subject", "choices", "answer"]) + writer.writeheader() + for question, subject, choices_repr, answer in mmlu_raw_rows: + writer.writerow({"question": question, "subject": subject, "choices": choices_repr, "answer": answer}) + _write( + freeze_root / "raw/local_model_generalization/data/raw/mmlu_raw.csv", mmlu_buffer.getvalue(), + ) + _write( + freeze_root / "raw/local_model_generalization/data/processed/mmlu_normalized.csv", + _normalized_csv(mmlu_normalized_rows), + ) + _write( + freeze_root / "raw/local_model_generalization/data/splits/benchmark/robustness_ids.json", + json.dumps(mmlu_selected_ids), + ) + _write( + freeze_root / "raw/local_model_generalization/data/splits/benchmark/robustness_metadata.json", + json.dumps({"seed": 7, "actual_size": len(mmlu_selected_ids)}), + ) + + ledger_lines = [] + for relative_path in _DATASET_RELATIVE_PATHS: + data = (freeze_root / relative_path).read_bytes() + digest = sha256(data).hexdigest() + if tamper_checksum_for == relative_path: + digest = "f" * 64 + ledger_lines.append(f"{digest} {relative_path}") + _write(freeze_root / "checksums" / "checksums.sha256", "\n".join(ledger_lines) + "\n") + return freeze_root + + +def test_build_stage1_expected_datasets_happy_path(tmp_path): + freeze_root = _build_dataset_freeze(tmp_path) + datasets = build_stage1_expected_datasets(freeze_root) + assert set(datasets) == {"arc_challenge", "mmlu"} + arc = datasets["arc_challenge"] + assert arc.reference_kind == "independent_input_snapshot" + assert arc.trust_label == "checksum_verified_freeze_internal" + assert set(arc.selected_question_ids) == {"arc_q1", "arc_q2", "arc_q3", "arc_q4"} + mmlu = datasets["mmlu"] + # the duplicate mmlu_q2 row collapses to one selected ID (keep-first) + assert set(mmlu.selected_question_ids) == {"mmlu_q1", "mmlu_q2"} + + +def test_build_stage1_expected_datasets_handles_variable_arc_option_count(tmp_path): + freeze_root = _build_dataset_freeze(tmp_path) + datasets = build_stage1_expected_datasets(freeze_root) + arc_frame = datasets["arc_challenge"].artifact_frame + row = arc_frame[arc_frame["question_id"] == "arc_q3"].iloc[0] + choices = json.loads(row["choices_json"]) + assert len(choices) == 3 # variable option count preserved, not padded + + +def test_build_stage1_expected_datasets_truncates_five_option_arc_row(tmp_path): + freeze_root = _build_dataset_freeze(tmp_path) + datasets = build_stage1_expected_datasets(freeze_root) + arc_frame = datasets["arc_challenge"].artifact_frame + row = arc_frame[arc_frame["question_id"] == "arc_q4"].iloc[0] + choices = json.loads(row["choices_json"]) + assert len(choices) == 4 # 5th raw option dropped, matching the archived normalizer + assert row["correct_option"] == "B" + + +def test_build_stage1_expected_datasets_verifies_checksums(tmp_path): + freeze_root = _build_dataset_freeze( + tmp_path, + tamper_checksum_for="raw/local_model_generalization/data/processed/arc_challenge_normalized.csv", + ) + with pytest.raises(Stage1ProfileError, match="checksum mismatch"): + build_stage1_expected_datasets(freeze_root) + + +def test_build_stage1_expected_datasets_rejects_revalidation_mismatch(tmp_path): + tampered = _default_arc_normalized_rows() + tampered[0] = {**tampered[0], "correct_option": "B"} # disagrees with raw answerKey "A" + freeze_root = _build_dataset_freeze(tmp_path, arc_normalized_rows=tampered) + with pytest.raises(Stage1ProfileError, match="revalidation failure"): + build_stage1_expected_datasets(freeze_root) + + +def test_build_stage1_expected_datasets_rejects_disagreeing_duplicate(tmp_path): + # The mutated 3rd row must still positionally revalidate against its OWN + # raw row (answer index 0 -> "A"/"x2"), so the disagreement is caught at + # the duplicate-question_id stage, not the raw/normalized revalidation + # stage -- isolating the specific invariant this test targets. + raw_rows = _default_mmlu_raw_rows() + raw_rows[2] = ("MQ2?", "history", "['x2', 'y2', 'z2', 'w2']", 0) + mismatched = _default_mmlu_normalized_rows() + mismatched[2] = {**mismatched[2], "correct_option": "A", "correct_answer_text": "x2"} + freeze_root = _build_dataset_freeze(tmp_path, mmlu_raw_rows=raw_rows, mmlu_normalized_rows=mismatched) + with pytest.raises(Stage1ProfileError, match="disagreeing field values"): + build_stage1_expected_datasets(freeze_root) + + +def test_build_stage1_expected_datasets_rejects_selected_id_missing_from_source(tmp_path): + freeze_root = _build_dataset_freeze(tmp_path, arc_selected_ids=["arc_q1", "arc_qXX"]) + with pytest.raises(Stage1ProfileError, match="absent from the"): + build_stage1_expected_datasets(freeze_root) + + +def test_build_stage1_expected_datasets_rejects_raw_normalized_row_count_mismatch(tmp_path): + freeze_root = _build_dataset_freeze(tmp_path, arc_raw_rows=_default_arc_raw_rows()[:2]) + with pytest.raises(Stage1ProfileError, match="row counts disagree"): + build_stage1_expected_datasets(freeze_root) diff --git a/tests/importing/test_transaction.py b/tests/importing/test_transaction.py new file mode 100644 index 0000000..fd07690 --- /dev/null +++ b/tests/importing/test_transaction.py @@ -0,0 +1,127 @@ +"""Tests for atomic staged publication (Task 11, reduced scope).""" + +from __future__ import annotations + +import pytest + +from choicebench.infra.artifacts import LockHeldError +from choicebench.importing.transaction import ( + ImportTransaction, + TransactionError, + atomic_publish_directory_no_replace, + fsync_tree, +) +import choicebench.importing.transaction as transaction_module + + +def test_fsync_tree_does_not_raise_on_a_populated_directory(tmp_path): + (tmp_path / "a").mkdir() + (tmp_path / "a" / "file.txt").write_text("hello") + fsync_tree(tmp_path) # must not raise + + +def test_atomic_publish_directory_no_replace_moves_staged_tree(tmp_path): + staged = tmp_path / "staged" + staged.mkdir() + (staged / "file.txt").write_text("content") + destination = tmp_path / "runs" / "run-1" + atomic_publish_directory_no_replace(staged, destination) + assert destination.is_dir() + assert (destination / "file.txt").read_text() == "content" + assert not staged.exists() + + +def test_atomic_publish_directory_no_replace_refuses_existing_destination(tmp_path): + staged = tmp_path / "staged" + staged.mkdir() + destination = tmp_path / "runs" / "run-1" + destination.mkdir(parents=True) + (destination / "existing.txt").write_text("already here") + with pytest.raises(TransactionError): + atomic_publish_directory_no_replace(staged, destination) + assert (destination / "existing.txt").read_text() == "already here" # untouched + assert staged.exists() # never moved + + +def test_renameat2_wrapper_fails_closed_when_unavailable(tmp_path, monkeypatch): + """If renameat2 cannot be located, this must refuse -- never silently + fall back to os.replace (which would race an existence check).""" + monkeypatch.setattr(transaction_module.ctypes.util, "find_library", lambda name: None) + staged = tmp_path / "staged" + staged.mkdir() + destination = tmp_path / "runs" / "run-1" + with pytest.raises(TransactionError, match="libc"): + atomic_publish_directory_no_replace(staged, destination) + assert staged.exists() + assert not destination.exists() + + +def test_import_transaction_publishes_a_new_run(tmp_path): + runs_dir = tmp_path / "runs" + runs_dir.mkdir() + validated_paths = [] + + def validator(path): + validated_paths.append(path) + (path / "manifest.json").write_text("{}") + + with ImportTransaction(runs_dir=runs_dir, run_id="run-1") as txn: + (txn.staged_run / "evidence.txt").write_text("data") + final_path = txn.publish(validator) + + assert final_path == runs_dir / "run-1" + assert (final_path / "evidence.txt").read_text() == "data" + assert (final_path / "manifest.json").read_text() == "{}" + # the validator ran against the staging directory, before the rename + assert len(validated_paths) == 1 + assert validated_paths[0].name.startswith(".staging-run-1") + # staging directory is gone (renamed into place, not left behind) + leftover = [p for p in runs_dir.iterdir() if p.name.startswith(".staging-run-1")] + assert leftover == [] + + +def test_import_transaction_verifies_existing_run_without_staging(tmp_path): + runs_dir = tmp_path / "runs" + final_run = runs_dir / "run-1" + final_run.mkdir(parents=True) + (final_run / "manifest.json").write_text("{}") + + validated_paths = [] + + def validator(path): + validated_paths.append(path) + + with ImportTransaction(runs_dir=runs_dir, run_id="run-1") as txn: + with pytest.raises(TransactionError, match="already exists"): + _ = txn.staged_run + result = txn.publish(validator) + + assert result == final_run + assert validated_paths == [final_run] + leftover = [p for p in runs_dir.iterdir() if p.name.startswith(".staging-run-1")] + assert leftover == [] # no staging directory was ever created + + +def test_import_transaction_cleans_up_staging_on_validator_failure(tmp_path): + runs_dir = tmp_path / "runs" + runs_dir.mkdir() + + def failing_validator(path): + raise ValueError("synthetic validation failure") + + with pytest.raises(ValueError, match="synthetic validation failure"): + with ImportTransaction(runs_dir=runs_dir, run_id="run-1") as txn: + txn.publish(failing_validator) + + assert not (runs_dir / "run-1").exists() + leftover = [p for p in runs_dir.iterdir() if p.name.startswith(".staging-run-1")] + assert leftover == [] + + +def test_import_transaction_holds_the_manifest_lock(tmp_path): + runs_dir = tmp_path / "runs" + runs_dir.mkdir() + with ImportTransaction(runs_dir=runs_dir, run_id="run-1"): + with pytest.raises(LockHeldError): + with ImportTransaction(runs_dir=runs_dir, run_id="run-1"): + pass diff --git a/tests/importing/test_validation.py b/tests/importing/test_validation.py new file mode 100644 index 0000000..a280a49 --- /dev/null +++ b/tests/importing/test_validation.py @@ -0,0 +1,284 @@ +"""Reduced-scope tests for importing.validation: question-identity join, +evidence-status computation with declaration cross-check, and evaluable-row +derivation. Reuses tests/importing/test_identity.py's condition/dataset +fixtures instead of re-deriving them here. +""" + +from __future__ import annotations + +from dataclasses import replace + +import pytest + +from choicebench.importing.csv_adapter import AdaptedTable, LogicalRecordSpan, SourceRow +from choicebench.importing.validation import ( + ImportValidationError, + normalize_realization_rows, + prepare_realization_validation_artifact, + validate_realization_validation_artifact, + validate_source_rows, + write_realization_validation_artifact, +) +from tests.importing.test_identity import _condition, _dataset + +_MAPPING = { + "question_id": "qid", + "question_text": "question", + "correct_option": "gold_answer", + "prediction": "predicted_letter", +} + + +def _row(index, qid, question, gold, a, b, predicted): + values = { + "qid": qid, + "question": question, + "gold_answer": gold, + "opt_a": a, + "opt_b": b, + "predicted_letter": predicted, + } + return SourceRow( + values=values, + span=LogicalRecordSpan(index=index, start=index * 10, end=index * 10 + 9, terminator=b"\n"), + raw_sha256=f"{index:064x}", + ) + + +def _valid_rows(): + return ( + _row(0, "q1", "One?", "A", "x", "y", "a"), + _row(1, "q2", "Two?", "B", "m", "n", "b"), + ) + + +def _table(rows): + return AdaptedTable(columns=tuple(_MAPPING.values()), rows=rows, source_sha256="1" * 64) + + +def test_validate_source_rows_accepts_matching_rows(): + condition = _condition() + expected = _dataset() + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=expected, + ) + assert validated.findings == () + assert set(validated.rows_by_question_id) == {"q1", "q2"} + + +def test_validate_source_rows_flags_null_question_id(): + rows = (_row(0, None, "One?", "A", "x", "y", "a"), _valid_rows()[1]) + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + codes = {f.code for f in validated.findings} + assert "NULL_QUESTION_ID" in codes + assert "MISSING_QUESTION_ID" in codes # q1 never arrived + + +def test_validate_source_rows_flags_duplicate_question_id(): + rows = (*_valid_rows(), _row(2, "q1", "One?", "A", "x", "y", "a")) + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + duplicate_findings = [f for f in validated.findings if f.code == "DUPLICATE_QUESTION_ID"] + assert len(duplicate_findings) == 2 # both occurrences flagged + assert "q1" not in validated.rows_by_question_id # ambiguous, not usable + + +def test_validate_source_rows_flags_unexpected_question_id(): + rows = (*_valid_rows(), _row(2, "q9", "Nine?", "A", "x", "y", "a")) + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + assert any(f.code == "UNEXPECTED_QUESTION_ID" and f.question_id == "q9" for f in validated.findings) + + +def test_validate_source_rows_flags_missing_question_id(): + validated = validate_source_rows( + _table(_valid_rows()[:1]), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + assert any(f.code == "MISSING_QUESTION_ID" and f.question_id == "q2" for f in validated.findings) + + +def test_validate_source_rows_flags_gold_mismatch(): + rows = (_row(0, "q1", "One?", "B", "x", "y", "a"), _valid_rows()[1]) + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + assert any(f.code == "CORRECT_OPTION_MISMATCH" and f.question_id == "q1" for f in validated.findings) + + +def test_validate_source_rows_flags_invalid_prediction(): + rows = (_row(0, "q1", "One?", "A", "x", "y", "z"), _valid_rows()[1]) + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + assert any(f.code == "INVALID_PREDICTION" and f.question_id == "q1" for f in validated.findings) + + +def test_validate_source_rows_flags_missing_prediction(): + rows = (_row(0, "q1", "One?", "A", "x", "y", None), _valid_rows()[1]) + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + assert any(f.code == "MISSING_PREDICTION" and f.question_id == "q1" for f in validated.findings) + + +def test_validate_source_rows_never_drops_or_reorders_rows(): + rows = (_valid_rows()[1], _valid_rows()[0]) # reversed input order + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=_condition(), expected=_dataset(), + ) + assert len(validated.rows_by_question_id) == 2 + + +def test_normalize_realization_rows_computes_complete_status(): + condition = _condition() + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + assert result.computed_evidence_status == "complete" + assert result.evaluable is True + assert len(result.evaluable_rows) == 2 + assert {row["question_id"] for row in result.evaluable_rows} == {"q1", "q2"} + assert all( + row["prediction_origin"] == "external_historical_inference" for row in result.evaluable_rows + ) + + +def test_normalize_realization_rows_refuses_declaration_mismatch(): + condition = replace(_condition(), evidence_status="malformed") + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + with pytest.raises(ImportValidationError, match="evidence status"): + normalize_realization_rows(validated, condition=condition, expected=_dataset()) + + +def test_normalize_realization_rows_produces_no_evaluable_rows_for_partial(): + condition = replace( + _condition(), + expected_question_ids=("q1", "q2", "q3"), + evidence_status="partial", + ) + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + assert result.computed_evidence_status == "partial" + assert result.evaluable is False + assert result.evaluable_rows == () + + +def test_normalize_realization_rows_produces_no_evaluable_rows_for_malformed(): + rows = (_row(0, "q1", "One?", "A", "x", "y", "z"), _valid_rows()[1]) # invalid prediction + condition = replace(_condition(), evidence_status="malformed") + validated = validate_source_rows( + _table(rows), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + assert result.computed_evidence_status == "malformed" + assert result.evaluable is False + assert result.evaluable_rows == () + assert result.defect_question_ids == ("q1",) + + +def test_normalize_realization_rows_computes_recoverable_status(): + condition = replace( + _condition(), + expected_question_ids=("q1", "q2", "q3"), + evidence_status="recoverable", + recoverable_question_ids=("q3",), + ) + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + assert result.computed_evidence_status == "recoverable" + assert result.evaluable is False + + +def test_normalize_realization_rows_computes_failed_status_for_zero_rows(): + condition = _condition() + validated = validate_source_rows( + _table(()), source_id="results", mapping=_MAPPING, + condition=replace(condition, evidence_status="failed"), expected=_dataset(), + ) + result = normalize_realization_rows( + validated, condition=replace(condition, evidence_status="failed"), expected=_dataset() + ) + assert result.computed_evidence_status == "failed" + assert result.evaluable is False + + +def test_normalize_realization_rows_excludes_from_paper_matrix_is_not_evaluable(): + condition = replace(_condition(), scope_disposition="excluded_from_paper_matrix") + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + assert result.computed_evidence_status == "complete" # row content is fine + assert result.evaluable is False # but scope excludes it from evaluation + assert result.evaluable_rows == () + + +def _realization_stub(digest="a" * 64): + return {"realization_id": "real_" + "b" * 16, "realization_digest": digest} + + +def test_validation_artifact_round_trips(tmp_path): + condition = _condition() + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + realization = _realization_stub() + prepared = prepare_realization_validation_artifact( + result, realization=realization, evidence_records=() + ) + path, reused = write_realization_validation_artifact(tmp_path, prepared) + assert reused is False + payload = validate_realization_validation_artifact( + path, manifest={}, realization=realization, expected_sha256=prepared.file_sha256 + ) + assert payload["realization_id"] == realization["realization_id"] + + _, second_reused = write_realization_validation_artifact(tmp_path, prepared) + assert second_reused is True + + +def test_validation_artifact_rejects_tampered_bytes(tmp_path): + condition = _condition() + validated = validate_source_rows( + _table(_valid_rows()), source_id="results", mapping=_MAPPING, + condition=condition, expected=_dataset(), + ) + result = normalize_realization_rows(validated, condition=condition, expected=_dataset()) + realization = _realization_stub() + prepared = prepare_realization_validation_artifact( + result, realization=realization, evidence_records=() + ) + path, _ = write_realization_validation_artifact(tmp_path, prepared) + path.write_bytes(prepared.json_bytes + b"\n// tampered") + with pytest.raises(RuntimeError, match="content integrity"): + validate_realization_validation_artifact( + path, manifest={}, realization=realization, expected_sha256=prepared.file_sha256 + ) diff --git a/tests/test_wheel_smoke.py b/tests/test_wheel_smoke.py index ecf990d..ba89a9d 100644 --- a/tests/test_wheel_smoke.py +++ b/tests/test_wheel_smoke.py @@ -94,13 +94,16 @@ def test_built_wheel_and_sdist_run_outside_repository(tmp_path): cwd=tmp_path, ) - # Functional CLI smoke reuses the test environment's already-installed - # dependencies. The release verification separately performs a true clean - # dependency install from the built wheel. + # A fully isolated venv with the wheel's declared dependencies installed + # alongside it. --system-site-packages previously stood in for this, but + # under a truly clean venv/container it inherits an empty (or unrelated) + # site-packages rather than this test's own dependencies, so the CLI + # commands below would crash on missing imports (e.g. pandas) outside + # this one machine's incidentally-populated user site-packages. venv = tmp_path / "workflow-venv" - subprocess.run([sys.executable, "-m", "venv", "--system-site-packages", str(venv)], check=True) + subprocess.run([sys.executable, "-m", "venv", str(venv)], check=True) pip = venv / "bin" / "pip" - _run_checked([str(pip), "install", "--no-deps", str(wheel)]) + _run_checked([str(pip), "install", str(wheel)]) for command in ("choicebench-prepare-toy", "choicebench-run", "choicebench-evaluate", "choicebench-prepare"): _run_checked([str(venv / "bin" / command), "--help"]) _run_workflow(venv, tmp_path / "wheel workspace é", "wheel")