Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
47 commits
Select commit Hold shift + click to select a range
65d9567
docs: specify external results importer
cotenthusiast Jul 18, 2026
b26a46e
docs: refine external importer identity design
cotenthusiast Jul 18, 2026
712beb5
docs: plan external results importer implementation
cotenthusiast Jul 18, 2026
fe5d90f
test: freeze manifest v2 compatibility
cotenthusiast Jul 18, 2026
9d98ce1
test: exercise first manifest v2 reader call
cotenthusiast Jul 18, 2026
9d51ed6
test: align manifest v2 IDs with package version
cotenthusiast Jul 18, 2026
582dac5
feat: add strict external import specification
cotenthusiast Jul 18, 2026
439220e
fix: tighten external import specification
cotenthusiast Jul 18, 2026
6598496
fix: accept external implementation identities
cotenthusiast Jul 18, 2026
732f8c3
feat: add deterministic CSV source adapter
cotenthusiast Jul 18, 2026
c23c239
fix: enforce declared CSV line terminators
cotenthusiast Jul 18, 2026
4edb977
feat: add trusted import dataset references
cotenthusiast Jul 18, 2026
b9bb0b7
fix: harden imported dataset identity handoff
cotenthusiast Jul 18, 2026
68f1892
fix: separate dataset artifact and selection identity
cotenthusiast Jul 18, 2026
3b3664a
fix: preserve normalized dataset integer semantics
cotenthusiast Jul 18, 2026
b05ba09
feat: separate condition and realization identity
cotenthusiast Jul 18, 2026
a0f789e
fix: close importer identity schemas
cotenthusiast Jul 18, 2026
930dae2
fix: harden importer identity trust boundaries
cotenthusiast Jul 18, 2026
280c5ba
fix: derive active importer identity from callables
cotenthusiast Jul 18, 2026
c095334
fix: bind importer lineage and dataset ownership
cotenthusiast Jul 19, 2026
9064d85
fix: preserve partial provenance coverage
cotenthusiast Jul 19, 2026
268422a
fix: close importer identity dependency graph
cotenthusiast Jul 19, 2026
fd03a2f
fix: bind importer identity to trusted runtime inputs
cotenthusiast Jul 19, 2026
e694eac
fix: close derived importer lineage edges
cotenthusiast Jul 19, 2026
175e4ee
fix: require reachable transformation lineage
cotenthusiast Jul 19, 2026
ee2108b
fix: close final Task 5 identity-contract gaps
cotenthusiast Jul 19, 2026
1c6625c
fix: bind cross-module helper globals into callable identity
cotenthusiast Jul 19, 2026
3a69eda
perf: cache implementation_identity for resolved-global bindings
cotenthusiast Jul 19, 2026
965dc77
feat: add additive manifest-v3 schema alongside v2
cotenthusiast Jul 19, 2026
c401864
feat: publish non-circular v2 result artifacts for v3 realizations
cotenthusiast Jul 19, 2026
6a3cf1b
feat: validate imported rows by question identity and compute evidenc…
cotenthusiast Jul 19, 2026
bdbffd6
feat: validate typed authorization and derive immutable repair overlays
cotenthusiast Jul 19, 2026
5b3ec6d
feat: stage external imports atomically with verified evidence storage
cotenthusiast Jul 19, 2026
7590e1e
feat: orchestrate generic dry-run and real imports
cotenthusiast Jul 19, 2026
fc408c9
feat: read and evaluate v3 imported result sets
cotenthusiast Jul 19, 2026
0bb6304
feat: translate Stage 1 queue/status/matrix accounting
cotenthusiast Jul 19, 2026
08e2c52
feat: build the Stage 1 ARC/MMLU expected-dataset trust chain
cotenthusiast Jul 19, 2026
80ee933
feat: overlay merge-into-new-run orchestration + expected_datasets ov…
cotenthusiast Jul 19, 2026
3f6e455
feat: full Stage 1 ImportSpec/AuthorizationSpec assembly
cotenthusiast Jul 19, 2026
79fffee
fix: real-freeze validation findings (containment root, perf, identit…
cotenthusiast Jul 19, 2026
b9167c3
fix: overlay lineage_components order was nondeterministic across pro…
cotenthusiast Jul 19, 2026
e63037b
fix: offline_transformation overlay digests use dataset order (Fix B)
cotenthusiast Aug 2, 2026
d79cf8d
fix: distinguish overlayable (structural) from evaluable (scoring) ba…
cotenthusiast Aug 2, 2026
ff27f10
fix: NaN-safe retained rows for overlay authorization (Fix A, importe…
cotenthusiast Aug 2, 2026
fb24ed2
fix: fourth-cell import against real choices_json format (Fix D, impo…
cotenthusiast Aug 2, 2026
0c7af03
fix(test): stop wheel-smoke test relying on ambient --system-site-pac…
cotenthusiast Aug 2, 2026
55c7fb7
docs: point README at PAPER.md on trunk
cotenthusiast Aug 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,7 @@
# choicebench

This branch is a standalone result-importer tool evaluated separately from choicebench's associated paper — see [PAPER.md](https://github.com/cotenthusiast/choicebench/blob/choicebench/PAPER.md) on `choicebench` (trunk) for details on how this branch relates to the other paper-supporting branches.

ChoiceBench is a lightweight framework for MCQ evaluation-method research on LLMs, with built-in support for answer-order bias analysis and mitigation methods.

[![Tests](https://github.com/cotenthusiast/choicebench/actions/workflows/test.yml/badge.svg)](https://github.com/cotenthusiast/choicebench/actions/workflows/test.yml)
Expand Down
3,260 changes: 3,260 additions & 0 deletions docs/superpowers/plans/2026-07-18-external-results-importer.md

Large diffs are not rendered by default.

1,344 changes: 1,344 additions & 0 deletions docs/superpowers/specs/2026-07-18-external-results-importer-design.md

Large diffs are not rendered by default.

2 changes: 2 additions & 0 deletions src/choicebench/identity.py
Original file line number Diff line number Diff line change
Expand Up @@ -167,6 +167,8 @@ def canonicalize(value: Any, *, redact_secrets: bool = True) -> Any:
return sorted(items, key=lambda item: json.dumps(item, sort_keys=True, separators=(",", ":"), ensure_ascii=True))
if isinstance(value, Path):
return value.as_posix()
if isinstance(value, (bytes, bytearray)):
return bytes(value).hex()
if isinstance(value, str):
text = unicodedata.normalize("NFC", value)
return redact_text(text) if redact_secrets else text
Expand Down
51 changes: 51 additions & 0 deletions src/choicebench/importing/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
"""Public types for strict external result import declarations."""

from choicebench.importing.schema import (
AuthorizationSpec,
CsvDialectSpec,
DatasetReferenceSpec,
DerivationOrigin,
EvidenceStatus,
ImportConditionSpec,
ImportMethodSpec,
ImportModelSpec,
ImportPromptSpec,
ImportSpec,
ImportSpecError,
ImportState,
NumericColumnSpec,
OptionMappingSpec,
OverlaySpec,
PredictionOrigin,
ResultOriginSpec,
ScopeDisposition,
SourceArtifactSpec,
import_spec_digest,
load_import_spec,
stable_import_projection,
)

__all__ = [
"AuthorizationSpec",
"CsvDialectSpec",
"DatasetReferenceSpec",
"DerivationOrigin",
"EvidenceStatus",
"ImportConditionSpec",
"ImportMethodSpec",
"ImportModelSpec",
"ImportPromptSpec",
"ImportSpec",
"ImportSpecError",
"ImportState",
"NumericColumnSpec",
"OptionMappingSpec",
"OverlaySpec",
"PredictionOrigin",
"ResultOriginSpec",
"ScopeDisposition",
"SourceArtifactSpec",
"import_spec_digest",
"load_import_spec",
"stable_import_projection",
]
191 changes: 191 additions & 0 deletions src/choicebench/importing/authorization.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,191 @@
"""Validate typed repair/offline-transformation authorization bundles.

An authorization bundle grants specific (condition_digest, question_id) pairs
permission to be replaced, typed as either `inference_repair` (executable --
a real model may be run) or `offline_transformation` (non-executing semantic
rematching only). This module never runs inference or opens a filesystem
path itself; it validates an already-opened source and already-known
condition digests/expected datasets.
"""

from __future__ import annotations

from dataclasses import dataclass
from typing import Mapping

from choicebench.identity import integrity_digest, short_id
from choicebench.importing.csv_adapter import OpenedSource
from choicebench.importing.dataset_reference import ExpectedDataset
from choicebench.importing.schema import AuthorizationSpec


class AuthorizationError(ValueError):
"""Raised when an authorization bundle or its use is invalid or unsafe."""


@dataclass(frozen=True)
class ValidatedAuthorizationBundle:
authorization_id: str
authorization_digest: str
authorization_type: str
grants: Mapping[str, Mapping[str, str]] # condition_digest -> question_id -> reason
authority: str
purpose: str
executable: bool
source_sha256: str
input_evidence_digests: Mapping[str, str]
expected_snapshot_digests: Mapping[str, str]


@dataclass(frozen=True)
class ValidatedAuthorization:
bundle_id: str
bundle_digest: str
authorization_type: str
condition_digest: str
question_reasons: Mapping[str, str]
authority: str
purpose: str
executable: bool
source_sha256: str
input_evidence_digests: Mapping[str, str]
expected_snapshot_digest: str


def validate_authorization_bundle(
declaration: AuthorizationSpec,
*,
opened_source: OpenedSource,
condition_digests: Mapping[str, str],
expected: Mapping[str, ExpectedDataset],
) -> ValidatedAuthorizationBundle:
"""Validate one authorization artifact against an independently opened
source and the caller's own known condition digests/expected datasets.

Fails closed on: source tampering, an authorization_type/executable
mismatch, a grant referencing an unknown condition or an out-of-selection
question, an empty reason, or a self-assigned authorization_id that does
not match the bundle's own recomputed digest.
"""
if opened_source.source_id != declaration.source_id:
raise AuthorizationError(
f"Authorization {declaration.authorization_id!r} declares source "
f"{declaration.source_id!r} but was validated against a differently "
f"declared opened source {opened_source.source_id!r}."
)
if declaration.authorization_type == "inference_repair" and not declaration.executable:
raise AuthorizationError(
f"Authorization {declaration.authorization_id!r} declares inference_repair "
"but executable=false."
)
if declaration.authorization_type == "offline_transformation" and declaration.executable:
raise AuthorizationError(
f"Authorization {declaration.authorization_id!r} declares "
"offline_transformation but executable=true."
)

grants: dict[str, dict[str, str]] = {}
for condition_digest, question_reasons in declaration.condition_question_reasons.items():
if condition_digest not in condition_digests.values():
raise AuthorizationError(
f"Authorization {declaration.authorization_id!r} grants a condition "
f"{condition_digest!r} that is not among the caller's known conditions."
)
dataset = expected.get(condition_digest)
if not isinstance(question_reasons, Mapping) or not question_reasons:
raise AuthorizationError(
f"Authorization {declaration.authorization_id!r} has an empty grant for "
f"condition {condition_digest!r}."
)
selected = set(dataset.selected_question_ids) if dataset is not None else None
normalized_reasons: dict[str, str] = {}
for question_id, reason in question_reasons.items():
if not isinstance(reason, str) or not reason.strip():
raise AuthorizationError(
f"Authorization {declaration.authorization_id!r} grant for "
f"{condition_digest!r}/{question_id!r} has no explicit reason."
)
if selected is not None and question_id not in selected:
raise AuthorizationError(
f"Authorization {declaration.authorization_id!r} grants question "
f"{question_id!r} that is not in condition {condition_digest!r}'s "
"expected selection."
)
normalized_reasons[question_id] = reason
grants[condition_digest] = normalized_reasons
if not grants:
raise AuthorizationError(
f"Authorization {declaration.authorization_id!r} grants no conditions."
)

bundle_payload = {
"schema_version": "choicebench.authorization-bundle.v1",
"authorization_type": declaration.authorization_type,
"grants": grants,
"authority": declaration.authority,
"purpose": declaration.purpose,
"executable": declaration.executable,
"source_sha256": opened_source.sha256,
"input_evidence_digests": dict(declaration.input_evidence_digests),
"expected_snapshot_digests": dict(declaration.expected_snapshot_digests),
}
authorization_digest = integrity_digest(bundle_payload)
expected_id = short_id("auth", bundle_payload)
if declaration.authorization_id != expected_id:
raise AuthorizationError(
f"Authorization ID {declaration.authorization_id!r} does not match its own "
f"recomputed digest (expected {expected_id!r}); it cannot self-assign an ID."
)
if set(declaration.expected_snapshot_digests) != set(grants):
raise AuthorizationError(
f"Authorization {declaration.authorization_id!r} expected_snapshot_digests "
"do not exactly cover its granted conditions."
)

return ValidatedAuthorizationBundle(
authorization_id=expected_id,
authorization_digest=authorization_digest,
authorization_type=declaration.authorization_type,
grants=grants,
authority=declaration.authority,
purpose=declaration.purpose,
executable=declaration.executable,
source_sha256=opened_source.sha256,
input_evidence_digests=dict(declaration.input_evidence_digests),
expected_snapshot_digests=dict(declaration.expected_snapshot_digests),
)


def authorization_for_condition(
bundle: ValidatedAuthorizationBundle, *, condition_digest: str
) -> ValidatedAuthorization:
"""Slice one condition's grant out of a bundle without weakening its
aggregate identity: bundle_id/bundle_digest are the whole bundle's, so a
slice can be traced back to (and cannot silently diverge from) the
exact bundle it came from, and it can never borrow another condition's
grants."""
question_reasons = bundle.grants.get(condition_digest)
if question_reasons is None:
raise AuthorizationError(
f"Authorization {bundle.authorization_id!r} grants no questions for "
f"condition {condition_digest!r}; this condition cannot self-authorize."
)
expected_snapshot_digest = bundle.expected_snapshot_digests.get(condition_digest)
if expected_snapshot_digest is None:
raise AuthorizationError(
f"Authorization {bundle.authorization_id!r} has no expected snapshot digest "
f"for condition {condition_digest!r}."
)
return ValidatedAuthorization(
bundle_id=bundle.authorization_id,
bundle_digest=bundle.authorization_digest,
authorization_type=bundle.authorization_type,
condition_digest=condition_digest,
question_reasons=dict(question_reasons),
authority=bundle.authority,
purpose=bundle.purpose,
executable=bundle.executable,
source_sha256=bundle.source_sha256,
input_evidence_digests=dict(bundle.input_evidence_digests),
expected_snapshot_digest=expected_snapshot_digest,
)
Loading
Loading