Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
6eec2d9
Add two-phase held-out chemistry validation harness
erinepshovel-code Sep 29, 2026
579e036
Test held-out freeze/reveal validation boundary
erinepshovel-code Sep 29, 2026
7fa4815
Add empty held-out chemistry oracle envelope
erinepshovel-code Sep 29, 2026
0e9591d
Document held-out chemistry validation protocol
erinepshovel-code Sep 29, 2026
1cc98fe
Package held-out validation harness
erinepshovel-code Sep 29, 2026
1281ddd
Repair held-out validation review and replay gates
erinepshovel-code Sep 29, 2026
c82325a
Close exact-head validation review findings
erinepshovel-code Sep 29, 2026
a5d7dd5
Stabilize held-out JSON boundary
erinepshovel-code Sep 29, 2026
7d8620a
Complete held-out contract evidence graph
erinepshovel-code Sep 29, 2026
3736139
fix: seal held-out evidence reloads and reconcile main
erinepshovel-code Oct 4, 2026
e85fd0f
fix: refresh join evidence after package inventory reconciliation
erinepshovel-code Oct 4, 2026
d6f10cc
fix: preserve unresolved provenance and verify persisted receipts
erinepshovel-code Oct 4, 2026
4ffbdfd
fix: reject unresolved prediction source identities
erinepshovel-code Oct 4, 2026
ddfef26
Reject JSON numeric literals that lose decimal value during evidence …
erinepshovel-code Oct 4, 2026
b1053fc
Keep unknown and non-finite evidence unresolved before every comparator
erinepshovel-code Oct 4, 2026
d70a200
Compare numeric evidence by its committed JSON decimal value
erinepshovel-code Oct 4, 2026
805072b
Reject unresolved sentinels in all case and domain identities
erinepshovel-code Oct 4, 2026
db6fc5c
Reject unpaired Unicode surrogates before loading or hashing evidence
erinepshovel-code Oct 4, 2026
91a3dcd
Declare held-out runtime contracts and preserve unresolved boundary m…
erinepshovel-code Oct 4, 2026
500c6e4
Declare the comparison-side packaged oracle resource dependency
erinepshovel-code Oct 4, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion data/epac-join-term-v0-receipt.json

Large diffs are not rendered by default.

22 changes: 22 additions & 0 deletions data/heldout_chemistry_oracle.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
{
"schema": "epac.heldout-chemistry-oracle",
"version": "v1",
"policy": {
"role": "comparison-only",
"construction_access": "forbidden",
"selection_binding": "Case IDs, domains, and comparator rules must be frozen and persisted in an epac.heldout-validation-plan commitment before predictions are frozen or revealed.",
"empty_corpus_semantics": "Packaging default only; anti-selection comes from the frozen validation plan, not corpus emptiness.",
"discovery_sources": [
"Periodic Table app or other chemistry references"
],
"acceptance": "Important values require an independently identifiable authoritative source before scoring."
},
"domains": [
"isotope",
"oxidation-state",
"valence",
"reaction",
"element-property"
],
"cases": []
}
6 changes: 3 additions & 3 deletions docs/boundary_probe_completeness.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,13 +35,13 @@ construction, evidence, provenance and visualization modules. The inventory cove

| item | count |
|---|---:|
| declared operations classified | 136 |
| declared operations classified | 173 |
| boundary-relevant operations | 58 |
| omitted boundary-relevant operations | 14 |
| omitted operations that distinguish same-B frozen states | 3 |
| ambiguous operations | 39 |
| ambiguous operations | 76 |

Thirty-nine public callable addresses, including the representation, probe-relativity, minimal-refinement and probe-completeness audits, have unresolved boundary relevance. Unknown names are not classified by substrings or treated as internal/non-boundary. The ledger records these uncertainties explicitly. Existing counterexamples still falsify completeness; unresolved operations do not erase that evidence.
Seventy-six public callable addresses, including the representation, probe-relativity, minimal-refinement, probe-completeness, join-term, and held-out validation surfaces, have unresolved boundary relevance. Each of the nine held-out operations has an explicit unresolved mapping reason: selection rules, prediction envelopes, oracle data, and comparison receipts have not been measured against the frozen 27-state construction or same-B discrimination. Integrity verification does not establish that mapping. Unknown names are not classified by substrings or treated as internal/non-boundary. The ledger records these uncertainties explicitly. Existing counterexamples still falsify completeness; unresolved operations do not erase that evidence.

## Partition Result

Expand Down
10 changes: 5 additions & 5 deletions docs/epac-join-term-v0-audit.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,7 @@ The public receipt recovers one complete typed S0-through-S6 join tree and its e
## Exact authority

- EPAC base: `1e5c999286f12221eff9870d1372203ac7935f2a` (tree `ce7781938ff684d826bd91f475b5423bd56b146a`)
- EPAC feature source manifest: `c3aad9a890e4ea42afcc8118a14c21b8ef10044ec1520a0d32f7daf668fbff0b`
- EPAC feature source manifest: `ad606b04310d1d04db2bb6b773b3e2d045bb0b10debc90e7bd568b3c2be54e22`
- METAPAT producer: `1cdfb09dd00a451cee30eec2e78624df8c682662` (tree `d946a5a18b0c53fd561dcd9d68fdf36d5dd11638`)
- METAPAT application: `metapat.application.epac_join_terms@epac-join-terms-application-v4`
- METAPAT application digest: `461ef7e059aa65b514017683ba9b57058ee9f582bc63748fabd021cdeb660b4b`
Expand All @@ -20,7 +20,7 @@ The containing feature commit cannot be embedded in a file inside itself. The ev

- Candidate digest: `11504eb61c4952fc3da04a51d3779484194292db2c8a5c3174b3c3ea540e6442`
- Candidate JSON SHA-256: `821d35a1e42c8d22423f78fdce8449066af40392d47f3e79e5b30b8f993fcbfd`
- Evidence digest: `215844eda0b1b1f3edf95bf7b360df84deecc97382a88f777c150fa50e359839`
- Evidence digest: `0a11bd666927926ab57fd0f5f7625034cbfa8670f4f1f3dc4c4c6f3905c2d638`
- Origins: `7`
- Authored joins (no Transformation or Time claim): `6`
- Slots: `18` (`6` members, `6` holes, `6` leftovers)
Expand Down Expand Up @@ -56,10 +56,10 @@ This is information loss. It is not a secrecy, computational-hardness, or confid
- `docs/research-status.md`: `1f5194f0dc44521b6dc81d9a9494c72ff385fd879b8eb2b2849824640c5c2d07`
- `docs/work-graph.json`: `4edb37ace798511ff4b5f667afd1f0a3852ce78b7335349aaa9662171de0b9e9`
- `docs/work-graphs/epac-join-term-v0.json`: `7487441ec7e3ff3203635e245fccba5d293192cef5201c657946a15788166f16`
- `epac_boundary_probe_completeness.py`: `61df1624adc2970fe2e7d9d94ee15359decd42ddb7285ac223bd46eb748541ec`
- `epac_boundary_probe_completeness.py`: `96e37f87e62c6346f8bd4f888e6151487636853d8c9309f199574b5c60da9600`
- `epac_join_term.py`: `4697da5c76d5dfc77aedc5c84c71b7175180550824b6a1a2c1bf23a8548839af`
- `pyproject.toml`: `414cdc71cdb2e0ff859fcb59a359d57a5327e6b105e0ff010a522216dcd3881f`
- `tests/test_boundary_probe_completeness.py`: `254590ac58a54979f0f05534cea357e2369577e96304dda0ae337f3114d13ac4`
- `pyproject.toml`: `44936ff75383ab25d3e295c6574b1759709fef7a072773d4c5774dd00243bcdd`
- `tests/test_boundary_probe_completeness.py`: `8cd71e6f708949ce03a6750ac690a563c7cb236e9145d33e583539f0af46f93d`
- `tests/test_epac_join_term.py`: `ed060a6ef7903f178e3da6b94890f088de8b37dc7b94e6e41b4ccc15a3b7f5c8`
- `tools/generate_join_term_v0.py`: `1836734e0ba692f23f604838ed3146c811528266a341979366eff6b19447e331`

Expand Down
167 changes: 167 additions & 0 deletions docs/heldout-chemistry-validation.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,167 @@
# Held-out chemistry validation

Purpose: test EPAC against chemistry facts without feeding those facts into EPAC construction.

Runtime effects are limited to reading caller-supplied evidence and the packaged oracle and returning in-memory results. The module writes no files or network messages; callers control persistence as shown below. The source declares its runtime boundary, documentation, capabilities, dependency edges, and existing owner.

The operation inventory retains all nine public held-out functions as `AMBIGUOUS`, with a separate mapping reason for each operation. Their relation to the frozen 27-state construction and ability to distinguish same-B states have not been measured. Validation integrity does not establish boundary non-relevance or transfer chemistry standing.

## Boundary

The selection boundary is committed before EPAC predictions are produced or revealed:

1. Choose stable case IDs, domains, and comparator rules only; do not include expected values or provenance.
2. Call `freeze_validation_plan(...)` and persist the returned plan commitment.
3. Run EPAC construction and produce predictions only for IDs in that frozen plan.
4. Call `freeze_predictions(..., source_identity=<exact EPAC head/receipt>, validation_plan=<persisted plan>)` and persist the prediction commitment.
5. Only then reveal/load a held-out oracle containing the same case IDs, domains, and comparator rules plus expected values and provenance.
6. Call `compare_after_freeze(commitment, oracle, validation_plan=<persisted plan>)`.
7. Preserve the receipt and its `SURVIVED / FALSIFIED / UNRESOLVED` results.

The flow is:

`frozen validation plan → EPAC derivation → frozen prediction → held-out oracle → comparator → evidence receipt`

The chemistry app is a **discovery/reference surface**, not an EPAC construction dependency and not the sole authority. Facts used for scoring carry an authority plus locator. Missing predictions, absent expected values, missing/null provenance, unsupported comparisons, and non-finite numeric comparisons are `UNRESOLVED`; they are never silently counted as success. Duplicate oracle IDs or drift in the frozen case inventory/domain/comparator contract are rejected before scoring.

The checked-in `data/heldout_chemistry_oracle.json` is intentionally empty as a distributable comparison-side resource and smoke-test default. Its emptiness does **not** prove independent case selection. The anti-selection control is the persisted validation-plan commitment that binds case identity and comparator rules before the prediction commitment exists. Stronger chronology or independent custody can be supplied by an external evidence store without changing EPAC construction.

Recommended stable IDs include `element:H:valence`, `element:O:oxidation`, `isotope:H-1:mass`, and `reaction:<canonical-id>:products`. The corpus can grow across isotope behavior, oxidation/valence, reaction outcomes, and element properties without changing EPAC construction code.

## Comparator and evidence input rules

Validation-plan fields are deliberately narrow so oracle-shaped data cannot cross into the prediction side:

- `domain` is either `null` or one nonempty canonical string; objects, arrays, numbers, blank strings, and padded strings are rejected;
- `exact` and `set-equality` comparators accept only `kind`;
- `numeric-tolerance` accepts only `kind` and `absolute_tolerance`; a supplied tolerance must be a number or `null`, so nested oracle fields, strings, and booleans cannot cross through that slot;
- an unknown comparator kind may be preregistered with `kind` only and will score `UNRESOLVED`;
- any additional comparator field is rejected before the plan is committed or verified.

Prediction and programmatic oracle values must use the JSON container/type model: `null`, booleans, numbers, strings, arrays/lists, and string-keyed objects/mappings. Python-only coercible values such as tuples are rejected before commitment or evidence hashing, including when nested. This keeps an in-memory commitment semantically identical to the same commitment after documented JSON persistence/reload. Non-finite Python floats remain admissible only so the comparator can classify them `UNRESOLVED`; they never produce a numeric success.

Numeric-tolerance comparison keeps integers as integers and uses exact rational arithmetic over accepted finite JSON decimal values, so large integer distinctions are not collapsed by an intermediate float conversion. Non-finite operands or tolerances remain `UNRESOLVED`.

Provenance `authority` and `locator` must both be nonempty strings after trimming; blank or null identities are `UNRESOLVED`. All evidence loaders reject duplicate JSON object keys at every nesting level, including identical duplicates and escaped spellings of the same key. Reload plans with `load_validation_plan(path)` and commitments with `load_prediction_commitment(path, validation_plan=plan)`: these verify envelopes and digests, and the latter also verifies the plan binding and prediction-ID membership before returning. Missing predictions remain admissible and score `UNRESOLVED`; extra predictions are rejected. `verify_commitment()` alone checks only the envelope, not plan membership. Raw `json.loads` discards duplicate-key evidence and must not be used to reload externally supplied plans or commitments.

In a source checkout, `load_packaged_oracle()` prefers the sibling `data/heldout_chemistry_oracle.json`; in an installed distribution it reads the `epac_data` package resource. Loading a plan or commitment never loads oracle data.

The reserved `hmmm` provenance identity (ignoring surrounding whitespace and case) is unresolved, just like a blank or null identity. It cannot produce a scored success.

The same reserved sentinel is rejected as a prediction `source_identity` during freezing, commitment verification, reloading, and comparison, including envelopes with a recomputed matching digest. Supply a resolved EPAC head or receipt identity before freezing predictions.

Case identifiers and non-null domains must also be resolved identities: the reserved `hmmm` sentinel is rejected in plans, predictions, oracle comparisons, and preserved receipts, including externally reconstructed envelopes with matching digests. An omitted domain remains `null`; it is not replaced by an unknown identity string.

Within prediction or expected evidence, a reserved `hmmm` string (ignoring case and surrounding whitespace) or a non-finite number makes the case `UNRESOLVED` before any comparator runs. This applies recursively to arrays and object keys/values, even when the other operand has a different type, length, or keys. Explicitly unknown evidence cannot establish agreement or disagreement, and the receipt loader enforces the same classification.

Every loader also requires decimal/exponent JSON numbers to preserve their exact decimal value when parsed and serialized back to JSON. Ordinary round-trip values such as `0.1` and `1.007825` remain supported; precision-losing literals such as `9007199254740993.0`, finite-number overflow, and underflow are rejected before hashing or scoring. Integer tokens retain Python's exact integer handling. This is a decimal-value round-trip check, not a claim that decimal fractions have exact binary encodings. Explicit non-finite Python JSON extensions retain the existing `UNRESOLVED` comparison policy.

All numeric comparison uses exact rational interpretations of the committed JSON decimal values, including nested exact/set equality and tolerance arithmetic. For example, `1e23` means exactly `100000000000000000000000`, not the nearby integer represented by its in-memory binary float. Programmatic floats use their canonical JSON decimal representation, so direct and persisted scoring agree.

Strings and object keys must contain valid Unicode scalar values. Loaders and programmatic normalization reject unpaired surrogates with `ValueError` before hashing or returning loaded evidence. Valid non-ASCII text, including JSON-escaped surrogate pairs that decode to a Unicode scalar, retains the same canonical UTF-8 digest.

Reload preserved receipts with `load_validation_receipt(path)`. It rejects duplicate keys, invalid envelopes and digest fields, repeated case IDs, inconsistent result statuses or counts, and a mismatched `receipt_sha256`. It also checks scored statuses against the evidence carried in each result. These checks establish internal integrity; an unkeyed checksum does not authenticate a custodian or prevent replacement of the entire evidence chain. Preserve externally trusted digests or independently controlled custody when that stronger claim matters.

## Runnable end-to-end example

This runs from a source checkout or the installed distribution. It persists and reloads both commitments before revealing a synthetic held-out value, exercises packaged-oracle loading, then persists the evidence receipt.

```bash
python - <<'PY'
import json
from pathlib import Path
from tempfile import TemporaryDirectory

from epac_heldout_validation import (
compare_after_freeze,
freeze_predictions,
freeze_validation_plan,
load_oracle,
load_packaged_oracle,
load_prediction_commitment,
load_validation_receipt,
load_validation_plan,
)

with TemporaryDirectory() as tmp:
root = Path(tmp)

plan = freeze_validation_plan([
{
"id": "example:scalar",
"domain": "example",
"comparison": {"kind": "exact"},
}
])
(root / "plan.json").write_text(json.dumps(plan), encoding="utf-8")
plan = load_validation_plan(root / "plan.json")

# Replace this literal with an EPAC-derived prediction in a real run.
predictions = {"example:scalar": 7}
commitment = freeze_predictions(
predictions,
source_identity="epac@example-exact-head-or-receipt",
validation_plan=plan,
)
(root / "prediction.json").write_text(
json.dumps(commitment), encoding="utf-8"
)
commitment = load_prediction_commitment(
root / "prediction.json", validation_plan=plan
)

# The package ships an empty comparison-side resource; loading it works
# identically from a checkout or installed wheel.
packaged = load_packaged_oracle()
assert packaged["schema"] == "epac.heldout-chemistry-oracle"

# Reveal/write held-out evidence only after the two commitments above.
oracle = {
"schema": "epac.heldout-chemistry-oracle",
"version": "v1",
"cases": [{
"id": "example:scalar",
"domain": "example",
"comparison": {"kind": "exact"},
"expected": 7,
"provenance": {
"authority": "example-authority",
"locator": "example:7",
},
}],
}
(root / "oracle.json").write_text(json.dumps(oracle), encoding="utf-8")
oracle = load_oracle(root / "oracle.json")

receipt = compare_after_freeze(
commitment,
oracle,
validation_plan=plan,
)
(root / "receipt.json").write_text(json.dumps(receipt), encoding="utf-8")
preserved = load_validation_receipt(root / "receipt.json")

assert preserved["counts"] == {
"SURVIVED": 1,
"FALSIFIED": 0,
"UNRESOLVED": 0,
}
assert preserved["validation_plan_sha256"] == plan["plan_sha256"]
assert preserved["prediction_commitment_sha256"] == commitment["commitment_sha256"]
assert len(preserved["oracle_sha256"]) == 64
assert len(preserved["receipt_sha256"]) == 64
print(preserved["counts"])
PY
```

## Failure definition

- **SURVIVED**: frozen prediction satisfies the preregistered comparator.
- **FALSIFIED**: frozen prediction is comparable and disagrees.
- **UNRESOLVED**: the comparison cannot honestly decide.

The receipt binds the frozen validation plan, frozen prediction commitment, revealed oracle digest, and detached result evidence.

## hmmm

The module proves its own ordering and digest relationships; it does not by itself prove that a human or external curator lacked access to predictions before creating the plan. If that stronger claim matters, persist the plan digest in an independently controlled/timestamped evidence system before EPAC derivation.
20 changes: 20 additions & 0 deletions epac_boundary_probe_completeness.py
Original file line number Diff line number Diff line change
Expand Up @@ -113,6 +113,7 @@
ObservableFn = Callable[[StateContext], Observable]

OPERATION_SOURCE_FILES = (
"epac_heldout_validation.py",
"epac_evidence_cache.py",
"viz/__init__.py",
"subatomic/__init__.py",
Expand Down Expand Up @@ -355,7 +356,22 @@ def _declared_operations() -> tuple[dict[str, str], ...]:
return tuple(sorted(operations, key=lambda item: item["operation"]))


_HELDOUT_OPERATION_RELEVANCE = {
"freeze_validation_plan": (AMBIGUOUS, "hmmm: case selection and comparator rules have no measured mapping to the frozen 27-state construction"),
"freeze_predictions": (AMBIGUOUS, "hmmm: externally derived predictions have no measured mapping to the frozen 27-state construction"),
"verify_commitment": (AMBIGUOUS, "hmmm: envelope integrity does not measure the prediction mapping to the frozen 27-state construction"),
"compare_after_freeze": (AMBIGUOUS, "hmmm: comparison results have no measured same-B discrimination on the frozen 27-state construction"),
"load_validation_plan": (AMBIGUOUS, "hmmm: persisted selection rules have no measured mapping to the frozen 27-state construction"),
"load_prediction_commitment": (AMBIGUOUS, "hmmm: reloaded prediction evidence has no measured mapping to the frozen 27-state construction"),
"load_validation_receipt": (AMBIGUOUS, "hmmm: reloaded comparison evidence has no measured mapping to the frozen 27-state construction"),
"load_oracle": (AMBIGUOUS, "hmmm: external reference cases have no measured mapping to the frozen 27-state construction"),
"load_packaged_oracle": (AMBIGUOUS, "hmmm: packaged reference cases have no measured mapping to the frozen 27-state construction"),
}


def _classify_operation(module: str, name: str) -> str:
if module == "epac_heldout_validation":
return _HELDOUT_OPERATION_RELEVANCE.get(name, (AMBIGUOUS, "hmmm: unclassified held-out operation"))[0]
if module == "epac_join_term":
# The join-tree candidate is not part of the frozen 27-state boundary
# quotient surface. Its possible boundary relevance remains unresolved
Expand Down Expand Up @@ -680,6 +696,10 @@ def declared_operation_ledger() -> tuple[OperationRecord, ...]:
),
}
)
if raw["module"] == "epac_heldout_validation":
records[-1]["boundary_relevance_reason"] = _HELDOUT_OPERATION_RELEVANCE.get(
raw["name"], (AMBIGUOUS, "hmmm: unclassified held-out operation")
)[1]
return tuple(records)


Expand Down
Loading
Loading