Scientific/product gap
Dynamic evaluation work now has Draft owner contracts for exact per-run item identity and provenance:
TEPP should later consume released/digest-pinned event projections from those owners to detect temporal change without copying their item-bank, provider, adjudication, or scoring authority.
Owner boundary
TEPP owns event/valid/assertion/availability/knowledge-cutoff semantics, temporal composition, longitudinal/multilevel/multiple-membership structure, drift/change-point evidence, and time-indexed invariance monitoring.
TEPP does not:
- generate evaluation items;
- call providers or store provider credentials;
- adjudicate observations;
- promote anchors or activate instruments;
- refit reusable static psychometric kernels already owned by fast-mlsirm;
- read another service database or consume mutable PR heads.
Required input contract
Do not implement against sibling source types. Wait for immutable released contracts and consume a versioned Anti-Corruption Layer carrying at minimum:
- exact evaluation run/item snapshot refs and contract digests;
- blueprint/rubric/criterion revisions;
- event, assertion, availability, and knowledge-cutoff clocks;
- generator and rater configuration/invocation refs;
- failed/abstained denominator states;
- adjudication case/resolution refs kept separate from source observations;
- calibration/validation/anchor-promotion/linking evidence refs;
- population/language/domain/testlet/context/multiple-membership refs and weights where owned upstream;
- explicit no-anchor/within-run/linked comparability status;
- source-text-free provenance by default.
Raw prompts, responses, source documents, generated item text, provider output, credentials, and tenant authorization records are not required by the canonical monitoring input.
Monitoring families
After sufficient longitudinal connectedness and sample evidence exists, support evidence-gated monitoring for:
- item/reference drift by exact item and successor lineage;
- rubric/criterion revision drift without conflating revisions on one scale;
- generator configuration and generation-failure drift;
- rater/facet severity, abstention, failure, and disagreement drift;
- adjudication-rate and resolution-pattern drift while retaining source observations;
- anchor stability and anchor-direction failures;
- DIF/invariance change across governed populations, languages, domains, and occasions;
- exposure/leakage or production-sample distribution change;
- blueprint coverage/composition change;
- comparability transitions from unavailable → within-run-only → linked, without back-casting unsupported historical linkage.
Fail-closed rules
- No fixed anchor set is required for pilot data collection, but zero-anchor runs cannot yield cross-version linked drift claims.
- A seed, model name, prompt revision, or nominal rubric label does not establish stable item identity.
- An adjudicated item is not automatically a validated anchor.
- Changed item/rubric/generator versions are explicit event strata or linked only through supplied evidence; they are never silently pooled.
- Availability time and knowledge cutoff gate every analysis so future adjudication/calibration/promotion evidence cannot leak into an earlier monitoring decision.
- Missing denominator events, incomplete membership weights, disconnected panels, absent anchors, unsupported linking, or insufficient repeated occasions produce typed unavailable evidence, not a passing flag.
- Statistical thresholds, window sizes, change-point penalties, and pooling structures require an explicit model/evidence contract; do not encode repository-authored universal heuristics.
Rust-first acceptance
Any production numerical path remains Rust-owned and requires:
- true-parameter or controlled synthetic recovery where a known truth exists;
- bias/RMSE/MAE and interval coverage as appropriate;
- convergence/failure-rate accounting including failed and abstained calls;
- deterministic seed manifests for simulation only, without claiming provider regeneration determinism;
- CPU
f64 reference behavior, bounded allocation, multithreading, and GPU parity when computationally material;
- permutation/order invariance where the estimand requires it;
- realistic irregular timing, missingness, panel turnover, item revision, anchor loss, and multiple-membership designs;
- comparison against static/no-change and misspecified baselines;
- no skip/xfail/source rewriting or threshold weakening to hide failed recovery.
Delivery gate
Keep this issue in design/evidence-gathering status until upstream immutable releases exist and Psychometrics Commons has durable run/item/panel/adjudication snapshots. The first implementation slice should be a versioned input/event ACL plus no-anchor/no-linking fail-closed tests—not a drift score or dashboard.
This issue does not authorize production monitoring from current Draft PR heads.
Scientific/product gap
Dynamic evaluation work now has Draft owner contracts for exact per-run item identity and provenance:
TEPP should later consume released/digest-pinned event projections from those owners to detect temporal change without copying their item-bank, provider, adjudication, or scoring authority.
Owner boundary
TEPP owns event/valid/assertion/availability/knowledge-cutoff semantics, temporal composition, longitudinal/multilevel/multiple-membership structure, drift/change-point evidence, and time-indexed invariance monitoring.
TEPP does not:
Required input contract
Do not implement against sibling source types. Wait for immutable released contracts and consume a versioned Anti-Corruption Layer carrying at minimum:
Raw prompts, responses, source documents, generated item text, provider output, credentials, and tenant authorization records are not required by the canonical monitoring input.
Monitoring families
After sufficient longitudinal connectedness and sample evidence exists, support evidence-gated monitoring for:
Fail-closed rules
Rust-first acceptance
Any production numerical path remains Rust-owned and requires:
f64reference behavior, bounded allocation, multithreading, and GPU parity when computationally material;Delivery gate
Keep this issue in design/evidence-gathering status until upstream immutable releases exist and Psychometrics Commons has durable run/item/panel/adjudication snapshots. The first implementation slice should be a versioned input/event ACL plus no-anchor/no-linking fail-closed tests—not a drift score or dashboard.
This issue does not authorize production monitoring from current Draft PR heads.