You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A single CEFR-aligned assessment result can support placement, but repeated results cannot be interpreted as language growth until form, task, rater, scoring, mode, population, and language-profile comparability have been evaluated. TEPP is the correct owner for temporal, multilevel, relational, and multiple-membership analysis of repeated language profiles.
This issue owns longitudinal and contextual language-profile analysis artifacts. It does not own assessment sessions, raw responses, task content, CEFR descriptor prose, cross-sectional scoring, cut scores, or LMS decisions.
multiple membership when one learner/occasion belongs to multiple relevant contexts;
missing-by-design domains and selective retesting;
correction/supersession without historical mutation.
Do not treat ordinal CEFR labels as equally spaced numeric observations. Use the underlying versioned domain-score/posterior evidence or appropriate latent-state measurement model.
Rust arithmetic ownership
All result-affecting temporal and multilevel arithmetic is Rust-owned:
longitudinal state transition;
random intercept/slope and context effects;
continuous-time or interval-aware transition where adopted;
invariance/drift tests and uncertainty;
change-point/onset detection where scientifically supported;
trajectory estimation and prediction intervals;
true-state simulation and recovery;
CPU multithreading and GPU execution where benchmark evidence justifies it.
Python may validate contracts, orchestrate jobs, and render reports only.
First bounded vertical
Implement a research-only English A1–B2 repeated-placement analysis with:
four domains: reading, listening, written production, spoken production;
at least three occasions per simulated participant;
irregular time gaps;
two form releases;
human and AI-rater-version drift in productive domains;
course/cohort/instructor cross-classification;
one multiple-membership sponsor/context dimension;
delayed result availability and one later correction;
explicit no-change verdict when invariance or linking evidence fails.
Every artifact pins exact source result identities, model/engine versions, data cutoff, membership revision, failure denominator, uncertainty, and limitations.
Recovery evidence
The simulator must inject known truth for:
stable person differences and within-person change;
domain correlations;
form/linking shifts;
item/task/rater drift;
occasion effects;
context random effects;
multiple-membership weights;
irregular lags;
missingness and delayed availability;
correction events.
Report bias/RMSE, interval coverage, state/transition recovery, drift-detection precision/recall, false change declarations, backend parity, worker-count determinism, failure denominator, and Monte Carlo uncertainty.
Acceptance
No trajectory or growth claim is emitted when required invariance/linking evidence fails.
No future-available result or correction enters an earlier cutoff.
Ordinal CEFR labels are never averaged or treated as an interval scale.
Within-person change is separated from stable between-person/context differences.
Cross-classified and multiple-membership design is explicit; no primary_group shortcut.
Historical results and trajectories are immutable; correction creates superseding artifacts.
Every number is reproducible from exact input and model artifacts.
Production statement/branch coverage and public API docstrings are 100%.
True-state recovery, CPU/GPU parity, memory/throughput evidence, and no skipped required GPU/recovery tests.
Visual outputs include exact-value accessible tables, uncertainty, provenance, data cutoff, and no causal language unless a separate causal design supports it.
No Council of Europe endorsement, validation, certification, or logo-use claim.
Standards and research baseline
CEFR Companion Volume (2020).
Council of Europe examination-linking and test-development manuals.
AERA, APA, & NCME Standards (2014).
Longitudinal measurement-invariance, DSEM/continuous-time, multilevel, cross-classified, and multiple-membership primary literature documented in APA 7th form.
TEPP may measure change in a governed language-profile system; it must not turn repeated labels into growth by assumption.
Temporal measurement gap
A single CEFR-aligned assessment result can support placement, but repeated results cannot be interpreted as language growth until form, task, rater, scoring, mode, population, and language-profile comparability have been evaluated. TEPP is the correct owner for temporal, multilevel, relational, and multiple-membership analysis of repeated language profiles.
This issue owns longitudinal and contextual language-profile analysis artifacts. It does not own assessment sessions, raw responses, task content, CEFR descriptor prose, cross-sectional scoring, cut scores, or LMS decisions.
Dependencies
Blocked by:
cwl_cefr_language_assessment/v1fromContextualWisdomLab/learning-interoperability-contractsPR feat(temporal): add typed six-clock values and uncertain intervals #5;Do not build this as a Python post-hoc regression over copied result JSON.
Analysis boundary
Consume only versioned, immutable result references and approved contextual membership references. Preserve at least:
Raw responses, audio, transcripts, task content, rater payloads, and PII remain in their owning systems.
Multi-clock and leakage contract
For every analysis cutoff
T, require:Separate:
A later correction, rescoring run, cut-score revision, or RLD update must not leak into an earlier historical analysis.
Measurement structure
Evaluate, rather than assume:
Do not treat ordinal CEFR labels as equally spaced numeric observations. Use the underlying versioned domain-score/posterior evidence or appropriate latent-state measurement model.
Rust arithmetic ownership
All result-affecting temporal and multilevel arithmetic is Rust-owned:
Python may validate contracts, orchestrate jobs, and render reports only.
First bounded vertical
Implement a research-only English A1–B2 repeated-placement analysis with:
Outputs:
Every artifact pins exact source result identities, model/engine versions, data cutoff, membership revision, failure denominator, uncertainty, and limitations.
Recovery evidence
The simulator must inject known truth for:
Report bias/RMSE, interval coverage, state/transition recovery, drift-detection precision/recall, false change declarations, backend parity, worker-count determinism, failure denominator, and Monte Carlo uncertainty.
Acceptance
primary_groupshortcut.Standards and research baseline
TEPP may measure change in a governed language-profile system; it must not turn repeated labels into growth by assumption.