feat: backtest the forecast against every recorded session, with its sampling bias controlled - #35
Merged
Merged
Conversation
Three model changes in one evening were each checked by replaying one or two hand-picked moments, and each found its counter-example within the hour. This scores every way the forecast can be made against what the handle actually did, over the whole history and per scenario, so a model change is judged on all of it at once. Every session is split into steady-current runs. A run that held its current for >= 3 tau yields a truth: an exponential fitted over the whole run with tau free, or (--truth last) the handle's own final three minutes - the two bracket the answer. Against it, two things are scored: the live forecast tick by tick as the current holds (trajectory basis vs model basis), and at every current change and session start the plateau that would have been predicted at the new current from each ambient on offer (sensor, previous run's trajectory, the warmer of the two, the idle handle) under each current law (I^2 prior vs fitted exponent), with the scored session left out of the fit. Errors are predicted minus actual: negative is optimistic, the direction that trips the charger. First run, 74 sessions, 40 scorable runs: the model-basis forecast is optimistic by 3-4 C on median and by more than 2 C in 79% of ticks; the mature trajectory by ~1 C; every cross-current method by 2.5-5 C at a step-down, under both truth definitions. The error grows with ambient - ~1 C on mild days, 3-5 C above 30 C - which is the ambient coefficient the drift regression had been reporting (0.35 C/C) and that its confidence interval alone did not make convincing. The fitted exponent beats I^2 everywhere scored. Whole-run tau runs 1-2 min longer than the 30 min fit windows' 11.25, which is where the trajectory's residual optimism comes from. No model change here. The next one gets scored by this before it ships. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ainst their own projection The history is not a neutral sample. The current and ambient every run happened at were chosen by the controller and the charger on the strength of this same model, which contaminates the truths three ways: a run whose current the charger trimmed has a plateau that is a setpoint, not an equilibrium; a run whose true plateau lay above the trip point tripped and ended before it could become truth, censoring exactly the optimistic errors the tool exists to find; and because the controller caps on hot days, hot and low-current arrive together, so an ambient effect and a current-law error wear the same signature. So: regulated runs (sag > 1.5%) are excluded from truth and counted; every run that tripped is listed with what each method predicted for it; and every free-running run cut short after >= 1 tau - capped, tripped, or ended - is scored against its own trajectory projection, SE alongside, as a proxy truth. The projection needs no ambient and no current law. On 74 sessions that reverses the first pass. Clean truths alone read the model as 3-4 C optimistic, worst on hot days, and an ambient term fitted to that looked like the fix. De-censored: at full-rate cold starts, hot days included, the deployed model is unbiased (+0.3 C from the sensor, |med| 0.7, 0% optimistic by > 2 C over 7 runs; +0.7 from the idle proxy over 18). The optimism is a history effect - a step-down right after a hot full-rate run reads 2.8-3.6 C optimistic from the sensor, and 0.7-1.4 C from the previous run's implied ambient, which carries the heat-soaked cable - and the deployed configuration (caps from the implied ambient, restores from the sensor) already sits on the right side of it. The ambient term is parked on wip/ambient-term with the evidence against it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
PRs #31–#33 were each validated by replaying one or two hand-picked moments from the production DB, and each found its counter-example within the hour. There was no way to score a model change against the whole history;
contrib/backtest_derate_amp_control.pyreplays the controller against forecasts the server recorded, not whether those forecasts were right.A first cut of this tool then produced a wrong conclusion of its own: scored on the runs that reached a plateau, the model read 3–4 °C optimistic, worst on hot days, and an ambient term fitted to that looked like the fix. The sample was the problem. The current and ambient every run happened at were chosen by the controller and the charger on the strength of this same model:
Code touched
contrib/backtest_forecast.py(new)--truth fit(whole-run exponential, τ free) or--truth last(handle's own last 3 min, model-free); the two bracket._current_sag_a/_free_plateau, the fitter's own gate) are never truth and are counted; runs with an alert 40 raised inside them are never truth and are listed with every method's prediction (a prediction under 65 °C is a miss the tables cannot show); every free-running run cut short after ≥ 1 τ — capped, tripped, ended — is scored against its own trajectory projection with the projection's SE (median 0.03 °C), as a proxy truth. The projection needs no ambient and no current law.--verbose,--json.tests/test_backtest_forecast.py: synthetic history with known truth (cold start then step-down, both ≥ 3 τ) — runs found and classified, plateaus recovered within 1 °C, trajectory converges, every cross-current method within 1.5 °C at the step, start scored from sensor + idle; run splitter on the session-80 ramp and a contactor drop.docs/thermal-model.md: "Measuring the forecast" — usage, the three contamination channels and what is done about each, and the findings.No model change. The ambient-term model built on the first (contaminated) pass is parked on
wip/ambient-termwith the evidence against it in this PR.Risk
Read-only tool; nothing in the running system changes.
Verification
uv run pytest: 154 passed (2 new).🤖 Generated with Claude Code