Skip to content

feat: backtest the forecast against every recorded session, with its sampling bias controlled - #35

Merged
zebraengine merged 2 commits into
mainfrom
feat/forecast-backtest-clean
Sep 15, 2026
Merged

zebraengine merged 2 commits into
mainfrom
feat/forecast-backtest-clean

Conversation

@zebraengine

Copy link
Copy Markdown
Owner

Problem

PRs #31–#33 were each validated by replaying one or two hand-picked moments from the production DB, and each found its counter-example within the hour. There was no way to score a model change against the whole history; contrib/backtest_derate_amp_control.py replays the controller against forecasts the server recorded, not whether those forecasts were right.

A first cut of this tool then produced a wrong conclusion of its own: scored on the runs that reached a plateau, the model read 3–4 °C optimistic, worst on hot days, and an ambient term fitted to that looked like the fix. The sample was the problem. The current and ambient every run happened at were chosen by the controller and the charger on the strength of this same model:

  1. Regulated runs. The charger trims current as the handle nears 65 °C and holds it there — a setpoint, not an equilibrium. 20 of 52 long-enough runs.
  2. Censoring. A run whose true plateau was above 65 °C tripped and ended before it could become truth. The optimistic errors that matter never reach the tables. 14 runs.
  3. Confounding. The controller caps on hot days, so hot and low current arrive together; an ambient effect and a current-law error wear the same signature.

Code touched

contrib/backtest_forecast.py (new)

  • Splits sessions into steady-current runs (same band rule as the live forecast). Truth = plateau of a run held ≥ 3 τ: --truth fit (whole-run exponential, τ free) or --truth last (handle's own last 3 min, model-free); the two bracket.
  • Contamination controls: runs whose current sagged > 1.5 % (_current_sag_a / _free_plateau, the fitter's own gate) are never truth and are counted; runs with an alert 40 raised inside them are never truth and are listed with every method's prediction (a prediction under 65 °C is a miss the tables cannot show); every free-running run cut short after ≥ 1 τ — capped, tripped, ended — is scored against its own trajectory projection with the projection's SE (median 0.03 °C), as a proxy truth. The projection needs no ambient and no current law.
  • Scores: in-run forecast tick by tick (trajectory vs model basis) by minutes-at-current and scenario; cross-current prediction at every change and session start from each ambient (sensor / previous run's implied / warmer of the two / idle handle) × law (I² / fitted n), scored session left out of the fit. Scenario tags: cold_start, warm_start, step_down, step_up, probe, hot. --verbose, --json.

tests/test_backtest_forecast.py: synthetic history with known truth (cold start then step-down, both ≥ 3 τ) — runs found and classified, plateaus recovered within 1 °C, trajectory converges, every cross-current method within 1.5 °C at the step, start scored from sensor + idle; run splitter on the session-80 ramp and a contactor drop.

docs/thermal-model.md: "Measuring the forecast" — usage, the three contamination channels and what is done about each, and the findings.

No model change. The ambient-term model built on the first (contaminated) pass is parked on wip/ambient-term with the evidence against it in this PR.

Risk

Read-only tool; nothing in the running system changes.

Verification

  • uv run pytest: 154 passed (2 new).
  • Production DB, 74 sessions, 312 runs. De-censored findings (fit truth / model-free truth):
    • Full-rate cold starts, hot days included: the deployed model is unbiased. Sensor: +0.3 °C median, |med| 0.7, 0 % optimistic by > 2 °C (7 runs); idle proxy: +0.7, |med| 1.0 (18 runs). e.g. s47 at 31 °C ambient: predicted 67.8, projected 67.8.
    • The optimism is a history effect. Step-downs after a hot full-rate run: −2.8 / −3.6 °C from the sensor; −0.7 / −1.4 from the previous run's implied ambient (carries the heat-soaked cable). The two worst clean truths (39.6 A "cold starts" on hot evenings) and the July-heat-wave probes (−7 °C, right after a trip) fit the same story.
    • The deployed configuration is on the right side of it: caps from the implied ambient (right for step-downs), restores from the sensor — at the 6 de-censored step-ups the sensor is unbiased-to-pessimistic (+2.6), and at session 80's own step-up predicted 64.9 vs projected 64.8. fix: work the sustainable current from whichever ambient leaves less headroom #33's "warmer" rule would have cost amps on restores.
    • Fitted exponent beats I² in every de-censored cell below 48 A (step-downs |med| 2.4 vs 3.9 from the implied ambient; probes 2.5 vs 3.7). Keep.
    • Trip table: of 6 real full-rate trips (all July, idle-proxy ambient), the deployed model predicted a trip for 4; the two misses were within 1.5 °C (63.5, 62.9).
  • Next, if the history effect is worth chasing: a second, slow time constant for the cable — which needs designed data (probes at one current with a cold cable and a warm one), not more history.

🤖 Generated with Claude Code

zebraengine and others added 2 commits September 14, 2026 20:59
Three model changes in one evening were each checked by replaying one or
two hand-picked moments, and each found its counter-example within the
hour. This scores every way the forecast can be made against what the
handle actually did, over the whole history and per scenario, so a model
change is judged on all of it at once.

Every session is split into steady-current runs. A run that held its
current for >= 3 tau yields a truth: an exponential fitted over the whole
run with tau free, or (--truth last) the handle's own final three minutes -
the two bracket the answer. Against it, two things are scored: the live
forecast tick by tick as the current holds (trajectory basis vs model
basis), and at every current change and session start the plateau that
would have been predicted at the new current from each ambient on offer
(sensor, previous run's trajectory, the warmer of the two, the idle handle)
under each current law (I^2 prior vs fitted exponent), with the scored
session left out of the fit. Errors are predicted minus actual: negative is
optimistic, the direction that trips the charger.

First run, 74 sessions, 40 scorable runs: the model-basis forecast is
optimistic by 3-4 C on median and by more than 2 C in 79% of ticks; the
mature trajectory by ~1 C; every cross-current method by 2.5-5 C at a
step-down, under both truth definitions. The error grows with ambient -
~1 C on mild days, 3-5 C above 30 C - which is the ambient coefficient the
drift regression had been reporting (0.35 C/C) and that its confidence
interval alone did not make convincing. The fitted exponent beats I^2
everywhere scored. Whole-run tau runs 1-2 min longer than the 30 min fit
windows' 11.25, which is where the trajectory's residual optimism comes
from.

No model change here. The next one gets scored by this before it ships.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ainst their own projection

The history is not a neutral sample. The current and ambient every run
happened at were chosen by the controller and the charger on the strength
of this same model, which contaminates the truths three ways: a run whose
current the charger trimmed has a plateau that is a setpoint, not an
equilibrium; a run whose true plateau lay above the trip point tripped and
ended before it could become truth, censoring exactly the optimistic
errors the tool exists to find; and because the controller caps on hot
days, hot and low-current arrive together, so an ambient effect and a
current-law error wear the same signature.

So: regulated runs (sag > 1.5%) are excluded from truth and counted;
every run that tripped is listed with what each method predicted for it;
and every free-running run cut short after >= 1 tau - capped, tripped, or
ended - is scored against its own trajectory projection, SE alongside, as
a proxy truth. The projection needs no ambient and no current law.

On 74 sessions that reverses the first pass. Clean truths alone read the
model as 3-4 C optimistic, worst on hot days, and an ambient term fitted
to that looked like the fix. De-censored: at full-rate cold starts, hot
days included, the deployed model is unbiased (+0.3 C from the sensor,
|med| 0.7, 0% optimistic by > 2 C over 7 runs; +0.7 from the idle proxy
over 18). The optimism is a history effect - a step-down right after a hot
full-rate run reads 2.8-3.6 C optimistic from the sensor, and 0.7-1.4 C
from the previous run's implied ambient, which carries the heat-soaked
cable - and the deployed configuration (caps from the implied ambient,
restores from the sensor) already sits on the right side of it. The
ambient term is parked on wip/ambient-term with the evidence against it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@zebraengine
zebraengine merged commit 88ab36e into main Sep 15, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant