Skip to content

feat: backtest the forecast, then fix what it found — an ambient term and 4 τ fit windows - #34

Closed
zebraengine wants to merge 2 commits into
mainfrom
feat/forecast-backtest
Closed

zebraengine wants to merge 2 commits into
mainfrom
feat/forecast-backtest

Conversation

@zebraengine

@zebraengine zebraengine commented Sep 15, 2026 •

Copy link
Copy Markdown
Owner

Problem

PRs #31–#33 were each validated by replaying one or two hand-picked moments, and each found its counter-example within the hour. There was no way to score a model change against the whole history — contrib/backtest_derate_amp_control.py replays the controller against forecasts the server recorded at the time, not whether those forecasts were right.

Built the tool first, then fixed what it found. Over 74 sessions the model-basis forecast was optimistic by 3–4 °C on median and by more than 2 °C in 79 % of ticks; every cross-current prediction (the number a restore is decided on) by 2.5–5 °C. The error grows with ambient: ~1 °C on mild days, 3–5 °C above 30 °C. The sensor reads the garage air; the handle's environment runs hotter than the air by an amount that scales with the heat, and the model had no term for it. The drift watch's regression had been reporting that slope (0.35 °C/°C) as an unexplained "ambient coefficient" for weeks. Separately, whole-run τ ran 1–2 min longer than the 30-min fit windows' 11.25 — the trajectory forecast's residual optimism.

Code touched

contrib/backtest_forecast.py (new) — splits every session into steady-current runs; a run held ≥ 3 τ yields a truth (--truth fit: exponential over the whole run, τ free; --truth last: the handle's own last 3 min, model-free — the two bracket). Scores the live forecast tick by tick (trajectory vs model basis) and, at every current change and session start, the plateau at the new current from each ambient (sensor / previous run's implied / warmer of the two / idle handle) under each law (I2, n, n+k), each fitted with the scored session left out. Scenario tags: cold_start, warm_start, step_down, step_up, probe, hot. --verbose, --json.

wallmonitor/thermal.py

  • Rise law gains an ambient term: rise(I, a) = rise_ref·(I/48)^n + k·(a − 25). CurrentLaw + _fit_current_law() fit n and k together — n profiled over a grid, k and rise_ref by OLS at each n, n's SE from the SSE curvature; each term only when the history spans enough (≥ 6 A / ≥ 4 °C), else its prior (2 / 0). Either can be pinned (the backtest compares laws that way). params_from_fits() re-normalizes every fit under the law and takes medians; fit_history() wraps it.
  • ThermalParams: ambient_coef, ambient_coef_se, plateau_at(), ambient_from_plateau(), safe_ambient_max_c(). Every site that turned an ambient into a plateau or back (_recent_steady_ambient, model basis, implied ambient, idle hypothetical, sustainable_max_current) goes through them. as_dict() exposes ambient_coef, ambient_coef_se, ambient_ref_c.
  • PREFIX_SPAN_TAU 2.5 → 4.0 (the 30-min floor had made 2.5 moot).
  • rise_ref_c now means rise at 48 A and 25 °C.

wallmonitor/static/app.js — model note states k when fitted. docs/thermal-model.md — ambient-term bullet; "Measuring the forecast" section with usage and the before/after table. Tests: backtest on a synthetic history with known truth + run splitter; k and n recovered together from a seeded k = 0.3 across 7 sessions, hot-day prediction lands, sustainable/safe-ambient shrink; seed helper takes ambient_coef.

Left alone: the controller (sensor-first restore ambient confirmed, see below), suggested_max_a, the drift regression.

Risk

  • Forecasts get more conservative on hot days, by design: on this install a full-rate charge now reads 66 °C at 28 °C ambient (was 64.8) and safe_ambient_max_c 28.7 → 27.6. Sustainable current 45 / 43 / 41 / 37 A at 28 / 30 / 32 / 35 °C. The remaining 1–2 °C optimism at a step-down is a history effect (cable still warm from the higher current); SUGGEST_MARGIN_C = 2 covers its median.
  • Longer windows: two fits that read free-running over 30 min read regulated over 4 τ (the sag shows), so the law is fitted from 12 free fits, and fit_rmse_c 0.40 → 0.50 with two fits at 0.59 against the 0.6 gate. Fit count unchanged at 19.
  • Drift watch residual scatter 1.44 → 1.85 °C, threshold self-calibrates to 4.75 °C: the three off-reference free fits (39.6 A ×2 on hot days, 32.5 A probe) still disagree about the law even with k (n = 1.52 ± 0.25). Verdict unchanged (delta −0.42, flat, not compromised). Monthly probes are what pin it.
  • Installs without a sensor: k is fitted from idle-handle ambients, whose error the (1 + k) term amplifies slightly; the backtest can be run per install. Fresh installs: priors unchanged (n = 2, k = 0).
  • API: additive fields only. Older controllers unaffected.

Verification

  • uv run pytest: 155 passed (3 new).
  • Backtest on production, before → after, as (whole-run-fit truth / model-free truth):
bias, °C optimistic by > 2 °C
model basis, in-run −3.3 / −3.7 → −2.1 / −1.2 79 % / 80 % → 52 % / 20 %
step-down, from the sensor −4.6 / −4.0 → −2.2 / −1.3 92 % / 100 % → 64 % / 20 %
probe end, from the sensor −3.8 / −3.9 → −0.6 / −0.8 67 % / 100 % → 0 % / 0 %
trajectory, 10–20 min in −1.0 / −0.6 → −0.6 / +0.4 31 % / 31 % → 21 % / 18 %

Deploy: pull + sudo systemctl restart wallmonitor.service. Server-only; the controller reads the new numbers on its next tick.

🤖 Generated with Claude Code

Three model changes in one evening were each checked by replaying one or
two hand-picked moments, and each found its counter-example within the
hour. This scores every way the forecast can be made against what the
handle actually did, over the whole history and per scenario, so a model
change is judged on all of it at once.

Every session is split into steady-current runs. A run that held its
current for >= 3 tau yields a truth: an exponential fitted over the whole
run with tau free, or (--truth last) the handle's own final three minutes -
the two bracket the answer. Against it, two things are scored: the live
forecast tick by tick as the current holds (trajectory basis vs model
basis), and at every current change and session start the plateau that
would have been predicted at the new current from each ambient on offer
(sensor, previous run's trajectory, the warmer of the two, the idle handle)
under each current law (I^2 prior vs fitted exponent), with the scored
session left out of the fit. Errors are predicted minus actual: negative is
optimistic, the direction that trips the charger.

First run, 74 sessions, 40 scorable runs: the model-basis forecast is
optimistic by 3-4 C on median and by more than 2 C in 79% of ticks; the
mature trajectory by ~1 C; every cross-current method by 2.5-5 C at a
step-down, under both truth definitions. The error grows with ambient -
~1 C on mild days, 3-5 C above 30 C - which is the ambient coefficient the
drift regression had been reporting (0.35 C/C) and that its confidence
interval alone did not make convincing. The fitted exponent beats I^2
everywhere scored. Whole-run tau runs 1-2 min longer than the 30 min fit
windows' 11.25, which is where the trajectory's residual optimism comes
from.

No model change here. The next one gets scored by this before it ships.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…indows

What the backtest found, fixed, and scored again before shipping.

The sensor reads the garage air. The handle's environment runs hotter than
the air by an amount that grows with the heat - sun on the wall, a
heat-soaked structure and cable - and the model had no term for it. Over
74 sessions the model-basis forecast was optimistic by ~1 C on mild days
and 3-5 C above 30 C, in 79% of ticks by more than 2 C, and every
cross-current prediction (the number a restore is decided on) by 2.5-5 C.
The degradation watch's regression had been reporting the same slope as an
unexplained "ambient coefficient" (0.35 C/C) for weeks; its confidence
interval alone had not made it convincing. Forty runs did.

So rise(I, a) = rise_ref * (I/48)^n + k * (a - 25), with n and k fitted
together from the free-running fits: n profiled over a grid, k and
rise_ref by least squares at each n, each term only when the history
spans enough to identify it (>= 6 A, >= 4 C). rise_ref_c now means the
rise at 48 A and 25 C; every fit is re-normalized under the law; every
site that turned an ambient into a plateau or back goes through
plateau_at / ambient_from_plateau. On this install: n = 1.52 +/- 0.25,
k = 0.25 +/- 0.13; a full-rate charge settles at 66 C at 28 C ambient and
75 C at 35 C, and the sustainable current reads 45 / 43 / 41 / 37 A at
28 / 30 / 32 / 35 C.

The fit window also runs to 4 tau instead of 30 min. Whole-run fits read
tau 1-2 min longer than the 30 min windows did, and that gap was the
trajectory forecast's residual optimism: a window that ends at 2.7 tau
has seen 93% of the rise and still trades a little tau for a little rise.
tau moves 11.25 -> 12.0 min.

Before -> after on production, as (whole-run-fit truth / model-free
truth), bias in C and share optimistic by more than 2 C:
  model basis, in-run       -3.3/-3.7 -> -2.1/-1.2   79%/80% -> 52%/20%
  step-down, from sensor    -4.6/-4.0 -> -2.2/-1.3   92%/100% -> 64%/20%
  probe end, from sensor    -3.8/-3.9 -> -0.6/-0.8   67%/100% -> 0%/0%
  trajectory at 10-20 min   -1.0/-0.6 -> -0.6/+0.4   31%/31% -> 21%/18%

The same run settled the restore-ambient question the anecdotes could
not: at a step-down the sensor beats the ambient the previous run's
trajectory implies, and "the warmer of the two" ties the sensor within
0.1 C. The sensor stays; nothing cleverer ships.

Two fits that read free-running over 30 min windows read regulated over
4 tau ones (the sag shows), so the law is fitted from 12 free fits rather
than 14, and the drift regression's residual scatter rises to 1.85 C - the
three off-reference free fits still disagree with each other about the
law, and the watch's threshold self-calibrates to 4.75 C until monthly
probes pin it. Verdict unchanged: flat.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@zebraengine zebraengine changed the title feat: backtest the forecast against every recorded session feat: backtest the forecast, then fix what it found — an ambient term and 4 τ fit windows Sep 15, 2026
@zebraengine

Copy link
Copy Markdown
Owner Author

Superseded by the contamination-controlled version. The ambient-term model in this PR was fitted to a biased sample (surviving and regulated runs); once cut-short runs are scored against their own projections the deployed model is unbiased at full-rate cold starts, hot days included, and the optimism is a heat-soak history effect. Model change parked on wip/ambient-term with the evidence.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant