feat: backtest the forecast, then fix what it found — an ambient term and 4 τ fit windows - #34
Closed
zebraengine wants to merge 2 commits into
Closed
zebraengine wants to merge 2 commits into
zebraengine wants to merge 2 commits into
Conversation
Three model changes in one evening were each checked by replaying one or two hand-picked moments, and each found its counter-example within the hour. This scores every way the forecast can be made against what the handle actually did, over the whole history and per scenario, so a model change is judged on all of it at once. Every session is split into steady-current runs. A run that held its current for >= 3 tau yields a truth: an exponential fitted over the whole run with tau free, or (--truth last) the handle's own final three minutes - the two bracket the answer. Against it, two things are scored: the live forecast tick by tick as the current holds (trajectory basis vs model basis), and at every current change and session start the plateau that would have been predicted at the new current from each ambient on offer (sensor, previous run's trajectory, the warmer of the two, the idle handle) under each current law (I^2 prior vs fitted exponent), with the scored session left out of the fit. Errors are predicted minus actual: negative is optimistic, the direction that trips the charger. First run, 74 sessions, 40 scorable runs: the model-basis forecast is optimistic by 3-4 C on median and by more than 2 C in 79% of ticks; the mature trajectory by ~1 C; every cross-current method by 2.5-5 C at a step-down, under both truth definitions. The error grows with ambient - ~1 C on mild days, 3-5 C above 30 C - which is the ambient coefficient the drift regression had been reporting (0.35 C/C) and that its confidence interval alone did not make convincing. The fitted exponent beats I^2 everywhere scored. Whole-run tau runs 1-2 min longer than the 30 min fit windows' 11.25, which is where the trajectory's residual optimism comes from. No model change here. The next one gets scored by this before it ships. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…indows What the backtest found, fixed, and scored again before shipping. The sensor reads the garage air. The handle's environment runs hotter than the air by an amount that grows with the heat - sun on the wall, a heat-soaked structure and cable - and the model had no term for it. Over 74 sessions the model-basis forecast was optimistic by ~1 C on mild days and 3-5 C above 30 C, in 79% of ticks by more than 2 C, and every cross-current prediction (the number a restore is decided on) by 2.5-5 C. The degradation watch's regression had been reporting the same slope as an unexplained "ambient coefficient" (0.35 C/C) for weeks; its confidence interval alone had not made it convincing. Forty runs did. So rise(I, a) = rise_ref * (I/48)^n + k * (a - 25), with n and k fitted together from the free-running fits: n profiled over a grid, k and rise_ref by least squares at each n, each term only when the history spans enough to identify it (>= 6 A, >= 4 C). rise_ref_c now means the rise at 48 A and 25 C; every fit is re-normalized under the law; every site that turned an ambient into a plateau or back goes through plateau_at / ambient_from_plateau. On this install: n = 1.52 +/- 0.25, k = 0.25 +/- 0.13; a full-rate charge settles at 66 C at 28 C ambient and 75 C at 35 C, and the sustainable current reads 45 / 43 / 41 / 37 A at 28 / 30 / 32 / 35 C. The fit window also runs to 4 tau instead of 30 min. Whole-run fits read tau 1-2 min longer than the 30 min windows did, and that gap was the trajectory forecast's residual optimism: a window that ends at 2.7 tau has seen 93% of the rise and still trades a little tau for a little rise. tau moves 11.25 -> 12.0 min. Before -> after on production, as (whole-run-fit truth / model-free truth), bias in C and share optimistic by more than 2 C: model basis, in-run -3.3/-3.7 -> -2.1/-1.2 79%/80% -> 52%/20% step-down, from sensor -4.6/-4.0 -> -2.2/-1.3 92%/100% -> 64%/20% probe end, from sensor -3.8/-3.9 -> -0.6/-0.8 67%/100% -> 0%/0% trajectory at 10-20 min -1.0/-0.6 -> -0.6/+0.4 31%/31% -> 21%/18% The same run settled the restore-ambient question the anecdotes could not: at a step-down the sensor beats the ambient the previous run's trajectory implies, and "the warmer of the two" ties the sensor within 0.1 C. The sensor stays; nothing cleverer ships. Two fits that read free-running over 30 min windows read regulated over 4 tau ones (the sag shows), so the law is fitted from 12 free fits rather than 14, and the drift regression's residual scatter rises to 1.85 C - the three off-reference free fits still disagree with each other about the law, and the watch's threshold self-calibrates to 4.75 C until monthly probes pin it. Verdict unchanged: flat. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Owner
Author
|
Superseded by the contamination-controlled version. The ambient-term model in this PR was fitted to a biased sample (surviving and regulated runs); once cut-short runs are scored against their own projections the deployed model is unbiased at full-rate cold starts, hot days included, and the optimism is a heat-soak history effect. Model change parked on wip/ambient-term with the evidence. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
PRs #31–#33 were each validated by replaying one or two hand-picked moments, and each found its counter-example within the hour. There was no way to score a model change against the whole history —
contrib/backtest_derate_amp_control.pyreplays the controller against forecasts the server recorded at the time, not whether those forecasts were right.Built the tool first, then fixed what it found. Over 74 sessions the model-basis forecast was optimistic by 3–4 °C on median and by more than 2 °C in 79 % of ticks; every cross-current prediction (the number a restore is decided on) by 2.5–5 °C. The error grows with ambient: ~1 °C on mild days, 3–5 °C above 30 °C. The sensor reads the garage air; the handle's environment runs hotter than the air by an amount that scales with the heat, and the model had no term for it. The drift watch's regression had been reporting that slope (0.35 °C/°C) as an unexplained "ambient coefficient" for weeks. Separately, whole-run τ ran 1–2 min longer than the 30-min fit windows' 11.25 — the trajectory forecast's residual optimism.
Code touched
contrib/backtest_forecast.py(new) — splits every session into steady-current runs; a run held ≥ 3 τ yields a truth (--truth fit: exponential over the whole run, τ free;--truth last: the handle's own last 3 min, model-free — the two bracket). Scores the live forecast tick by tick (trajectory vs model basis) and, at every current change and session start, the plateau at the new current from each ambient (sensor / previous run's implied / warmer of the two / idle handle) under each law (I2,n,n+k), each fitted with the scored session left out. Scenario tags: cold_start, warm_start, step_down, step_up, probe, hot.--verbose,--json.wallmonitor/thermal.pyrise(I, a) = rise_ref·(I/48)^n + k·(a − 25).CurrentLaw+_fit_current_law()fit n and k together — n profiled over a grid, k and rise_ref by OLS at each n, n's SE from the SSE curvature; each term only when the history spans enough (≥ 6 A / ≥ 4 °C), else its prior (2 / 0). Either can be pinned (the backtest compares laws that way).params_from_fits()re-normalizes every fit under the law and takes medians;fit_history()wraps it.ThermalParams:ambient_coef,ambient_coef_se,plateau_at(),ambient_from_plateau(),safe_ambient_max_c(). Every site that turned an ambient into a plateau or back (_recent_steady_ambient, model basis, implied ambient, idle hypothetical,sustainable_max_current) goes through them.as_dict()exposesambient_coef,ambient_coef_se,ambient_ref_c.PREFIX_SPAN_TAU2.5 → 4.0 (the 30-min floor had made 2.5 moot).rise_ref_cnow means rise at 48 A and 25 °C.wallmonitor/static/app.js— model note states k when fitted.docs/thermal-model.md— ambient-term bullet; "Measuring the forecast" section with usage and the before/after table. Tests: backtest on a synthetic history with known truth + run splitter; k and n recovered together from a seeded k = 0.3 across 7 sessions, hot-day prediction lands, sustainable/safe-ambient shrink; seed helper takesambient_coef.Left alone: the controller (sensor-first restore ambient confirmed, see below),
suggested_max_a, the drift regression.Risk
safe_ambient_max_c28.7 → 27.6. Sustainable current 45 / 43 / 41 / 37 A at 28 / 30 / 32 / 35 °C. The remaining 1–2 °C optimism at a step-down is a history effect (cable still warm from the higher current);SUGGEST_MARGIN_C= 2 covers its median.fit_rmse_c0.40 → 0.50 with two fits at 0.59 against the 0.6 gate. Fit count unchanged at 19.Verification
uv run pytest: 155 passed (3 new).Deploy: pull +
sudo systemctl restart wallmonitor.service. Server-only; the controller reads the new numbers on its next tick.🤖 Generated with Claude Code