Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
89 changes: 64 additions & 25 deletions contrib/derate_amp_control.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,14 +21,24 @@

**Restoring up is the risky direction** — it's what pushes the equilibrium
back toward the trip point — so it stays conservative on every axis: only
``trajectory`` basis (this session's own proven data, never ``model``), only
one ``--restore-step-a`` at a time rather than snapping straight back to
``--normal-amps``, and never while the handle is still within
``--restore-margin-c`` of the trip point even if the trajectory reads clear.
The same 2026-08-03 session showed why: capping straight back to 48A the
moment a trajectory read clear, twice, immediately restarted the climb both
times, converting a caught derate into three near-misses before the third
one wasn't caught in time.
``trajectory`` basis (this session's own proven data, never ``model``), never
while the handle is still within ``--restore-margin-c`` of the trip point
even if the trajectory reads clear, and never straight back to
``--normal-amps`` on trust. The same 2026-08-03 session showed why: capping
straight back to 48A the moment a trajectory read clear, twice, immediately
restarted the climb both times, converting a caught derate into three
near-misses before the third one wasn't caught in time.

Where it restores *to* is the server's ``sustainable_max_a``: the highest
current whose modelled plateau stays under the trip point at today's ambient
(the LAN sensor when one reports, else the ambient the live trajectory
implies). One move there, then the trajectory and the confidence guard trim
the last amp or two. The alternative — climbing
``--restore-step-a`` at a time — resets the trajectory window at every rung
and took half an hour to find the same number. The model is trusted once per
session: after a quick reversal it has already been wrong about today, and
the climb falls back to single ``--restore-step-a`` steps (and to that ladder
entirely against a server that doesn't report the field).

Earlier live testing (2026-08-01, a full 48A session) is why ``hypothetical``
basis is never trusted at all: it leans on the historical per-install
Expand All @@ -43,8 +53,7 @@
A cap fully lifts three ways, in order of how eagerly they should fire:
1. the trajectory forecast reports ``will_trip: false`` for
``--confirm-ticks`` consecutive polls *and* the handle has real margin
below the trip point, stepped up ``--restore-step-a`` at a time — see
above;
below the trip point, restored to the sustainable current — see above;
2. the charging session ends (``state`` leaves ``charging``) — the normal,
expected end of any cap, restored immediately since there's no more
climb to protect against;
Expand Down Expand Up @@ -346,8 +355,8 @@ def _decide_thermal(thermal: dict, state: State, cfg: Config) -> tuple[Action, S
)

# Restore path: deliberately narrower than the cap path. Only
# `trajectory` basis (never `model`), only a step at a time, gated on
# real thermal margin, and backed off exponentially after repeated
# `trajectory` basis (never `model`), never on trust to full rate, gated
# on real thermal margin, and backed off exponentially after repeated
# quick reversals — see the module docstring for the incidents that
# justify every one of these guards.
if basis == "trajectory" and will_trip is False:
Expand Down Expand Up @@ -382,7 +391,16 @@ def _decide_thermal(thermal: dict, state: State, cfg: Config) -> tuple[Action, S
f"(need {cfg.restore_margin_c:g}C): holding {state.cap_value:g}A"
),
)
next_value = min(cfg.normal_amps, state.cap_value + cfg.restore_step_a)
# One move to the model's sustainable current when the server
# reports one, instead of a 2 A ladder that resets the trajectory
# window at every rung. The model is trusted once per session:
# after a quick reversal it has already been wrong about this
# session, so the climb falls back to single steps.
step_value = state.cap_value + cfg.restore_step_a
sustainable = forecast.get("sustainable_max_a")
jump = isinstance(sustainable, (int, float)) and state.restore_attempts == 0
next_value = min(cfg.normal_amps, max(step_value, float(sustainable)) if jump else step_value)
how = "jumping" if jump and next_value > step_value else "stepping up"
step_state = replace(next_state, clear_streak=0, last_step_up_ts=now_ts)
if next_value >= cfg.normal_amps:
final_state = replace(step_state, capped=False, cap_value=None)
Expand All @@ -396,7 +414,7 @@ def _decide_thermal(thermal: dict, state: State, cfg: Config) -> tuple[Action, S
Action("cap", next_value),
final_state,
(
f"trajectory clear, {margin_c:.1f}C of margin: stepping up to {next_value:g}A "
f"trajectory clear, {margin_c:.1f}C of margin: {how} to {next_value:g}A "
f"(still under {cfg.normal_amps:g}A)"
),
)
Expand Down Expand Up @@ -470,17 +488,37 @@ def _apply_probe(
return Action("none"), new_state, "probe holding (no timestamp to age it against)"
held_min = (now_ts - prev.probe_started_ts) / 60.0
if held_min >= cfg.probe_hold_min:
done = replace(
new_state, probe_started_ts=None, last_probe_ts=now_ts, clear_streak=0, trip_streak=0
)
# The probe's own plateau is the best measurement of today's
# conditions the session will ever have, and the server has
# already turned it into the sustainable current. Go there.
# Snapping to normal_amps instead, on a day the model already
# knew full rate would trip, cost a 32 -> 48 -> 42 A oscillation.
sustainable = (thermal.get("forecast") or {}).get("sustainable_max_a")
if isinstance(sustainable, (int, float)) and sustainable < cfg.normal_amps:
target = float(sustainable)
if target <= cfg.probe_amps:
return (
Action("none"),
replace(done, capped=True, cap_value=cfg.probe_amps),
(
f"probe complete: held {cfg.probe_amps:g}A for {held_min:.0f}min; "
f"sustainable {target:g}A is no higher, holding"
),
)
return (
Action("cap", target),
replace(done, capped=True, cap_value=target, last_step_up_ts=now_ts),
(
f"probe complete: held {cfg.probe_amps:g}A for {held_min:.0f}min, "
f"restoring to the sustainable {target:g}A"
),
)
return (
Action("restore", cfg.normal_amps),
replace(
new_state,
probe_started_ts=None,
last_probe_ts=now_ts,
capped=False,
cap_value=None,
clear_streak=0,
trip_streak=0,
),
replace(done, capped=False, cap_value=None),
(
f"probe complete: held {cfg.probe_amps:g}A for {held_min:.0f}min, "
f"restoring to {cfg.normal_amps:g}A"
Expand Down Expand Up @@ -643,8 +681,9 @@ def main(argv: list[str] | None = None) -> int:
"--restore-step-a",
type=float,
default=2.0,
help="raise the cap by at most this much per confirmed-clear cycle, "
"instead of snapping straight back to --normal-amps (default %(default)s)",
help="raise the cap by this much per confirmed-clear cycle when the server reports no "
"sustainable current, or after the model has already been wrong once this session "
"(default %(default)s)",
)
parser.add_argument(
"--restore-margin-c",
Expand Down
35 changes: 26 additions & 9 deletions docs/amp-control.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,12 +34,25 @@ reality. `model` and `trajectory` are trusted, but **not symmetrically**:
alert 40 fired inside exactly that gap during live testing.
- **Restoring up is the risky direction** (it's what pushes the equilibrium
back toward the trip point), so it stays conservative on every axis: only
`trajectory` basis, one `--restore-step-a` at a time rather than snapping
straight back to `--normal-amps`, and never while the handle is within
`--restore-margin-c` of the trip point even if the trajectory reads clear.
Snapping straight back to full current, twice, immediately restarted the
climb both times during live testing — turning a caught derate into
repeated near-misses before a third one wasn't caught in time.
`trajectory` basis, never straight back to `--normal-amps` on trust, and
never while the handle is within `--restore-margin-c` of the trip point
even if the trajectory reads clear. Snapping straight back to full
current, twice, immediately restarted the climb both times during live
testing — turning a caught derate into repeated near-misses before a
third one wasn't caught in time.
- **Where it restores *to* is the model's answer, in one move.** The server
reports `sustainable_max_a` on every forecast: the highest current whose
modelled plateau stays under the trip point at today's ambient (the LAN
sensor when one reports, else the ambient the live trajectory implies —
the sensor is preferred because a plateau measured at a low current
carries the current law's extrapolation error, and it comes back doubled
when rescaled to a high one). The daemon restores straight to that (or to full rate
when that is what it says), and lets the trajectory and the confidence
guard trim the last amp or two. It used to climb `--restore-step-a` at a
time instead — but every rung resets the trajectory window, so a climb
from 32 A took half an hour to find the same number. The model is trusted
once per session: after a quick reversal it has already been wrong about
today, and the climb falls back to single `--restore-step-a` steps.

Either direction needs a signal held for `--confirm-ticks` consecutive polls
(default 3) before acting — a single noisy fit can't flip a real amp change.
Expand Down Expand Up @@ -81,8 +94,8 @@ untrustworthy forecast is. As a window matures its SE shrinks, and the
guard relaxes tick by tick on its own.

A cap fully lifts three ways: the trajectory forecast reports the risk has
passed *and* the handle has real thermal margin (stepped up gradually, see
above), the charging session ends (restored immediately — no more climb to
passed *and* the handle has real thermal margin (restored to the
sustainable current, see above), the charging session ends (restored immediately — no more climb to
protect against), or — a safety net — a new session starts while the
daemon's on-disk state still says "capped" from a run that never saw its
session close out (crash, restart, etc.). That last case always restores
Expand Down Expand Up @@ -125,7 +138,11 @@ sudo ./deploy/install-derate-amp-control.sh --tesla-ble http://<esp32-host> --pr

Once every `--probe-interval-days` (default 30), the first charging session
to come along is held at `--probe-amps` for `--probe-hold-min` (default 40)
minutes, then released. Pick a current low enough that neither this daemon
minutes, then restored to the sustainable current its own plateau implies —
the probe is the best measurement of today's conditions the session will
get, so its end is the one moment a restore target is most trustworthy.
(Releasing to full rate instead, on a day the model already knew full rate
would trip, produced a 32 → 48 → 42 A oscillation.) Pick a current low enough that neither this daemon
nor the vehicle wants to reduce it — on a 48 A install where foldback starts
around 61 °C, 32 A plateaus near 53 °C with room to spare. Hold it for more
than ~3x the install's time constant (`model.tau_min` in `/api/thermal`) so
Expand Down
22 changes: 17 additions & 5 deletions docs/thermal-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,17 @@ the trip happens.
the default has a standing cost — the fitter judges a charge's window
against the default τ, so only charges of ~22 min or more at steady
current teach the model there.
- **How rise scales with current is fitted too.** Joule heating says rise
∝ I², and that is the prior — but a real handle carries heat that does
not scale with current (the charger's own electronics, cable heat soak),
so measured plateaus fall off more gently as current drops. On one
install a 32 A probe settled 4 °C above what I² predicted, and every
forecast at an off-reference current inherited the error — including
the current the amp controller was told to restore to. Once the
free-running fits span ≥ 6 A of current, the exponent *n* in
rise = rise₄₈ · (I/48)ⁿ is fitted by log-log regression (`current_exp`
in `/api/thermal`, with its standard error), every fit's `rise_ref_c` is
re-normalized with it, and the model note says so.
- **The charger is its own thermometer.** Idle, the handle sits ~1–2 °C above
ambient (an ambient-dependent offset), so ambient can be read without any
extra sensor. The offset model ships as a seed from one install and is
Expand Down Expand Up @@ -165,7 +176,7 @@ coefficient. The reported Δ is that slope times the observed span.
It did once compare a recent median against a baseline median, and that
asks the wrong question. "Are the last few fits higher?" is answered for
you by anything that moved with the calendar: a garage that cooled between
the two halves, or a vehicle capped to a lower current whose (48/I)²
the two halves, or a vehicle capped to a lower current whose (48/I)ⁿ
normalization then lifts every recent fit at once. On one install the
split reported **+7.2 °C with a 95 % CI of [5.4, 9.1]** — "statistically
confirmed" — for a connector whose rise, regressed on time with ambient and
Expand Down Expand Up @@ -223,10 +234,11 @@ row. More sessions either confirm it or dissolve it.
typical and pooled the operating current out of its own comparison).
- **Pooled across a wide current band when the fits are clean.**
Ambient-bracketed fits join from a wider band, and the regression's own
current term then *adjusts* them: residual error in the I² normalization
lands on that coefficient instead of masquerading as a trend. On the
install above that coefficient read −0.99 °C per amp, which is the whole
of the phantom +7.2 °C. The band is wide enough on purpose to admit a
current term then *adjusts* them: residual error in the current
normalization lands on that coefficient instead of masquerading as a
trend. On the install above, under the I² prior, that coefficient read
−0.99 °C per amp, which is the whole of the phantom +7.2 °C — and is
what the fitted exponent now removes at the source. The band is wide enough on purpose to admit a
[calibration probe](amp-control.md#the-calibration-probe).
- **Never only stale sessions.** If none of the newest few free-running
charges make it into the comparison, the install has moved to a current
Expand Down
Loading
Loading