Skip to content

feat: automate a monthly calibration probe, and stop discarding its fits - #30

Merged
zebraengine merged 1 commit into
mainfrom
feat/calibration-probe
Sep 14, 2026
Merged

zebraengine merged 1 commit into
mainfrom
feat/calibration-probe

Conversation

@zebraengine

Copy link
Copy Markdown
Owner

Problem

Follow-up to #29, and it starts with a correction to that PR.

#29 attributed the in-window current sag to the charger's own thermal
regulation. That was wrong. The charger's internal foldback is counted by
lifetime.thermal_foldbacks, and on this install it has not incremented
since 08-03 — through every one of the 18 fitted sessions:

foldback counter 34 -> 42 over the recorded history
last increment:  session 33, 08-03 20:44
fitted sessions: 36 .. 73, 08-05 .. 09-08   <- zero foldbacks

The sag is mostly the monitor's own doing. The amp controller caps on
this model's forecast — 285 amp_capped events in a month — and 6 of the 7
windows the free-plateau gate excluded have controller events inside them
(the seventh looks like vehicle taper). So it is a feedback loop: the
forecast caps the current, the cap contaminates the fit, the fit feeds the
forecast. The gate and the verdict in #29 are unaffected — a window whose
current moved is contaminated whoever moved it — but the comments and docs
explaining why were wrong and are corrected here.

That also reframes the fix. Steady windows on this install are scarce (11 of
18) and land at whatever current the capping happened to stop at, split
across 39.6 / 44.4 / 48.6 A — a moving target, and one correlated with the
weather, since a hot garage triggers capping sooner.

The component that moves charge current is the one that can hold it still on
purpose.

Code touched

contrib/derate_amp_control.py — --probe-amps holds a chosen current
for --probe-hold-min (default 40) every --probe-interval-days (default
30). Off unless set.

Implemented as _apply_probe, a layer over the renamed _decide_thermal,
rather than a branch inside it: that logic is tuned by a string of live
incidents and the probe has no business reaching into it. Precedence is the
entire contract:

  • a thermal cap below the probe current wins, is applied, and abandons
    the probe — a window whose current just moved teaches nothing;
  • the probe outranks restoring, since stepping back toward full rate is
    exactly what would ruin the measurement (it also pins clear_streak at
    zero, so the restore path cannot bank confirming polls while the probe
    runs and spend them the instant it ends);
  • only a completed hold updates the cadence, so an early unplug or a
    session that needed a real cap retries next time rather than consuming the
    month;
  • a probe never starts from a current already at or below the probe value.

--probe-amps is validated against --min-amps/--normal-amps. State
gains two fields; load_state already ignores unknown keys and defaults
missing ones, so an existing state file upgrades silently (tested).

wallmonitor/thermal.py — a probe is useless if its fits are discarded,
and they were: 32 A against a 48 A install fell outside the pooling band.
Three changes, each safe because #29 replaced the median split with a
regression, and none of which the median split could have afforded:

  • pooling band 0.25 → 0.45. It was narrow to stop one off-current fit
    swinging a median; there is no median now, every fit reaching the
    comparison held its current steady, and the current term adjusts the rest.
  • typical_current_a is now the median of the whole comparable history, not
    the newest three fits. Chasing the newest current was necessary when only
    like could be compared with like. Now it is actively harmful: two probes
    landing in the newest three made the probe current "typical" and pooled
    the operating current out of its own comparison.
    Caught by the new test,
    not by inspection.
  • not chasing the newest current needs its own guard against stale verdicts,
    so a comparison must reach into the newest DRIFT_RECENCY_N (3)
    free-running charges or report nothing — which also lets a stale alert
    clear. Deliberately "the newest few", not "the newest": one odd charge
    must not blank an otherwise current watch.

Assumption stated inline: 32 A is this install's safe probe current
(foldback starts ~61 °C; 32 A plateaus ~53 °C). The docs tell other
installs how to choose rather than presenting 32 as universal.

Deliberately left alone: the controller is currently stuck holding 40 A
("giving up after 3 quick reversals") and runs with
--forecast-confidence-k 4 against a default of 2, which is why it caps so
often. Both are live-tuning questions, not code defects, and changing either
would alter derate protection — out of scope here, flagged for a decision.
Also still open from #29: regulated fits continue to feed the live forecast,
biasing it toward under-predicting derates.

Risk

  • The probe deliberately charges slower, once a month, for ~40 minutes.
    That is the cost, and it is why the feature is off by default.
  • It can be abandoned and retried, so in a month of short sessions it may
    not complete. That is preferred to banking a window it did not hold.
  • Safety is unchanged. The probe cannot raise current above what the
    thermal logic allows, cannot start below an active cap, and any cap under
    the probe current takes precedence. _decide_thermal is byte-for-byte the
    old decide(); a --probe-amps 0 config is asserted identical to the
    pre-change behavior.
  • Pooling changes which fits enter existing installs' verdicts — wider
    band, different typical_current_a. On the production database the verdict
    is unchanged (Δ 0.30 °C, CI [−2.83, 3.43]); typical_current_a moves
    39.6 → 48.6 A, which is the install's actual operating current.
  • The staleness guard can return None where a verdict previously appeared,
    after a sustained current change. That clears the alert rather than
    latching a stale one.
  • No schema change, no migration. State gains fields; old state files load.

Verification

142 tests pass (uv run pytest), 39 of them in
tests/test_derate_amp_control.py. The 2 pre-existing ruff E731 warnings
are untouched; bash -n clean on the installer.

Nine tests added. Probe: disabled-is-identical, starts when due, respects the
interval, holds against the restore path (and cannot bank clear_streak), completes
and restores, safety cap wins and abandons, unplug does not consume the
slot, never starts from a lower cap, old state file upgrades and is then due.
Analysis: a 32 A probe interleaved among 48 A charges reaches the regression
(n == len(fits), off_current_n == 0) and keeps current separable from the
calendar (collinear_with_time empty) — the test that caught the
typical-current inversion.

Two existing tests updated where the mechanism changed (typical is now the
usual current, not the newest; the "can't judge yet" case is now enforced by
the staleness guard rather than by chasing the current). Their intent is
preserved and re-stated in the comments.

Re-run against the production database with this branch's code:

delta_c            0.30      delta_ci95_c  [-2.83, 3.43]
drifting           False     n / regulated_n  11 / 7
typical_current_a  48.6      (was 39.6 — the controller's cap, not the install's current)
cross_current_n    2         (was 9)
current_coef       -0.986    ambient_coef  -0.434 -> compromised_by: ['ambient']

Deploy: --probe-amps 32 has not been applied to the mini-PC — that is a
separate install-derate-amp-control.sh run once this merges. Note also that
wallmonitor.service there is still running the pre-#29 code; the restart is
pending and best done between charging sessions.

🤖 Generated with Claude Code

The degradation watch can only compare charges whose current held steady
through the ramp. On this install those are scarce and land wherever the
capping stopped — and the reason is the monitor itself: the amp controller
caps on this model's own forecast, 285 times in a month, contaminating 7 of
18 fitted windows. The charger's internal foldback (lifetime
thermal_foldbacks) had not fired once in the same period, so the previous
commit's attribution of the current sag to the charger was wrong; the
comments and docs are corrected. It is a feedback loop: the forecast caps
the current, the cap contaminates the fit, the fit feeds the forecast.

The component that moves charge current is also the one that can hold it
still on purpose. --probe-amps holds a chosen current for --probe-hold-min
every --probe-interval-days, giving the watch a plateau nothing trimmed at a
repeatable operating point. It lands as a layer over decide() rather than a
branch inside it: that logic is tuned by a string of live incidents and the
probe has no business reaching into it. Precedence is the contract — a
thermal cap below the probe current wins and abandons the probe, the probe
outranks restoring, and only a completed hold updates the cadence, so an
early unplug retries next time instead of consuming the month.

The probe is useless if its fits are then discarded, and they were: at 32 A
against a 48 A install they fell outside the pooling band. Three things fix
that, all of which the regression made safe and none of which the median
split could have afforded:

- the pooling band widens (0.25 -> 0.45). It was narrow to stop one
  off-current fit swinging a median; there is no median now, every fit that
  reaches the comparison held its current steady, and the current term
  adjusts what is left.
- "typical current" becomes the median of the whole comparable history
  rather than the newest three fits. Chasing the newest current was
  necessary when only like could be compared with like; now it just lets an
  occasional off-current charge become typical and invert the band. Two
  probes landing together did exactly that, pooling the operating current
  out of its own comparison.
- not chasing the newest current needs its own guard against judging from
  stale data, so a verdict must now reach into the newest few free-running
  charges or report nothing. Deliberately "the newest few", not "the
  newest": one odd charge must not blank an otherwise current watch.

Verdict on the production database is unchanged (delta 0.30 C, CI
[-2.83, 3.43]); typical_current_a now reads 48.6 A, the install's actual
operating current, rather than 39.6 A.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@zebraengine
zebraengine merged commit 7181e44 into main Sep 14, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant