Day-ahead (24h) hourly load forecasting, benchmarked from naive baselines to gradient boosting, with an emphasis on honest evaluation and error analysis grounded in power system behaviour.
Benchmarked 11 approaches — from naive baselines to gradient boosting with observed weather — for 24h-ahead hourly load forecasting on Great Britain's national demand (NESO, 2022–2026). The deployable model (LightGBM on known-ahead calendar/solar features and demand lags, no weather) reached 5.75% MAPE on a test set touched exactly once, a 30% reduction on the best baseline (persistence, 8.26%). The largest and most persistent source of error is not holidays or weather but embedded solar: midday MAPE runs 9.23% against 3.4–4.9% elsewhere, and SHAP confirms the model learned the actual PV-suppression mechanism behind that gap rather than merely correlating with it.
Short-term load forecasting drives unit commitment, reserve sizing, and energy procurement. Errors are not symmetric in cost: under-forecasting risks reserve shortfall, over-forecasting wastes committed capacity.
I have spent fifteen years in electrical engineering, and planning is the part of the work I enjoy most. Planning keeps running into forecasting. You size equipment for load that does not exist yet, and you commit money now for demand that arrives in five years. When the forecast is off, the plan fails quietly, and usually long after anyone can reverse the decision.
The arithmetic is rarely what breaks. One method gets trusted too far, with nothing standing beside it for comparison. That is the reason this repository is a benchmark rather than a model: several approaches, the same data, the same evaluation rules, and the disappointing results published next to the good ones. A single number with no reference point cannot tell you whether 6% MAPE is competent or embarrassing.
That conviction shaped three specific choices here. Baselines stay in every results table, because a forecast is only as good as the simplest thing it beats. Configuration is chosen on a selection set rather than on the data the score is reported from, because a plan built on a tuned-to-death number is a plan built on nothing. And nothing that physics can compute is hardcoded to one country's clock, because a planning tool that only works at one latitude is not a tool.
AI is the new half of this for me. I am putting planning work I already know beside a field I am still learning, and finding out how much of it survives an honest test.
| Item | Value |
|---|---|
| Source | NESO Historic Demand Data (Great Britain) |
| License / terms of use | NESO Open Data Licence — commercial reuse permitted with attribution |
| Coverage | GB transmission system, 2019-01-01 → 2026-07-27 (2001–2026 available) |
| Native resolution | 30 min (settlement periods 1–48) |
| Resolution used | Hourly |
| Target | ND — National Demand, MW |
| Weather covariates | Open-Meteo Historical Weather API, no key required — temperature, dew point, cloud cover, shortwave radiation, at the GB demand-weighted centroid |
Contains data from the National Energy System Operator, used under the NESO Open Data Licence.
Data is not committed to this repository. Run make data (or
python -m src.data.download) to fetch it. See docs/DATA.md
for the full acquisition and preprocessing procedure.
Full audit: python -m src.data.check_completeness. Detail in
docs/DATA.md.
| Issue | Detection | Handling | Rationale |
|---|---|---|---|
| DST: 8 days have 46 settlement periods, 7 have 50 | Periods-per-day distribution | Timestamps built in UTC from each day's local midnight | Naive local-time arithmetic duplicates the autumn hour and invents the spring one |
FORECAST_ACTUAL_INDICATOR present only in the current-year file |
Column-set diff across yearly files | Null indicator treated as actual | The obvious == 'A' filter silently drops 2019–2025 (123k of 132k rows) — this bug was hit and fixed during the audit |
SCOTTISH_TRANSFER absent before 2023 |
Column-set diff | Excluded from features | A column that starts mid-series reads to a model as a real regime change |
| MW is average power, not energy | Unit check on the source schema | Resampled with mean, never sum |
sum inflates every value by exactly 2× and remains invisible in every distribution check |
| 2020 demand depressed ~7.5% by COVID | Annual means, overview plot | Training range starts 2022 | Lockdown load shape does not represent current behaviour |
| Data ends ~3 weeks behind today | Coverage check | Accepted; test set ends at the last settled day | NESO publishes in batches, not in real time |
Precise framing, because "load forecasting" is ambiguous:
- Target: GB National Demand (
ND), transmission-system level, MW — not per-client consumption. - Forecast horizon: 24 hours ahead
- Forecast origin: 00:00 local (Europe/London), for the whole of the following local day
- Update frequency: once per day, at the forecast origin — this benchmark does not model intraday re-forecasting
- Information available at forecast time: actuals strictly before the origin — calendar, UK holidays, computed solar geometry, installed PV/wind capacity, and demand lags/rolling stats (see Features). The honest ("operational") models use no weather at all. A third variant is scored with observed weather for the forecast period, which is leakage against a live deployment — no real weather-forecast feed exists in this benchmark — so it is always reported separately and labelled optimistic, never as the headline number.
Fixed before any modelling, and not changed afterwards.
- Split: chronological — never randomised. Random splits leak future
information into training and inflate results.
- Train: 2022-01-01 → 2024-12-31 (26,304 h)
- Validation: 2025-01-01 → 2025-12-31 (8,760 h) — everything below is scored here
- Test: 2026-01-01 → 2026-07-27 (4,991 h) — held out, untouched
- Reserved: 2019–2021 excluded from the benchmark entirely, see
docs/SHOCK_SIMULATION_PLAN.md
- Cross-validation: rolling-origin (expanding window)
- Leakage control: enforced by the harness, not by convention.
backtest()slices history at each forecast origin, so a model cannot see the day it is predicting even if its code tries to. A spy test asserts the harness never exposes a timestamp at or after the origin. - Metrics: MAPE, MAE, RMSE — reported together. MAPE is intuitive but unstable near low load; RMSE penalises the large peak-hour errors that matter operationally.
- Reported disaggregation: overall, by hour of day, by day type, and holidays separately.
Everything above this point was decided on train and validation only. The
full Stage 4 model roster was then refit on train+valid combined
(2022–2025) and scored against the test set exactly once, with
python -m src.evaluation.test_score:
| Model | MAPE (%) | MAE (MW) | RMSE (MW) | Bias (MW) | Notes |
|---|---|---|---|---|---|
| Ridge [calendar] | 9.28 | 2,324 | 2,905 | −852 | Known-ahead features only |
| Seasonal naive (t−168) | 8.96 | 2,237 | 3,068 | +231 | Baseline |
| Persistence (t−24) | 8.26 | 2,039 | 2,824 | +40 | Baseline — best simple rule |
| Ensemble w=0.6 (t−24, t−168) | 7.10 | 1,750 | 2,370 | +116 | Stage 2 best, weight chosen on 2024 selection |
| LightGBM [calendar] | 6.68 | 1,609 | 2,166 | +211 | Known-ahead features only |
| Ridge [operational], stationary PV | 6.41 | 1,565 | 2,078 | +76 | |
| Ridge [operational] | 6.31 | 1,558 | 2,046 | −70 | |
| LightGBM [operational] | 6.14 | 1,478 | 1,988 | +241 | Raw PV capacity included |
| LightGBM [operational], stationary PV | 5.75 | 1,399 | 1,894 | +149 | Deployable champion |
| Ridge [weather] | 5.49 | 1,350 | 1,734 | −246 | Observed weather — optimistic |
| LightGBM [weather] | 4.87 | 1,148 | 1,524 | +356 | Observed weather — optimistic |
Generalization held, no surprises. The champion moved from 5.46% (validation) to 5.75% (test), the optimistic weather variant from 4.48% to 4.87% — both roughly +0.3pp, and the full ranking order from validation reproduced exactly on test. All six models holding no external weather frame re-passed the leakage check against the test set. See Error analysis below for where the remaining 5.75% sits.
Everything below reproduces the selection process, on the validation period
used to choose it. Reproduce with python -m src.evaluation.report (baselines)
and python -m src.models.train (learned models).
| Model | MAPE (%) | MAE (MW) | RMSE (MW) | Bias (MW) | Notes |
|---|---|---|---|---|---|
| Hour × weekday mean (all train) | 15.41 | 3,964 | 4,957 | +168 | Baseline — pure calendar shape, no recency |
| Hour × weekday mean (rolling 8w) | 9.77 | 2,537 | 3,381 | −31 | Baseline — table rebuilt at every origin |
| Seasonal naive (t−168) | 8.44 | 2,181 | 2,915 | −42 | Baseline |
| Persistence (t−24) | 7.33 | 1,840 | 2,565 | −14 | Baseline — best so far |
| Ridge [calendar] | 8.86 | 2,221 | 2,839 | −247 | Known-ahead features only |
| LightGBM [calendar] | 7.53 | 1,881 | 2,480 | +457 | Known-ahead features only |
| LightGBM [operational] | 6.54 | 1,589 | 2,089 | +625 | Raw PV capacity included |
| Ridge [operational] | 5.82 | 1,452 | 1,909 | +72 | |
| LightGBM [operational], stationary PV | 5.46 | 1,366 | 1,826 | +80 | Best without observed weather |
| Ridge [weather] | 5.05 | 1,269 | 1,632 | −270 | Observed weather — optimistic |
| LightGBM [weather] | 4.48 | 1,097 | 1,452 | +273 | Observed weather — optimistic |
python -m src.evaluation.diagnose. Every configuration choice — window
length, ensemble weight, per-segment winner — is made on a selection set
carved out of training data (2024) and only then scored on validation.
Choosing on validation and reporting the result would be a subtler version of
the leak the harness exists to prevent: the number would look good and mean
nothing.
| Configuration | MAPE (%) | MAE (MW) | RMSE (MW) | P90 |err| (MW) | Within ±3% | Chosen on |
|---|---|---|---|---|---|---|
| Hour × weekday (all train) | 15.41 | 3,964 | 4,957 | 8,289 | 12.97% | — |
| Hour × weekday (rolling 8w) | 9.77 | 2,537 | 3,381 | 5,605 | 21.14% | — |
| Seasonal naive (t−168) | 8.44 | 2,181 | 2,915 | 4,635 | 24.54% | — |
| Hour × weekday (rolling 2w) | 7.84 | 2,030 | 2,786 | 4,498 | 26.63% | selection 2024 |
| Persistence (t−24) | 7.33 | 1,840 | 2,565 | 4,249 | 33.05% | — |
| Hybrid — best model per load period | 7.24 | 1,850 | 2,558 | 4,040 | 30.27% | selection 2024 |
| Ensemble, 3-member equal | 6.45 | 1,654 | 2,252 | 3,656 | 33.44% | — |
| Ensemble 50/50 (t−24, t−168) | 6.41 | 1,630 | 2,216 | 3,680 | 34.20% | — |
| Ensemble w=0.6 (t−24, t−168) | 6.34 | 1,607 | 2,197 | 3,661 | 35.96% | selection 2024 |
Within ±3% is the share of hours whose absolute percentage error falls inside 3%. It answers a question no mean can: how often is this forecast actually usable? Two models can share a MAPE while one is steadily mediocre and the other alternates between excellent and unusable. P90 is the 90th percentile of absolute error — the bad hour rather than the average one, which is what reserve margins are sized against.
Note that these two disagree with MAPE about the hybrid: it improves P90 (4,040 vs 4,249 MW) while making more hours fall outside ±3% than plain persistence (30.27% vs 33.05%). It trims the worst errors and loosens the middle. Whether that is an improvement depends on whether the cost of being wrong is linear — which for reserve procurement it is not.
- A two-line ensemble beats every single baseline — 7.33% → 6.34%, a 14% reduction in MAPE, from averaging two models that were already in the table. The weight curve is flat between w=0.5 and w=0.6, so the choice is not balanced on a knife edge.
- Per-segment switching did not pay. Choosing the best model separately for
each load period on 2024 bought 0.09 points on 2025 (7.33 → 7.24). The
per-segment winners did not transfer:
rolling 2wwon midday on the selection set and then lost it on validation. Reported because it is the result, not because it is flattering. - Rolling window: 2 weeks is optimal, and the curve is not monotonic — 26 weeks is worse than 52. A half-year window straddles two seasons at their most different; a full year at least averages symmetric ones.
| Segment | Share of hours | Persistence MAPE | Ensemble w=0.6 |
|---|---|---|---|
| Night (sun down) | 50% | 5.22% | 4.98% |
| Day (sun up) | 50% | 9.41% | 7.82% |
| High sun (>15°) | 33% | 10.68% | 8.92% |
| Dark | 50% | 5.22% | 4.98% |
| Overnight base 23–06 | 29% | 4.87% | 4.82% |
| Midday / solar 09–16 | 29% | 10.64% | 9.23% |
| Evening peak 16–20 | 17% | 6.03% | 5.32% |
| Triad window (winter wkdy 16–19) | 3% | 4.25% | 4.89% |
Day and night are split by computed solar elevation. No clock hour is
hardcoded anywhere in src/; a test enforces it. At 53°N the daylight window
already swings from 7.3 hours at the winter solstice to 16.7 at the summer one,
so a fixed 06:00–18:00 rule would file the same clock hour as "day" in December
and June while the sun is below the horizon in one of them.
The same code is portable without modification — verified across latitudes:
| Site | 21 Jun | 21 Dec | Behaviour |
|---|---|---|---|
| Jakarta (6.2°S) | 11.6 h | 12.4 h | near-constant, as expected on the equator |
| Great Britain (53.0°N) | 16.7 h | 7.3 h | strong seasonality |
| Tromsø (69.6°N) | 24.0 h | 0.0 h | polar day and polar night, no NaN |
| Santiago (33.4°S) | 9.8 h | 14.2 h | hemisphere correctly inverted |
Separating morning from afternoon exposes an asymmetry that any symmetric bucket would average away:
| Solar phase | Hours | Persistence MAPE | Ensemble w=0.6 |
|---|---|---|---|
| Night (below civil twilight) | 3,784 | 5.07% | 4.97% |
| Dawn twilight | 286 | 8.08% | 6.12% |
| Morning sun | 2,266 | 10.72% | 8.54% |
| Afternoon sun | 2,140 | 8.03% | 7.05% |
| Dusk twilight | 284 | 4.42% | 3.91% |
Dawn is nearly twice as hard as dusk, and morning is harder than afternoon, even though irradiance is roughly symmetric about solar noon. Load is not: the morning ramp is where occupancy, heating and industrial start-up all move at once, while the evening decline is habitual and slow. A twilight bucket that merged dawn with dusk would have reported ~6% and hidden both ends.
Bucketing by day length (computed, so it needs no month names and works in either hemisphere):
| Day length | Persistence | Seasonal naive | Ensemble w=0.6 |
|---|---|---|---|
| Short (<9 h) | 6.42% | 8.89% | 6.17% |
| 9–11 h | 6.25% | 8.39% | 5.90% |
| 11–13 h | 7.75% | 9.69% | 7.25% |
| 13–15 h | 7.93% | 8.02% | 6.39% |
| Long (>15 h) | 8.21% | 7.58% | 6.44% |
The two curves run in opposite directions and cross. Persistence degrades as days lengthen — more solar variance between one day and the next. Seasonal naive improves, because a week-old observation is a better guide once the level stops drifting. Averaging them is not a trick that happened to work: each covers the regime where the other is weakest, which is also why per-segment switching gained nothing while blending gained 14%.
- Daylight hours are ~1.8× harder than dark hours (9.41% vs 5.22%), and MAE
rises monotonically with solar elevation, from 1,270 MW in deep dark to
2,670 MW at 30–40°. Embedded PV is not metered at transmission level, so it
reaches the forecaster only as unexplained variance in
ND. - The triad window is the easiest segment, not the hardest — 4.25% MAPE in the winter weekday evenings that set network charges. Demand there is high, stable and strongly habitual. Note that the ensemble is worse here (4.89%) than plain persistence: averaging in a weekly model dilutes exactly the recency that peak hours reward.
- Weekends are harder than weekdays (9.11% vs 6.62%) and winter is easier than spring for persistence — a level-following model does well when the level is high and stable.
The headline number depends on a choice, not on a model. LightGBM reaches 4.48% MAPE with observed weather and 5.46% without. An operational forecaster has a weather forecast, not an observation, so the honest figure for deployment sits closer to the second. Both are reported; only one is defensible as a claim about live performance.
Gradient boosting does not automatically win. On the operational feature set Ridge scored 5.82% against LightGBM's 6.54%. The reason turned out to be a feature, not the algorithm.
Installed PV capacity is a trap for a tree. The column rises monotonically,
so 99.7% of validation hours and 100% of test hours sit above the largest
value the model saw in training. A tree cannot extrapolate; every split
saturates at the top bin, the model under-counts PV suppression, and it
over-forecasts. That shows up as a +625 MW bias. Replacing raw capacity with
pv_potential_share, which divides potential PV by the recent demand level,
took LightGBM from 6.54% to 5.46% and cut the bias from +625 to +80 MW.
The choice between the three PV treatments was made on the 2024 selection set, not on validation:
| Model | PV treatment | Selection MAPE | Validation MAPE |
|---|---|---|---|
| Ridge | all PV features | 6.18 | 5.82 |
| Ridge | no PV at all | 5.63 | 5.80 |
| Ridge | stationary PV only | 5.62 | 5.95 |
| LightGBM | all PV features | 5.16 | 6.54 |
| LightGBM | no PV at all | 5.14 | 5.48 |
| LightGBM | stationary PV only | 5.13 | 5.46 |
The selection set picked the stationary form for both. It transferred for LightGBM and did not transfer for Ridge, where the plain no-PV set scored better on validation. Reported as it happened.
The PV features barely earn their place. Stationary PV beats no PV at all by
0.02 points for LightGBM, which is noise. The lags already carry the signal:
yesterday's demand at the same hour has PV suppression baked into it. Adding an
explicit PV term mostly re-states what lag_24h already knows. The engineering
value here was diagnostic, not predictive.
Reading this table before building anything clever:
- Persistence beats seasonal naive. Yesterday carries more information than last week, so day-to-day level (weather, in practice) dominates weekly shape. Any model that ignores recent level will lose.
- Recency is worth ~5.6 points of MAPE — the same calendar model goes from 15.41% to 9.77% purely by rebuilding its table from the last 8 weeks.
- Calendar shape alone is not enough. Even the rolling calendar model loses to plain persistence, so a pure hour×weekday feature set will not be competitive without lags.
- 7.33% is the number to beat. A gradient boosting model that lands at 7% would not be worth reporting as a success.
python -m src.features.build → 34 features in three groups, and the group
decides what a model may see.
| Group | Count | Available at the origin? |
|---|---|---|
| Known ahead — calendar, holidays, solar geometry, installed capacity | 18 | Yes |
| History only — lags and rolling stats, all shifted ≥ 24 h | 10 | Yes, from the past |
| Weather observed — temperature, dew point, cloud, radiation | 6 | No. See below |
A test corrupts demand from the forecast origin forward and asserts that no column in the first two groups moves. Weather is excluded from that test on purpose: those columns are observations of the forecast period, which is what the group name records.
| Demand | |
|---|---|
| Coldest bin (−6 to −4 °C) | 31,491 MW |
| Minimum (16 to 18 °C) | 23,395 MW |
| Warmest bin (30 to 32 °C) | 24,750 MW |
The cold arm rises 8,096 MW above the minimum. The warm arm rises 1,356 MW, a gap of roughly six to one, and a single linear temperature term cannot bend at both ends at once. The feature set carries heating and cooling degree hours instead, which keeps each arm separate and each coefficient physically readable.
Most of that six-to-one gap is not heating versus cooling. Regressing demand
on degree hours with hour and weekday controls, then adding solar terms one at
a time (python -m src.visualization.temperature_sensitivity):
| Controls | Heating (MW/°C) | Cooling (MW/°C) | Ratio | R² |
|---|---|---|---|---|
| Hour + weekend only | +737 | −358 | 2.1 | 0.749 |
| + installed PV × sun height | +498 | +221 | 2.3 | 0.861 |
| + measured radiation | +483 | +399 | 1.2 | 0.877 |
| + cloud cover | +492 | +409 | 1.2 | 0.878 |
Without a solar term the cooling coefficient is negative, which would mean air conditioning reduces electricity use. Hot hours in Britain are sunny hours, so embedded PV suppresses transmission demand at exactly the times cooling load appears, and an uncontrolled regression charges the PV dip to temperature. Once solar is in the model the sign flips and the true asymmetry is 1.2 to 1, not 6 to 1.
What remains of the asymmetry comes from duration and depth rather than from per-degree response: 17.7% of hours sit below 5 °C, against 1.2% above 25 °C, and a British winter runs a 25 K gap to a 20 °C indoor target where a hot day runs 12 K.
Cooling sensitivity also appears to be rising:
| Year | Heating | Cooling | Hours >20 °C |
|---|---|---|---|
| 2022 | +534 | +286 | 518 |
| 2023 | +489 | +321 | 386 |
| 2024 | +456 | +332 | 294 |
| 2025 | +469 | +421 | 648 |
| 2026 | +361 | +494 | 563 |
Cooling moved +73% while heating fell 32%. Treat this as a hint rather than a measurement: each year rests on a few hundred hot hours, 2026 stops in July so its heating figure comes from winter months alone, and the model carries no interaction terms.
Installed PV grew 58% across the modelling period (15,036 → 23,805 MW) while wind capacity moved −2%. None of that PV is metered at transmission level, so it reaches the forecaster only as demand that failed to appear.
Summer (May–Jul) midday depression, the gap between the 08:00 shoulder and the 12:00–14:00 trough:
| Year | 08:00 | 12–14h | Depression | Installed PV |
|---|---|---|---|---|
| 2022 | 24,842 MW | 23,922 MW | 3.7% | 14,363 MW |
| 2023 | 24,013 MW | 22,388 MW | 6.8% | 15,781 MW |
| 2024 | 23,948 MW | 22,519 MW | 6.0% | 17,664 MW |
| 2025 | 23,184 MW | 20,575 MW | 11.3% | 20,786 MW |
| 2026 | 22,937 MW | 19,758 MW | 13.9% | 23,701 MW |
Depression correlates with installed PV at +0.97, which on its own proves nothing. Five points, both series rising with time, and almost any pair would score near 1. Two controls carry the argument instead:
Wind capacity stayed flat while PV grew 58%, so this is not renewables in general.
The placebo: embedded PV cannot act at 02:00. Between 2022 and 2026 midday demand fell 17.4% while overnight demand rose 7.7%. A 25-point divergence inside the same series, appearing only while the sun is up, is much harder to explain away than a correlation across five years.
The feature pv_potential_MW is installed capacity multiplied by the sine of
solar elevation: a physical proxy for how much load embedded PV is hiding at
that instant. Neither term says it alone, since capacity is flat within a day
and the sun knows nothing about how much plant is out there. Across sunlit
hours it correlates −0.51 with demand.
The most useful part of this repo. Computed once, on the champion's already-
scored test predictions — python -m src.evaluation.error_analysis describes
errors that are already fixed, it does not re-select or retune anything.
1. When does the model fail? Daily MAPE on test ranges from 1.58% (2026-02-08) to 20.41% (2026-07-02), median 4.98%. But 8 of the 10 worst days were also above-median for the plain persistence baseline — most "bad days" are genuinely hard demand days for every model, not a champion-specific failure.
2. Holidays degrade the forecast by +1.45pp (7.16% on holiday hours vs 5.71% otherwise, 120 vs 4,871 hours). It is not uniform: Easter Monday (11.52%) and New Year's Day (10.19%) are the worst, while May Day (3.45%) is actually easier than an average day — a single "holiday" feature is carrying at least two different demand patterns.
3. The one DST transition day in the test window (2026-03-29, 23h spring-forward) is a sharp, rare failure: 12.48% MAPE vs 6.48% in the surrounding week, a +6.00pp jump. The autumn transition falls outside the test window (which ends 2026-07-27), so only the spring case is measured here.
4. Temperature extremes. MAPE rises from 4.11% in the second-coldest quintile to 7.26% in the warmest, but this test window is Jan–Jul only, so the "warm" tail is really the "sunny" tail — see the next point. It restates the solar confound from Load and temperature: the U is lopsided, not a genuine cold/heat effect.
5. Peak vs off-peak — the consistent finding across every stage of this project. Midday still dominates the error budget: 9.23% MAPE in the midday load period against 3.4–4.9% everywhere else, and 8.63% in the "high sun (>15°)" solar regime against 4.04% in the dark. The honest model narrowed this gap a great deal from the Stage 2 baselines (10.64% midday / 5.22% dark) but never closed it. The triad window — winter weekday 16:00–19:00, which sets GB transmission charges — remains the easiest segment at 3.45%.
6. Bias. The champion over-forecasts by +148.6 MW overall, but the sign flips by month: February is −460.7 MW (under), July is +1,065.2 MW (over). By day/night: +358.5 MW over-forecast in daylight, −102.1 MW under-forecast at night — the model still does not fully price in PV suppression at the moments it matters most, the same direction of error found back in Stage 4's raw-capacity ablation.
SHAP explanations for the deployed champion (python -m src.evaluation.shap_analysis), computed on the same test predictions above —
opening up the model already selected, not choosing a different one.
| Feature | Mean |SHAP| (MW) | Share |
|---|---|---|
lag_24h |
2,885 | 36.5% |
lag_168h |
1,085 | 13.7% |
dayofweek |
796 | 10.1% |
pv_potential_share |
320 | 4.0% |
lag_336h |
290 | 3.7% |
solar_gen_lag_24h |
227 | 2.9% |
lag_24h dominates, which is exactly what the results table already implies
— persistence alone reaches 8.26% MAPE, so the model leans hard on
yesterday's level and refines it. dayofweek's dependence plot is a clean
step function: positive SHAP on weekdays, negative on weekends, no
surprises.
The interesting one is pv_potential_share: SHAP value decreases
monotonically as the feature rises, and the effect is strongest exactly
when solar elevation (the colour axis) is highest. That is the model
internals confirming, independently of any correlation analysis, the same
PV-suppression mechanism argued for in
Embedded solar: the variable the grid cannot see —
this is not a spurious association, the tree has genuinely encoded "more
sun and more installed PV → suppress the forecast."
git clone https://github.com/Strooms/Electrical-Load-Forecasting-Benchmark.git
cd Electrical-Load-Forecasting-Benchmark
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
cp .env.example .env # add API credentials if using ENTSO-E
python -m src.data.download # fetch raw data
python -m src.data.download_weather # fetch Open-Meteo weather
python -m src.data.check_completeness # audit the raw download
python -m src.data.make_dataset # clean + resample to hourly
python -m src.evaluation.report # Stage 2: baselines on validation
python -m src.evaluation.diagnose # Stage 2: configuration search
python -m src.features.build # Stage 3: feature matrix
python -m src.models.train # Stage 4: Ridge/LightGBM on validation
python -m src.evaluation.test_score # Stage 5: final test-set score (once)
python -m src.evaluation.error_analysis # Stage 5: where the error sits
python -m src.evaluation.shap_analysis # Stage 5: SHAP interpretability
python -m src.visualization.header_figure # regenerate the header image
pytest # 34+ tests, incl. leakage/mutation checksRandom seeds are fixed in configs/. Exact package versions that produced
the reported numbers are pinned in requirements.txt (Python 3.13.14).
├── configs/ # YAML configs — no hardcoded parameters in code
├── data/ # gitignored
│ ├── raw/ # immutable download, never edited
│ ├── interim/ # cleaned, resampled
│ └── processed/ # model-ready feature matrices
├── docs/
│ ├── DATA.md # acquisition + preprocessing detail
│ └── DECISIONS.md # design decisions and why
├── notebooks/ # exploration only — logic lives in src/
├── reports/figures/
├── src/
│ ├── data/ # download, clean, resample
│ ├── features/ # calendar, lags, weather
│ ├── models/ # baselines + ML models, shared interface
│ ├── evaluation/ # metrics, backtesting, error analysis
│ └── visualization/
└── tests/
- Single grid, single climate. Every number here is GB transmission
demand. The physics-vs-convention split in
region_profile(seeconfigs/config.yaml) is built for portability, but it has never actually been run against a second country's data. - The optimistic figure needs a weather forecast feed this benchmark does not have. 4.87% MAPE uses observed weather for the forecast period, which does not exist at 00:00 the day before. The 0.88pp gap to the 5.75% honest number is the real cost of that missing feed, not modelling error.
- Point forecasts only. No quantile or interval output, so there is no direct answer to "how much reserve margin does a 95% confidence forecast imply" — the question this whole project exists to eventually serve.
- No exogenous economic drivers. Price, industrial output, and EV charging load are all absent. Some of that signal likely already leaks in through the lag features, uncredited.
- The test window is 7 months, not a full year. 2026-01-01 to 2026-07-27 has no autumn DST transition, no true summer heat extreme, and no Christmas/New Year holiday cluster beyond New Year's Day itself. The headline 5.75% is real but was not tested against a full seasonal cycle.
- DST handling is structurally correct but not specifically modelled. The harness produces 23/25-hour days correctly (Stage 2), but no feature flags "this is a transition day" — which is very likely why it is the sharpest single-day failure mode found in Stage 5 (+6.00pp).
- England-only holiday calendar.
holidays.UnitedKingdom(subdiv="England")is a known simplification; Scotland keeps a different calendar and is folded into the same national series.
- A DST-transition flag. The sharpest single-day failure found in Stage 5 (+6.00pp on 2026-03-29) has an obvious, cheap fix that was never tried: a binary feature for "today has 23 or 25 hours," letting the model treat it as the irregular case it is instead of silently forcing 24-hour lag logic onto it.
- Quantile regression (pinball loss) alongside the point forecast, so reserve sizing has an actual interval to work from instead of a single number and an MAE that says nothing about the tails that matter.
- A real weather forecast feed, to find out how much of the 0.88pp gap between the honest and optimistic models survives contact with an actual forecast API instead of observed weather standing in for one.
- Hierarchical reconciliation down to regional or feeder level, where the embedded-PV problem this project spent three stages on gets worse, not better — capacity is lumpier and less diversified at smaller scale.
- National Energy System Operator (NESO). Historic Demand Data (GB). neso.energy/data-portal/historic-demand-data, used under the NESO Open Data Licence.
- Open-Meteo. Historical Weather API. open-meteo.com/en/docs/historical-weather-api.
- Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q. and Liu, T.Y., 2017. LightGBM: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems, 30.
- Lundberg, S.M. and Lee, S.I., 2017. A unified approach to interpreting model predictions (SHAP). Advances in Neural Information Processing Systems, 30.
dr-prodigy/python-holidays— UK bank holiday calendar (holidays.UnitedKingdom).
Author: Raditya Yoga K. · LinkedIn License: MIT (code). Data remains under its original terms — see above.



