Skip to content

Latest commit

 

History

19 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Short-Term Electricity Load Forecasting — A Benchmark Study

Day-ahead (24h) hourly load forecasting, benchmarked from naive baselines to gradient boosting, with an emphasis on honest evaluation and error analysis grounded in power system behaviour.

Forecast vs actual


TL;DR

Benchmarked 11 approaches — from naive baselines to gradient boosting with observed weather — for 24h-ahead hourly load forecasting on Great Britain's national demand (NESO, 2022–2026). The deployable model (LightGBM on known-ahead calendar/solar features and demand lags, no weather) reached 5.75% MAPE on a test set touched exactly once, a 30% reduction on the best baseline (persistence, 8.26%). The largest and most persistent source of error is not holidays or weather but embedded solar: midday MAPE runs 9.23% against 3.4–4.9% elsewhere, and SHAP confirms the model learned the actual PV-suppression mechanism behind that gap rather than merely correlating with it.


Why this project

Short-term load forecasting drives unit commitment, reserve sizing, and energy procurement. Errors are not symmetric in cost: under-forecasting risks reserve shortfall, over-forecasting wastes committed capacity.

I have spent fifteen years in electrical engineering, and planning is the part of the work I enjoy most. Planning keeps running into forecasting. You size equipment for load that does not exist yet, and you commit money now for demand that arrives in five years. When the forecast is off, the plan fails quietly, and usually long after anyone can reverse the decision.

The arithmetic is rarely what breaks. One method gets trusted too far, with nothing standing beside it for comparison. That is the reason this repository is a benchmark rather than a model: several approaches, the same data, the same evaluation rules, and the disappointing results published next to the good ones. A single number with no reference point cannot tell you whether 6% MAPE is competent or embarrassing.

That conviction shaped three specific choices here. Baselines stay in every results table, because a forecast is only as good as the simplest thing it beats. Configuration is chosen on a selection set rather than on the data the score is reported from, because a plan built on a tuned-to-death number is a plan built on nothing. And nothing that physics can compute is hardcoded to one country's clock, because a planning tool that only works at one latitude is not a tool.

AI is the new half of this for me. I am putting planning work I already know beside a field I am still learning, and finding out how much of it survives an honest test.


Data

Item Value
Source NESO Historic Demand Data (Great Britain)
License / terms of use NESO Open Data Licence — commercial reuse permitted with attribution
Coverage GB transmission system, 2019-01-01 → 2026-07-27 (2001–2026 available)
Native resolution 30 min (settlement periods 1–48)
Resolution used Hourly
Target ND — National Demand, MW
Weather covariates Open-Meteo Historical Weather API, no key required — temperature, dew point, cloud cover, shortwave radiation, at the GB demand-weighted centroid

Contains data from the National Energy System Operator, used under the NESO Open Data Licence.

Data is not committed to this repository. Run make data (or python -m src.data.download) to fetch it. See docs/DATA.md for the full acquisition and preprocessing procedure.

Known data issues and how they were handled

Full audit: python -m src.data.check_completeness. Detail in docs/DATA.md.

Issue Detection Handling Rationale
DST: 8 days have 46 settlement periods, 7 have 50 Periods-per-day distribution Timestamps built in UTC from each day's local midnight Naive local-time arithmetic duplicates the autumn hour and invents the spring one
FORECAST_ACTUAL_INDICATOR present only in the current-year file Column-set diff across yearly files Null indicator treated as actual The obvious == 'A' filter silently drops 2019–2025 (123k of 132k rows) — this bug was hit and fixed during the audit
SCOTTISH_TRANSFER absent before 2023 Column-set diff Excluded from features A column that starts mid-series reads to a model as a real regime change
MW is average power, not energy Unit check on the source schema Resampled with mean, never sum sum inflates every value by exactly 2× and remains invisible in every distribution check
2020 demand depressed ~7.5% by COVID Annual means, overview plot Training range starts 2022 Lockdown load shape does not represent current behaviour
Data ends ~3 weeks behind today Coverage check Accepted; test set ends at the last settled day NESO publishes in batches, not in real time

Problem definition

Precise framing, because "load forecasting" is ambiguous:

  • Target: GB National Demand (ND), transmission-system level, MW — not per-client consumption.
  • Forecast horizon: 24 hours ahead
  • Forecast origin: 00:00 local (Europe/London), for the whole of the following local day
  • Update frequency: once per day, at the forecast origin — this benchmark does not model intraday re-forecasting
  • Information available at forecast time: actuals strictly before the origin — calendar, UK holidays, computed solar geometry, installed PV/wind capacity, and demand lags/rolling stats (see Features). The honest ("operational") models use no weather at all. A third variant is scored with observed weather for the forecast period, which is leakage against a live deployment — no real weather-forecast feed exists in this benchmark — so it is always reported separately and labelled optimistic, never as the headline number.

Evaluation protocol

Fixed before any modelling, and not changed afterwards.

  • Split: chronological — never randomised. Random splits leak future information into training and inflate results.
    • Train: 2022-01-01 → 2024-12-31 (26,304 h)
    • Validation: 2025-01-01 → 2025-12-31 (8,760 h) — everything below is scored here
    • Test: 2026-01-01 → 2026-07-27 (4,991 h) — held out, untouched
    • Reserved: 2019–2021 excluded from the benchmark entirely, see docs/SHOCK_SIMULATION_PLAN.md
  • Cross-validation: rolling-origin (expanding window)
  • Leakage control: enforced by the harness, not by convention. backtest() slices history at each forecast origin, so a model cannot see the day it is predicting even if its code tries to. A spy test asserts the harness never exposes a timestamp at or after the origin.
  • Metrics: MAPE, MAE, RMSE — reported together. MAPE is intuitive but unstable near low load; RMSE penalises the large peak-hour errors that matter operationally.
  • Reported disaggregation: overall, by hour of day, by day type, and holidays separately.

Results

Final result — test set (2026), touched once

Everything above this point was decided on train and validation only. The full Stage 4 model roster was then refit on train+valid combined (2022–2025) and scored against the test set exactly once, with python -m src.evaluation.test_score:

Model MAPE (%) MAE (MW) RMSE (MW) Bias (MW) Notes
Ridge [calendar] 9.28 2,324 2,905 −852 Known-ahead features only
Seasonal naive (t−168) 8.96 2,237 3,068 +231 Baseline
Persistence (t−24) 8.26 2,039 2,824 +40 Baseline — best simple rule
Ensemble w=0.6 (t−24, t−168) 7.10 1,750 2,370 +116 Stage 2 best, weight chosen on 2024 selection
LightGBM [calendar] 6.68 1,609 2,166 +211 Known-ahead features only
Ridge [operational], stationary PV 6.41 1,565 2,078 +76
Ridge [operational] 6.31 1,558 2,046 −70
LightGBM [operational] 6.14 1,478 1,988 +241 Raw PV capacity included
LightGBM [operational], stationary PV 5.75 1,399 1,894 +149 Deployable champion
Ridge [weather] 5.49 1,350 1,734 −246 Observed weather — optimistic
LightGBM [weather] 4.87 1,148 1,524 +356 Observed weather — optimistic

Generalization held, no surprises. The champion moved from 5.46% (validation) to 5.75% (test), the optimistic weather variant from 4.48% to 4.87% — both roughly +0.3pp, and the full ranking order from validation reproduced exactly on test. All six models holding no external weather frame re-passed the leakage check against the test set. See Error analysis below for where the remaining 5.75% sits.

Validation period (2025) — how the model was selected

Everything below reproduces the selection process, on the validation period used to choose it. Reproduce with python -m src.evaluation.report (baselines) and python -m src.models.train (learned models).

Model MAPE (%) MAE (MW) RMSE (MW) Bias (MW) Notes
Hour × weekday mean (all train) 15.41 3,964 4,957 +168 Baseline — pure calendar shape, no recency
Hour × weekday mean (rolling 8w) 9.77 2,537 3,381 −31 Baseline — table rebuilt at every origin
Seasonal naive (t−168) 8.44 2,181 2,915 −42 Baseline
Persistence (t−24) 7.33 1,840 2,565 −14 Baseline — best so far
Ridge [calendar] 8.86 2,221 2,839 −247 Known-ahead features only
LightGBM [calendar] 7.53 1,881 2,480 +457 Known-ahead features only
LightGBM [operational] 6.54 1,589 2,089 +625 Raw PV capacity included
Ridge [operational] 5.82 1,452 1,909 +72
LightGBM [operational], stationary PV 5.46 1,366 1,826 +80 Best without observed weather
Ridge [weather] 5.05 1,269 1,632 −270 Observed weather — optimistic
LightGBM [weather] 4.48 1,097 1,452 +273 Observed weather — optimistic

Configuration search

python -m src.evaluation.diagnose. Every configuration choice — window length, ensemble weight, per-segment winner — is made on a selection set carved out of training data (2024) and only then scored on validation. Choosing on validation and reporting the result would be a subtler version of the leak the harness exists to prevent: the number would look good and mean nothing.

Configuration MAPE (%) MAE (MW) RMSE (MW) P90 |err| (MW) Within ±3% Chosen on
Hour × weekday (all train) 15.41 3,964 4,957 8,289 12.97% —
Hour × weekday (rolling 8w) 9.77 2,537 3,381 5,605 21.14% —
Seasonal naive (t−168) 8.44 2,181 2,915 4,635 24.54% —
Hour × weekday (rolling 2w) 7.84 2,030 2,786 4,498 26.63% selection 2024
Persistence (t−24) 7.33 1,840 2,565 4,249 33.05% —
Hybrid — best model per load period 7.24 1,850 2,558 4,040 30.27% selection 2024
Ensemble, 3-member equal 6.45 1,654 2,252 3,656 33.44% —
Ensemble 50/50 (t−24, t−168) 6.41 1,630 2,216 3,680 34.20% —
Ensemble w=0.6 (t−24, t−168) 6.34 1,607 2,197 3,661 35.96% selection 2024

Within ±3% is the share of hours whose absolute percentage error falls inside 3%. It answers a question no mean can: how often is this forecast actually usable? Two models can share a MAPE while one is steadily mediocre and the other alternates between excellent and unusable. P90 is the 90th percentile of absolute error — the bad hour rather than the average one, which is what reserve margins are sized against.

Note that these two disagree with MAPE about the hybrid: it improves P90 (4,040 vs 4,249 MW) while making more hours fall outside ±3% than plain persistence (30.27% vs 33.05%). It trims the worst errors and loosens the middle. Whether that is an improvement depends on whether the cost of being wrong is linear — which for reserve procurement it is not.

  • A two-line ensemble beats every single baseline — 7.33% → 6.34%, a 14% reduction in MAPE, from averaging two models that were already in the table. The weight curve is flat between w=0.5 and w=0.6, so the choice is not balanced on a knife edge.
  • Per-segment switching did not pay. Choosing the best model separately for each load period on 2024 bought 0.09 points on 2025 (7.33 → 7.24). The per-segment winners did not transfer: rolling 2w won midday on the selection set and then lost it on validation. Reported because it is the result, not because it is flattering.
  • Rolling window: 2 weeks is optimal, and the curve is not monotonic — 26 weeks is worse than 52. A half-year window straddles two seasons at their most different; a full year at least averages symmetric ones.

Where the error actually sits

Segment Share of hours Persistence MAPE Ensemble w=0.6
Night (sun down) 50% 5.22% 4.98%
Day (sun up) 50% 9.41% 7.82%
High sun (>15°) 33% 10.68% 8.92%
Dark 50% 5.22% 4.98%
Overnight base 23–06 29% 4.87% 4.82%
Midday / solar 09–16 29% 10.64% 9.23%
Evening peak 16–20 17% 6.03% 5.32%
Triad window (winter wkdy 16–19) 3% 4.25% 4.89%

Day and night are split by computed solar elevation. No clock hour is hardcoded anywhere in src/; a test enforces it. At 53°N the daylight window already swings from 7.3 hours at the winter solstice to 16.7 at the summer one, so a fixed 06:00–18:00 rule would file the same clock hour as "day" in December and June while the sun is below the horizon in one of them.

The same code is portable without modification — verified across latitudes:

Site 21 Jun 21 Dec Behaviour
Jakarta (6.2°S) 11.6 h 12.4 h near-constant, as expected on the equator
Great Britain (53.0°N) 16.7 h 7.3 h strong seasonality
Tromsø (69.6°N) 24.0 h 0.0 h polar day and polar night, no NaN
Santiago (33.4°S) 9.8 h 14.2 h hemisphere correctly inverted

Splitting at solar noon: dawn is not dusk

Separating morning from afternoon exposes an asymmetry that any symmetric bucket would average away:

Solar phase Hours Persistence MAPE Ensemble w=0.6
Night (below civil twilight) 3,784 5.07% 4.97%
Dawn twilight 286 8.08% 6.12%
Morning sun 2,266 10.72% 8.54%
Afternoon sun 2,140 8.03% 7.05%
Dusk twilight 284 4.42% 3.91%

Dawn is nearly twice as hard as dusk, and morning is harder than afternoon, even though irradiance is roughly symmetric about solar noon. Load is not: the morning ramp is where occupancy, heating and industrial start-up all move at once, while the evening decline is habitual and slow. A twilight bucket that merged dawn with dusk would have reported ~6% and hidden both ends.

Why the ensemble works — the two baselines fail in opposite regimes

Bucketing by day length (computed, so it needs no month names and works in either hemisphere):

Day length Persistence Seasonal naive Ensemble w=0.6
Short (<9 h) 6.42% 8.89% 6.17%
9–11 h 6.25% 8.39% 5.90%
11–13 h 7.75% 9.69% 7.25%
13–15 h 7.93% 8.02% 6.39%
Long (>15 h) 8.21% 7.58% 6.44%

The two curves run in opposite directions and cross. Persistence degrades as days lengthen — more solar variance between one day and the next. Seasonal naive improves, because a week-old observation is a better guide once the level stops drifting. Averaging them is not a trick that happened to work: each covers the regime where the other is weakest, which is also why per-segment switching gained nothing while blending gained 14%.

  • Daylight hours are ~1.8× harder than dark hours (9.41% vs 5.22%), and MAE rises monotonically with solar elevation, from 1,270 MW in deep dark to 2,670 MW at 30–40°. Embedded PV is not metered at transmission level, so it reaches the forecaster only as unexplained variance in ND.
  • The triad window is the easiest segment, not the hardest — 4.25% MAPE in the winter weekday evenings that set network charges. Demand there is high, stable and strongly habitual. Note that the ensemble is worse here (4.89%) than plain persistence: averaging in a weekly model dilutes exactly the recency that peak hours reward.
  • Weekends are harder than weekdays (9.11% vs 6.62%) and winter is easier than spring for persistence — a level-following model does well when the level is high and stable.

Stage 4 in four findings

The headline number depends on a choice, not on a model. LightGBM reaches 4.48% MAPE with observed weather and 5.46% without. An operational forecaster has a weather forecast, not an observation, so the honest figure for deployment sits closer to the second. Both are reported; only one is defensible as a claim about live performance.

Gradient boosting does not automatically win. On the operational feature set Ridge scored 5.82% against LightGBM's 6.54%. The reason turned out to be a feature, not the algorithm.

Installed PV capacity is a trap for a tree. The column rises monotonically, so 99.7% of validation hours and 100% of test hours sit above the largest value the model saw in training. A tree cannot extrapolate; every split saturates at the top bin, the model under-counts PV suppression, and it over-forecasts. That shows up as a +625 MW bias. Replacing raw capacity with pv_potential_share, which divides potential PV by the recent demand level, took LightGBM from 6.54% to 5.46% and cut the bias from +625 to +80 MW.

The choice between the three PV treatments was made on the 2024 selection set, not on validation:

Model PV treatment Selection MAPE Validation MAPE
Ridge all PV features 6.18 5.82
Ridge no PV at all 5.63 5.80
Ridge stationary PV only 5.62 5.95
LightGBM all PV features 5.16 6.54
LightGBM no PV at all 5.14 5.48
LightGBM stationary PV only 5.13 5.46

The selection set picked the stationary form for both. It transferred for LightGBM and did not transfer for Ridge, where the plain no-PV set scored better on validation. Reported as it happened.

The PV features barely earn their place. Stationary PV beats no PV at all by 0.02 points for LightGBM, which is noise. The lags already carry the signal: yesterday's demand at the same hour has PV suppression baked into it. Adding an explicit PV term mostly re-states what lag_24h already knows. The engineering value here was diagnostic, not predictive.

Reading this table before building anything clever:

  • Persistence beats seasonal naive. Yesterday carries more information than last week, so day-to-day level (weather, in practice) dominates weekly shape. Any model that ignores recent level will lose.
  • Recency is worth ~5.6 points of MAPE — the same calendar model goes from 15.41% to 9.77% purely by rebuilding its table from the last 8 weeks.
  • Calendar shape alone is not enough. Even the rolling calendar model loses to plain persistence, so a pure hour×weekday feature set will not be competitive without lags.
  • 7.33% is the number to beat. A gradient boosting model that lands at 7% would not be worth reporting as a success.

Features (Stage 3)

python -m src.features.build → 34 features in three groups, and the group decides what a model may see.

Group Count Available at the origin?
Known ahead — calendar, holidays, solar geometry, installed capacity 18 Yes
History only — lags and rolling stats, all shifted ≥ 24 h 10 Yes, from the past
Weather observed — temperature, dew point, cloud, radiation 6 No. See below

A test corrupts demand from the forecast origin forward and asserts that no column in the first two groups moves. Weather is excluded from that test on purpose: those columns are observations of the forecast period, which is what the group name records.

Load and temperature: the U is lopsided

Demand
Coldest bin (−6 to −4 °C) 31,491 MW
Minimum (16 to 18 °C) 23,395 MW
Warmest bin (30 to 32 °C) 24,750 MW

The cold arm rises 8,096 MW above the minimum. The warm arm rises 1,356 MW, a gap of roughly six to one, and a single linear temperature term cannot bend at both ends at once. The feature set carries heating and cooling degree hours instead, which keeps each arm separate and each coefficient physically readable.

Most of that six-to-one gap is not heating versus cooling. Regressing demand on degree hours with hour and weekday controls, then adding solar terms one at a time (python -m src.visualization.temperature_sensitivity):

Controls Heating (MW/°C) Cooling (MW/°C) Ratio R²
Hour + weekend only +737 −358 2.1 0.749
+ installed PV × sun height +498 +221 2.3 0.861
+ measured radiation +483 +399 1.2 0.877
+ cloud cover +492 +409 1.2 0.878

Without a solar term the cooling coefficient is negative, which would mean air conditioning reduces electricity use. Hot hours in Britain are sunny hours, so embedded PV suppresses transmission demand at exactly the times cooling load appears, and an uncontrolled regression charges the PV dip to temperature. Once solar is in the model the sign flips and the true asymmetry is 1.2 to 1, not 6 to 1.

What remains of the asymmetry comes from duration and depth rather than from per-degree response: 17.7% of hours sit below 5 °C, against 1.2% above 25 °C, and a British winter runs a 25 K gap to a 20 °C indoor target where a hot day runs 12 K.

Cooling sensitivity also appears to be rising:

Year Heating Cooling Hours >20 °C
2022 +534 +286 518
2023 +489 +321 386
2024 +456 +332 294
2025 +469 +421 648
2026 +361 +494 563

Cooling moved +73% while heating fell 32%. Treat this as a hint rather than a measurement: each year rests on a few hundred hot hours, 2026 stops in July so its heating figure comes from winter months alone, and the model carries no interaction terms.

Embedded solar: the variable the grid cannot see

Installed PV grew 58% across the modelling period (15,036 → 23,805 MW) while wind capacity moved −2%. None of that PV is metered at transmission level, so it reaches the forecaster only as demand that failed to appear.

Summer (May–Jul) midday depression, the gap between the 08:00 shoulder and the 12:00–14:00 trough:

Year 08:00 12–14h Depression Installed PV
2022 24,842 MW 23,922 MW 3.7% 14,363 MW
2023 24,013 MW 22,388 MW 6.8% 15,781 MW
2024 23,948 MW 22,519 MW 6.0% 17,664 MW
2025 23,184 MW 20,575 MW 11.3% 20,786 MW
2026 22,937 MW 19,758 MW 13.9% 23,701 MW

Depression correlates with installed PV at +0.97, which on its own proves nothing. Five points, both series rising with time, and almost any pair would score near 1. Two controls carry the argument instead:

Wind capacity stayed flat while PV grew 58%, so this is not renewables in general.

The placebo: embedded PV cannot act at 02:00. Between 2022 and 2026 midday demand fell 17.4% while overnight demand rose 7.7%. A 25-point divergence inside the same series, appearing only while the sun is up, is much harder to explain away than a correlation across five years.

The feature pv_potential_MW is installed capacity multiplied by the sine of solar elevation: a physical proxy for how much load embedded PV is hiding at that instant. Neither term says it alone, since capacity is flat within a day and the sun knows nothing about how much plant is out there. Across sunlit hours it correlates −0.51 with demand.


Error analysis

The most useful part of this repo. Computed once, on the champion's already- scored test predictions — python -m src.evaluation.error_analysis describes errors that are already fixed, it does not re-select or retune anything.

Error analysis

1. When does the model fail? Daily MAPE on test ranges from 1.58% (2026-02-08) to 20.41% (2026-07-02), median 4.98%. But 8 of the 10 worst days were also above-median for the plain persistence baseline — most "bad days" are genuinely hard demand days for every model, not a champion-specific failure.

2. Holidays degrade the forecast by +1.45pp (7.16% on holiday hours vs 5.71% otherwise, 120 vs 4,871 hours). It is not uniform: Easter Monday (11.52%) and New Year's Day (10.19%) are the worst, while May Day (3.45%) is actually easier than an average day — a single "holiday" feature is carrying at least two different demand patterns.

3. The one DST transition day in the test window (2026-03-29, 23h spring-forward) is a sharp, rare failure: 12.48% MAPE vs 6.48% in the surrounding week, a +6.00pp jump. The autumn transition falls outside the test window (which ends 2026-07-27), so only the spring case is measured here.

4. Temperature extremes. MAPE rises from 4.11% in the second-coldest quintile to 7.26% in the warmest, but this test window is Jan–Jul only, so the "warm" tail is really the "sunny" tail — see the next point. It restates the solar confound from Load and temperature: the U is lopsided, not a genuine cold/heat effect.

5. Peak vs off-peak — the consistent finding across every stage of this project. Midday still dominates the error budget: 9.23% MAPE in the midday load period against 3.4–4.9% everywhere else, and 8.63% in the "high sun (>15°)" solar regime against 4.04% in the dark. The honest model narrowed this gap a great deal from the Stage 2 baselines (10.64% midday / 5.22% dark) but never closed it. The triad window — winter weekday 16:00–19:00, which sets GB transmission charges — remains the easiest segment at 3.45%.

6. Bias. The champion over-forecasts by +148.6 MW overall, but the sign flips by month: February is −460.7 MW (under), July is +1,065.2 MW (over). By day/night: +358.5 MW over-forecast in daylight, −102.1 MW under-forecast at night — the model still does not fully price in PV suppression at the moments it matters most, the same direction of error found back in Stage 4's raw-capacity ablation.

Feature importance

SHAP explanations for the deployed champion (python -m src.evaluation.shap_analysis), computed on the same test predictions above — opening up the model already selected, not choosing a different one.

SHAP summary

Feature Mean |SHAP| (MW) Share
lag_24h 2,885 36.5%
lag_168h 1,085 13.7%
dayofweek 796 10.1%
pv_potential_share 320 4.0%
lag_336h 290 3.7%
solar_gen_lag_24h 227 2.9%

lag_24h dominates, which is exactly what the results table already implies — persistence alone reaches 8.26% MAPE, so the model leans hard on yesterday's level and refines it. dayofweek's dependence plot is a clean step function: positive SHAP on weekdays, negative on weekends, no surprises.

SHAP dependence, top 4 features

The interesting one is pv_potential_share: SHAP value decreases monotonically as the feature rises, and the effect is strongest exactly when solar elevation (the colour axis) is highest. That is the model internals confirming, independently of any correlation analysis, the same PV-suppression mechanism argued for in Embedded solar: the variable the grid cannot see — this is not a spurious association, the tree has genuinely encoded "more sun and more installed PV → suppress the forecast."


Reproducing

git clone https://github.com/Strooms/Electrical-Load-Forecasting-Benchmark.git
cd Electrical-Load-Forecasting-Benchmark

python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -r requirements.txt

cp .env.example .env             # add API credentials if using ENTSO-E

python -m src.data.download           # fetch raw data
python -m src.data.download_weather   # fetch Open-Meteo weather
python -m src.data.check_completeness # audit the raw download
python -m src.data.make_dataset       # clean + resample to hourly

python -m src.evaluation.report       # Stage 2: baselines on validation
python -m src.evaluation.diagnose     # Stage 2: configuration search

python -m src.features.build          # Stage 3: feature matrix
python -m src.models.train            # Stage 4: Ridge/LightGBM on validation

python -m src.evaluation.test_score      # Stage 5: final test-set score (once)
python -m src.evaluation.error_analysis  # Stage 5: where the error sits
python -m src.evaluation.shap_analysis   # Stage 5: SHAP interpretability
python -m src.visualization.header_figure  # regenerate the header image

pytest                                # 34+ tests, incl. leakage/mutation checks

Random seeds are fixed in configs/. Exact package versions that produced the reported numbers are pinned in requirements.txt (Python 3.13.14).


Repository layout

├── configs/            # YAML configs — no hardcoded parameters in code
├── data/               # gitignored
│   ├── raw/            # immutable download, never edited
│   ├── interim/        # cleaned, resampled
│   └── processed/      # model-ready feature matrices
├── docs/
│   ├── DATA.md         # acquisition + preprocessing detail
│   └── DECISIONS.md    # design decisions and why
├── notebooks/          # exploration only — logic lives in src/
├── reports/figures/
├── src/
│   ├── data/           # download, clean, resample
│   ├── features/       # calendar, lags, weather
│   ├── models/         # baselines + ML models, shared interface
│   ├── evaluation/     # metrics, backtesting, error analysis
│   └── visualization/
└── tests/

Limitations

  • Single grid, single climate. Every number here is GB transmission demand. The physics-vs-convention split in region_profile (see configs/config.yaml) is built for portability, but it has never actually been run against a second country's data.
  • The optimistic figure needs a weather forecast feed this benchmark does not have. 4.87% MAPE uses observed weather for the forecast period, which does not exist at 00:00 the day before. The 0.88pp gap to the 5.75% honest number is the real cost of that missing feed, not modelling error.
  • Point forecasts only. No quantile or interval output, so there is no direct answer to "how much reserve margin does a 95% confidence forecast imply" — the question this whole project exists to eventually serve.
  • No exogenous economic drivers. Price, industrial output, and EV charging load are all absent. Some of that signal likely already leaks in through the lag features, uncredited.
  • The test window is 7 months, not a full year. 2026-01-01 to 2026-07-27 has no autumn DST transition, no true summer heat extreme, and no Christmas/New Year holiday cluster beyond New Year's Day itself. The headline 5.75% is real but was not tested against a full seasonal cycle.
  • DST handling is structurally correct but not specifically modelled. The harness produces 23/25-hour days correctly (Stage 2), but no feature flags "this is a transition day" — which is very likely why it is the sharpest single-day failure mode found in Stage 5 (+6.00pp).
  • England-only holiday calendar. holidays.UnitedKingdom(subdiv="England") is a known simplification; Scotland keeps a different calendar and is folded into the same national series.

What I would do next

  • A DST-transition flag. The sharpest single-day failure found in Stage 5 (+6.00pp on 2026-03-29) has an obvious, cheap fix that was never tried: a binary feature for "today has 23 or 25 hours," letting the model treat it as the irregular case it is instead of silently forcing 24-hour lag logic onto it.
  • Quantile regression (pinball loss) alongside the point forecast, so reserve sizing has an actual interval to work from instead of a single number and an MAE that says nothing about the tails that matter.
  • A real weather forecast feed, to find out how much of the 0.88pp gap between the honest and optimistic models survives contact with an actual forecast API instead of observed weather standing in for one.
  • Hierarchical reconciliation down to regional or feeder level, where the embedded-PV problem this project spent three stages on gets worse, not better — capacity is lumpier and less diversified at smaller scale.

References

  • National Energy System Operator (NESO). Historic Demand Data (GB). neso.energy/data-portal/historic-demand-data, used under the NESO Open Data Licence.
  • Open-Meteo. Historical Weather API. open-meteo.com/en/docs/historical-weather-api.
  • Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q. and Liu, T.Y., 2017. LightGBM: A highly efficient gradient boosting decision tree. Advances in Neural Information Processing Systems, 30.
  • Lundberg, S.M. and Lee, S.I., 2017. A unified approach to interpreting model predictions (SHAP). Advances in Neural Information Processing Systems, 30.
  • dr-prodigy/python-holidays — UK bank holiday calendar (holidays.UnitedKingdom).

Author: Raditya Yoga K. · LinkedIn License: MIT (code). Data remains under its original terms — see above.

About

No description, website, or topics provided.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages