Selective Calibration for Probabilistic Forecast Enhancement
SCOPE is a post-processing methodology for improving the probabilistic forecasts of frozen time-series foundation models. Its central premise is that a foundation model already contains useful predictive structure, but a revision should be applied only where its expected benefit is supported by information available at forecast time.
The working paper title is:
SCOPE: Label-Safe Selective Calibration for Frozen Time-Series Foundation Models
Frozen generative forecasters can provide well-shaped forecast distributions without task-specific retraining. In renewable power forecasting and other non-stationary settings, however, a globally applied correction can improve a subset of windows while degrading another subset. Point-error optimization alone also risks distorting forecast intervals.
SCOPE treats post-processing as a selective decision problem. It separates two questions that are often conflated:
- Should a candidate revision replace the base distribution for this forecast block?
- Once a distribution is selected, how should its uncertainty be calibrated without moving the central point forecast unnecessarily?
For an input context (x), a frozen generation model produces base quantiles
[ Q^0(x) = {q^0_{\tau}(x)}_{\tau \in \mathcal{T}}, ]
where (\mathcal{T}) is an ordered set of quantile levels. An optional upstream adaptation path produces a candidate distribution (Q^1(x)). SCOPE does not require a particular foundation model or candidate generator: Chronos, TimeFM, TimeMoE, and other generators can supply the same quantile interface.
The goal is to produce a final distribution (Q^{\mathrm{SCOPE}}(x)) with lower point and probabilistic loss, while controlling harmful revisions and interval miscalibration.
SCOPE contains two public stages.
BSAP groups forecast horizons into local blocks and learns a selector from forecast-time features only. These features describe the base and candidate quantile distributions, their relative displacement and width, horizon phase, persistence disagreement, and any upstream routing weights. They do not include the future target.
During validation, a block receives a positive label only when the candidate improves mean absolute error and does not violate a pinball-loss guard relative to the base distribution. A threshold selected on a held-out validation subset then controls the final action:
[ Q^{\mathrm{BSAP}}(x) = \begin{cases} Q^1(x), & s(x) \geq \gamma, \ Q^0(x), & s(x) < \gamma, \end{cases} ]
where (s(x)) is the selector score and (\gamma) is the validation-fitted threshold. Block-level routing prevents isolated horizon decisions from creating an incoherent forecast trajectory.
MLQC calibrates the selected distribution using validation residual structure. It considers intercept and conformal-style interval adjustments, then selects the candidate that balances pinball loss and coverage error. By default, the median is locked:
[ q_{0.5}^{\mathrm{MLQC}}(x) = q_{0.5}^{\mathrm{BSAP}}(x). ]
This makes uncertainty calibration explicit: MLQC can repair interval placement or width without silently claiming point-forecast gains that it did not create. Monotonicity is enforced after calibration so quantile crossing cannot occur.
All fitting decisions use only labeled validation windows. The validation set is internally partitioned into selector-training and threshold/calibration holdout windows. Test windows are never used to train the selector, choose a threshold, or select a calibration rule.
At deployment time, SCOPE accepts a label-free forecast frame. The inference
interface rejects y_true by design, which makes accidental target use visible
rather than implicit. The fitted artifact stores only model state, feature
schema, threshold, and calibration parameters; it excludes validation rows,
labels, and selection sweeps.
Quantile Distribution Alignment Layer (Q-DAL) is an optional upstream candidate generator. It can transform a routed base distribution before BSAP, but it is not an always-on public SCOPE stage. This preserves the methodological claim: SCOPE evaluates and selectively calibrates candidate forecast distributions, regardless of how a valid candidate was produced.
SCOPE is evaluated with complementary point, distributional, and revision-risk criteria: MAE, MSE, pinball loss, interval coverage error, interval width, improvement rate, and harmful revision rate. The method is intended to show not only whether a correction can help on average, but whether the decision to apply it is stable, calibrated, and causally valid at inference time.