Risk: Conformal, ACI, EVT & Epistemic Uncertainty¶
Navigation:
Theory introduction: See the Intro
Synthesis: One Atom, Many Statistics
Related implementation: Risk (architecture)
This chapter establishes the mathematical foundation for the framework’s uncertainty quantification. The risk layer bounds predictions after the fit. It operates on raw arrays / duck-typed model outputs, so any meta-model’s predictions flow in identically.
One source of truth: the static engine SafetyTAM holds the finite-sample quantile and p-value; the layers below build on it.
Split Conformal Prediction & CQR¶
Unlike traditional Bayesian confidence intervals (which depend heavily on prior distributional assumptions) or bootstrap methods, Conformal Prediction offers a finite-sample marginal coverage guarantee without assuming anything about the underlying error distribution, provided the data points are exchangeable [Angelopoulos and Bates, 2023, Vovk et al., 2005].
The framework implements the Split Conformal method [Lei et al., 2018] to ensure computational efficiency. The mathematical procedure is defined as follows:
Data Splitting: The historical data is divided into a disjoint Training set and a Calibration set of size \(n\).
Non-Conformity Scores: We define a non-conformity measure. For standard regression, this is the absolute residual \(s_i = \left| Y_i - \Phi_i \hat{\theta} \right|\).
ConformalDistributionalTAMuses a Conformalized Quantile Regression (CQR) score [Romano et al., 2019] defined as \(s_i = \max(\hat{Q}_{\alpha/2}-Y_i,\ Y_i-\hat{Q}_{1-\alpha/2})\) and supports Mondrian (per-stratum) calibration.Finite-Sample Quantile: To guarantee exact marginal coverage of \(1 - \alpha\), we compute the empirical quantile \(\hat{q}\) of the calibration scores at a corrected level: $\( \hat q = \mathrm{Quantile}\!\left(\{s_i\},\ (1-\alpha)\big(1+\tfrac1n\big)\right) \)$
Prediction Interval: For any new observation at time \(t\), the prediction interval is symmetrically defined around the baseline by \(\hat{q}\).
Adaptive Conformal Inference (ACI)¶
The fundamental vulnerability of the Static mode in industrial time series is that the exchangeability assumption is frequently violated. Continuous Concept Drift causes the underlying data distribution to shift [Principato and Stoltz, 2025].
To maintain valid coverage in non-stationary environments, TAM implements the Adaptive Conformal Inference (ACI) algorithm [Gibbs and Candes, 2021]. Instead of relying on a fixed historical quantile, ACI dynamically adjusts the risk parameter \(\alpha_t\) at every time step using an online feedback loop:
If the model makes an error (\(\text{err}_t = 1\)), the effective \(\alpha_{t+1}\) decreases, forcing the model to widen the interval for the next step. If safely covered (\(\text{err}_t = 0\)), \(\alpha_{t+1}\) increases, tightening the interval. Tracking cumulative miscoverage errors guarantees asymptotic long-term validity even under arbitrary, adversarial distribution shifts [Principato and Stoltz, 2025].
Extreme Value Theory (EVT)¶
A parametric location-scale fit constrains the whole shape and may misfit the extreme tail. By the Pickands-Balkema-de Haan theorem [Pickands III, 1975, Balkema and De Haan, 1974], exceedances of a high threshold \(u\) follow a Generalized Pareto Distribution (GPD). fit_gpd_tail models these upper tails, giving asymptotically-justified tail probabilities and an EVT surprisal score \(-\log P(M>m)\) for severe anomalies.
Epistemic (Parameter) Uncertainty¶
The penalized fit is a MAP estimate under a Gaussian prior [Wahba, 1983]. The posterior covariance is \(V_\beta=\sigma^2(\Phi^\top\Phi+nS)^{-1}\), and the epistemic standard error \(\sqrt{\phi V_\beta \phi^\top}\) widens where the training design is data-poor.
Note: This is computed via an efficient Cholesky solve, never requiring a dense matrix inverse.
Hierarchical Limitations & Future Upgrades¶
The current SafetyTAM module operates marginally—meaning it calculates safety bounds for each node independently. However, Conformal Prediction intervals do not obey linear summation. Summing the marginal intervals of the children (via Minkowski sum) does not mathematically yield the correct joint interval for the aggregated parent [Principato et al., 2024]. Applying ACI independently level-by-level inevitably breaks the physical realities of the hierarchy.
Roadmap: Future iterations of the framework must implement interval reconciliation via polytope projection to strictly enforce hierarchical aggregation constraints across the entire uncertainty band [Principato et al., 2024].