Uncertainty Quantification in Forecast Comparisons

May 5, 20262605.03997

Marc-Oliver Pohle, Tanja Zahn, Sebastian Lerch

stat.MEecon.EMphysics.ao-ph

TLDR

This paper introduces simultaneous confidence bands to rigorously quantify uncertainty in multi-dimensional forecast comparisons, addressing the multiple comparison problem.

Key contributions

Introduces simultaneous confidence bands for expected and skill scores.
Provides a statistically rigorous framework for multi-dimensional forecast evaluation.
Addresses the multiple comparison problem, preventing inflated Type I error rates.
Applicable to diverse forecast types, from mean to full distributional.

Why it matters

Quantifying uncertainty in forecast comparisons is crucial for valid inference, especially in complex, multi-dimensional scenarios. This framework offers a robust solution, preventing inflated Type I error rates and enabling reliable joint evaluations. It's vital for fields like economics and meteorology.

Original Abstract

Skill scores, which measure the relative improvement of a forecasting method over a benchmark via consistent scoring functions and proper scoring rules, are a standard tool in forecast evaluation, yet their sampling uncertainty is rarely rigorously quantified. With modern forecasting applications being increasingly multivariate and involving evaluations across multiple horizons, variables, spatial locations, and forecasting methods, standard tools like the pairwise Diebold-Mariano forecast accuracy test or pointwise confidence intervals fail to account for the multiple comparison problem, leading to inflated Type I error rates and invalid joint inference. To address the lack of a coherent, statistically rigorous framework for quantifying uncertainty across these multi-dimensional evaluation problems, we introduce simultaneous confidence bands for expected scores and skill scores. Our framework provides a versatile tool for joint inference that is applicable to any forecast type from mean and quantile to full distributional forecasts. We develop a bootstrap implementation and show that our bands are valid under multivariate extensions of the classical Diebold-Mariano assumptions. We demonstrate the practical utility of the approach in two case studies by quantifying the benefits of time-varying parameter models for macroeconomic forecasting, and by comparing data-driven and physics-based models in probabilistic weather forecasting.

View on arXiv Download PDF

📬 Weekly AI Paper Digest

Get the top 10 AI/ML arXiv papers from the week — summarized, scored, and delivered to your inbox every Monday.

TLDR

Key contributions

Why it matters

Original Abstract

📬 Weekly AI Paper Digest

Related papers