Football probability calibration: making numbers mean what they say
The essential promise of a probability forecast is simple: among many comparable events assigned 60%, roughly six in ten should happen over time.
Updated 2026-06-22
Probability is not decorative confidence
A 60% forecast is not 'safe'; it leaves roughly a 40% alternative. That distinction keeps a model output from being mistaken for certainty.
How to test calibration
Bucket historical forecasts by probability and compare their average prediction with observed frequency. A reliable model keeps them close and reports uncertainty from finite samples.
When to trust less
A new season, sparse competitions, major squad changes or missing data can break historical calibration. Forecasts should show refresh time and coverage.
FAQ
Are higher probabilities always more reliable?
Not automatically. A high forecast can still be overconfident; inspect its historical calibration for comparable cases.
Can calibration guarantee a single match?
No. Calibration is a long-run group property, not a promise about any one result.