How model calibration improves sports forecasts
Calibration makes a model's percentages trustworthy. What it is, how it is measured with reliability charts, the methods used to fix it, and why it can backfire on small samples.
By Patrick C · Founder & Editor, EdgeIQ
In this article
A calibrated model is one whose probabilities can be taken at face value. When it says 70%, events like that happen about 70% of the time. Calibration does not make a model smarter at telling good teams from bad ones; it makes the numbers honest. For anyone reading probabilities, that honesty is what matters most.
Ranking versus scaling
It helps to separate two jobs a model does. The first is ranking: putting stronger favorites ahead of weaker ones. The second is scaling: attaching the right number to each. Many machine-learning methods are good at ranking but produce poorly scaled outputs — some are systematically too cautious, pushing everything toward 50%, and others are too bold, pushing probabilities toward the extremes. Calibration is a separate step that fixes scaling without disturbing ranking.
Seeing calibration: the reliability chart
The standard picture is a reliability chart. Sort past forecasts into bins by stated probability, then plot each bin's average forecast on one axis and the share that actually happened on the other. A perfectly calibrated model sits on the diagonal line.
- Points below the diagonal at high probabilities mean the model is overconfident: its 80% calls win less than 80% of the time.
- Points above the diagonal at high probabilities mean it is underconfident.
- Bins with few games bounce around; always look at the count behind each point.
Model Lab shows this chart for EdgeIQ's models, with the number of games in each band, so you can judge the evidence rather than a summary.
Common calibration methods
Temperature or Platt scaling
A simple, smooth adjustment: one or two parameters that stretch or shrink every probability toward or away from 50%. Because it has so few moving parts it is hard to overfit and works well when the model's error is uniform — for example, consistently a bit too bold.
Isotonic regression
A flexible, step-shaped mapping that can fix uneven errors — for instance, a model that is fine at 60% but overconfident at 85%. The price of flexibility is data hunger: with too few games per step, isotonic calibration memorizes noise and makes probabilities worse.
| Method | Strength | Risk |
|---|---|---|
| No calibration | Nothing to overfit | Leaves known bias in place |
| Temperature / Platt | Stable with modest data | Cannot fix uneven errors |
| Isotonic | Fixes complex patterns | Overfits small samples |
How EdgeIQ chooses
EdgeIQ's engine does not assume calibration helps. For each league and market, it holds out a validation window of games the model did not train on, then scores three options on that window: no calibration, temperature scaling and isotonic regression. A method is only adopted if it improves log loss on the held-out games by a meaningful margin, and isotonic steps require a minimum number of games behind each one. If nothing clears that bar, the raw probabilities are kept.
This matters because the tempting mistake is to calibrate on the same games you evaluate on. That always looks like an improvement and often is not.
A real example of miscalibration
During development, EdgeIQ's college football engine showed overconfidence on heavy favorites: in the highest probability bands, favorites won less often than stated. College football has enormous gaps between the strongest and weakest programs, so raw models are tempted into extreme numbers. Identifying that pattern is exactly what reliability charts are for, and it is noted openly in our evaluation.
By contrast, the NFL engine's backtest bands landed close to target — 65–70% forecasts won 69% of the time and 70–80% forecasts won 79% of the time across 768 unseen games.
What calibration cannot fix
- A model with no real signal. Calibrating a coin flip produces a well-calibrated 50% for every game — honest, and useless.
- Drift. A mapping learned last season may not fit this one; it needs regular re-checking.
- Tiny samples. With 30 games, a reliability chart is mostly noise.
Calibration turns a ranking into a trustworthy probability. Without it, a percentage is just a score with a % sign attached.
Why readers should care
If you ever compare a model's probability with another number — your own judgment, a friend's opinion, or a market price — the comparison only makes sense if the model's percentages are calibrated. An uncalibrated 75% cannot be meaningfully compared with anything. That is why calibration sits alongside accuracy and Brier score on every EdgeIQ model report.
Sources and further reading
- Niculescu-Mizil, A. & Caruana, R. (2005). Predicting good probabilities with supervised learning. Proceedings of ICML 2005.
- Murphy, A. H. (1973). A new vector partition of the probability score. Journal of Applied Meteorology, 12(4).
- Gneiting, T. & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477).
About the author
Patrick C founded EdgeIQ and edits its Learning Center. He oversees how EdgeIQ collects sports data, how its prediction models are evaluated and how results are explained to readers.
EdgeIQ is a sports analytics platform. Projections are statistical estimates based on historical data, carry uncertainty and are never guarantees. EdgeIQ takes no wagers and is not a sportsbook. Read the methodology and analytics disclaimer.