Prediction accuracy vs probability calibration
Accuracy asks whether the favorite won. Calibration asks whether the percentages were honest. Why you need both, how they can disagree, and which to trust when they do.
By Patrick C · Founder & Editor, EdgeIQ
In this article
"How accurate is it?" is the first question people ask about any prediction model. It is a reasonable question with a surprisingly incomplete answer. Accuracy captures one dimension of forecast quality. Calibration captures another. A model can be strong on one and weak on the other, and knowing the difference protects you from misleading claims.
Definitions
| Measure | Question it answers | What it ignores |
|---|---|---|
| Accuracy | How often did the favored team win? | How confident the forecast was |
| Calibration | When the model said X%, did it happen about X% of the time? | Whether the model separated winners from losers well |
| Brier score / log loss | Overall quality of the probabilities | Harder to explain in one sentence |
How they can disagree
Accurate but miscalibrated
A model picks the winner 65% of the time but labels every favorite 90%. Its picks are decent, but its numbers are badly overconfident: its "90%" calls win only 65% of the time. Anyone taking those percentages literally will be misled.
Calibrated but not useful
A model says 55% for the home team in every game. If home teams win 55% of the time, it is perfectly calibrated — and almost worthless, because it never distinguishes one game from another.
The goal
A good model does both: it gives strong favorites high numbers and weak favorites low numbers (discrimination), and those numbers match reality (calibration). Proper scoring rules such as the Brier score and log loss reward both at once, which is why serious evaluations lead with them.
Why accuracy alone is easy to game
- Cherry-picking: reporting accuracy only on high-confidence games, or only on a good stretch.
- Easy slates: accuracy looks great in weeks full of lopsided matchups and poor in weeks full of close ones, regardless of model quality.
- No baseline: 60% sounds impressive until you learn a simple rating system gets 61% on the same games.
An example from EdgeIQ
On its 768-game NFL walk-forward backtest, EdgeIQ's engine picked 60.3% of winners — while the ratings reference picked 60.8%. On accuracy alone, the engine looks slightly behind. On calibration, its bands landed close to target (65–70% forecasts won 69%; 70–80% forecasts won 79%). Neither number alone tells the full story; together they say the engine is roughly as good as a strong baseline and its percentages are trustworthy on that sample.
Always ask for accuracy, calibration and a baseline — on the same games, tested chronologically.
Which to trust when they conflict
If you use a model's percentages — comparing them to your own judgment or to other numbers — calibration and Brier score matter more than raw accuracy. If you only care which team is favored, accuracy is a reasonable summary, but it should still be compared with a baseline on the same games.
How EdgeIQ reports both
Model Lab shows accuracy, Brier score, log loss and reliability charts for each league and model version, with sample sizes. Live-season results are kept separate from chronological backtests, and figures are suppressed when the sample is too small to mean anything.
Sources and further reading
About the author
Patrick C founded EdgeIQ and edits its Learning Center. He oversees how EdgeIQ collects sports data, how its prediction models are evaluated and how results are explained to readers.
EdgeIQ is a sports analytics platform. Projections are statistical estimates based on historical data, carry uncertainty and are never guarantees. EdgeIQ takes no wagers and is not a sportsbook. Read the methodology and analytics disclaimer.