Understanding Brier score in sports predictions
The Brier score measures how good probability forecasts are, not just whether picks were right. Worked examples, benchmarks and the common traps when comparing models.
By Patrick C · Founder & Editor, EdgeIQ
In this article
Accuracy — the share of games where the favorite won — is the number everyone quotes. It is also crude. It treats a 51% forecast and a 95% forecast the same way as long as the same team was favored. The Brier score fixes that by grading the probability itself.
The definition
The Brier score, introduced by meteorologist Glenn Brier in 1950 to grade weather forecasts, is the average squared difference between a forecast probability and what happened. The outcome is recorded as 1 if the event happened and 0 if it did not.
Brier score = average of (forecast probability − outcome)², where outcome is 1 or 0. Lower is better. 0 is perfect.
Worked examples
| Forecast for Team A | Result | Calculation | Score |
|---|---|---|---|
| 70% | A wins (1) | (0.70 − 1)² = 0.30² | 0.09 |
| 70% | A loses (0) | (0.70 − 0)² = 0.70² | 0.49 |
| 50% | Either | (0.50 − 1)² or (0.50 − 0)² | 0.25 |
| 95% | A wins (1) | (0.95 − 1)² | 0.0025 |
| 95% | A loses (0) | (0.95 − 0)² | 0.9025 |
Two things jump out. First, confident forecasts are rewarded heavily when right and punished severely when wrong. Second, always saying 50% produces exactly 0.25 every time. That makes 0.25 the natural "knows nothing" benchmark for a two-outcome game.
Why squaring matters
Squaring the error makes the score a proper scoring rule: a forecaster gets the best expected score by reporting what they actually believe. If you think a team wins 70% of the time, reporting 90% to look bold will hurt your average over many games. Reporting 55% to look cautious will also hurt it. Honest probabilities win in the long run, which is exactly the incentive you want.
What counts as a good Brier score?
There is no universal target, because it depends on how predictable the sport is. In a league where many games are close to coin flips, even an excellent model will score near 0.23–0.24. In a league with large gaps between teams, a good model can go meaningfully lower. The useful comparison is always relative:
- Against the 0.25 coin-flip baseline.
- Against a simple rule such as "home team 57%".
- Against a basic rating model.
- Against the previous version of the same model on the same games.
A new model that scores 0.215 when the rating reference scores 0.218 on the same games has improved slightly. A model that scores 0.215 on one season and is compared with a rival's 0.225 on a different season tells you nothing, because the seasons differ in predictability.
Accuracy and Brier score can disagree
Imagine two models over four games where the favorite wins three times. Model A says 55% for every favorite. Model B says 80%. Both have 75% accuracy. Their Brier scores differ:
| Model | Three wins | One loss | Average Brier |
|---|---|---|---|
| A (55%) | 3 × 0.2025 | 1 × 0.3025 | 0.2275 |
| B (80%) | 3 × 0.04 | 1 × 0.64 | 0.19 |
Model B scores better because its confidence matched reality: favorites won 75% of the time and it said 80%, while A said 55%. Accuracy could not see that difference at all.
Breaking the score apart
Statisticians decompose the Brier score into three parts: reliability (does 70% really mean 70%? — this is calibration), resolution (does the model separate likely winners from unlikely ones, or does it say 55% for everything?), and uncertainty (how unpredictable the games were to begin with, which the model cannot control). A model can improve its Brier score either by becoming better calibrated or by becoming more discriminating. The best models do both.
Traps to avoid
- Comparing Brier scores computed on different sets of games.
- Reporting a Brier score on training data. It will always look better than on unseen games.
- Treating tiny differences as meaningful on small samples. On 100 games, a 0.003 difference is well within noise.
- Forgetting draws or ties in sports where they happen; the outcome coding must match the question being asked.
How EdgeIQ uses it
EdgeIQ reports Brier score alongside accuracy, log loss and calibration for every model version, both in chronological backtests and in the live record. When a candidate model is considered for promotion, it has to hold up on these probability-quality measures on games it never saw — not just pick more winners.
Sources and further reading
- Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1).
- Murphy, A. H. (1973). A new vector partition of the probability score. Journal of Applied Meteorology, 12(4).
- Gneiting, T. & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477).
About the author
Patrick C founded EdgeIQ and edits its Learning Center. He oversees how EdgeIQ collects sports data, how its prediction models are evaluated and how results are explained to readers.
EdgeIQ is a sports analytics platform. Projections are statistical estimates based on historical data, carry uncertainty and are never guarantees. EdgeIQ takes no wagers and is not a sportsbook. Read the methodology and analytics disclaimer.