Statistics & Probability
Published · 8 min read

Understanding Brier score in sports predictions

The Brier score measures how good probability forecasts are, not just whether picks were right. Worked examples, benchmarks and the common traps when comparing models.

By Patrick C · Founder & Editor, EdgeIQ

Accuracy — the share of games where the favorite won — is the number everyone quotes. It is also crude. It treats a 51% forecast and a 95% forecast the same way as long as the same team was favored. The Brier score fixes that by grading the probability itself.

The definition

The Brier score, introduced by meteorologist Glenn Brier in 1950 to grade weather forecasts, is the average squared difference between a forecast probability and what happened. The outcome is recorded as 1 if the event happened and 0 if it did not.

Brier score = average of (forecast probability − outcome)², where outcome is 1 or 0. Lower is better. 0 is perfect.

Worked examples

Forecast for Team AResultCalculationScore
70%A wins (1)(0.70 − 1)² = 0.30²0.09
70%A loses (0)(0.70 − 0)² = 0.70²0.49
50%Either(0.50 − 1)² or (0.50 − 0)²0.25
95%A wins (1)(0.95 − 1)²0.0025
95%A loses (0)(0.95 − 0)²0.9025

Two things jump out. First, confident forecasts are rewarded heavily when right and punished severely when wrong. Second, always saying 50% produces exactly 0.25 every time. That makes 0.25 the natural "knows nothing" benchmark for a two-outcome game.

Why squaring matters

Squaring the error makes the score a proper scoring rule: a forecaster gets the best expected score by reporting what they actually believe. If you think a team wins 70% of the time, reporting 90% to look bold will hurt your average over many games. Reporting 55% to look cautious will also hurt it. Honest probabilities win in the long run, which is exactly the incentive you want.

What counts as a good Brier score?

There is no universal target, because it depends on how predictable the sport is. In a league where many games are close to coin flips, even an excellent model will score near 0.23–0.24. In a league with large gaps between teams, a good model can go meaningfully lower. The useful comparison is always relative:

  • Against the 0.25 coin-flip baseline.
  • Against a simple rule such as "home team 57%".
  • Against a basic rating model.
  • Against the previous version of the same model on the same games.

A new model that scores 0.215 when the rating reference scores 0.218 on the same games has improved slightly. A model that scores 0.215 on one season and is compared with a rival's 0.225 on a different season tells you nothing, because the seasons differ in predictability.

Accuracy and Brier score can disagree

Imagine two models over four games where the favorite wins three times. Model A says 55% for every favorite. Model B says 80%. Both have 75% accuracy. Their Brier scores differ:

ModelThree winsOne lossAverage Brier
A (55%)3 × 0.20251 × 0.30250.2275
B (80%)3 × 0.041 × 0.640.19

Model B scores better because its confidence matched reality: favorites won 75% of the time and it said 80%, while A said 55%. Accuracy could not see that difference at all.

Breaking the score apart

Statisticians decompose the Brier score into three parts: reliability (does 70% really mean 70%? — this is calibration), resolution (does the model separate likely winners from unlikely ones, or does it say 55% for everything?), and uncertainty (how unpredictable the games were to begin with, which the model cannot control). A model can improve its Brier score either by becoming better calibrated or by becoming more discriminating. The best models do both.

Traps to avoid

  • Comparing Brier scores computed on different sets of games.
  • Reporting a Brier score on training data. It will always look better than on unseen games.
  • Treating tiny differences as meaningful on small samples. On 100 games, a 0.003 difference is well within noise.
  • Forgetting draws or ties in sports where they happen; the outcome coding must match the question being asked.

How EdgeIQ uses it

EdgeIQ reports Brier score alongside accuracy, log loss and calibration for every model version, both in chronological backtests and in the live record. When a candidate model is considered for promotion, it has to hold up on these probability-quality measures on games it never saw — not just pick more winners.

Sources and further reading

About the author

Patrick C founded EdgeIQ and edits its Learning Center. He oversees how EdgeIQ collects sports data, how its prediction models are evaluated and how results are explained to readers.

EdgeIQ is a sports analytics platform. Projections are statistical estimates based on historical data, carry uncertainty and are never guarantees. EdgeIQ takes no wagers and is not a sportsbook. Read the methodology and analytics disclaimer.

Related articles