Why historical performance doesn't guarantee future results
Backtests describe the past. Regression to the mean, changing conditions, small samples and selection effects all explain why a strong record can fade — and how to set realistic expectations.
By Patrick C · Founder & Editor, EdgeIQ
In this article
Every model evaluation looks backward. Even a carefully run chronological backtest answers one question: how would this model have done on games that already happened? Whether it keeps performing depends on things the backtest cannot see. Understanding why protects you from over-trusting any record, good or bad.
Regression to the mean
Any measured performance is partly skill and partly luck. When a record is unusually good, some of that is likely luck, and luck does not persist. Over time, extreme results tend to drift back toward the average. This applies to teams — a team that wins many close games one season usually wins fewer the next — and it applies equally to prediction models and the people who pick games.
The world changes
- Rules change: overtime formats, pace-of-play rules and scoring rules shift how games unfold.
- Rosters turn over: star players retire, get traded or get hurt.
- Strategy evolves: leagues adapt to analytics themselves, eroding old edges.
- Environments change: crowd effects, travel and scheduling vary between eras.
A model trained on older seasons encodes the relationships of those seasons. When conditions move, its assumptions go stale — a problem statisticians call drift.
Small samples mislead in both directions
A short great run can make a mediocre model look brilliant, and a short bad run can make a good model look broken. The table below shows how wide the plausible range of results is for a model whose true winner-pick rate is 60%.
| Games evaluated | Rough range of observed accuracy (about 95% of the time) |
|---|---|
| 20 | about 38% – 82% |
| 100 | about 50% – 70% |
| 400 | about 55% – 65% |
| 1,000 | about 57% – 63% |
With 20 games, a genuinely 60% model can plausibly show anything from below 40% to above 80%. That is why EdgeIQ hides accuracy figures on thin samples and shows the count behind every number.
Selection effects
When many people or models make predictions, some will have excellent records by chance alone. If you only hear about the winners, the past looks far more predictive than it is. The same happens inside a single model's development: trying many variations and keeping the best-looking one on the same historical data produces an optimistic result. The defense is to hold back data the model never touched and evaluate on it only once.
What a backtest can legitimately tell you
- Whether the model has real signal compared with simple baselines.
- Whether its probabilities were calibrated on unseen games.
- Where it is weak — particular leagues, markets or confidence ranges.
What it cannot tell you is that next season will look the same. EdgeIQ's NFL engine, for instance, matched a strong rating baseline on 768 unseen games. That is evidence the model is sound; it is not a promise about any upcoming week.
Past performance is evidence, not a guarantee. Treat it as a reason for measured trust, and keep checking the live record.
How EdgeIQ handles this
EdgeIQ keeps live-season grades separate from backtests so you can see whether real-world results are tracking historical ones. Model versions are recorded permanently, new versions are only promoted after passing held-out tests, and results remain visible when performance is weak. Projections are always presented as estimates, never guarantees.
Sources and further reading
About the author
Patrick C founded EdgeIQ and edits its Learning Center. He oversees how EdgeIQ collects sports data, how its prediction models are evaluated and how results are explained to readers.
EdgeIQ is a sports analytics platform. Projections are statistical estimates based on historical data, carry uncertainty and are never guarantees. EdgeIQ takes no wagers and is not a sportsbook. Read the methodology and analytics disclaimer.