How to evaluate a sports prediction model
A practical checklist for judging any sports prediction source: sample size, chronological testing, baselines, probability quality, transparency and the red flags that should make you walk away.
By Patrick C · Founder & Editor, EdgeIQ
In this article
Sports prediction is crowded with bold claims and very little evidence. You do not need a statistics degree to separate the credible from the rest; you need a handful of questions and the willingness to ask them. This checklist works for any source, including EdgeIQ.
1. Is there a full, public record?
The first test is simple: can you see every prediction, including the misses, recorded before the games were played? A record that only shows highlights, or that is assembled after the fact, cannot be evaluated. Look for timestamps and a policy that predictions are never edited after kickoff. EdgeIQ stores every projection with its model version and data cutoff and does not overwrite it.
2. How big is the sample?
A 10–2 week proves nothing. Random chance alone produces hot streaks. As a rough guide, dozens of games start to hint at a pattern, and hundreds are needed for confident conclusions — more in high-randomness sports like baseball and hockey. Be wary of any accuracy figure presented without a count of games behind it.
3. Was it tested on unseen games, in time order?
Ask whether the model's historical performance comes from games it had never seen, predicted in chronological order using only information available at the time. If a source tested on data it trained on, or shuffled seasons together, its historical numbers will be optimistic. This is called walk-forward or out-of-sample testing.
4. What is it compared against?
| Baseline | What beating it shows |
|---|---|
| Always pick the home team | The model knows more than venue |
| Pick the team with the better record | The model adds value beyond standings |
| A simple rating system (e.g. Elo) | The model's complexity is earning its keep |
| The previous model version | Changes are real improvements |
EdgeIQ reports its NFL engine against these references. On 768 unseen games it picked 60.3% of winners versus 48.3% for always-home and 60.8% for a ratings reference — clearly better than naive rules, not demonstrably better than the rating system. That kind of comparison is what you should expect from any source.
5. Are the probabilities honest?
If a source gives percentages, check calibration: do its 70% calls win about 70% of the time? Ask for a Brier score or log loss and a reliability chart. A source that only publishes picks cannot be checked this way at all.
6. Does it explain its inputs and limits?
You should be able to find out what data the model uses, what it does not use, and where it is known to be weak. Vague references to "proprietary AI" with no detail are a warning sign. EdgeIQ's methodology page lists its inputs and known gaps, including the absence of real-time injury data.
7. Is performance separated by context?
Good evaluation breaks results down by league, market and confidence level. A model can be solid on NFL winners and weak on totals. EdgeIQ publishes per-league, per-market results and notes, for example, that spread and total questions are close to coin flips.
Red flags
- Guarantees, "locks", or claims of near-perfect accuracy.
- Records that start at a convenient date or omit losing periods.
- No sample sizes, no baselines, no definition of accuracy.
- Results that change retroactively.
- Pressure tactics tied to spending money.
A trustworthy model publishes its misses, its sample sizes and its baseline comparisons — and never promises an outcome.
Applying the checklist to EdgeIQ
You can run every check above on EdgeIQ in Model Lab: live records separated from backtests, sample sizes on every figure, calibration charts, baseline comparisons and version history. If something is missing or unclear, the contact page is there so you can ask.
Sources and further reading
- Hyndman, R. J. & Athanasopoulos, G. Forecasting: Principles and Practice (3rd ed.), section on time-series cross-validation.
- Gneiting, T. & Raftery, A. E. (2007). Strictly proper scoring rules, prediction, and estimation. Journal of the American Statistical Association, 102(477).
- Tetlock, P. E. & Gardner, D. (2015). Superforecasting: The Art and Science of Prediction. Crown.
About the author
Patrick C founded EdgeIQ and edits its Learning Center. He oversees how EdgeIQ collects sports data, how its prediction models are evaluated and how results are explained to readers.
EdgeIQ is a sports analytics platform. Projections are statistical estimates based on historical data, carry uncertainty and are never guarantees. EdgeIQ takes no wagers and is not a sportsbook. Read the methodology and analytics disclaimer.