1. Data collection
Schedules, scores and team records are collected from sports data providers and updated on a schedule. Team-level history powers the current model; player statistics may be shown for context but are not automatically treated as a reliable model input. When a feed is delayed or unavailable, a displayed timestamp helps readers distinguish a fresh result from an older snapshot.
2. Features without leakage
Each projection uses information available before the game started: team ratings, scoring and margin trends, recent form, strength of schedule, rest, home and away splits and data-quality signals. Historical games are processed in date order: a team's rating and record are updated only after the features for that game have been recorded. That prevents a finished game's result from leaking into its own projection.
Recent form summarizes a team's preceding games, not its future schedule. Home/away splits compare how each team has played in those settings, while strength of schedule helps contextualize a record assembled against weaker or stronger opponents. Small early-season samples are less informative; an unfamiliar team or missing game history can lower the quality of the available estimate.
3. Models
For each league and market, EdgeIQ trains several candidates — a simple baseline, an Elo-style rating model and a gradient-boosted model — and compares them. A new version only replaces the current one if it performs better on games it has never seen.
Projected home and away scores are estimates from learned scoring relationships in prior games. Their difference produces a projected margin; their sum produces a projected total. The winner probability is learned and calibrated separately rather than treating a projected score as a certain final score. Where a newer model is not eligible or a data feed is incomplete, the existing eligible model may be used instead of silently substituting invented numbers.
4. Calibration
Raw probabilities are adjusted using a separate validation window. Calibration asks whether, over many predictions, events labeled 70% occurred about seven times in ten. That is a diagnostic goal, not a promise about the next game. Brier score averages the squared distance between the stated probability and the binary result: a lower score is better, and a confident wrong estimate is penalized more heavily than a cautious one. Read the historical record in the Performance page and model versions in the Model Lab.
5. Chronological testing
Models are evaluated with walk-forward tests: train on earlier seasons, test on later ones, and repeat. We report accuracy, Brier score, calibration and score error, always with the sample size. Small samples are hidden rather than shown as misleading percentages.
Win accuracy is the share of evaluated games where the preferred winner was correct. Average score error describes how far a score estimate was from the final result; probability scoring captures information that accuracy alone misses. Results from a held-out historical test and results from live pregame predictions are labeled separately because they answer different questions. Neither should be mistaken for an all-time guarantee.
6. Live tracking and immutability
Projections are saved before kickoff and cannot be edited afterwards. When games finish they are graded automatically, and the live record sits alongside the backtest.
Before kickoff, a projection can change when newly synced schedules, scores or team form change its inputs, or when a new validated version becomes eligible. The model does not retroactively rewrite the recorded pregame prediction after a result is known. If a game lacks enough completed history or its status is unclear, a metric may be withheld instead of presenting a misleading rate.
7. Confidence labels
Confidence reflects both the strength of the estimate and the quality and amount of team-history data behind it. An early-season sample or incomplete historical data can lower a label even if the estimated probability looks strong. A label is not a guarantee; compare groups with enough completed games before inferring reliability.
How We Measure Ourselves
Completed pregame predictions are graded against actual results. EdgeIQ publishes sample sizes, win accuracy and available probability and score-error measures, rather than relying on a marketing claim. The public performance record distinguishes recent live grading from chronological historical testing. Weak periods remain part of the record. A reader should be able to inspect both the number and what it measures.
8. Known limitations
- Late injury news, weather and lineup changes may not be reflected.
- Sports feeds can be delayed or incomplete; missing data and small samples limit what can be inferred.
- College football projections have been overconfident on heavy favorites in testing.
- Point spread and total projections perform close to chance and should be treated as descriptive.
- Past performance does not guarantee future results.
No prediction is guaranteed: a well-calibrated forecast still expects some favored teams to lose. The analytics disclaimer explains the limits of every estimate, and our responsible-use page puts market comparisons in context.