Behind the model
Here is what our game-training pipeline uses: historical game results, prior team performance, schedule context, and available pregame market data. Learn what goes in, what the model estimates, and how to read the evidence.
Completed games and information available before each prediction.
Current odds, news, and injuries are distinct from verified historical inputs.
This explains the training code. Exact data dates, counts, and coverage require the deployed model's training record.
The Pipeline
Build historical context
- Completed game scores, results, and available team statistics
- Prior team form and schedule context, calculated before the target game
- Historical market snapshots where pregame coverage is available
- Sport-specific pitcher or goalie history when usable records exist
Train, then evaluate
- Earlier games for training; later games for calibration and evaluation
- Separate league models for win probabilities and score estimates
- Probability calibration is checked alongside prediction error
- Missing market history limits what a betting backtest can establish
Put predictions in context
- Compare model estimates with available market prices
- Keep current injuries and news separate from verified historical inputs
- Treat confidence and estimated edge as uncertain signals
- Keep a complete record of published picks and settled outcomes
Live Model Status
| League | Status | Training Samples | Test Accuracy | Log Loss | Calibration (ECE) | Backtest ROI | Last Trained |
|---|---|---|---|---|---|---|---|
These are reported model evaluation metrics, not customer returns. Accuracy alone does not establish profitability. ROI needs a test period, sample size, recorded prices, and staking assumptions. Calibration (ECE) measures probability error; lower is better.
What data goes in?
These are supported input categories, not a promise that every source is available for every game. Provider adapters include ESPN game records, MLB and NHL data sources, and historical sportsbook snapshots. Coverage and freshness vary; model artifacts need their own provenance.
Stored completed games provide outcomes and prior scoring context.
Recent team performance is calculated from earlier games, without using the target result.
Possession and scoring data support pace and efficiency estimates where records exist.
Game dates describe time off and consecutive-day games.
Historical game records can identify postseason play.
Only timestamped context available before the prediction cutoff belongs in a historical example.
Historical pregame snapshots supply market context when coverage exists.
Available opening and later pregame snapshots provide price context.
An absent price is not evidence of an available betting opportunity. Coverage varies by league and period.
The MLB training path can incorporate prior pitcher performance when usable player records exist.
The NHL training path can incorporate prior goalie performance when usable player records exist.
Live context is a separate layer. Today's report is not proof of what was known before a historical game.
A forecast and a market price answer different questions. The model estimates outcomes; the available odds determine the price of acting on them. A strong favorite can still be a poor price.
We explain the data categories and evaluation approach publicly. Model weights, feature transformations, and tuning remain proprietary. Futures are rating-based simulations, separate from the trained game-prediction models.
Check the matchup, the estimate, and the reasoning. Missing information changes how much a forecast can tell you.
Odds move. An estimated edge is tied to the price and information available when the pick was generated.
Losses and passes are part of the process. Evaluate a complete record over a meaningful period, not a highlight reel.