Methodology
Sources supply facts. Bumeral calculates the ratings, form, efficiency and predictions itself.
Data sources
| Source | Sport | What we take from it |
|---|---|---|
| ESPN | All sports | Schedules, scores, teams, standings, summaries, news and some odds |
| API-Football | Football | Fixtures, lineups, events, player stats, injuries, transfers, odds |
| nflverse | NFL | Play-by-play, rosters, snap counts, depth charts, injuries, Next Gen Stats |
| MLB Stats API | MLB | Box scores, play-by-play, batting and pitching lines, transactions |
| NBA public data | NBA | Box scores, advanced stats, lineups, shooting |
| EuroLeague public data | EuroLeague | Box scores, play-by-play, shot coordinates, lineups |
We add a source only when it brings data the others do not have. Candidates for later: The Odds API for multi-bookmaker prices, Baseball Savant for Statcast, StatsBomb and Understat for football event data and xG.
From raw data to a prediction
- Collect raw responses from each source and store every one
- Normalise into one schema with a shared ID for each team, player and game
- Calculate metrics: form, ratings, efficiency, strength of schedule
- Train and validate models on point-in-time data only
- Publish tables, probabilities and projections, then grade them against results
How we keep models honest
Features are built only from information available before the game starts. Backtests use walk-forward splits, never random shuffles. Every published probability is stored with its model version and graded once the result is in, so calibration and ROI history are visible to subscribers.