How good is the model?

A predictor that never reports its own accuracy is a horoscope. This page is the scoreboard for the one on this site.

Read this first

Three models are scored here, and they are not all comparable:

  • elo_pregame — a genuine forecast. Its rating going into a game was built only from games already played, so nothing about the result leaked into the prediction.
  • market_consensus — the betting market's closing line, converted to a probability. This is the benchmark, not a model of ours. Beating it is hard, and most models do not.
  • blended_season_ratings — the full model, but scored against end-of-season efficiency ratings. Predicting a January game with a rating that describes the whole season means using March information. Its numbers look excellent and mean nothing as a forecast. It is shown because it is the model that runs the bracket, and hiding it would be worse than labelling it.

Once the pipeline has run through a full season, the ratings table carries a rating snapshot for every date, and the blended model can be scored honestly too.

No Results

Accuracy is the least informative column. A model that picks the favourite in every game scores well on it and has told you nothing. Log loss and Brier score are proper scoring rules: they reward confidence only when it is justified, and punish confident mistakes hard. A coin flip scores 0.693 and 0.25 respectively.

Calibration

The question that decides whether the tournament odds mean anything: when the model says 70%, does it happen 70% of the time? A model can rank teams perfectly and still be badly calibrated, and a bracket built on overconfident numbers looks far more certain than the tournament actually is.

The diagonal is perfect calibration. Above it the model is underconfident; below it, overconfident.

Loading...
No Results

Where the model struggles

Not every game is equally predictable. Conference tournament games pit teams from the same league against each other at neutral sites, which strips out most of what a rating knows — and the numbers show it.

No Results

Accuracy by season

Loading...