Did the model beat the market? Charts home Rankings Scorecard Matchup Model Season Tournament Dashboard Three predictors scored on the same season, and the calibration curve that decides which of them means what it says. Each predictor is scored only on the games it actually priced. Read Forecast? in the table first: a model allowed to see the season it is predicting is not forecasting, it is remembering. Only Elo and the market consensus are honest on that column, and the game counts differ because each model prices a different slice of the schedule. Every model, every scoring ruleRegular season. A coin flip scores 0.250 on Brier and 0.693 on log loss.ModelAccuracyBrierLog LossSkill vs Coin FlipMargin ErrorMargin BiasGamesForecast?Blended season ratings76%0.1610.48336%8.6−0.225,435NoMarket consensus720.1780.524299.0+0.175,431YesElo, pregame730.1860.5532612.4−4.295,968Yes 0.0000.0500.1000.1500.200Blended season ratingsMarket consensusElo, pregame0.1610.1780.186Brier score, lower is betterMean squared error on the probability itself, not on the pick2022202320242025202670%75%Regular-season accuracyElo, season by seasonThe only model here that never sees the season it is predicting Reliability Accuracy rewards a model for being right. Calibration asks the harder question: whether its confidence means anything. A line above the diagonal is a model underselling itself, and one below it is a model charging for certainty it does not have. 0%20%40%60%80%100%Probability the model stated0%50%100%Share that actually wonWhen a model says 70%, how often does it happen?One point per probability bucket. A model on the diagonal means what it says.Perfect calibrationElo, pregameBlended season ratingsMarket consensus Regular seasonOther postseasonNCAA tournament0%20%40%60%80%AccuracyBlended season ratingsElo, pregameMarket consensusThe tournament is where confidence goes to dieSame models, same season, harder games Data as of 01:14 UTC on 16 Sep 2026 made withdbt Charts