Model Performance
Lower Brier and log loss are better. The floor is what you'd score knowing nothing but home-field advantage; the market is the ceiling — it is measured against, deliberately not chased. Landing just below it with honest calibration is the goal.
Buckets are the model's predicted probability that a player scores, against how often players in that bucket actually did. Graded over the active roster — a player who dressed and never touched the ball counts as a miss, while one who was inactive, on the practice squad or on IR is left out, because a touchdown prop voids rather than losing when its player never plays.
The scoreboard above and the last-graded slate cover the moneyline market, measured on graded predictions only — ties excluded. The market comparison is the published closing moneyline with the vig removed. It is never a model input: the model is built without any market feature, so that whether it is accurate stays a question this page can actually answer. The touchdown markets are scored against real per-player results, and are headlined by the model's top-ranked player in each game rather than a hit rate: scoring a touchdown is rare enough across a full roster that “nobody scores” would look over 90% accurate.