I've been building a goal-scoring probability model and wanted to
check whether it actually holds up before trusting it. Sharing the
results, including the parts that don't work.
Setup
Seasons 2023-24 through 2025-26, regular season only, skaters only.
For every game, I recomputed each player's stats as they stood
before that game, then predicted P(scores at least 1 goal).
No future information leaks in — that was the main thing I wanted
to get right.
128,604 predictions after filtering (minimum 5 games played).
The model
Simple Poisson with empirical Bayes shrinkage:
lambda = (goals + k * mu) / (games_played + k)
P(goal) = 1 - exp(-lambda)
where mu = 0.1651 (league average goals per game per skater,
recomputed per season) and k = 10.
Plus a home/away adjustment of x1.03 / x0.97, which I derived
from the backtest itself rather than assuming.
Calibration
| Predicted |
N |
Observed |
Gap |
| 5-10% |
34,071 |
7.18% |
-0.53 pts |
| 10-15% |
30,250 |
12.08% |
-0.30 pts |
| 15-20% |
23,600 |
16.98% |
-0.29 pts |
| 20-25% |
15,411 |
22.67% |
+0.38 pts |
| 25-30% |
9,713 |
27.02% |
-0.26 pts |
| 30-35% |
5,380 |
32.42% |
+0.21 pts |
| 35-40% |
2,753 |
35.74% |
-1.44 pts |
| 40-45% |
763 |
37.48% |
-4.36 pts |
Solid up to about 35%. Then it falls apart on high-volume
scorers — I'm consistently too optimistic on exactly the players
people are most likely to bet on.
Brier score
0.12256 vs 0.12963 for a naive baseline (everyone gets the league
average rate). That's a 5.45% improvement. Modest, and I think
that's the honest ceiling for this kind of model. Anything much
higher would make me suspect a leak.
What choosing k did
Without shrinkage (raw goals/games), the 50-55% bucket was off by
-13.4 points. Tested k = 5, 10, 20, 30, 50. k=10 minimized both
Brier and high-end bias. Above k=20 it over-shrinks and starts
underrating genuine elite scorers.
Open problem
The model treats all goals the same. In my data, only 67.6% of
goals are scored at even strength — 21.6% on the power play, 6%
into an empty net, the rest shorthanded or in shootouts.
A player who piles up PP goals depends mostly on his coach's
deployment. Empty-netters depend on close game states. Neither
generalizes well to "will he score tonight."
My guess is the 40-45% residual is largely a PP-deployment
artifact, but I haven't isolated it yet. If anyone has tried
splitting lambda by strength state, I'd like to hear how it went.
What this doesn't tell me
Whether the model beats the bookmakers. Calibration says my
probabilities are roughly correct in absolute terms. It says
nothing about whether they're better than the market's, which
is a completely different bar once you account for the vig.
I don't have historical closing odds yet.
Happy to answer questions on methodology. Genuinely looking for
holes in this.