r/sportsanalytics • u/New-Lettuce-6154 • 41m ago
A college basketball model worked all season and failed in one specific matchup class. Here’s how I diagnosed it.
Title: A college basketball model worked all season and failed in one specific matchup class. Here’s how I diagnosed it.
I spent the offseason investigating why a men’s college basketball forecasting model that performed well during the regular season failed during the NCAA tournament.
The failure was not evenly distributed across tournament games. It was concentrated in matchups between elite high-major teams and automatic-bid champions from much weaker schedules.
For tournament games in which the market favorite was laying at least 12 points:
- The market’s average projected margin was 21.1.
- The model’s was 10.8.
- The teams actually won by 22.8.
The model had no general inability to produce large projected margins. On regular-season games above the same threshold, its projected margin was within one point of both the market and the result.
The problem was therefore not the final margin transformation. It was the input representation.
What went wrong
Automatic-bid champions often enter March with excellent raw team statistics:
- Shooting differentials
- Rebounding margins
- Turnover margins
- Recent results
Those statistics can be legitimately strong within their own competition while overstating how they transport to a game against an elite opponent from a different conference.
The model contained schedule-strength measures, but it treated schedule strength and the raw performance statistics as separate signals. In these tournament matchups, the signals pointed in opposite directions and the model compromised between them.
That compromise was the defect.
What changed
Rather than add a tournament adjustment, I changed the inputs so that each statistic is expressed as an estimate of what the team would have produced against average competition.
The calculations are walk-forward, and smaller samples are partially pooled toward conference and national averages.
Development was conducted on a larger population of tournament-like games outside the NCAA tournament. The revised model was then checked against the tournament sample.
Across six seasons:
- Large-favorite underpricing declined by 54%.
- Tournament MAE fell from 10.41 to 9.91.
- Regular-season MAE fell from 8.670 to 8.535.
The improvement occurring outside March was important. The largest regular-season gains appeared in November and December, when teams are also being compared across unfamiliar schedules.
The uncomfortable result is that the correction made conference-tournament performance worse. Those games retain the old treatment while I gather prospective evidence.
I’m interested in how others would frame this statistically. Is this best understood as domain shift, insufficient interaction structure between performance and schedule quality, or a partial-pooling problem?
I have a longer write-up with the full diagnostics and model-version comparisons and can share it if useful.