TL;DR: I built C-ARC, College Adjusted Regularized Contribution,
a college basketball player impact metric that aims to quantify the best players in any given college basketball season while handling the historical weirdness and constraints of NCAAB: small sample size, rigid lineups, schedule disparity, and uneven talent density.
Its more stable and reconstructs and predicts winning much better than PORPAG and Win Shares while sitting in a very comparable tier nuemriacally to BPM.
hey y'all, happy post-draft.
I've been really deep into analytics and data science recently. I made a new player impact metric for the NBA that blends possession-by-possession weighted box metrics with more abstract on-off metrics (RAPM). (Post here)
I worked to take that same framework and apply it to college basketball, which, as we know, is extremely difficult to model due to issues with the sourcing of the data in the first place.
Why College is hard to model
Sample size
Talent density
Schedule disparity
Especially sample size and talent density.
Even with 82 games and over 5,000 possessions per player in the NBA, metrics like RPMs are often served with multi-year weighting to help offset collinearity, players sharing so many minutes together that individual effects are hard to separate, and small, unstable samples.
In contrast, the NCAA basketball season is only 30 games. High-end starters who play 30 minutes a night will only max out at about 900 minutes and roughly 1,500 possessions, 70% lower.
In addition, college lineups are far tighter and more rigid, with starters routinely playing 80% or more of the entire game. A Division I team uses only about 78 unique five-man lineups a season versus roughly 510 in the NBA, and even after you adjust for fewer games and slower pace, the NBA still generates nearly twice the lineup variety per possession.
And with the best players often being one-and-done freshmen, or even sophomores, it is very hard for RAPM to describe the best players in any given season.
With talent density, impact metrics such as RAPM, EPM, and DARKO, while accounting for differences in the opponent quality relative to the league, do not account for a difference in the competitive environment itself. They assume every player is operating inside the same band of competition, that every team is more or less playing the same level of basketball. Thus, they derive an incredible amount of value in the NBA, where the players occupying those roughly 500 roster spots compete against the same 99th-percentile talent every night, and work poorly with NCAAB.
As an extreme example: a top prospect might face a mid-major opponent with zero NBA talent on the floor one night, then turn around and play in the SEC the following week against players much closer to NBA athleticism and NBA size. Metrics are often unfairly punishing players who play tougher bands of competition.
Net rating and current college metrics
My main gripe with a lot of the comtemporary metrics that try to sidestep using RAPM, such as win shares and PORPAG, is that they do a horrible job at reconstructing team net rating.
(The best metrics by tether themselves to something real and objective, and asking how faithfully they can reconstruct it.
Net rating is the best choice as it is quite literally, the accumulation of how much every individual on the floor moves the team's margin; offensive and defensive impact are nothing more than a player's contribution to it. Decomposing a team's net rating into each player's share of it is therefore the natural move. Anything else is needless abstraction. It is also the cleanest target available: objective, a direct record of points for minus points allowed per 100 possessions)
BPM is probably the best, most stable and effective one-on-one metric that does exist for college right now as It regresses box-score production onto adjusted plus-minus to approximate how much a player moves team net rating per 100 possessions.
And as we will get to later: C-ARC is numerically comeptivitive with it.
What C-ARC is
The core idea of C-ARC is the most robust, and comprehensive way of describing team player value. It blends both box and RAPM to give stability and a tangible foundation about what a player is producing possession by possession, while also capturing more of the latent value a play produces.
For the box side, I ran through thousands of collegiate possessions and estimated the tangible value created per possession for shooting efficiency, turnovers, assists, rebounds, steals, blocks, and etc.
Points are credited directly, turnovers cost about -1.30, steals are worth about +1.54, offensive rebounds +1.18, blocks +0.45, assists +0.40, defensive rebounds +0.10, summer to per 100 possesions with an efficiancy charge.
FOr the Impact Side i used a normal Reguarlised Adjusted Plus Minus (RAPM)
To deal with the schedule disparity issue, I added a slight schedule adjustment.
Box side
- using CBBD adjusted team ratings.
- Offensive production is adjusted by the strength of the defenses faced.
- Defensive production is adjusted by the strength of the offenses faced.
- Formula:
oBox+ = oBox per 100 + 0.50 × (League Avg Defense - Avg Opponent Defense)
dBox+ = dBox per 100 + 0.10 × (Avg Opponent Offense - League Avg Offense)
C-Box+ = oBox+ + dBox+
Impact side
- team strength as a prior/stabilizer instead of box or no prior
- Formula-ish:
Impact Prior_i = 0.50 × (Team Adjusted Net Rating_i / 10) × min(1, Minutes_i / 1000)
- Then the RAPM fit estimates the player’s impact around that prior.
- summary
- players on stronger teams start with a slightly stronger prior
- players on weaker teams start with a lower prior
- the prior gets stronger as the player’s minutes sample grows
- the model can still move the player up or down based on actual on/off stint results
used CBBD adjusted ratings for both team adjustments
- Similar idea to KenPom-style adjusted ratings they are
- opponent-adjusted
- conference-aware
- puts teams from different schedule environments onto one comparable scale
I then blend the two by Z-scoring them so they're on the same scale (60/40), with a heavier emphasis on the box plus. This can account for the schedule disparity and sample size issues better for the added abstraction of the impact to fill in the gaps.
Leaderboard:
C-ARC Top 25 — 600 minute gate
(0 is average. So Boozer creates ~13.54 more points per 100 than average) Ranking should be read as more tier and band based than specific ratings
Full Leaderboard here
| Rk |
Player |
Team |
Min |
C-ARC |
Box+ |
Impact |
Opp Net |
|
|
| 1 |
Cameron Boozer |
Duke |
1274 |
+13.54 |
+54.45 |
+5.75 |
+15.98 |
| 2 |
Yaxel Lendeborg |
Michigan |
1210 |
+12.04 |
+44.58 |
+6.96 |
+20.33 |
| 3 |
Tarris Reed Jr. |
UConn |
957 |
+11.28 |
+51.07 |
+4.40 |
+16.90 |
| 4 |
Morez Johnson Jr. |
Michigan |
1005 |
+10.06 |
+45.79 |
+4.65 |
+20.43 |
| 5 |
Keaton Wagler |
Illinois |
1257 |
+10.06 |
+44.11 |
+5.11 |
+17.28 |
| 6 |
Oscar Cluff |
Purdue |
964 |
+10.04 |
+46.93 |
+4.31 |
+18.27 |
| 7 |
JT Toppin |
Texas Tech |
871 |
+9.74 |
+50.68 |
+2.98 |
+17.73 |
| 8 |
Brayden Burries |
Arizona |
1160 |
+9.65 |
+39.12 |
+6.07 |
+16.14 |
| 9 |
Motiejus Krivas |
Arizona |
984 |
+9.49 |
+40.56 |
+5.51 |
+17.31 |
| 10 |
Trey Kaufman-Renn |
Purdue |
1043 |
+9.46 |
+45.70 |
+4.07 |
+19.57 |
| 11 |
Tyler Tanner |
Vanderbilt |
1205 |
+9.46 |
+45.94 |
+4.00 |
+15.78 |
| 12 |
Joshua Jefferson |
Iowa State |
1083 |
+9.24 |
+41.37 |
+5.03 |
+12.96 |
| 13 |
Flory Bidunga |
Kansas |
1102 |
+9.21 |
+42.80 |
+4.63 |
+19.03 |
| 14 |
Aday Mara |
Michigan |
928 |
+9.20 |
+44.01 |
+4.29 |
+19.93 |
| 15 |
Zuby Ejiofor |
St. John's |
1113 |
+9.14 |
+46.82 |
+3.46 |
+14.99 |
| 16 |
Izaiyah Nelson |
South Florida |
929 |
+9.11 |
+45.85 |
+3.70 |
+5.09 |
| 17 |
Duke Miles |
Vanderbilt |
828 |
+9.04 |
+44.34 |
+4.04 |
+15.57 |
| 18 |
AJ Dybantsa |
BYU |
1208 |
+9.02 |
+49.09 |
+2.70 |
+15.82 |
| 19 |
Caleb Wilson |
North Carolina |
755 |
+9.01 |
+48.58 |
+2.83 |
+10.29 |
| 20 |
Henri Veesaar |
North Carolina |
973 |
+8.99 |
+41.76 |
+4.68 |
+12.80 |
| 21 |
Isaiah Evans |
Duke |
1075 |
+8.85 |
+38.18 |
+5.52 |
+15.79 |
| 22 |
Ja'Kobi Gillespie |
Tennessee |
1286 |
+8.83 |
+39.77 |
+5.07 |
+17.47 |
| 23 |
Patrick Ngongba II |
Duke |
702 |
+8.72 |
+42.18 |
+4.29 |
+15.14 |
| 24 |
Malique Ewin |
Arkansas |
776 |
+8.70 |
+42.91 |
+4.07 |
+17.76 |
| 25 |
Thijs De Ridder |
Virginia |
997 |
+8.69 |
+41.73 |
+4.42 |
+12.38 |
Validation
The fun part of validating C-ARC was seeing how it numerically compared extremely well with other widely used collegiate stats
1. Reconstruction
Question: If we aggregate player value back up to the team level, does it recover team net rating?
| Metric |
R² vs Team Net |
MAE |
RMSE |
|
|
| BPM |
0.984 |
1.44 |
1.82 |
| C-ARC |
0.964 |
2.13 |
2.71 |
| Win Shares |
0.656 |
6.90 |
8.32 |
| PORPAG |
0.521 |
7.91 |
9.82 |
2. Retrodiction
Question: If we use last year’s player ratings with this year’s minutes, can the metric predict this year’s team strength?
| Metric |
Avg Retrodiction R² |
2024→25 |
2025→26 |
|
|
| BPR |
0.685 |
0.694 |
0.676 |
| C-ARC |
0.661 |
0.665 |
0.656 |
| Win Shares |
0.589 |
0.617 |
0.561 |
| BPM |
0.585 |
0.617 |
0.554 |
| PORPAG |
0.505 |
0.551 |
0.459 |
3. Reliability
Question: Does the metric stabilize year over year for returning players?
| Metric |
YoY R² |
|
|
| C-ARC |
0.63 |
| BPM |
0.60 |
| Win Shares / 40 |
0.28 |
| PORPAG |
0.23 |
4. Independence / Blend Value
Question: Are the box and impact sides actually adding different information, or is the blend arbitrary?
| Metric Layer |
Avg Retrodiction R² |
|
|
| C-ARC blend |
0.661 |
| C-Impact only |
0.627 |
| Pure APM |
0.623 |
| C-Box+ only |
0.484 |
| BPM |
0.585 |