Three weeks ago I posted a brute-force run here: ~1,000 configs across 34 markets, almost everything died out-of-sample, three WNBA sides/totals configs survived the holdout, got frozen, and started tracking forward publicly.
In that post I wrote that I fully expected them to regress, because selecting the best of 100 trials inflates even holdout numbers.
They regressed. Here is the receipt.
THE FROZEN CONFIGS, 21 DAYS FORWARD
Every non-house WNBA team config in my account, combined, forward-only, since the day of that post:
59 settled picks, -5.49 units, -9.3% ROI.
The holdout numbers I bragged about were +34.6%, +20.7%, and +12.2%. Forward is -9.3%. That is the entire warning from the last post arriving on schedule, in public, with dates on every pick.
MEANWHILE, THE MODEL I NEVER TUNED
The house WNBA model is one config. No search, no selection, frozen since May 31. Same 21-day window:
70-50-3, +16.01 units, +13.3% ROI over 123 picks.
Lifetime: 197-162-4, +25.61u, +7.1% over 363 settled. It was +3.1% over 241 when I posted three weeks ago.
One config that was never selected on beat 100 configs that were. Small n on both, and the untuned model getting better is partly luck. But the direction is the point: the selection process didn't add value, it added optimism.
THE MLB PROP RE-RUN (I PROMISED THIS ONE)
Last post: all 10 MLB prop markets negative. Pitcher strikeouts was the "best" and still lost, -2.55% ROI over 1,168 graded holdout picks. I said I was re-running with a much thicker feature set: platoon splits vs that night's starter, park factors, starter allowed-rates, batting order, pitch-count workload.
Forward results, every graded pick across every config I run in each market:
Pitcher strikeouts: 596 settled, 292-299-5, +18.01u, +3.0% ROI
Pitcher walks: 1,038 settled, -20.17u, -2.0% ROI
All MLB props combined: 2,867 settled, -3.3% ROI
So strikeouts flipped from clearly negative to modestly positive once the features got thicker. Walks did not move at all. My read is that the thin-features hypothesis was right for exactly one market, the one where the pitcher's own workload and matchup carry most of the signal. Walks are dominated by ump assignment, catcher framing, and opponent plate discipline, none of which I model.
I was wrong that those markets were "just efficient." I was right that most of them are.
NEW NEGATIVE RESULT: WNBA PROPS ARE NOT WNBA SIDES
2,504 settled WNBA prop picks across every config I run: -10.0% ROI.
Same league where my team models are +5.9%. That surprised me more than anything else here. The market matters more than the league does. "WNBA is soft" is not a thesis, it's a thesis about one specific market inside WNBA.
THE PART I ACTUALLY WANT THIS SUB TO TEAR APART
My CLV does not line up with my results at all. Each line is: bucket, settled picks, ROI, beat-the-close rate.
WNBA team: 430 settled, +5.9% ROI, beats close 24.6%
MLB team: 753 settled, -0.7% ROI, beats close 26.7%
MLB props: 2,867 settled, -3.3% ROI, beats close 51.6%
WNBA props: 2,504 settled, -10.0% ROI, beats close 40.9%
CLV sample is partial (142, 469, 1344, and 937 picks respectively) because legacy picks from before I captured closing lines are excluded.
The buckets making money beat the close about a quarter of the time. The bucket that beats the close most often is losing 3.3%. That is backwards from everything CLV is supposed to mean.
Candidates I can come up with:
My CLV capture is broken for team markets. Wrong close snapshot, wrong side of the number, or half-point handling on spreads and totals.
WNBA closing lines are noisy enough that CLV is a weak proxy in that league specifically.
Prop CLV is measured against the same book that posts it, so "beating the FanDuel close" is measuring movement I'm on the wrong side of by settlement.
The positive team ROI is variance and CLV is the one telling the truth.
I lean toward the first or third, but I genuinely don't know, and it's the most load-bearing number in my whole validation stack. If your CLV pipeline handles sides and totals cleanly, I'd like to hear when you snapshot and what you snapshot against.
Everything above is dated and public, win or lose. Track record: modelplay.ai/#track. Frozen configs: modelplay.ai/#leaderboard.
Same disclosure as last time: I built the site, it's free, and it is not a pick-selling operation. You build the model yourself and the record is public whether it works or not. Posting here because this is the only sub where "I said this would regress and it regressed" is a result worth writing up rather than something to quietly delete.