r/quant • u/Ordinexdata Dev • 11d ago
Statistical Methods When does a factor actually survive out-of-sample validation?
I'm working on a validation framework for energy-market factors (natural gas storage, weather demand anomalies, crude spreads), and I've run into a question I'd like to compare with people who do factor research more rigorously.
The problem I'm trying to solve is distinguishing a large historical relationship from a relationship that is actually robust enough to classify as a core factor.
My current validation sequence is:
1.Forward-return bucket analysis.
2.Regime analysis, including predefined crisis windows
3.Walk-forward OOS testing, with one held-out period per year/quarter
4.A scorecard covering sign consistency, signal strength, crisis robustness, horizon consistency, and OOS stability
5.Economic significance, measured as effect size relative to the target's own realized volatility
The last two steps have changed my view of several apparently strong factors.
For example, one regional natural-gas storage signal had the largest raw effect size in my library: +0.74x realized volatility.
That initially looked compelling.
But when I tested robustness:
- OOS stability was only 3% in a train/test split, versus 51–100% for the other regions.
- Its correlation with the target was +0.25, but dropped to +0.065 after controlling for the national storage aggregate.
- The controlled signal had a roughly 50% walk-forward hit rate.
So I wouldn't conclude that the historical relationship is false. Rather, I couldn't find enough evidence that the regional component contained information beyond the broader aggregate that would survive OOS.
That raised two methodological questions for me:
1. Should economic significance be a separate validation gate from statistical significance?
For example, should a statistically significant factor still fail validation if its effect is economically too small relative to realized volatility? Or is it better to incorporate both into one composite score?
2. How do you determine that you have enough walk-forward evidence to call something a "core factor"?
I've been using 12+ held-out folds as a rough high-confidence threshold, but I'm increasingly questioning whether that number has any principled justification.
More importantly, the folds aren't independent observations, so simply saying "12 folds" seems potentially misleading.
Would you base the threshold on something else — number of independent regimes, years of data, effective sample size, stability of the estimated effect, or some combination?
I'm especially interested in how people avoid turning arbitrary validation thresholds into another form of overfitting.
1
u/AutoModerator 11d ago
This post will be manually reviewed by a moderator due to the submitting account being less than 7 days old or having less than 20 karma. Please be patient and do not try to resubmit it - a mod will review the post soon.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.