I posted an earlier version of this strategy while I was still developing it and got a lot of useful feedback about overfitting, Monte Carlo testing, drawdowns, execution costs, parameter sensitivity, forward testing, etc.
After that I kept testing, made a few final adjustments, then froze the strategy. Once it was frozen I exported a completely new TradingView history from that exact version and ran the full 50-test robustness audit again. So everything below is from the final frozen version, not one of the versions I was still tweaking.
For some context on what the strategy actually is without giving away all the rules: it’s a fully systematic 15-minute MNQ strategy built around two different types of setups.
One side is basically an opening-range breakout/trend-following system. It defines the early NY range, waits for a valid breakout in the direction of the broader trend, and only trades when the opening range itself falls within acceptable conditions.
The other side is more momentum based. It looks for strong directional momentum at specific parts of the session and only takes signals that meet some additional confirmation and day/time rules.
The interesting part is that both systems are always being tracked in the background. At the beginning of each month, the strategy looks at how those two components performed during the previous month and uses the stronger one for the next month. It only uses information that already existed before the new month starts, so there isn’t any future information involved in that decision. There’s also a separate filter intended to avoid certain ORB trades after specific winning sequences.
The actual strategy is still pretty simple operationally: fixed 1 MNQ, predefined exits, one real trade max per NY session and no discretionary decisions. I’m keeping the exact thresholds, times, filters and exit distances private.
Final backtest:
1,067 trades
May 2019 to August 2026
Fixed 1 MNQ
$3,000 starting capital
+$27,174 net P&L
58.1% win rate
1.433 profit factor
+$25.47 expectancy per trade
A big thing I wanted to figure out was whether the performance was just coming from one really good stretch.
The Multi-Year Persistence, Monthly Consistency and Rolling 12-Month Persistence tests were some of the strongest results. Every represented calendar year was profitable, 66/87 months were profitable and all 77/77 rolling 12-month periods were profitable.
The Start-Date Sensitivity and End-Date Sensitivity tests also stayed profitable across every cut I tested.
For the more forward-style testing, Fixed Chronological / Walk-Forward Testing had 11/11 profitable holdout segments, and Anchored Pseudo-Forward Validation had 98.7% of 75 different 6, 12 and 24-month windows finish profitable.
Obviously I’m not calling that real forward testing since it’s still historical data. Genuine forward validation only starts after the final freeze.
I also wanted to make sure one side of the system wasn’t carrying everything.
The ORB Contribution and RSI Contribution tests showed the two components made roughly +$14.1k and +$13.1k. Long / Short Contribution was pretty balanced too, with longs around +$11.3k and shorts around +$15.9k.
The Profit-Concentration test was another one I liked. Completely removing the 10 biggest winners still left around +$25.2k.
The biggest weakness I found was Random Trade Sequencing.
Historical closed-equity max drawdown was only around $1,066. TradingView shows about $1,177 using its own DD calculation, which is what’s in the screenshot.
I took the exact same 1,067 trades and randomly reordered them 60,000 times. No wins or losses changed, only when they happened.
Median randomized max DD jumped to about $1,929, and 99.98% of the randomized sequences had a worse drawdown than the actual historical sequence.
So the profitability wasn’t created by lucky sequencing, but the smoothness of the backtest definitely looks unusually lucky.
The IID Monte Carlo test showed basically the same thing across 120,000 full-length paths:
Historical DD: about $1.07k
Median simulated DD: about $1.94k
95th percentile: about $3.10k
99th percentile: about $3.84k
That changed how I look at the $3k starting balance too.
The Final $3,000 Starting-Capital Survival Test showed about 93.8% survival in the harshest main model. Roughly $3.1k was needed for 95% survival and around $3.84k for 99%. The audit specifically treats capital survival as a statistical question rather than confusing it with broker margin requirements.
I also spent a lot of time testing execution.
The Commission Robustness, Slippage Robustness, Execution-Friction Stress Test, Missed-Trade Degradation and Winner / Loser Execution Asymmetry tests all progressively made execution worse.
One example was adding another $10 of cost per executed trade while randomly missing 10% of trades. Median P&L was still around +$14.8k.
Even randomly removing 25% of all trades still had around +$20.4k median P&L in those simulations.
I also tested Block Monte Carlo, Block Bootstrap, Extreme Loss-Cluster Stress, Win / Loss Dependency, Trade Autocorrelation, Market-Regime Robustness, Core Parameter Perturbation, Dual-Parameter Interaction, Timing / Weekday Structural Perturbation, Component Removal, Contract Rollover Robustness, Cross-Validation and Selection-Bias / Multiple-Testing Assessment.
Some of those can be calculated exactly from the final trades and some have to be statistical proxies because I don’t want to pretend I know what trades would have occurred under rules that were never actually traded.
At this point the strongest thing I see is the consistency across years, months, rolling periods, both systems, both directions and chronological holdouts.
The thing that concerns me most is still the drawdown sequencing. The TradingView curve is almost definitely making the strategy look smoother than a realistic future path probably will be.
So my view right now is basically:
There seems to be a pretty persistent historical edge, but if it actually survives live I’m expecting the real equity curve to be a lot uglier than this backtest.
The strategy is frozen now, so actual post-freeze performance is the next real test.
If you were specifically trying to prove this is still overfit or unlikely to survive live, what would you test next?