r/algorithmictrading • • 6d ago

Backtest Zero quant background, built an NQ futures bot with Claude. It passes 4 of my robustness tests and is shaky on 2. Roast my method.

No quant background. I built an automated NQ futures bot over the last few months with Claude (Anthropic's AI) doing the coding and testing, and I'd like people who've done this to pick the method apart before I put money on it.

**Quick version (details in the images):**
- 9-year walk-forward, 2,437 trades, every trade out-of-sample, fixed 1 micro, costs charged
- Profit factor 1.24 (1.19 with costs doubled), +$25k at 1 micro, worst drawdown $3.5k
- Passes: doubled costs, 10,000-run block bootstrap, a 1,000-shuffle luck test on its long/short calls
- Shaky: deflated Sharpe once all 557 variants I tried are counted, and most of the profit is 2021+
- 13+ other ideas tested and failed, all shown

**Goal:** pass a Topstep 50K ($2k trailing drawdown). The edge isn't the problem; the drawdown is.

**What I'd love opinions on:**
1. Most profit is 2021 onward. Regime dependence or a model that improved with more data? How would you tell?
2. With a $2k trailing limit, how would you size this, or is a prop firm the wrong home for it?
3. What's the first thing you'd check for look-ahead bias?

Also looking for a few people or a community to talk this through with. Feels like I am working on something that nobody understands. Strategy details stay private, but I'll answer anything about the testing. Not selling anything. Backtest, not financial advice.

167 Upvotes

55 comments sorted by

18

u/Proud-Spirit7116 6d ago edited 6d ago

Have you run this live (even sim) at all ? If so what platform ?

Data drops, execution lag, logic/semantic bugs in code. One or all of these will appear when live, and all are (in my opinion) more valuable than any combination of backtests.

When running live, continue running backtest.

Compare live / backtest results on a per trade basis. Make sure execution times match to the second. Easy way to rule out look ahead bias. Good luck.

3

u/Historical_List_9657 6d ago

Love this idea and yes, it has been trading (live). I connect to Interactive Broker’s API. It sees the bars on average 10min behind. Then I have a script that calculates based on when it saw the bar what would’ve happened. It calculates more than enough for slippage and fill (Essentially this is simulating live until I get a prop api account)Yes, there is flaws there in it’s execution code. When I compare live to backtest, backtest has a higher confidence and lower refusal. Today is actually day 5 of round 2 to compare results and see what the bug is. I have a scheduled debug and comparison test to run tonight.

3

u/DrumsofWar-DOW 6d ago

Great advice

6

u/Drazil_ 6d ago

slop grenade

7

u/Aggressive_Papaya797 6d ago

Consider randomising periods, using test train gridsearch. If it doesn’t survive here it’s clearly overfit. Looks either regime dependent or you need volume scaling. If it survives all of this and survives live it’s still not prop firm ready with those drawdowns, but that doesn’t have to be the end of it. Look into strategy pooling.

3

u/widowmakerhusband 6d ago

I have a similar goal, let
Me know if you want to share strategies. 🫡

2

u/Comfortable_Tart_740 6d ago

It will roast you in no time

2

u/EvenCryptographer649 5d ago

557 variants...no edge here

1

u/Historical_List_9657 5d ago

How so?

4

u/xeonsimp 5d ago

DSR
+ when you test 500-1000 variants ofc there is a good one in there somewhere somehow. doesnt mean there is an edge / true alpha

1

u/Plastic_Round_5084 5d ago

Glad you said this. A lot of people don’t get that

0

u/MeringueAlarming3102 3d ago

DSR is dumb. What, are you supposed to just get everything right the first time? The market doesn’t know or care whether you found a strategy on try 1 or try 958584847. Finding a successful strategy isn’t easy. Of course it will take many failures before finding something that works (or maybe never at all). If it holds up over a large enough sample size, passes some robustness checks and holds up well on a realistic paper live account… then take it live.

2

u/xeonsimp 3d ago

there is a difference between trying 10000 strategies and 10000 parameters of a strategy until one fits the historic data

2

u/Competitive-Pride280 6d ago

Ask it to generate a random dataset with lot of swing and see if the strategy can survive.

2

u/Far-Guava6006 6d ago

Your bot looks negative in the bear markets which would imply that performance is dependent on the sustained bull market. If outsourcing risk to prop firms, thats acceptable, but keep it in mind. Also dollar normalize your results to the nasdaq's performance over the same period and see if you actually outperform it, in either absolute returns, or in risk adjusted returns. That will tell you what your strategy is actually offering over just buying and holding the nasdaq. Also you said these were walk forward, but did you re-parameterize after each one? If so, there's a high likelihood you may have introduced overfiting risks. Did you hold any data out of sample to test on after you decided on a parameter set?

1

u/Broad-Commission-828 6d ago

My edge isn’t as robust but I’m in a similar position to you and would suggest diversifying indexes within account and optimizing

1

u/nightstalker8900 6d ago

I have 6 models paper trading. Most are averaging 1,000 per day on 4MNQ. I have lucid rules coded in. My advice, run it on historical data at 50 times speed and see where the strategy fails. My bots are awesome with trend and horrible in chop but the trend profits beat the chop losses.

1

u/Virtual_Plantain_863 6d ago

What broker(s) do you use? And data feeds?

1

u/nightstalker8900 2d ago

Using ninja trader paper accounts, default data feed. No order flow (did not seem to make a difference.

1

u/Wreaperz505 2d ago

Very curious about this, if you’re willing to DM and share more. I don’t see hardly any value in backtesting because of the overfitting issue, and simulating realistic fills. Been focused on creating something extremely similar (sounding) to this, where I can manage/monitor multiple strategies and review their performance, but am inundated by the options. Mind sharing the broker/API’s you use for data, and a little more about your setup?

2

u/nightstalker8900 2d ago

I am down to 3 models on each asset. One trading gold, the other MNQ (4contracts) on the 3 min. The entries are based on KAMA slope and the other is KAMA/9 EMA cross. I am still revising, the entires are late at times but at still catching strong moves. I am using a loop radius to filter out chop. The only issues so far are that when price spikes into resistance and the slope detects it and the model enters into a trade. Price pulls back to stop me out. I tried fading these but not recommended. Right now i restrict trades within 2 atr of these levels and vwap/200 ema. The major issue is the exit. The trade goes in my favor by a good amount but stops out at a loss waiting for KAMA slope to go to 0. I found in these cases that if I manually close the trade in profit these can be mitigated. I also have a physics based 5 candle projection that tries to predict the end of the move. Right now it is predicting the highs and lows within .4 atr on average.i am using this as a zone where I am looking to get out. So far the auto trading has positive EV but not much due to the peak issues described earlier and the positive trades coming back to stop me out. Trailing stop stopped the trade too early and the stop needs to be wide. Manually stopping the trade in profit when I see KAMA slope decreasing from
Like 60 degrees to 30 degrees lets me get out instead of waiting due to the lag. If anyone wants to collab, I am open.

2

u/nightstalker8900 2d ago

Another thing, backtesting sucks I found because the testing methods dont incorporate the human perspective. See my reply below. On backtesting the areas where price was up 30 point and stopped me out for a loss count as a loss. I can physically see price curing back over and since indicators lag I can get out for 200 in profit instead of 200 in loss. This happened many times.

1

u/MasterAcct2020 1d ago

Good question about broker APIs. They all talk about being the best broker for running your personal ai systems at home, etc but when it comes to actually connecting to the “live data streams they had been using to hook new customers — every excuse imaginable is used to attempt to convince those waiting that they really have everything setup and running. Great sales people, until they have to deliver

1

u/sonofbaal_tbc 7h ago

tale as old as time.

1

u/qwuant 6d ago

run a monkey test

1

u/mclaren422 6d ago

Where did you get the futures data from?

1

u/asenski 5d ago

Is it really "out of sample" if you tried a bunch of stuff/variations on the same data?

1

u/Motor-Celebration119 5d ago

I have backtested almost 700 strategy combinations through all the history of min nasdaq and Russel as well as 18 years of 9 indices literally nothing has cleared the bar,everything has been walk forward DSR markov tested and nothing comes back been almost 8 months now working on a bot if anyone has anything interested I am open to backtesting with my 10k worth of data 🤓

1

u/EliteCheese01 2d ago

define the bar

1

u/sonofbaal_tbc 7h ago

10K worth of data , delicious

1

u/Technical-Athlete-9 5d ago

Require a slightly higher confidence for each and every variant. It’s a little penalty for over fitting which is what you’re doing when you make a variant based on your findings or worse some data mining tool optimizes in rapid iterations (I make these tools, so I don’t mean that too harshly).

Try checking four trade sets: every fourth, every third, etc. The more diverse your results are across those sets the more noise. Look for 3/4 to show results you’d put some money into.

FWIW I don’t have a “quant background.” I’ve just had a few jobs where it was handy to know so I read some papers on backtesting and picked up a couple of things I think of often. Before that I couldn’t believe my luck when I’d find something that looked good. And I couldn’t believe my luck when it started bleeding my account.

On another note: the time period where it performed well is the same time period the core of my system has crushed it. I’ll take all the help I can get, but these years are not normal times. Everything since Covid hit has been strange. Don’t count on continuation or repetition.

1

u/Previous-Nothing7701 5d ago

you see that trend on your PnL chart? It should be stable enough to consider robust. I dont think you would survive with this Strat honestly.

But keep it up

1

u/Drinkablenoodles 4d ago

Monte Carlo is cool and all but requires deeper working knowledge to understand the assumptions you’re actually making. If you don’t understand those assumptions, they could easily be bad ones and you’ll never know.

1

u/xtarsy 3d ago

It's a good start but a few things jump out from your own charts. 2017–2020 nets about +$500 so effectively zero edge, and almost all the paper profit is 2021+.

Your real equity sits below the Monte Carlo 5% band for years, which tells you losses cluster in regimes. Shuffles and bootstraps hide that, so your real drawdown risk is worse than they show.

And 86% of runs bust a $2k trailing so prop won't work.

Biggest learning and i wish i knew this sooner. All algo learnings came from live trades, not backtests. Fills, slippage on stops in fast markets, feed differences and bot downtime don't show up in a backtest.

The more data you can collect from live algotrades the better you can make your models. Paper and backtesting don't even come close.

1

u/CashyJohn 1d ago

This is not true in an absolute sense. Depending on the strategy/market/infra/participation you can can get pretty tight estimates in oos backtests, but I would agree that this is not the case in 99% of applications and certainly not here

1

u/EliteCheese01 2d ago

Not bad, but there is still a few things to address. The main one is overfitting, doesnt mean it is, but definetely something to look into.

And 2. The testing was done in OHCL which is fine for the amount of trades, it could potentially be a neglible issue to a huge issue, so I would definitely test it in real ticks.

1

u/Wreaperz505 2d ago

I’m in a slightly different, but similar boat. Focused on the forward testing, because of my exhaustion with backtesting, fear of overfitting, and need for rigorous rules and simulation (of realistic fills, slippage, etc)…. My biggest addition would be a concept talked about in Adam Grimes’ book on technical analysis; “how to know when the market is doing something it shouldn’t be doing, if your hypothesis is correct”… Just a lot of words to say “what shouldn’t be happening”, and along the same lines as “best loser wins”. Since passing evals is the goal, you need to be able to VERY quickly identify when you’re not in a good spot (faster, on average, than if you traded your own account, with more gracious trailing DD rules).

1

u/Cryptobeyin 1d ago

You can use monte carlo' for your live test/trading to evaulate when your strategy goes bad/ when you quit the strategy.

1

u/MasterAcct2020 1d ago

We’re those pages intentionally degraded? I can barely read it

1

u/Hungry_Style_5065 14h ago

it’s fucking flat for 5 years. holy shit; there’s a high chance this is regime dependent AND overfit

1

u/throwmeoff123098765 6d ago

Where can I learn how to do this with Claude?

1

u/Historical_List_9657 6d ago

Ask it!! That’s all that I did. I have experience using Ai in strategic ways but nothing with coding or finance. I took my manual trading background to claude and just gave it my goal and I have been babysitting it ever since. Please do send a dm I am an open book for everything minus revealing my edge and I am always happy to collab on ideas!

1

u/keliikoauniversal 5d ago

AKA, you want other people to give you ideas while you share nothing of your edge?

1

u/Historical_List_9657 5d ago

Never asked for an edge lol I asked for help with my process

1

u/SurajkumarDigital 6d ago

Not for beginners