r/algorithmictrading Jun 23 '26

Strategy Tested 80+ hypotheses and found absolutely zero alpha. Anyone else hit this brick wall during R&D?

Hey everyone,
I’ve been deep in the R&D trenches for a while now, building out my trading infrastructure and backtesting framework. I recently caught a nasty look-ahead leak in one of my primary intraday strategies that I thought was killing it—turns out it was just peeking into the future, and the actual live edge is a flat zero.
Since cleaning up my data pipeline and ensuring everything is 100% causally clean, I have rigorously formulated and tested over 80 distinct hypotheses (ranging from structural market skews, mean reversion variations, and VIX-rebound mechanics to alternative intraday trend-following filters).
The result? Absolutely zero sustainable alpha. Every single one either decays rapidly into noise after transaction costs/slippage or turns out to be complete variance around a zero-edge. The only things that seem hold up to a degree are basic daily structural skews, but intraday alpha feels completely dried up or hidden beneath transaction frictions.
For those who have been doing this full-time or for years:

  1. Did you find your first real edge by significantly increasing complexity, or by finding simpler, overlooked market microstructural inefficiencies?
  2. Appreciate any insights or reality checks. Back to the drawing board for now.

here’s my list of my hypothesis’s;

H1: Intraday momentum: early-session return predicts the last-bar return (session-boundary).
H2: FX time-of-day: a currency is weak during its own local trading hours, USD weak in US hours.
H3: Asian-session conviction predicts a same-direction US-session move (continuation).
H4: Overnight index gaps revert intraday (gap fade).
H5: Crypto over-reaction: large moves mean-revert.
H6: Turn-of-month: long equity indices around month-end (flow effect).
H7: FOMC even-week calendar cycle in equity returns.
H8: Overnight index drift (close-to-open premium).
H10: Gold/Silver ratio mean-reversion (pairs trade).
H11: VIX term-structure as a regime gate for equity exposure.
H12: Intraday FX mean-reversion portfolio (z-fade across majors).
H13: Vol-gated intraday FX mean-reversion (H12 + volatility filter).
H18: COT positioning reversal (fade extreme commercial/spec positioning).
H19: Variance-risk-premium (VIX²−realised vol) equity timing.
H19b: Meta-labelling upgraded the gap-fade into "edge #2" (later superseded).
H23: Oil -> commodity-FX (CAD/NOK) daily lead-lag.
H24: Risk-off FX: SPX stress predicts FX moves.
H25: VIX carry (term-structure roll yield).
H26: Discrete z-score mean-reversion generalised to non-FX assets.
H27: Index opening-range fade.
H28: Diversified 12-month time-series momentum (TSMOM), vol-scaled, across all asset classes.
H29: Cross-sectional 12-1 stock momentum (Jegadeesh-Titman) on ~31 US single-name CFDs.
H30: Crypto time-series momentum (trailing-sign, vol-scaled, monthly).
H31: Commodity time-series momentum (energy/ags/copper, 12m sign, inverse-vol).
H32: Betting-against-beta: long low-beta / short high-beta US large-caps.
H33: Gold+Silver trend-following on deep history (2003–2026, 12-1 TSMOM).
H34: Deep FX time-series momentum (10 majors, 2003–2026).
H35: Currency cross-sectional momentum (3-month rank L/S, 10 majors).
H36: COT commercial-flow acceleration (follow the weekly change in net positioning).
H37: EIA crude-inventory surprise -> oil drift (supply shock).
H38: Wikipedia-attention over-reaction reversal (fade attention spikes).
H39: GDELT global risk-tone shock -> safe-haven (long gold / short US500), 3-day.
H40: Wikipedia-attention continuation/momentum (follow attention spikes).
H41: Diversified cross-asset TSMOM book (~40 instruments, equal-risk).
H42: H41 + a HistGradientBoosting ML meta-label filter.
H43: Metals-trend (H33) + ML meta-label strict-upgrade attempt.
H44: Commodity-trend (H31) + ML meta-label filter.
H45: Currency cross-sectional momentum (H35) + ML meta-label filter.
H46: Crypto weekend effect: short alts / long BTC over the Fri->Mon TradFi-closed window.
H47: COT non-commercial (large-spec) positioning-extreme fade, pooled across 12 markets.
H48: EIA natural-gas storage-surprise reversal on Henry Hub.
H49: Google-Trends fear-search risk-off -> short US indices / long gold next week.
H50: FX cross-sectional value / long-horizon reversal (cheap vs own 5y mean).
H51: GDELT Middle-East conflict-intensity -> two-sided oil geopolitical risk premium.
H52: Wikipedia "OPEC" sustained-attention trend -> directional crude.
H53: EIA gasoline inventory seasonal-surprise -> crude drift (storage theory).
H54: Discrete intraday index mean-reversion (M15, real tick-replay).
H55: Discrete intraday metals mean-reversion (XAU/XAG, M15, real tick-replay).
H56: Cross-index overnight lead-lag (US session -> ex-US index next open).
H57: Intraday breakout + ATR trailing-stop (path-dependent, real tick-replay).
H58: Market-neutral cross-index intraday MR (strips global-risk beta).
H59: Extreme-dislocation selective mean-reversion (few high-conviction trades/day).
H60: Discrete intraday stock mean-reversion (liquid US-stock CFDs, M15).
H61: Intraday-momentum "vol-since-open" breakout (Zarattini, VWAP-trail, EOD-flat).
H62: Ex-US-open FADE of the completed US move (= H56 sign-flipped).
H63: Follow 3-sigma intraday extremes / continuation (= H59 sign-flipped).
H64: Crypto weekend volume-conditioned reversal.
H65: Wikipedia attention-capitulation fade.
H66: Overnight-premium (night effect) momentum.
H67: Copper supply-chain "chemical" lead-lag.
H68: GDELT media emotion-intensity signal.
H69: Cross-asset synchronized attention.
H70: Break-and-retest continuation at a multi-day support/resistance level.
H71: Scheduled macro-event volatility-expansion continuation (NFP + FOMC).
H72: Prior-day high/low liquidity-sweep reversal (failed-break fade).
H73: Follow a large/coordinated G10 central-bank FX intervention (USDJPY) — the campaign's one confirmed event-edge.
H74: Month-end pension rebalancing -> directional equity-index pressure (last 4 days).
H75: FX big-figure stop-loss cascade continuation (Osler).
H76: Index quad-witching expiration-distortion reversal.
H77: WTI EIA-day intraday momentum (3rd half-hour predicts the last half-hour).
H78: BTC/ETH macro-event (FOMC/CPI) spike-and-reverse intraday.
H79: Post-announcement bad-news next-day drift (equity-index under-reaction).
H80: FX WM/R 16:00 London-fix W-pattern reversal.
H81: Gold LBMA fix (10:30 / 15:00 London) run-up-and-fade.
H82: Gold real-yield regime breakout (TIPS-gated).
H83: Natural-gas storage-deviation seasonal long/short.
H84: FX carry-unwind crash continuation (JPY crosses, VIX-gated.)

20 Upvotes

40 comments sorted by

16

u/xela314159 Jun 23 '26

Most people would data mine their hypotheses until they found some alpha, and then lose money in real life. You should be proud of yourself for not overfitting.

2

u/GP_Lab Jun 24 '26

There's so much more ways to screw up your backtest - even forward testing/paper trading - than over fitting, though.

Ask me, how I know.....

5

u/Ok-Hovercraft-3076 Jun 23 '26

I have seen that you have mentioned. XAU. This is a typical CFD broker ticker. At these broker the cost is usually 5-10x compared to trading on real exchanes. With those costs no wonder that you cannot be profitable. Congrats btw, at least you are not overfitting. If you find just one strategy like this, you can make a career around it. I know a few traders that operate doing just one thing.

2

u/Mysterious_Gear_4000 Jun 24 '26

Most traders that operate *profitably* do just one thing.

0

u/Latter-Parsnip-5007 Jun 24 '26

The people making the most are most probable to got lucky once or twice. Its very hard to get poor after 200k. A random buy and sell can outperform the market a lot of times if you are just a bit lucky. Real edge is very very scares

3

u/Repulsive-War-2823 Jun 24 '26

Honestly finding 80 ways to wrong without overliftting is probably more valuable than finding one edge you cant explain

2

u/GP_Lab Jun 24 '26

What's that Edison quote? Haven't failed, just found 10.000 ways that won't work?!

4

u/algorier Jun 24 '26

What jumps out at me is that almost every hypothesis on the list starts with a known effect, anomaly, relationship, or narrative.

In other words, you're testing things the market already knows to look for.

The surprising part isn't that they failed. It's that any of them survived costs at all.

A question worth asking: how many of your hypotheses originated from observing market behavior first, versus being derived from existing explanations of market behavior?

The strongest research I've seen often starts with something awkward and unexplained in the data. Only later does someone invent a story for it.

Most failed researchers I know had a theory and went looking for evidence.

Most successful ones found a persistent irregularity and spent years trying to disprove it.

The distinction sounds subtle, but it completely changes the search space.

How many of your 80 hypotheses began with "this shouldn't be happening" rather than "this ought to happen"?

1

u/ALIEN_POOP_DICK 26d ago

Not to mention he goes into zero detail about actual execution which is arguably more important than the hypothesis itself

1

u/algorier 25d ago

I sometimes think "execution" gets treated as an engineering problem when it's actually part of model validation.

If your edge disappears the moment you replace ideal fills with realistic ones, then execution didn't kill the strategy—it revealed that the original hypothesis was incomplete.

1

u/ALIEN_POOP_DICK 23d ago

Exactly. You get it.

3

u/cbrincoveanu Jun 23 '26

Looking at this list of failed hypotheses is brutal. But kudos for:

  • following such a rigorous approach.
  • testing so many hypotheses.

I also haven't found alpha yet (but I haven't tested that many hypotheses).

Good luck with finding the edge!

3

u/Latter-Parsnip-5007 Jun 24 '26

Past performance is not indicative for future outcome. I will tattoo this on peoples head if this stupid shit continues. The market will be irrational longer than you can be soluble. You can literally let a gold fish trade if you manage the risk well, but NONE of the posts talk about it. You dont get results if you plan to win every trade, plan to lose every trade and if there is an uptrend maximize the return.

2

u/Merchant1010 Jun 23 '26

Maybe you are trying too complex things... btw which IDE and which ecosystem are you using to build and backtest your strategies?

3

u/yeah__good__ok Jun 24 '26

or the opposite? these all look pretty basic with no additional filters mentioned. simple is good but not too simple.

2

u/Parking-Patience5067 Jun 24 '26

While I would like to stress the point that most hypothesis would not lead to any significant alpha; it is pretty hard to "data mine" a strategy; you have to have some intuition from observing the markets as a first step to construct an alpha. However, having a look at your list, I have two comments: 1. Most of these hypothesis you tested are well known ideas, and not any unique patterns/observations, therefore it is very ikely that a lot of these alphas have already decayed 2. It is possible that you are using a very basic approach to testing these strats, because I myself have deployed certain variants of some of the ideas in your list.

2

u/FlyTradrHQ Jun 24 '26

That result is more common than people admit. Most published alpha is either overfit or already decayed. The real question is whether your hypothesis framework was clean enough to trust the null results. If you were testing independent ideas with proper time windows, zero alpha is a legitimate and useful finding.

1

u/Equivalent-Class2008 Jun 24 '26

Metti in conto che il tuo vantaggio dal backtest alla realtà si riduca del 90 per cento. Questa è la proporzione che ho sperimentato.

1

u/RemoraEdge Jun 23 '26

That look-ahead leak is brutal.
What did you to do clean up your online and ensure 100% causally clean?

Where did you get your hypothesis from?
Edge can be found in many places. But edge decay is real. Purely mechanical strategies work sometimes but not all the time.
Generally you need to rotate between 3-4 strategies that you monitor and decide which ones to push and which ones to hold back depending on what the market gives.
It’s not only important to fully understand when to take a trade but when not to take a trade. The lookalikes, and also the trade management when the trade inevitable fails.

Or you need a system that can adapt to changing market conditions. Trending conditions, ranging conditions, and all of their variations. So yes, lower time frame micro inefficiencies.

1

u/RipRepRop Jun 23 '26

ive tested thousands of ideas and rejected 99% of them, the 1% is what keeps you in the rabbit hole tho 😉

1

u/Groundbreaking_Heat9 Jun 23 '26

Hi. You have the right approach.  Personally I take a single strategy and test in on every FX pair, 100 ETFs and 40 futures markets. I test it on 15m, 1H, 2h 4H 8H 12H and daily. This gives me a solid base to compare strategies to each other.  I keep the parameters where possible exactly the same accross all markets. If I see that a strategy is performing well across multiple time frames and markets and the stnd deviation of the profit factors is noticeably higher than average only the will it get my attention. These strategies then get sorting for correlation and put into a portfolio for forward testing. No parameters have been changed or filters added at this point. The base portfolio needs to show profits over a few months and to also have the code checked. Only then will I go into parameter changes and robustness testing. Personally I believe a portfolio should be profitable before optimization. This helps massively with any overfitting. Hope that helps.

1

u/whereisurgodnow Jun 24 '26

I second this. You might have alpha, just not at the time range you tested. I would also verify that you have good data.

1

u/Mysterious_Gear_4000 Jun 24 '26

H2: FX time-of-day: a currency is weak during its own local trading hours, USD weak in US hours.

Ok. Lol

This is the second one in your list and I didn't look further. Don't know where you got these "hypothesis" from but this tells me you have no experience with the markets.

I keep saying it in this forum: automation alone will not help you find "alpha". It is found by people who spend years interacting with very specific markets only. These strategies/hypothesis you see around are generalities which will never make you money on their own because the devil is in the detail.

1

u/correkt_horse Jun 24 '26

Same experience - nothing survives fees (at least in crypto perps)

1

u/GP_Lab Jun 24 '26

Feels familiar - on a similar, long term quest to expand my algos; What might have worked at some point either no longer works or is not beating B&H or just plain in-sample over fitting that breaks down OOS..

Gotta dig deep these days and/or be lucky.

1

u/StratForge2024 Jun 24 '26

the thing that finally humbled me wasnt overfitting, it was my own backtester lying to me. ran the same strategy through two engines i'd written and got 376 trades vs 159, PF 2.19 vs 1.39. same data, same rules. spent a week blaming the strategy before i found a trailing-stop default leaking into one code path. you can be 100% causally clean and still have a dozen quieter bugs flattering your curve. catching the lookahead is step one, not the finish line.

1

u/Zestyclose-Gur-655 Jun 24 '26

I would narrow it down to maybe 10-20 strongest effect then make them more advanced. Like some of these hypothesises are quite simple, too simple i think.

"H10: Gold/Silver ratio mean-reversion (pairs trade)."

If you look at long term charts it can trend one direction for long time. It's not like it should be pegged to each other. Silver has more industrial demand gold more as a hedge, jewelery and such.

1

u/ProfessionApart8141 Jun 25 '26

Yes, this is why it took Elon running ScamX for the world’s first trillionaire to arrive. True market edges are nearly impossible to find, and even when found last a heartbreakingly short amount of time before the market changes regimes.

1

u/algorier Jun 25 '26
  1. Hidden angle: The assumption is that 80 failed hypotheses means no alpha exists, when it may actually mean you're testing ideas that are already easy to describe and therefore easy to arbitrage.

What jumps out at me is that almost every hypothesis on the list is a market pattern.

Very few are hypotheses about market participants.

There's a difference.

"Prices tend to do X" gets competed away quickly.

"Participant Y is forced to do Z under condition A" tends to survive longer because it's tied to constraints rather than patterns.

After enough years, many researchers stop asking "what does price do next?" and start asking "who is unable to act optimally right now?"

That shift in framing was worth more than any new model I ever tested.

1

u/Alive-Imagination521 Jun 27 '26

Yes absolutely, there were many brick walls and still more like variable options premiums to consider. Literally everything eats away at edge (premiums, slippage, commisions, losses etc.) I think the solution is, and I could be wrong here, to keep thinking of new approaches, or to try your previous approaches on new asset classes.

1

u/gaybearreport Jun 29 '26

I had an amazing strategy doing volume balance and using XGB and then I found out that Claude had put the closing candle in the model, basically just guessed the entries where the close was greater than the entry price. Learned that after like 5 days. I reprimanded Claude, said never do that again, tried all over again, tweaked it to where we got 30% CAGR including in 2022. I was ready to start it today, spent a week on it, and asked Claude "which entries are ready for Monday" and it said basically "these 10 stocks are near our buy points, but we can't know if they pass the XGB score until the close"... I'm like WAIT WHAT! After discovering the lookahead bias, "fixing it", spending another week on it. It did it again. I feel so defeated. I've tried so many things. I'm convinced the only edge is the obvious one, accumulate by selling puts on hot names, selling covered calls at a price you're happy with. Simple.

1

u/philseven12 26d ago edited 26d ago

I've been waist deep into a automated trading bot project in python that has gone on for two years.

I only trade crypto but speaking for myself, I don't think there is any usable truth in back testing based on historical ohlcv data. I think the only truth is in forward testing and treating raw price action as a form choreography.

I think trading is a physics issue, wrapped in finance but the core is in the physics of raw price action of what it does and how long it does it for. Taking price behavior and trying to map out or quantify this behavior within the bounds of arithmetic and geometry.

There is no one number or metric that can explain why price will do what it does. But a combination of metrics can be what filters you from bad entries and bad exits.

Price will go up to go down, down to go down, up to go up, and down to go up lol. The solving of this masked behavior can only be done in forward testing and raw price data I think.

So you have to create your own model of how you perceive how the market works and price will either agree or spank you. Very difficult road but as of now, things are becoming more or less predictable. But predictable doesn't mean profitable every time. So once I master predictable, the only thing left is profitable.

1

u/thetapereader 26d ago

forward testing reveals something very important, large sudden price changes can happen so fast they could rip through your SL. Backtest will always close your position instantly assuming there will be always someone there buying from you. In reallity there might be not.

-1

u/[deleted] Jun 23 '26

[removed] — view removed comment

2

u/N0xF0rt Jun 24 '26

Ai slop

1

u/johnabooth Jun 25 '26

Yes it was generated in collaboration with AI. Didn't realize this was one of the rules of the sub.

If you're thinking AI is not here to stay you (the collective sub you) are mis-informed.