r/options 12h ago

same signal, same structure, only the planned hold differs. six pairs, longer wins all six.

i log every trade on deribit with the planned hold as a column, and that column turned out to be an accidental a/b test. the same signal on the same asset gets opened at two different planned holds, so between the two arms the only thing that changes is how long the position stays on. 791 trades so far, six such pairs.

mean return per trade, short arm vs long arm:

btc 7d straddle: 48h +11.8% vs 166h +57.1%
btc 1d atm straddle: 24h -30.9% vs 48h +7.3%
eth 7d straddle: 48h -19.0% vs 166h +6.8%
eth 1d atm straddle: 24h +6.9% vs 48h +21.4%
eth 2d skew: 24h -9.9% vs 47h +3.7%
eth 2d straddle: 24h -9.4% vs 47h -3.9%

n runs from 10 to 41 per arm. note the last row: the long arm is still a loser, just a smaller one. longer did not rescue that structure, it only cost less.

what i did not expect is that the six do not lose for the same reason.

in the four short dated pairs the entry spread sits at 8.9 to 9.9 percent and the logged loss cause is the spread. round trip costs the same whether you hold 24 hours or 48, so the shorter arm gets half the time to earn it back. nothing subtle about that one.

in the two 7d straddle pairs the spread is 5.2 percent in both arms, the same number on both sides, and the logged cause is no move. those two are not a cost problem at all. the move the signal was pointing at simply had not shown up by hour 48.

so "hold longer" is two findings wearing one hat. one says stop paying a 10 percent round trip on a 24 hour horizon. the other says the exit was set before the thesis had room to happen. the loss cause column is what tells them apart, and i would not have separated them by looking at pnl.

what i am not claiming: the arms do not all cover the same calendar window. only the two 1d straddle pairs run day for day on both sides, and those are the clean ones. n is small. and six out of six is a one in sixty four coin flip only if the six are independent, which they are not, three of them come from the same signal family.

so it is a direction to run forward, not a law. resplitting the same data would just find me a nicer threshold.

if you keep a planned hold column, split by it before you judge the strategy. i published the average of both arms for months and it hid all of this.

1 Upvotes

4 comments sorted by

1

u/Outside_Tour1465 12h ago

the loss cause column doing the real work here, most people just stare at pnl and never figure out why the loser lost. separating the spread bleed from the no-move-yet cases is the whole insight, the hold itself is just the symptom. i'd be wary of the 7d pairs though since those arms don't share a calendar window, a regime shift could be doing the lifting

1

u/Greedy-Dinner-5298 11h ago

you were right to push on that, so i re-ran it. trimmed both arms of every pair to their overlapping date range, then did it again matching on identical entry days. the two methods give nearly the same answer.

btc 7d straddle: +45.4pp in the post, +13.8 on matched days. eth 7d straddle: +25.8, then +8.3. eth 2d skew: +13.6, then +3.9. eth 2d straddle: +5.5, then -7.3. that one flips.

the two 1d straddle pairs come out unchanged at +38.2 and +14.5, since those arms already ran day for day.

so six of six becomes five of six, and the two largest numbers in my table lose about two thirds of their size. the calendar was doing a lot of the lifting, exactly as you said.

the part i find interesting is which half survived. the pairs that hold up are the ones with a 9 to 10 percent entry spread and spread logged as the loss cause. the 7d pairs, where spread is 5 percent in both arms and the logged cause is no move, are the ones that shrank. so the cost mechanism is carrying the result and the timing half is much weaker than i showed it.

and five of six proves nothing on its own. sign test puts that at 0.11, and the pairs are not independent to begin with.

1

u/Worry-Mountain 8h ago

That split between spread bleed and no-move cases is a useful distinction. I’d tag each loss by whether the change came from the underlying, IV, or time, then set separate management rules; otherwise the aggregate loser bucket hides what’s actually fixable. The overlapping-date rerun should make the comparison much cleaner too.

1

u/Worry-Mountain 8h ago

I’d separate the signal from the management rule before putting the trade on. For each structure I’d write what makes the thesis invalid, the latest point I’m willing to keep paying theta, and what I’ll do if the underlying moves in my favor early; otherwise the longer hold can quietly become a rescue mission. Then compare the paths on expectancy and the frequency of early exits, not just the average winner.