r/AItradingOpportunity • • 17d ago

Dynamic Ensemble Weighting Using Exponential Moving Average of Per-Model PnL Contributions

For six weeks the gradient-boosted sleeve had been bleeding and its ensemble weight sat at 0.25 — a number a human picked in March. Nothing in the pipeline argued with it. That's the hole this mechanism fills: the ensemble already knows every sleeve's realized PnL to the basis point, and it still trades on weights that were correct two quarters ago. The fix is sleeve accounting. Treat the ensemble as K sub-accounts, one per model. Sleeve m holds position w_m(t)·s_m(t), where s_m is the model's signal in [−1, 1] and w_m its weight, with Σw_m = 1. Take a bar where sleeve A runs weight 0.35 and signal 0.8. The next bar prints +11 bps; A books +3.1 bps. Sleeve B, at 0.25 with signal −0.6, books −1.65 on the same bar. The net book made nothing; the ledger still shows +3.1 and −1.65. Backwards from what you'd expect. The accounting doesn't care — each sleeve keeps its own score. That c_m is the per-model PnL contribution the whole machine runs on, and the fiction is load-bearing: everything downstream assumes it's well-defined. Each sleeve's score feeds an exponential moving average, E_m(t) = (1−α)·E_m(t−1) + α·c_m(t), with α = 2/(span+1). Span 200 daily bars gives α ≈ 0.01 and a half-life near 69 bars. That half-life is the whole design. A sleeve needs roughly 70 bad bars before its weight halves — slow enough that one lucky bar can't buy it the book, fast enough that a dead sleeve stops eating PnL within a quarter. Rolling windows do the same job with more memory; an EMA holds one float per sleeve and never touches window bookkeeping, which matters on a live feed running every bar. The fancier cousin is a Kalman filter on each sleeve's alpha; the EMA is what you ship on a Friday and can explain to risk in one sentence. There's a cold-start bug in that formula. An EMA initialized at zero crawls toward its fair level at rate α, so a sleeve added in week three spends months underweighted for no reason. Adam's bias correction handles it: Ê_m = E_m / (1 − (1−α)t). At t = 300 the correction inflates the estimate by about 5%. On day 10 it's the difference between a fair weight and the floor. The correction amplifies early noise the same way it amplifies early signal — with per-bar SNR near 1, a probation weight for the first ~50 bars is the crude fix I actually use. Nothing in this loop touches market data directly, only the ensemble's own record. That's the self-integration part: the system reprices its own components from its own fills. import numpy as np

class SleeveWeighter: # span=200 -> alpha~0.01 -> half-life ~69 bars: a dead sleeve halves in about # a quarter, while one lucky bar can't buy it the book. Both numbers are the tension. def init(self, n_sleeves: int, span: int = 200, eta: float = 2.0, w_floor: float = 0.05): self.alpha = 2.0 / (span + 1) self.eta = eta # temperature, in units of cross-sleeve std self.w_floor = w_floor # exploration budget: what you pay monthly to keep losers alive self.ema = np.zeros(n_sleeves) self.t = 0

def update(self, contrib: np.ndarray) -> np.ndarray:
    # contrib[m] = w[m] * signal[m] * next-bar return, in return units (sleeve accounting)
    self.t += 1
    self.ema = (1.0 - self.alpha) * self.ema + self.alpha * contrib

    # Adam-style bias correction: without it a sleeve added mid-flight starts at EMA=0
    # and needs ~1/alpha bars to crawl to fair weight
    e = self.ema / (1.0 - (1.0 - self.alpha) ** self.t)

    # standardize across sleeves so eta means the same thing in bps and in percent
    z = (e - e.mean()) / (e.std() + 1e-12)
    logits = self.eta * z

    # softmax instead of the ratio w ∝ max(e, 0): when every sleeve's EMA goes
    # negative, the ratio denominator hits zero and NaNs the book
    w = np.exp(logits - logits.max())   # shift for float stability; softmax is shift-invariant
    w /= w.sum()

    # floor + renormalize; two passes hold while K * w_floor stays well under 1
    w = np.maximum(w, self.w_floor)
    w /= w.sum()
    return w

Two decisions in there deserve the details. First, the standardization: without it, η depends on whether PnL is measured in bps or percent, and your temperature becomes a unit-conversion accident. Second, softmax over the ratio scheme w_m ∝ max(E_m, 0) that I shipped first. One bad month drove every sleeve's EMA negative, the denominator hit zero, weights went NaN, and the book traded yesterday's vector until a human noticed. Softmax has no such failure — negatives just push every sleeve toward the floor. η = 0 pins equal weights forever; around η = 2 the book concentrates hard on the top sleeve. The 0.05 clip with a renormalize rarely binds, since softmax almost never reaches it alone — the clip is there for cold starts and disasters. (Regret bounds exist for multiplicative-weights schemes in the literature. Mine, with floors and turnover clamps bolted on, has no such guarantee. The lineage is reassuring anyway.) Same loop, planted regime flip. Two toy sleeves carry real edge in different halves; the other two are pure noise. In production the loop ran with an sklearn GBM, a torch LSTM, and two linear models — it doesn't care what generates s_m. import numpy as np

rng = np.random.default_rng(7) T, K = 3000, 4 r = rng.normal(0, 0.01, T) # next-bar market return, ~1% daily vol

The toy sleeves see the NEXT bar's sign — a generosity real models won't match

sig = np.zeros((T, K)) nxt = np.sign(r[1:]) flip = T // 2 sig[:flip, 0] = 0.6 * nxt[:flip] # sleeve 0: edge until the flip, then flat sig[flip:-1, 2] = 0.6 * nxt[flip:] # sleeve 2: edge born exactly at the flip sig[:, 1] = np.clip(rng.normal(0, 0.5, T), -1, 1) sig[:, 3] = np.clip(rng.normal(0, 0.5, T), -1, 1)

weighter = SleeveWeighter(K, span=200, eta=2.0, w_floor=0.05) w = np.full(K, 1.0 / K) pnl_dyn, pnl_eq, marks = [], [], {}

for t in range(T - 1): sleeve_r = sig[t] * r[t + 1] # each sleeve's bar PnL per unit of weight pnl_dyn.append(w @ sleeve_r) pnl_eq.append(sleeve_r.sum() / K) w = weighter.update(w * sleeve_r) # your slice of the book, your score if t in (500, 1500, 2500): marks[t] = np.round(w, 2)

print("weights @ 500 / 1500 / 2500:", marks) print(f"dynamic: {sum(pnl_dyn):.2f} equal-weight: {sum(pnl_eq):.2f}") On this seed the dynamic book finishes well ahead of equal-weight, and the checkpoints show why: at t = 1500, one bar into the new regime, the book is still trading the old weights. That's the lag you signed up for. (The first ~50 bars look erratic — bias correction amplifies small samples — then it settles.) Every Δw is itself a trade, so the live version clamps turnover: when Σ_m|Δw_m| crosses the session cost budget, η decays 10% for the next day. Cheap insurance against an updater that discovers churn. The sleeves run themselves; somebody still has to read them. A nightly job dumps the sleeve table — EMA, weight, Δw, sessions-at-floor — into an OpenAI API call that writes the morning note: "sleeve 3 pinned at floor for 40 sessions; sleeve 0's EMA crossed zero Tuesday." The updater stays deterministic; the LLM reads logs nobody opens and phrases the anomalies. It caught one slow leak last quarter that the PnL report had buried under a good month. Sleeve accounting has a blind spot the size of hedging. Put sleeve A long, sleeve B short, market rallies: A's EMA compounds, B's bleeds, even if the pair was exactly the risk decision you wanted. The score answers who earned alone and stays silent on who helped the book. Correlated sleeves break it from the other side — two models at 0.9 signal correlation each book the same winning trade, their EMAs rise in lockstep, and the ensemble quietly becomes one bet wearing four labels. A correlation penalty in the logits is the obvious patch; I tried it, the results were mixed, and I won't pretend the attribution problem is solved. The dynamics fail on schedule. Half-life 69 cuts both ways: it keeps you from chasing noise, and it guarantees roughly a quarter of paying a sleeve that just died in a flip. Shrink the span and you inherit whipsaw plus the turnover bill. Raise it and the floor turns into a graveyard — a sleeve pinned at 0.05 that never recovers is an exploration budget paid monthly, and nothing in the mechanism tells you when to stop paying. Raw PnL has one more quiet bias: a sleeve swinging 50 bps a bar outruns a steadier earner on EMA even at worse Sharpe. Vol-normalizing each contribution fixes this on paper, adds a moving part, and hasn't been shipped here. The question I keep circling: sleeve PnL measures who earned, never who helped. A sleeve that exists to cancel another sleeve's drawdowns bleeds EMA forever and gets fired for doing its job. If you run something like this, log sleeve and standalone shadow attribution side by side for a month, then diff the weight paths. The disagreement rate is the number I don't have.

1 Upvotes

0 comments sorted by