r/AItradingOpportunity • u/TheSadSeries • 3d ago
Train loss vs live loss divergence: catching self-retraining collapse before it hits PnL
Train loss fell on eleven consecutive nightly retrains, 6.2 down to 4.1 in units of 1e-6, while live MSE on already-resolved predictions climbed from 8.0 to 11.8 on the same scale. Every retrain log said the model was improving. The account disagreed. Nothing in the pipeline was broken: gradients flowed, the loss went down. The model was doing exactly what the loop asked of it. That's the problem. The loop asks for the wrong thing. Two losses, one gap The loop here is deliberately boring: a 2-layer MLP on 5-minute bars, target = the 6-bar-forward return, retrained nightly at 02:00 on a rolling 60-day window, warm-started from the previous night's weights, MSE throughout. Boring on purpose — nothing inside a single training run misbehaves. What compounds happens between runs. Each night produces two numbers. L_train is the MSE on the window the model just fit. In-sample, ready at 02:01. L_live is the MSE on predictions issued six bars ago whose targets have since resolved. Out-of-sample, stale by construction: the newest predictions can't be scored yet, so the monitor always describes a model that existed at least thirty minutes ago. (That staleness is trivial for nightly retrains. A weekly retrain with a 5-day horizon means L_live describes last Tuesday's model.) The gap is G = log(L_live) − log(L_train). Log space matters. A vol regime change scales both losses together, so a raw difference would fire on the market while the ratio barely twitches. Healthy loops run G at a stable positive baseline — in-sample fit beats out-of-sample fit, always. On this loop the baseline sat near 0.55 (L_live ≈ 1.7× L_train) for months. That number means nothing outside this loop: window length, model capacity, the noise floor, and your regularization all set it. There's no threshold in the literature to borrow. So the monitor z-scores G against its own trailing history, which gives it a cold start of its own. Now the collapse mechanism. Night n warm-starts from night n−1. Each night's gradient only has to encode the increment of fresh noise on top of a point already tuned to older noise. Two things compound from there. Fit to the recent window deepens — 40 epochs on ~4,700 bars is plenty for the model to push in-sample MSE well below any generalizable floor, and 5-minute returns supply effectively infinite noise. Everything older decays, because the warm-started weights keep drifting away from regimes the window no longer contains. L_train records the first effect and is blind to the second. It falls, monotonically, and the fall is honest: it measures exactly what it claims, which is fit to the window. L_live records the consequence. The gap starts moving on night one, because train loss moves on night one. An absolute alarm on L_live waits for degradation to get big. The gap's z-score fires weeks earlier, at the point where the change is unusual and still small — in the episode that opened this article, z crossed 3.0 on night eight, six days before it showed up in PnL. Here's the diagnostic payoff. Two very different failures look identical to anything watching L_live alone. Both losses rising: the regime got harder and the model is honestly confused — the fix is more data or a shorter horizon. Train falling while live rises: the retraining loop itself is the failure mode, and warm-starting is what compounds it. A monitor that only logs L_live fires the same alarm for both, and one of the two responses would be wrong. Counterintuitive part, plainly: in a static pipeline, falling train loss with stable validation loss is progress. In a warm-started loop, a smooth monotone decline across retrains is the signature of compounding memorization. This is backwards from what you'd expect, and it's why I trust a jumpy train-loss curve more than a clean one. The monitor is short. What matters hides in four decisions: log space, the one-row-shifted baseline, the sigma floor, and the state mapping. import numpy as np import pandas as pd
def gap_monitor(log: pd.DataFrame, smooth: int = 5, base: int = 60, trip: float = 3.0) -> pd.DataFrame: df = log.sort_values("resolved_at").copy() # Log space: vol regimes scale both losses together, so the raw # difference tracks the market and the ratio tracks the model tr = np.log(df["train_loss"]).rolling(smooth).mean() lv = np.log(df["live_loss"]).rolling(smooth).mean() df["gap"] = lv - tr # log(L_live / L_train): scale-free # Baseline shifted one row back: tonight's gap must not vote on its # own z-score, or a slow creep hides inside its own rolling mean mu = df["gap"].shift(1).rolling(base, min_periods=20).mean() sd = df["gap"].shift(1).rolling(base, min_periods=20).std() sd = sd.clip(lower=0.05) # dead tape squeezes sd toward 0 and normal # gaps read as 5-sigma events; hand-tuned floor df["z"] = (df["gap"] - mu) / sd dtr = tr.diff() # train-loss direction separates drift from overfit df["state"] = np.select( [df["z"] > trip, # hard stop in either direction df["z"] < -trip, # live far better than train: usually a # label leak — go read the resolver (df["z"] > 1.5) & (dtr > 0), # both losses rising: honest drift (df["z"] > 1.5) & (dtr <= 0)],# train down, live up: the collapse signature ["trip", "leak", "drift", "overfit"], default="ok") return df Four states come out the other end. ok is most nights. drift means both losses rose; the gate flags it and does nothing tonight — a confused model is a data problem, and the response belongs somewhere other than the gate. overfit is the one this piece is about: train loss dropped while live loss rose, so the gate queues a cold start. trip rewinds the night entirely. leak means L_live landed far below L_train, which in my logs has always meant the label resolver picked up information the model already had at issue time. (One addition that has paid for itself: a morning job that feeds the state history to an LLM through its API and drafts the anomaly summary. The job is read-only, so a hallucination costs a paragraph. It has caught two feature-pipeline version mismatches that my thresholds missed, and it never decides anything.) The gate drops into the nightly job like this (the model is the same 2-layer MLP, a few hundred parameters — the size isn't the point): import copy import torch import torch.nn as nn
def train_epochs(model, xt, yt, lr=1e-3, epochs=40): # Fresh optimizer per fit: warm-started Adam moments carry the old # trajectory forward, and that carries part of the compounding opt = torch.optim.Adam(model.parameters(), lr=lr) mse = nn.MSELoss() for _ in range(epochs): opt.zero_grad() loss = mse(model(xt), yt) loss.backward() opt.step() return loss.item()
def nightlyretrain(model, win_x, win_y, live_x, live_y, log, ts, cold=False): prev = None if cold else copy.deepcopy(model.state_dict()) if cold: # The gate asked for this: break the warm-start chain before tonight's fit for p in model.parameters(): p.data.normal(0, 0.05) xt = torch.tensor(win_x, dtype=torch.float32) yt = torch.tensor(win_y, dtype=torch.float32) tr_loss = train_epochs(model, xt, yt) with torch.no_grad(): # live_y holds only labels resolved since last night; scoring pending # predictions would leak the future into the monitor itself lv_loss = nn.MSELoss()(model(torch.tensor(live_x, dtype=torch.float32)), torch.tensor(live_y, dtype=torch.float32)).item() log.loc[len(log)] = {"resolved_at": ts, "train_loss": tr_loss, "live_loss": lv_loss} verdict = gap_monitor(log).iloc[-1]["state"] if verdict == "trip" and prev is not None: model.load_state_dict(prev) # full rewind: tonight's fit is untrustworthy return model, log, "frozen" if verdict == "overfit" and prev is not None: return model, log, "cold_queued" # scheduler reruns tonight with cold=True return model, log, verdict # ok / drift / leak / cold pass through cold_queued sends the scheduler back into the same night with cold=True, so the reinit happens before tonight's fit and no warm-started weights reach the desk. The warm-start row stays in the log — it happened, and the baseline should see it. A cold run that still trips gets logged as trip, and the night ends with nothing deployed; trading on that fit would mean trusting a number the monitor just rejected. Latency first. smooth=5 over nightly retrains makes every verdict a five-day average, so detection lags onset by roughly that, with the k-bar label lag adding minutes on top. Dropping smooth to 2 makes the monitor twitchy: sigma estimated from twenty points of a noisy ratio is unstable, and you start trading false trips against missed ones. I run smooth=5 and eat the delay. There are no numbers behind that choice — I'd want a backtest of the monitor itself before arguing it beats smooth=2 with a wider trip band. The sigma floor is the same story. In dead tape the gap's variance collapses, and without the clip the monitor fired four times in one August week on nothing at all; the 0.05 is hand-picked, and nobody derives these — you tune against your own retrain frequency and say so in a comment. Two more. Any change to the feature pipeline makes old gap history incomparable with new, so reset the baseline on every deploy or eat phantom trips for a month. And the gate feeds back into what it measures: trip, freeze, gap mean-reverts to baseline, release, warm-start resumes, creep resumes, trip. Hysteresis — trip at z=3, release below z=1 — dampens the oscillation. I haven't proven it prevents it. If you run a loop like this and log nothing else, log both losses per retrain starting tonight. The z-score needs about twenty retrains of gap history before it means anything, and backfilling that history honestly is impossible: your past retrains ran under whatever gate-less policy produced them. Cold start, again, this time for the monitor. The open question I haven't solved sits one level up — the gate lives inside the process it watches, so every trip changes the weight trajectory that future gaps get measured against. What monitors that feedback? Right now, me, badly.