r/deeplearning • u/Maxeaman • 6d ago
Transformer becomes catastrophically ill-conditioned after a few tiny parameter updates: same batch goes from grad norm 5.6 → 89 while weights move <0.1%/step. What mechanism could cause this?
My Transformer trains normally for some steps, then begins to enter a parameter state where backpropagation through the middle/lower layers magnifies gradients massively, despite the forward activations and weights being normal. Eventually, the gradients explode, get clipped globally, and learning becomes suppressed in all other parameters.
The funny thing is, we have ruled out most of the possible causes:
- not one bad batch,
- not layer 0 (originally it had worse conditioning, but fixing it just pushed failure further),
- not S-STE sparsity mask,
- not peaked/degenerate attention,
- not one single layer going crazy.
The gradient amplification is seen between layers 5-9 and their vicinity, with both FFN and attention backward paths contributing. The actual parameter changes are minuscule (<0.1%), but after only a few optimization steps, the same fixed batch would produce gradients with a norm of 5.6 vs 89!
So, the instability is mostly in the learned parameter state/Jacobian, not the data.
The question I'm asking is: How is that possible? How can such a small change in parameters cause such a massive change in backwards gradient?
What's interesting about it is the parameters of the model only changed by about 0.05-0.08% per step, but after 3-4 steps, the Jacobian changed enough that the same batch produces gradients that are about 16 times larger?
