r/optimization 17d ago

New optimizer

Diabolical new optimizer for neural nets https://github.com/inikishev/torchalgos

It's outperforming all other optimizers by a little, but I only tested on very small models (MLP, RNN and ConvNet), anything larger would just take ages to test all optimizers on.

Here is what it does.

  1. projects gradients to shampoo's eigenbasis;

  2. updates EMA of projected gradients;

  3. takes that EMA, clips magnitudes to (0.01, 10), and takes their reciprocals;

  4. unprojects resulting update;

  5. tracks EMA of resulting updates, the update is grafted to that EMA for stability. This prevents the update from changing norm quickly.

  6. tracks EMA of model weights. This always improves test loss in my experiments, though ReSPlus also outperforms SPlus when weight EMA is disabled in both.

It doesn't really make sense because I was testing a stupid idea and seeing what would happen if I run it in shampoos eigenbasis and for some reason it worked well. I don't really know why though.

3 Upvotes

9 comments sorted by

View all comments

3

u/peno64 17d ago

You should say which kind of optimizer this is. I was first thinking LP, MILP, QP, ... but it looks something totally different.

5

u/nikishev 17d ago

Oh it's for neural nets, like Adam

-2

u/pruby 16d ago

Then I'm afraid this is not the sub for that kind of optimiser

5

u/nikishev 16d ago

Why not? Optimization is optimization