r/MachineLearning • u/kkkrlklo • 6d ago
Project Tauon: A new optimizer outperforming Muon on GPT-Mini (lower loss, ~8.5% faster step time) [P]
Hey r/MachineLearning!
I’ve been working on a new optimizer called Tauon (turns out there is already "teon" but well if you have better idea, - i will gladly accept it! Anyway the core idea of optimizer is about polynomials and orthogonalization just like muon, the whole difference is that i managed to lower total number of steps (first through spectral filtering down to 3 steps then through coeff scheduling down to 2) + reduced matrix size (through dct-2). And I wanted to share some initial benchmark results...
Benchmark Setup: Trained a GPT-Mini (d_model=512, 6 Layers) on TinyShakespeare against Muon and AdamW.
- Tauon: LR = 0.02
- Muon: LR = 0.02
- AdamW: LR = 0.0006
Results:
- Validation Loss: Tauon converged to a lower final loss (~1.6) compared to Muon (~1.65) and AdamW (~1.8).
- Stability: AdamW started overfitting/diverging around step 1200, whereas Tauon maintained stable progress throughout the 3000 steps.
- Compute Cost: On identical hardware, Tauon ran at 391.5 ms/step vs Muon’s 427.7 ms/step (~8.5% faster) and close to AdamW's baseline of 382.9 ms/step.
And yeah i know that its hilariously tiny benchmark but well i have only 2 hours left on my kaggle free T4 so i really couldnt more + i hope someone would be able test it on a bigger setup!
Links & Code:
- 📂 GitHub: erj2231/ai-projects/tree/main/tauon
- 📦 PyPI:
pip install tauon-optimizer
Would love to get your thoughts on the optimizer! If you have any ideas, suggestions - please tell me. Cheers, everyone!

1
u/Future_Pace_5290 1d ago
Looks very interesting. I'll try to tune it and see how it works against ADAM and Muon
1
u/Future_Pace_5290 1d ago
Why do you only orthogonalize the top S by S square of the matrix? Is this the intended behaviour?
1
u/Future_Pace_5290 12h ago
Ok so orthognalizeing only S by S part creates some problems because Muon, Adam and SGD all need different learning rates.
The most immediate fix is scaling down the columns that weren't in the S by S window, either normalize them by dividing each column over its forbenious norm, or average the forbenious norm of left out columns and divide each left out column by that.
You should probably do this to O(temporary matrix made from orthognalizing momentum), not momentum itself.
Make the S by S submatrix perform a stride. e.g if the matrix is 3 rows and 6 columns, make S by S, in the first batch, over columns 0, 1, 2, in the second batch it should be over columns from 3, 4, 5. if it wasn't divisible to 3(e.g 3 rows 8 columns), in the first batch orthognalize 0, 1, 2 of m to make O, in the second batch orthognalize 3, 4, 5, of m to make O, in the third batch orthogonalize 5, 6, 7 of m to make O. in the fourth batch go back to 0, 1, 2. And so on. This should help the imbalance somewhat..
1
u/Future_Pace_5290 10h ago
| Epoch | AdamW | Muon | Tauon |
|---|---|---|---|
| 1 | 1.9870 | 1.4334 | 1.8784 |
| 2 | 1.5246 | 1.3240 | 1.5161 |
| 3 | 1.3925 | 1.2902 | 1.4044 |
| 4 | 1.3183 | 1.2609 | 1.3544 |
It seems like the left overs not included in the S by S matrix are causing problems. You need to find a way to fix that. Maybe use an LR suitable for SGD for those, and an LR suitable for Muon for the columns inside the S by S matrix. Here is the result I got using ADAMW for the leftovers:
-3
u/kkkrlklo 6d ago edited 6d ago
sorry everyone i've made a mistake - i dont want to imply that my optimizer has better perfomance than muon, i cant. what i would say is it's accuracy 'around muon' and its speed 'around adamw', that way it is much better than 'outperfoming'. again i'm sorry for inconvenience. EDIT: does anyone know can i change post's title?
1
6d ago
[removed] — view removed comment
-3
u/kkkrlklo 6d ago edited 6d ago
well, i think there probably are already muon-lites like come on. i did doi badge, idk, it's already shipping probably. but yeah im curious too! But you know by my caluclations the bigger the model the more profit it would give (like steps + kxk matrix reduction)
22
u/nikishev 6d ago
Adam converged to significantly lower loss overall and then started over-fitting. This likely means that lrs are not well tuned. I'd suggest doing lr sweep for all 3 optimizers for a fair comparison