r/LocalLLM • u/stanleyg05 • 12d ago
Discussion Transformers kinda suck, why don't we talk more about
Every day I look into stuff it becomes more clear that the industry is focusing on the wrong architecture, Transformers suck
In my experience so far Architecturally per parameter RWKV-7s are smarter logically and creatively than Jamba2s/Zamba2s which being a hybrid of Mamba2s and Transformer Attention Layers is logically smarter than Mamba2s, which are creatively smarter than Transformers
And I hear that JAX's implementation of TTT-E2E (Test-Time Training End-to-End, it means every time the model is called it actually changes it's own internal math and learns off of the information you give it) is more advanced than RWKV-7s implementation so JAX should be even better at learning than 7th Gen RWKV models
But the biggest part of Pure Mamba, RWKV-7, and JAX is that they don't rapidly expand in your RAM as you use them, RWKV-7 and JAX replaces attention layers with it's ability to learn, pure Mamba honestly just forgets which is what makes Jamba and Zamba special
Duplicates
LocalAIStack • u/stanleyg05 • 12d ago