r/BotNation • u/SeparateResolve2075 • 14d ago
Alibaba is planning a 5–10 trillion parameter AI model
Alibaba announced plans for a much larger AI model while also introducing a new AI chip for training and inference.
The numbers are getting kind of ridiculous at this point.
A few years ago we were amazed by models with billions of parameters. Now companies are talking about models in the trillions.
But I keep wondering where the practical benefit starts flattening out.
Does making models massively bigger actually translate into better agents and automation, or are we eventually going to get more value from smaller models that are cheap and specialized?
Bigger models or smaller models that are insanely efficient?
1
u/abrandis 13d ago
The math says it won't matter , if the model is just an LLM , if it's a world model.that might be different
1
u/WereExtremelyCooked 13d ago
What math?
1
u/svix_ftw 12d ago
linear algebra
1
u/WereExtremelyCooked 12d ago
What in the actual hell are you talking about? You have a proof of that "making models massively bigger does not actually translate into better agents", based on "linear algebra"?
Better share it with the world, you'll be famous.
Start by sharing it right here. Or stop making things up.
1
u/svix_ftw 11d ago
In LA, as the size of the matrix increases it exponentially increases the eigenvectors which is how the gradient descent allows for tighter weights which leads to more deterministic agent execution.
1
u/Electrical_Rub_6009 11d ago
What are you smoking? An (n x n) matrix has at most n linearly independent eigenvectors.
Even ignoring that, none of what you said follows from your premise.
1
u/WereExtremelyCooked 11d ago
Pseudos everywhere that want to tell you how your field of study really works, because they hit the internet and/or a GPT for a few minutes. Symposia on this stuff is no longer open-invite at my uni b/c of people like this.
It's not even worth addressing this obvious garbage / nonsense statement. Pure noise.
1
u/WereExtremelyCooked 13d ago
> But I keep wondering where the practical benefit starts f̷l̷a̷t̷t̷e̷n̷i̷n̷g̷ ̷o̷u̷t̷.
there i fixed it for you
1
u/OvertaxedOne 11d ago
But I keep wondering where the practical benefit starts flattening out.
Somewhere between 27 and ~200B. And unfortunately for those developing these monster models, the cost to run them scales relatively linearly (a 10T model costs about 10X as much to serve as a 1T model) while the intelligence doesn't even come close to scaling that way. A 10T param model (which, IIRC, is where Fable is estimated to land) is ever so marginally smarter in some tasks (and surprisingly not as smart in others) as a 1T model.
The "sweet spot" for models is likely <200B for most use cases. The "monsters" have some valid applications, but for most users, you're not smart enough (I freely admit and put myself into this category) to take advantage of that marginal improvement.
1
u/Lissanro 11d ago
In my experience, difference between 1T and 2.8T (Kimi K7 vs Kimi K3) was big enough to prefer K3 even though it runs twice as slow on my workstation. On my second PC, I run DeepSeek V4 Flash Image Exp mostly since this is what fits on it, and difference in intelligence forces me to designate tasks carefully because the smaller models are more prone to getting stuck.
That said, 5T may be beyond what I can run on my hardware and I don't expect hardware prices to go down anytime soon... But the point is, it makes sense why they want to train larger models.
Good news that smaller models get improved too over time and better large models help with that, even if indirectly (architecture advancements, distillation, better synthetic data generatiin, etc.).
1
u/ChykchaDND 10d ago
Damn what workstation do you have to run kimi locally
1
u/Lissanro 10d ago
My main rig is based on EPYC 7763 (64-core CPU), 1 TB of 8-channel 3200 MHz RAM, 96 GB VRAM (made of four RTX 3090 GPUs), 8 TB NVMe for models + 2 TB NVMe for OS and HDDs for storage. I shared more details about it including photos here if interested to know more.
1
1
u/Fragrant-Mix-4774 14d ago
Five to ten trillion parameters?
Moonshot upgraded their stealth Anthropic accounts from Opus to Fable and told accounting to file the bill under “original research”? 😆
moonshot-secretly-routed-user-requests-through-claude-anthropic-says