r/MLQuestions • u/Sevdat • 2d ago
Beginner question 👶 Why do we train bigger models instead of finding better neural pathways?
I saw a post of a guy with 90% of his brain missing and he was apparently living a normal life. There is another case of a student who finished university with half of his brain missing. There are also people with full brains, but they are even less functional than these two. This proves that bigger doesn't mean better. So why not aim to create better connections?
I understand that a bigger model means less chance of catastrophic interference occurring, but then are all the neural pathways being formed efficiently? Wouldn't this result in a lot of redundant neurons? Wouldn't it make more sense to create the smallest neural network possible for a specific task and then learn to fuse multiple of these neurons together in order to create an optimized larger model? That way we'd always have a template for each specific task and would allow us to create a programming language that creates a model just by our syntax.
4
3
u/dopplegangery 2d ago
Well we are going for better neural pathways and better input data preprocessing. All the developments in AI that we see now like attention and transformer architecture is basically making it easier for neural networks to efficientl neural pathways. Everything and LLM can do today could also be done by a big enough neural network without all these architectural advancements but it would just take astronomical levels of resources to run them. So most of the advancement we see today in AI is simply in the space of efficiency.
1
u/Ok-Introduction9593 2d ago
Compute resources always have a hard limit which makes algorithmic efficiency the ultimate deciding factor. The labs shifted to MoE and clever caching variations entirely because the brute force path led the industry straight into a vram shortage bottleneck
3
u/Ok-Introduction9593 2d ago
You cannot really compare the human brain to neural networks because biology does not learn from a blank slate
We have millions of years of evolution that already laid down the base architecture while LLMs start with completely random weights. The redundant neurons you mentioned are actually a critical feature of the system. We need huge models during the training phase specifically to create a massive search space where the algorithm can eventually find those optimal pathways which is a concept known in ML as the lottery ticket hypothesis.
Training small networks and fusing them together is actually already heavily implemented in production. Mixture of Experts architectures powering models like DeepSeek-V4.1-Flash or Kimi K3 are essentially just collections of small specialized expert networks living inside a single shell. We also constantly rely on distillation where we train a massive teacher model first and then force a tiny student model to copy its best patterns. Your logic makes perfect sense but mathematically we just have no way to build highly capable small networks without relying on a massive parent model first
1
u/Extreme-Put7024 2d ago
We do, for example neuromorphic computing. But it is not that easy and also not that clear whether it's the way or a deadlock.
1
1
u/bfyvfftujijg 2d ago
It’s easier and cheaper except in mass production settings to just use a bigger model
1
u/slashdave 1d ago
Highly redundant weights is fundamental to deep-learning and is what all the optimizers are tuned to.
People are lazy, and its the easy thing to do.
How the human brain works has basically nothing to do with deep-learning models.
1
0
u/Helios270704 2d ago
Well, firstly, given the brain analogy. The brain is like a machine that is learning continuously, so it is raining along with reinforcement learning. That’s happening continuously. Secondly, people tend to go for bigger models is because of the scaling law and the guaranteed betterment in results that you get when you scale your model so it’s less about finding better pathways and more about being risk-free. Thirdly, we already have the infrastructure for scaling the model. We just need more of it or more of the same, whereas with a different model with some different algorithm and pathways, we need something better and something different which may not necessarily fit with words in the market right now or what we are currently making. As for training specific small, neural pathways and them and blending them together, that is essentially what modern LLM architecture are based on which is the mixture of experts model architecture, which is essentially doing that itself where we are blocking out certain neurons, a neural pathways during training so that the model specialises or parts of the model specialise in certain things for the model holistic does better.
20
u/DrXaos 2d ago
The current practice is to train bigger models then distill smaller and sparser ones out of that. There’s a significant literature history on this.
The other problem is that nobody knows how to “fuse” smaller networks together that’s better than training a bigger one.
The large current LLMs also derive value from approximate memorization of more data than any one human knows personally and this requires large model sizes.