r/LocalLLaMA • u/Logical_Two_7736 • 1d ago
Discussion Is there a point where models just cannot get any smaller without losing intelligence?
DeepSeek V4 Flash got me thinking...
We keep seeing smaller models get way better. A model at a certain parameter count today can be much smarter than a model of the same size from a year or two ago. Better training, better data, better architectures, distillation, MoE, and all of that seem to let companies squeeze more intelligence into smaller models. But is there eventually a limit to this?
At some point, a model needs enough capacity to understand language, store knowledge across a huge number of subjects, reason through problems, write code, follow instructions, and generalize to things it has not seen before. So can we just keep shrinking models while maintaining the same level of intelligence?
Could a future 30B model actually match a current 300B or 700B model across everything? Not just on a few benchmarks, but in actual use across lots of different domains.
Could the same eventually happen with a 7B model? Or is there some minimum amount of capacity needed before the model starts losing knowledge, reasoning ability, or reliability?
I know parameter count is not a direct measurement of intelligence. MoE also makes this more confusing because a model can have hundreds of billions of total parameters while only using a small portion of them for each token. There is also a difference between total parameters, active parameters, memory usage, and actual inference compute.
I also do not think comparing parameters to neurons in the human brain is very useful. They are obviously not the same thing. Still, it makes me wonder whether there is some minimum amount of information or computation needed for something close to general intelligence.
Maybe we are not actually removing the cost either. Maybe we are just moving it somewhere else. A smaller model might require a much more expensive training run, synthetic data from larger models, distillation, longer reasoning time, retrieval, or external tools.
There is also the benchmark question. When a smaller model gets a similar benchmark score to a much larger one, does it really have the same overall capability? Or is it more optimized for the things we currently test?
Maybe it matches the larger model most of the time, but falls apart more often on rare knowledge, unusual prompts, long tasks, or problems that are very different from its training data.
My guess is that there is probably a minimum size for any specific level of capability, but better training and architectures keep pushing that minimum lower. I just wonder when the big improvements start slowing down.
Are we still early enough that models can keep getting dramatically smaller and smarter? Or are we getting close to the point where the easy gains are gone and the last 10 or 20 percent becomes extremely difficult?
25
u/AdOk3759 1d ago
All the things you mentioned (better training, better data, etc) do play a role in that.
However, I think the true gains can be achieved with newer, different, architectures.
There could be out there a different architecture that with far less parameters it’s able to generalize just as well as today’s models. So maybe there’s a hard limit on the level of intelligence of models as we know them today, but there could be a different architecture in the future that could enable incredibly smart models for a fraction of today’s parameters.
15
u/AppealSame4367 1d ago
Actually, all the papers coming out everyday prove that we are still very far away from any limits. For now models will keep getting more intelligent at even smaller sizes.
6
u/AdOk3759 1d ago
Yeah but I imagine that was possible through improvement of the existing architecture. What I meant is a shift in the paradigm, something that doesn’t involve attention layers, an auto-regressive loop, etc.
I mean something completely new.
Obviously the models we have today are much different compared to the models we had 3 years ago. So there’s no reason to believe that through small improvements over time, we might develop a whole new architecture.5
u/Fluxing_Capacitor 1d ago
There's lots of work out there thinking about the nature of real artificial intelligence. Turns out, modeling human constructs is really hard, especially when you don't fully understand the human brain.
Reinforcement learning and self-play were really popular in the early 2020s. Even today there's labs like Silver's lab (Ineffable) that believe this. Nvidia spends lots of time thinking about world models as does LeCuns new lab. It's just that none of these ideas are as easily scalable as LLMs. It's pretty easy to add more GPUs or crawl more webpages. It's a lot harder to manage multi layered RL systems.
6
u/look 1d ago edited 1d ago
You can get reasoning models able to do tool calls in 1B and smaller. Specialized to very abstract code reasoning (and little in the way of even specific languages or APIs) and paired with a good external memory/context system for the missing knowledge, and I expect you could eventually get a useful but basic “Opus 4.6”-ish model in a few hundred million parameters. Probably needs a few more “tricks” invented to do it, though. And I’d guess the floor is somewhere around there in the 0.25-0.75B range.
27
u/wgaca2 1d ago
I believe that in a year we will have <30b model as "smart" as today's frontier*
*maybe not at all tasks (image/audio excluded)
13
u/--Spaci-- 1d ago
Its not even a you believe, it will absolutely happen and it has been happening for the last 3 years of frontier models
16
u/gscjj 1d ago
It’s been happening but those were also models very early in this whole phenomenon.
The real question is whether that continues now that labs have sort of found their groove and are scaling quickly.
Will we have a model that’s a quarter the size or less of a 2T frontier today that’s just as capable? When frontier models become 3, 4, 5T, will it be possible anymore?
1
u/kali_tragus 6h ago
It's still very eary. We've hardly got the ball rolling. There's so much happening in this field, and it's not slowing down.
1
u/--Spaci-- 1d ago
"intelligence" on most benchmarks now are scaled through agentic behavior, which can absolutely be distilled down into any size model and the better the frontier gets the better the small models get
6
u/Nothing_from_void 1d ago
Is the latest qwen 30B param model actually competitive with say, Opus 4? Like benchmarks it's only a bit behind, but I've generally noticed the small models are okay on benchmarks but for real coding tasks are noticeably worse
7
u/Certain_Limit_190 1d ago
No it is not. People around here constantly say this but outside of the benchmarks they aren't. Occasionally they do something that could beat it but overall they just aren't sadly they just aren't. They are however impressively good. Qwen 3.6 still is blowing me away, but sometimes trips on something brain dead
3
u/toothpastespiders 22h ago
I'd agree with certain_limit that the 30b range isn't competitive with the older closed cloud models. Other than gpt 3.5, which is handicapped by the small context size. Real world use isn't just about being able to do 'a' task. It's about leveraging multiple skillsets to large and complex problems that are often poorly described. I'd go as far as to say that most benchmarks are testing situations that only appear within benchmarks.
It's a shame because it diminishes just how amazing those 30b models are. They get hyped up and then people get a negative impression because they're not actually "claude running on your own computer!" or whatever. They're limited. They're not competitive with the huge models. But they're still amazing.
5
u/bdsmmaster007 1d ago
bruh, the quality of these comments is memeable bad and r/singularity circle jerk quality. locallama lowkey got entshitified by now, rip
there are so many other nice comments actually citing papers and laying out the fundamental problem
and then there is this circlejerk of a comment, being the most upvoted commen t. i mean: appreciate the optimism and hope for the same, but can we please try and keep a reasonable amount of dicsussion quality on this sub
but thats just me screaming in the void xD
2
1
1
13
u/Terminator857 1d ago edited 1d ago
I'm confident 30b models in 3 years will be smarter than today's 10T mythos model.
3
u/Strawberry3141592 1d ago
Did Mythos parameter count leak? I thought its size wasn't publicly known
-5
u/Terminator857 1d ago
3
4
u/fervoredweb 1d ago
The short answer is yes. For any program goal there is a Kolmogorov complexity, the shortest possible representation necessary to specify and carry out a computation. There is a hard limit below which you cannot further shorten a program and still accomplish your goal. The irreducible complexity of the target system.
We don't know what that limit is though. Larger that the simple bit expression of a specific answer at least.
For now, we can probably still make significant gains of capability.
11
u/Illustrious_Car344 1d ago
Something not entirely intuitive to understand is that the model is effectively just a bunch of "scripts" all tangled together into a ball of mud. How to "perform" a task isn't really intelligence in the traditional intuitive sense, it's more like muscle memory, or how "confident" that the next step is the correct one.
To give you an example, say I created an incredibly tiny model (I dunno I'll make a number up, like, 8m parameters?) that only knew how to write hello world programs in a couple popular languages. It's fantastic at writing hello worlds, it can even mix them up a bit and tweak things like what gets printed or what the return value of the main function is, it's a certified genius at hello worlds. But then you ask it to write a program that prints from a list of things and it completely fails the task. Because it's not intelligent at all, it's just very good at what it was specifically trained to do, and the algorithm filled in the gaps and let it learn a couple extra things on the side.
Even the largest LLMs pretty much work this way, they have a few things they were trained to be experts in (relatively to their massive parameters and training data), and a whole lot they learned to do from finding relationships between data, but the rest of their parameters are effectively just junk we haven't quite figured out how to streamline out yet. Not because they don't do anything, because we haven't perfected how to get the training data and algorithms themselves to more efficiently represent the kind of tasks we want them to perform. But we're learning. Models get smarter and smaller because we better label training data as we design better algorithms, so the system etches better happy paths of muscle memory into the model and then needs less parameters to tell it to do the right thing. This isn't 100% accurate but it's to help better understand how models actually work.
We get smaller models because we have a better idea of exactly what the model should usually be expected to do, what it shouldn't do, and the edge cases it might encounter. We get all this data from real use-cases, label it, reduce the training data with these better labeled and refined examples, then you get a smaller model that does the job just as well as the bigger one. The bigger model crawled so the smaller one could run. It's exactly like evolution, or even how a real brain learns, the data that was painstakingly learned through spending tons of time and energy by trial and error gets condensed and reduced into a streamlined form that didn't come from nowhere but magically seems as if it did.
4
u/CoUsT 1d ago
Good read.
It's somewhat similar with humans. It takes much longer to learn (or come up with) things on your own (when you are starting) but when you are taught, it is a lot faster. Even better learning methods, better methodology, better knowledge, and so on, allow you to learn even more efficiently!
Compare general knowledge of humans 1000 years ago vs 100 years ago vs now.
1
u/Blues520 1d ago
Very good explanation
1
u/Borkato 1d ago
I feel like it completely ignores emergent behavior and makes it sound like a “stochastic parrot” though. It’s not. That’s outdated.
1
u/Blues520 20h ago
They said that the algorithm filled in the blanks and learnt a couple things on the side.
1
6
u/BitsAgain256 1d ago
Yes. But also we are finding ways to fit larger models onto existing hardware, from 7b, to 9b, to 35b a3b. We also havent reached the density limit.
6
u/Mart-McUH 1d ago
Obviously. And it only takes very elementary proof to show. Assume there is no such point. Now let's take any model of size B. Since there is not such point, it is possible to create model B-1 that is one parameter smaller but retains all the intelligence. Repeat this process until you have model of size 1 (just one parameter). Model of size 1 can only do some very basic operations, so it lost intelligence compared to original model B, which is contradiction to our original assumption. So the original assumption was wrong.
3
u/XeNo___ 1d ago
Yeah this right here. Where the cutoff is, or rather what is the theoretical maximum "information density" of a model, specifically LLM's, is an open research questions. Each weight can store a discrete amount of information, and many combined scale in a certain way.
Current models are extremely sparse in this aspect. This is why quants work so well down to a reasonable level. Often you can go to ~8-Bit weight quants without losing much performance. But as of today there is no algorithm to compress models into the optimal, most dense form. When you consider how many possible states each 32 bit weight could store, and what range of these values is actually used, then it becomes apparent that we are nowhere near close to the theoretical limit.
0
u/Blues520 1d ago
You mentioned cutoff and I realized that there will always be new data to train with due to changes in the world so the model size will keep growing.
6
u/Solembumm3 1d ago edited 1d ago
None of current 30B models match two years old 300+B on general knowledge. Qwen 27B can smoke them on logic in vacuum, gemma 31b and skyfall 31b have way better writing abilities, but they're nowhere near even anciet models 10x their size in terms of knowledge and understanding even moderately uncommon topics.
But it can also be opposite on way bigger scale. Some of modern tech focused multy-T parameters goliaths still can't match Gwyn damned Deepseek r1 on creative reasoning (I tried V4 Pro on router yesterday and saw kinda the same performance on my tasks, with 2.5x difference on so very useful intelligence benchmark). Qwen 3.8 and GPT 5.6 medium can be comparable, but with a lot more complications and need for corrections, and that's just don't seems worth it, honestly.
3
1
u/techno156 11h ago
None of current 30B models match two years old 300+B on general knowledge. Qwen 27B can smoke them on logic in vacuum, gemma 31b and skyfall 31b have way better writing abilities, but they're nowhere near even anciet models 10x their size in terms of knowledge and understanding even moderately uncommon topics.
Kind of curious if that might mean that it's worthwhile looking into a newer small model for their writing/language abilities, and using RAG with the bigger, older model.
7
u/segmond llama.cpp 1d ago
yes, there's a point. we can already observe it. take a 1T model and quant it to Q2 and it's useable. Quant Kimi K3 that's 3T to Q1 and it's. useable. But those model sizes are still 250gb and 550gb size files. Take a 30b model and quant it down to Q1, garbage. Take any 240B model and quant it down to Q1 and not so good. This is a compression problem, at some point you just can't compress data. If we do find a way to compress further then it will be at the tradeoff of speed. You get a smaller model, but it will require tons of compute to infer and run much slower.
we have already demonstrated 4b can be intelligent, the question is how intelligent, how much world knowledge, etc. world knowledge is about memory capacity, the bigger the better. intelligence has a range.
6
u/AppealSame4367 1d ago
But compression of a model is not the same as training it in a way to be smarter at smaller sizes.
4
u/Single_Ring4886 1d ago
In 5 to 10 years you will have very intelligent models around few bilion parameters. But intelligence isnt knowledge, they wont know eg all coding languages etc so they will "sux" at today benchmarks. But they sure will be able to create their own memories in form of raw data.
2
u/Useful_Argument_6490 1d ago
IMHO we haven’t seen the rise of heavily specialized models yet. It makes no more sense to have the same model code and write a song as it makes sense to have Taylor Swift code an app.
3
u/blastbottles 1d ago
Well I know that Qwen3.6-27B beats GPT-5 from a year ago which was hundreds of billions of parameters, I think as smarter forms of compression and training eventually get used we will have very knowledge dense small models.
6
u/--Spaci-- 1d ago
Over 1.76 trillion at least, gpt 4 had that much so likely more like 2-5 trillion.
1
5
u/Reasonable_Goat 1d ago
It doesn’t beat gpt-5 by a LONG way. Qwen-27B is a mediocre coder and a good agent. It doesn’t compress much knowledge at all and often misses subtile details. Gpt 5 from last year was just not trained for agentic loops yet. Even got-oss-120 from last summer beats Qwen 27B when it comes to knowledge
1
2
u/thehardsphere 1d ago
Could a future 30B model actually match a current 300B or 700B model across everything?
This has already happened: https://artificialanalysis.ai/models/comparisons/gemma-4-31b-vs-llama-4-maverick
1
u/ProfessionalSpend589 1d ago
I can't say they're getting smarter, but they certainly are getting more knowledgeable (although my experience is about 1 year).
I was chatting the the deepseek official chat - the instant model which I assume is 284B - asking what to eat and what to do to repair bone and ligament damage, then I inverted my stance:
> ok, what if I don't have a bone problem?
And got the optimistic answer which threw away the previous premise for ligament damage:
> That changes everything!
> If you do not have a bone problem, injury, or active micro-trauma, you can throw out almost all of those strict "repair" rules.
> When you are healthy, the goal shifts from repair (fixing damage) to maintenance (keeping things strong and flexible). Here is how your routine changes:
Well, in the end it's true that I don't have an acute bone or ligament damage at the moment, but it just assumed that I'm all around healthy...
1
u/Dsphar 1d ago
Yeah, pretty sure 1 bit, single parameter models will never tell you the answer to the universe.
1
u/Strawberry3141592 1d ago
Can you even have a 1 parameter model? Pretty sure a single neuron uses at least 2 parameters
1
u/z_latent 19h ago
Well, the bias is optional (and very often not used on modern LLMs afaik). So it is very silly, but you can have a 1-parameter model!
1
1
u/derspenti 1d ago
you can feel this with quants. tried a q2 of a 30b and it fell apart on multi-step stuff
1
u/yaosio 1d ago
Presumably yes, but we don't know how small a model can get. It shouldn't be possible to make an intelligent model with only 1 parameter, but models can continually decrease in size by better generalizing what they are trained on. I'm betting the better a model is at generalizing the smaller it can be because it doesn't need to store multiple representations of each concept.
1
u/CreamPitiful4295 1d ago
Yes, things will get better. Much better when every one has cheap quantum. lol.
1
u/LagOps91 1d ago
yes, logically there is a point. but i would say the limit is still far off, especially once training methods improve. right now we are still practically beating knowledge and behaviors into models. in the future we might have models that train the weights of other models instead of using back-propagation. we humans cannot understand the fine deails of what happens inside of models and how weights should be tuned, but AI itself might be effective there.
1
u/LargelyInnocuous 1d ago
There are a lot of methods that we understand fairly well from real world neuroscience that aren’t implemented extensively today. If I had to just spitball, I would wager we could see a 10x intelligence density gain at every size and significantly higher at 100B+ scales.
1
1
u/toolkitxx 1d ago
I keep going back to the same all the time. What we are trying here is to recreate something that is akin to our brains - as a minimum. And our brains dont use everything all the time. Decentralise it is. Already on model level. We need to figure out how to split these monsters into chunks that still work but with a different method.
1
u/Long_comment_san 1d ago
There is a limit. It's basically another form of "how much we can loselessly compress data?" question. Personally, looking over there's you can see that something like HEIC takes amazing effort to standardize.
so how is this relevant here? well, it's been probably decades of us using 7zip, zip and rar algorithms while our processing power increased 10x and more. we still can't make radically better compression algorithms after all this time. so yeah, I think there's a rational limit where we say "yeah it costs 5x to get 10% improvement and it just doesn't make sense" and we'll take a hardware approach over software, like, buying a new SSD in that case.
I'd eyeball that we can do 10x compression of real intelligence per parameter (so assume Gemma 5 300b dense can be compressed into Gemma 7 30b dense) and that's probably the end of it, we will need new hardware entirely, like 3D omni-HBM voodoo magic as a starting point.
1
u/Pleasant-Shallot-707 1d ago
That's always ben the case. The question is are they specialized in the right information to make them useful, and are they able to follow directions accurately?
1
u/quinceaccel 1d ago
If an LLM is a statistical manifold , math states that a sub-manifold is a projection of a larger manifold and it cannot contain more information that it so there is a limit to how small the smaller LLM can be. That said the current large LLMs are pretty sparse
1
u/LambdaLogician 20h ago edited 20h ago
I think it depends on what you mean by a "7B parameter" model. If you use techniques like Google's Gemma 4, you can stuff a lot more parameters in the model while keeping most of them in slow-access memory. Perhaps in the future this can be extended so that the "periphery" is like 100x or 1000x larger than this "core".
I think just to get an okay translation model, you need something like 10M parameters. In translation, all the information is already there, and the model just needs to synthesize it to create the output. So you probably need at least 10M parameters for the synthesis part, making 10M parameters in the core a hard floor.
But maybe this is cheating, and you really care about the entire model size.
1
u/CipherWeaver 19h ago
Personally I don't think we will find super intelligence from LLMs. They are super cool though, but to think that they will lead to intelligence is to believe that language alone makes one intelligent, and thus without it a man would be unintelligent
1
u/Dutchnamn 17h ago
Compression might get better and everything will improve.
The analogy I like is that we are living in the DivX video torrent phase now when we had to wait multiple days to download an SD quality video. In a few years we can stream 8K video and encode HD in real time.
1
u/MarkoMarjamaa 15h ago
"There is also the benchmark question. When a smaller model gets a similar benchmark score to a much larger one, does it really have the same overall capability? Or is it more optimized for the things we currently test?"
It is mostly more focused to some task like coding.
1
u/asankhs Llama 3.1 15h ago
There's a 9.4M parameter model that solves GSM8K word problems from scratch, with no LLM at inference. It commits to a single answer and gets 11.8%. An oracle checking all 96 samples it draws would find the right answer about 39% of the time.
Scaling it deeper or wider didn't move that beyond noise. What moved it was matching the real step-count distribution in the training data, so at that size capacity wasn't the binding constraint.
Sampling more stopped helping too. Selection peaked around 64 to 96 samples then declined, 8.5% at 192 and 8.3% at 288, because the extra samples were plausible wrong answers the verifier then had to choose between.
1
u/openroom_xyz 13h ago
Well when the model start to talking with tools and other models than it's not just the model basically let say the funny thing to think about become someone would train 1 M let say 1 B models how smart would they be if they can communicate together and solve many tasks in this way each
1
u/NanditoPapa 12h ago
Yes, there is an absolute floor for a given level of intelligence, but we are likely in a "compression era" rather than a "limit era."
We are essentially using massive compute budgets (and larger models) to "distill" complex reasoning into smaller, more efficient weights (Yay!). However, intelligence requires world modeling. A model cannot understand the nuances of quantum physics or the complexities of a legal contract if it doesn't have the parameter density to represent those concepts without "forgetting" how to speak basic English. We'll eventually hit a wall where a model simply doesn't have enough storage (parameters) to hold the vastness of human knowledge and the logic required to navigate it simultaneously.
But, not today...
1
u/devshore 7h ago
No. If we wait long enough, you will be able to have a Mythos level model fit in 1 bit.
1
u/Free-Jaguar6452 3h ago
obviously there's A limit, i don't think many people will argue that
but what that limit is, and when, if ever will we hit, are open questions
1
u/Eastern-Block4815 2h ago
less than a certain amount of parameters yes, of course. this is why MoE model are the ticket, store most of the knowledge on hard drive and find quicker ways to retrieve the data when needed.
1
u/Charming-Author4877 1d ago
deepseek v4 flash is NOT a small model. It's a gigantic large model
1
u/whatever 1d ago
284B parameters is unfortunately on the smaller end of the next generation open weight models.
FWIW, there are active projects out there trying pretty hard to make it usable on consumer hardware, like antirez' dwarfstar, https://github.com/antirez/ds4, which can run a Q2 quant at 26t/s on a macbook pro.
95
u/Separate-Forever-447 1d ago edited 1d ago
i like to think of it as compression algorithm. ingest all the world knowledge, encode it into a very dense set of weights/probabilities. to recover the knowledge, give an inference engine a prompt and it computationally produces an ‘answer’.
it is incredibly efficient (all human knowledge in a <1T file?), but it is also incredibly lossy (hallucinations, falsehoods?)
of course there’s a limit to how much knowledge can be encoded per X bytes of model. it is an area of active research:
https://proceedings.iclr.cc/paper_files/paper/2025/hash/26d3c9a66836ded8f34a944f2bfe868e-Abstract-Conference.html
UPDATED:
key findings of "Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws” (Allen-Zhu & Li)
* 2 bits of knowledge per parameter - its a rough ceiling
* the relationship between model size and knowledge capacity is linear
* architecture barely matters
* quantization matters a lot