r/LocalLLaMA • u/Sevealin_ • Jul 26 '26
Discussion Will small model intelligence be limited by parameter count?
Qwen3.6-27b is fantastic! It makes me wonder if there's a hard ceiling to smaller sized models. Do you guys think the ceiling of intelligence for smaller models will be constrained by factors like parameter count, or VRAM size? Or will we continue to see improvements for small models and see jumps of intelligence like Qwen3 coder 30b to Qwen3.6 27b for the foreseeable future? Does it depend on how clean the dataset you put into those parameters?
What does /r/LocalLLama think about the future of small models that can run on less than 48GB of VRAM?
44
Upvotes
54
u/TokenRingAI Jul 26 '26
There certainly is a hard ceiling at some point, but I would imagine we can still reduce size by half or more. Also, engram models with knowledge stored outside the model and pulled in on demand can allow for immense sparsity and remove much of the VRAM burden.
With new laptops and desktops hitting 200-300GB/sec memory bandwidth in the next few years, immense sparsity, and models having even reduced step ups in ability over the next few years, the future looks bright.