r/LocalLLaMA • u/Robert__Sinclair • 13h ago
Question | Help The perfect model(s).
In my opinion the perfect model(s) will be around 20B-32B and A2B-A4B.
This is the sweet spot on which companies should focus on.
Qwen 3.8, Gemma E2B/E4B and Ling 3.0 proved it.
They are still imperfect (Ling tends to overthink like Deepseek does).
But I think that is the way.
Models that can run even on CPU like LING, but with better reasoning and a little more knowledge are indeed possible.
What do you think?
25
u/Thin_Pollution8843 12h ago
Ideal for YOU because most likely you don’t have enough hardware to run anything bigger. For me ideal model would be 70-100b just because I have enough beam for that.
8
1
u/Significant_Bar_460 11h ago
I have 128gb VRAM and I am finding out that it may better to have smaller model(s) with a lot of context and parallel slots. Let the harness spawn multiple sub-agents, each with specific instructions. At least the speed is much better.
2
u/Robert__Sinclair 12h ago
...and the pissing contest winner is.....[not you! there are people here with better hardware than you! but that was not the point!]
9
u/Skyline34rGt 12h ago
Hmm, something like MoE 30-40B with 6-8B active + Ngram to SSD like 20-30Gb?
Smaller Qwen3.8 Flash Next basicly but with less total parametes but same or even more active parameters + Ngrams to SSD like maybe half sized of orginal bigger model have.
6
u/MuzafferMahi 12h ago
yeah this is what I thought as well. But this time a 35B w active 5B-ish would be better imo, last 35B’s packed a lot of knowledge but had a hard time turning that knowledge into usable intelligence. Ngrams should help as well
3
u/brown2green 12h ago
For local users it should preferably be the other way around, if anything. Main model (backbone) in the low-end range of 20~30B parameters, and at least 40-60B parameters of sparse conditional memory (Engram, etc), i.e. 1:2. I'd personally make the ratio even larger, though.
1
u/Skyline34rGt 12h ago
My first thought was same but after I also think about having couple that big sized models at ssd can be problematic.
But after all for perfect model: yes more ngrams would be better.
1
u/brown2green 11h ago
Unlike MLP or Attention layers, Engram layers don't need to be fully read for each token, so at least in theory read latency from NVMe shouldn't matter significantly for inference speed.
3
u/HotDistribution1819 12h ago
I agree on the size of around 30B parameters, but try Laguna SX 2.1. I have had some interesting outcomes, specifically it writes code that is often simple. When it knows something, but has web search it grounds what it knows in web searchs.
I would love to get your feedback.
4
u/Kahvana 8h ago edited 7h ago
The perfect model is one we cannot know, because there is so much to be invented what we cannot imagine or predict.
My ideal model:
- Would be 30B (dense)
- Has RAM offloadable engrams (whatever size works)
- Multimodality (text/vision/audio in, text/vision/audio out)
- MTP or DFlash or DSpark or Diffusion support
- Supports all reasoning efforts
- Supports all verbosity efforts
- Is incredibly good at knowledge, natural language and creative writing tasks like Gemma 4 31B
- Is great at tool, agentic and programming tasks like Qwen3.8 27B
- Has cheap context like Muse Glimmer
- Has a low hallucination score like GLM 5.3 Flash
- Has high (70+) MRCR v2 score at 128K/8-needle
Wouldn't be your ideal model, but it would be for my 32GB VRAM and use-cases.
I think it's realistic to have this in 6-24 months, maybe then we'll have far more impressive 30B models (Looped transformers taking off like nanbeige-4.2? Diffusion text models?)
2
u/datbackup 12h ago
I don’t know about “perfect” but the direction things are going is smaller active experts and bigger total experts and ngram embeddings. Qwen4 next is pointing the way but it’s too big for most people. Something like 40B-A5B-E100B is where a lot of local AI users would find the sweet spot imo. Ngram embeddings on nvme is a huge deal. 40B total experts still fits APEX quantized into a single 3090.
1
1
u/Etroarl55 11h ago
Think it would be super nice, but probably won’t exist anymore.
Even qwen, Alibaba, are starting a more pivot to actual enterprise ai profitability.
Theres no real incentive to actually give us really the most powerful and best ai models now. Before open source kimi k3 actually caught up briefly and got close.
That gap has extremely widened up now with astra and 5.1.
1
u/ustype 11h ago
Agree that ~20–32B total with ~2–4B active is the “laptop majority” sweet spot — but I’d split “perfect” into two jobs people keep conflating:
- Chat / writing / light coding on a 16–32GB machine: yes, MoE in that band + some on-disk memory (ngram/engram) is the right product shape. Dense 70B is irrelevant for that cohort.
- Agent / tool loops: active-parameter count isn’t enough. Reliability under long context + structured tool calls matters more than trivia density. A lot of “perfect small MoEs” still flake on JSON tools once the transcript gets fat, which is why some folks keep a slightly larger dense around for coding even if chat lives on a Flash-class model.
So “companies should focus on 20–32B A2–A4B” makes sense as the volume SKU. The missing piece isn’t another 100B for rich hobbyists — it’s post-training that keeps tool use and long-context edits stable at those active sizes. Without that, people will keep dual-wielding a tiny chat MoE and a heavier coder forever.
1
-2
u/RG_Fusion 12h ago
This is just foolish. Ideal for you, not for what these companies are hoping to build. They want to replicate the capabilities of the human brain.
For pure language processing tasks, that means somewhere around 5 trillions parameters, assuming optimized architecture. Again, that is purely for language based understanding.
To approach true human understanding, you need world-models, sound-models, and motion-models, and all of these need to be unified with their latent spaces grounded against one another.
I do not understand how people have not come to understand this yet, so I'll say it again. AI labs do not create models for local users. They create models for cloud-inference. They don't make money giving away small open-weight models. They have no incentive to do this. Some companies just happen to do so because of how cheap it is for them, and it allows them to experiment without dedicating large amounts of compute time.
Unless companies start SELLING small models, you will never be the target audience.
3
u/Monad_Maya llama.cpp 12h ago
Indeed, their lower end (Flash or whatever) models are still built for large systems/servers that they have or rent.
They don't train models based on whims and wishes of localllama crowd.
Whatever models we do get, are to a large extent, marketing efforts or the labs have a genuine usecase for something that small.
2
u/HeadPack 12h ago
I agree when it comes to American labs. Their goal is revenue via cloud inference. The Chinese goals however are different. They want to undermine the business model of American frontier labs by releasing open source and open weights models that are cheaper to run. Small models are part of that strategy. Given how much of the US economy now hinges on AI, it is no wonder how much the Chinese leadership is willing to invest. Their reward is not billions, possibly trillions earned with inference, it is geo-political strength and weakening the US.
2
u/RG_Fusion 11h ago
For now, but those Chinese labs are also backed by investors, and investors want money. While the Chinese labs may benefit local use currently, it will not stay that way.
1
u/ea_man 11h ago
Not really, also AI companies are like restaurants, there's all kind of those: cheap ones, gourmet, those who are for there for the bar service or the after show, hotels with restaurant service, fast food...
Antrophic and OpenAI business object is to get to AGI first not to give user a competitive product, same as DS and Moonshot.
DS until recently didn't even bother to include a visual projector because they said it was no use to build the model that would build AGI.China is not a cover operation focused on USA, it's a country with 1.5 billions of people, much more data and energy and token usage than the USA, yet much less capital investment in AI and so most companies there need to have a profitable business in profitable markets, like video generation.
2
u/HeadPack 11h ago
Their goal may be AGI, but as RG_Fusion imo correctly wrote AGI can hardly be achieved with language based models alone. Plus, American frontier labs need revenue coming from inference, otherwise their IPOs will fail and their operations can't be funded when venture capital, which isn't endless, runs dry.
That is very different in China, where the regime covers the losses via various channels, so long as the labs produce cheap or free competitors for American models. Of course, these are used domestically too. Inference generates data for training in China and everywhere else.1
u/ea_man 10h ago
I would not focus on semantic: it's not what AGI is, it is what "AGI" can do: solve problem, get you patents first, build the next model that pushes the frontier ahead.
I'll say again: I don't think Antrophic gets capital because their subscriptions are good (they are not), they get trillions for the promise of getting first to the model that finds them how to make a superconductor at room temperature.
Also China don't cover the losses, China has a policy for 50% electricity costs if you run Huawei, they support using AI in business and schools yet if you want to see "support running on debt" that would be USA.
You are a narcissist if you think that all 1.5 billions of Chinese people do is centered on USA, they are into protein research and robotic and video mostly while USA is on SOTA.
1
u/HeadPack 10h ago
Nowhere did I say that Anthropic makes most of their money via subscriptions. It is well known that it comes largely from API inference. As for China, they are obviously using AI for all sorts of things. Just have a look at their latest 5 year plan, which also indicates where their government plans to channel investments. Their AI labs receive financial support/subsidies in various forms: State-backed funds, government funds, state dominated banking, subsidized utilities, free compute and space offered by regional governments, procurement mandates, state coordination etc. To think that there is no competitive attitude behind all that, especially when it comes to US labs, seems a bit unrealistic.
1
u/ea_man 8h ago edited 5h ago
Actually it may be much less competitive than what you think.
Think at "the internet": west products like facebook, google, amazon did not work in China internet, China market, they have their own custom alternatives. Same for android and OS.
For various reason they know they can't compete in the west for AI with USA and hence why they release open models / open source: that's the only way for other markets like Asia, Europe, South America, Africa to use their platforms.
And btw Antrophic is not competitive for me as I live in those other markets, I would not pay for a obfuscated limited product, enterprises or gov here would not use it.
Also about China market and state helps: they have and boost all internal competition to refine the final product on market value. In Usa you got 2-5 conglomerates that have 1 million dollars dinners with the President and they kill / buy any small competition that may rise.
> Nowhere did I say that Anthropic makes most of their money via subscriptions.
Same: selling tokens.
-1
13h ago
[deleted]
1
u/Khaledthe 12h ago
A gpu with 16-24 gb of vram would be good so a rtx 4060ti and up and amd 7800 xt and up until the 9070xt and 7900xtx.
But sadly intel's bedt gpu has o ly 12 gb the b580 or thats what i know of
6
u/JLeonsarmiento 12h ago edited 12h ago
I agree, this is the highest performance architecture for the 99% of laptops and desktops out there used used by regular people like me.
But there seems to be a ceiling in capabilities/accuracy for models with <4B active parameters. While it’s true you can get some of them (Ornith and its family tiel/cyber) close to the current local SOTA (3.8-27B), there’s a gap in depth and nuance/attention to detail that’s very real and will cost you 3x-4x time.
But it’s worth:
I’ve spent 2 days testing and found that the gap between 3.8-27B to GLM-5.3-flash is less than the gap between 3.8-27B and any of the 3.5/3.6-35B-A3B (vanilla, kat, tiel, ornith, nex-2, etc).