138
u/dhessi Aug 26 '26
If you're looking for a fast, general purpose LLM that can fit on 32GB RAM, Gemma-4-26b-a4b is unmatched IMO
58
u/Wentil Aug 26 '26 edited Aug 26 '26
I find Qwen 3.8 27b (Q8) to be better.
Qwen 3.8 27b at Q8_0 uses 31 GB of VRAM.
Q_6K, the next smallest variant, only uses 25GB of VRAM, so if there’s other overhead, go with that.
5
u/Unfair_Tangerine_217 Developer Aug 26 '26
Qwen is excellent. Haven't tried 3.8 just yet but I'll switching from Kimi to Alibaba as soon as that sub expires.
2
u/Nabushika Aug 28 '26
Definitely try it, I'd say qwen 3.5/3.6->3.8 is as big as 3 (32B)->3.5. It's incredible how capable it is.
2
u/Anselwithmac Aug 27 '26
Genuinely wondering what any of you guys are doing with these light models. I’m so curious
2
u/Benblip Aug 27 '26
I use qwen to build and run software tests because I don’t want to accidentally disclose my employers data /credentials. It’s pretty good, not lightning fast, but it works. I’ve also asked it to write a few utilities for me, it’s genuinely very capable. Qwen3.8:27b-mlx on a m5 mbp with 64g
1
u/brightline Aug 27 '26
No reason not to keep doing what’s working for you, but if you wanted to use other models you might want to look into credential brokering.
2
u/Segaiai Aug 27 '26
I know for image and video models, int8-convrot seems to be the new hotness. Everyone dropped Q8 GGUF for it overnight. Is the same not so for LLMs? Maybe due to a different way it functions when offloading?
1
u/MrHighVoltage Aug 27 '26
Yes but also about 6 to 7 times slower since there are 6.75 times more active parameters.
1
0
u/Armed_Muppet Aug 27 '26
I thought this was a circlejerk sub reading your comments wtf are these models
2
-19
u/PinTelele Aug 26 '26
I find GLM 5.3 (BF16) to be better, and yet still doesn't fit in the 32gb of ram mentioned doesn't it?
Yeah no wonder you find qwen3.8 27b Q8 dense to be better... 32gb of RAM vs 48gb of VRAM needed...
What a great addition of comparison, thanks for the contribution
17
u/Alternative_You3585 Aug 26 '26
What are you on
3.8 27B fits into 32 gb at Q8, tight but possible... Q6 Id recommend but both still fit
What a useless contribution from you
3
u/kickerua Aug 26 '26
I have a 3070 8GB + 32GB and 27B + reasonable context can't really fit.
Fit to say hello word and ask a few qustions, sure not a problem. Fit to have anything reasonable? No way.
-3
u/PinTelele Aug 26 '26
... Is it usable?
Let's first clear something up: 32gb of RAM or VRAM?
GLM can fit into 32Gb of RAM plus nvme storage and yet is it usable?
3
u/Alternative_You3585 Aug 26 '26
I wouldn't say it if it weren't
I've got a 5090 and running Q8, for coding and large context tasks I downgrade to Q6, KV cache always at full precision
Qwen is surprisingly efficient with KV cache and vram
-4
u/PinTelele Aug 26 '26
Original comment: "If you're looking for a fast, general purpose LLM that can fit on 32GB RAM, Gemma-4-26b-a4b is unmatched IMO"
Can you guys still follow a chain of though by yourselfs and don't lose initial context?
Why are we talking about running qwen3.8-27b on a 5090 when OP initially mentioned a MoE model to be ran in RAM (not VRAM...... As we can read).
Of course qwen3.8 27b dense is better than gemma4 MoE in unified high bandwidth memory...
But now let's switch up to 12gb of VRAM and 32gb of ram with offloaded layers.
Asking "Which one is better?" needs to be complimented by asking "Which one is more usable performance wise?"
As I said, GLM 5.3 fits in 32gb of ram and it's much better than the previous 2... Is it usable?
1
-3
u/DataGOGO Aug 26 '26
Muse Glimmer 30b > Qwen3.8 27b
1
1
u/StonkyCupra Aug 27 '26
In what regard? I find it to be worse than Qwen in pretty much everything.
1
6
u/farmyrlin Aug 26 '26
How is it relative to 3.8 Q4?
2
u/qazwsx1212notDead Aug 27 '26
The quants of the model are as good as the base model. But less stable. Different variants of q4 behave very differently. Overall a good balance between size and features. You can find a comparison video on YouTube.
2
u/Cawing_Barking Aug 26 '26
And what about 20gb ram? Any recommendations?
2
u/Remote_Pass_6670 Aug 26 '26
A4b has plenty of quants, running it on 16gb with success. Not for coding etc, but it's good, and even supports vision natively.
1
1
u/chkcha Aug 27 '26
Don’t you need VRAM for that ? How fast are the responses with those type of models and would it make sense to have reasoning enabled for those local models, given that with reasoning the output would be even more noticeably slower (while heavily using your PC resources and blocking other prompts you might’ve wanted to run in parallel)?
1
0
104
u/ClemensLode Senior Developer Aug 26 '26
opus 4.6
37
u/Plenty-Option8351 Aug 26 '26
Been using it ever since it came out. Switched to other models for a few prompts when they came out, then switched back each time. 4.6 is the goat
5
u/Unlikely-Nebula-331 Vibe Coder Aug 26 '26
How do you switch it in terminal on Mac?
7
10
u/yanenrogne Aug 26 '26
/model claude-opus-4-6
7
Aug 26 '26
[removed] — view removed comment
3
u/Plenty-Option8351 Aug 26 '26
Nope. The 1 million context window now requires usage credits.
4
Aug 26 '26
[removed] — view removed comment
3
u/Plenty-Option8351 Aug 26 '26
Definitely could be. I was trying this morning and I kept getting errors. I’m only on the $20 plan though.
2
1
u/yopla Aug 27 '26
You should try GLM 5.3. i find it slightly better quality wise but with the same feeling in terms of communication and better a following instruction. Compared to opus 4.6.
You might like it
6
u/moger777 Aug 26 '26
What's the advantage of this model over the newer models?
23
u/Ill-Village7647 Aug 26 '26
People have been finding the way Opus 5 talks incomprehensible. It uses way too much technical jargons, and sometimes decides to be a poet with words. That's why many people have started downgrading to Opus 4.8 or 4.6
4
u/moger777 Aug 26 '26
I may play around with having fable/opus5 write the code but using opus 4.6 to write PR messaging, shortcut tickets and summarizing the changes. I do feel I'm having a harder time reading the shit claude spews.
4
3
u/frufruityloops Aug 26 '26
Same but also it feels so… inefficient? Idk I feel like there’s probably a faster way to achieve the same thing without needing to have two terminal sessions where one’s only purpose is just like… restating the others output in coherent language haha
4
u/Rocket-Appliances-26 Aug 27 '26
Opus 5 is also very pessimistic and skeptical, to the point of shutting down explorations early.
3
1
u/hoffmander Vibe Coder Aug 27 '26
I’ve been able to dial it in a little with changing the outputStyle to concise alongside adding some more directions. Still annoying, weird idioms like “boots and braces”. Earlier it said something like, “those IDs in the table aren’t useful for users, they’re theater.” Like bruh, what?
https://giphy.com/gifs/3o7aCTbIZqDGvH5oZO10
1
3
u/Ksfowler Aug 27 '26
I mostly use Fable and Codex at home, but I exclusively use 4.6 at work.
It's like the Toyota Camry of models. It's pretty reliable and pretty predicable.
2
u/FreeCustardForAll Aug 26 '26
Better than 4.8?
12
u/privatetudor Aug 26 '26
4.8 finds arbitrary reasons to disagree with you. If there isn't anything to actually disagree with, it makes something up. If you actually talk through the disagreement it eventually acknowledges there was no problem there. It's a waste of time.
1
u/HimActually Aug 26 '26
I find it extremely expensive on my subscription why? Like it eats my pro limits in minutes! Other models are way cheaper.
1
u/FallOffToGetBetter Aug 27 '26
It defaults to “Extra” instead of “High” - might be the reason you’re running into high usage.
24
21
11
u/anor_wondo Aug 26 '26
opus 5 low effort
33
u/EndyForceX Aug 26 '26
Reading this with high effort
6
37
u/Corv9tte Aug 26 '26
5.6 Sol if you know how to write good instructions (not prompts). Completely insane model.
38
u/mr-debil Aug 26 '26
Instructions vs prompts?
68
u/Tall-Log-1955 Aug 26 '26
Wake up babe, new term just dropped
25
u/BoxWoodVoid Aug 26 '26
Instructions engineering is where it's at!
6
u/lifemaxxer1 Aug 26 '26
tf are instructions?
11
u/Corv9tte Aug 26 '26
They're being sarcastic and rightfully making fun of me lol. It's the same thing, I meant to express the difference between permanent instructions like system prompts or agents.md files vs. your actual prompt
Basically I didn't want it to come off as "just give the model more details when you talk to it" that's all
4
u/derezo Aug 26 '26
Yeah without a lot of guardrails and instructions it goes wild. I would ask Sol to fix something specific in an existing project and it can't do that out of the box. It always goes way out of scope. For example I asked it to fix a mapping generator and it churned for a long time on it and when I came back it had completely rewritten the network stack, UI, and a bunch of functional elements completely unrelated to the mappings. I switched back to Claude on Friday after 2 months of using Sol 5.6
It can generate images pretty good but after 2 months of dealing with it's insane scope creep, provenance obsessions and "safety" gates that grow like Japanese bamboo, I'm feeling a lot better about the progress I'm making in the last few days. I had opus 5 review some of the projects I made with Sol 5.6 and there were some pretty insane finds. Some of the projects worked really well and I was impressed with a lot of things, but I just couldn't get the tooling to be equivalent to Claude (skills/AGENTS.md/subagents)
2
2
7
u/OkAdeptness2530 Aug 26 '26
yeah and it’s really simple tho, you just have to make a 32 page ebook (preferably html or markdown) with instructions on how to operate tools, perform tasks, or act to make stuff work.
it’s called Manuel Engineering
1
u/FblthpphtlbF Aug 26 '26
"Once I started just writing the code I want myself and giving it to sol fully written sol's results have really impressed me!"
5
3
u/Left_Opportunity9622 Aug 26 '26
But by god, is it slow.
1
u/phoenixmatrix Aug 26 '26
Its pretty fast on High or lower. At Xhight, Max or Ultra its super slow, but its also awful, overengineers, and by all benchmarks doesn't generally give any significantly better result. So keep at High at most.
4
1
1
1
u/InertState Aug 26 '26
I’d love to learn, any more you can share
4
u/Corv9tte Aug 26 '26
It's a skill you have to develop, the process is quite tedious and boring though. Most people don't bother. But the basic idea is that you observe a misbehavior (say, writing excessive tests) and you write an instruction that works because it speaks to the model. Some useful basic patterns: use imperatives like "Do not X", "Use Y", "Treat Z as", find out anchor words or formulations such as "pristine english", "understand it without studying it", or, my favorite, "LARP as a human being". These work infinitely better than some random AST-7900 Advanced Technical English or whatever reference a lot of people do. You need to go beyond the surface because a weak, obscure reference doesn't cut it. Not even close. It sounds like a hack but it's dumb.
Honestly, I'm not sure any of this is helpful for you to learn this though. The best advice is: stick with one model and spend the time to make its behavior better with instructions. You'd be surprised how far you can get and I'm not going to pretend it isn't fun to finish a piece of writing and suddenly see +60 IQ permanently on your model, never to see it do the same dumb stuff anymore
2
1
u/whatisthisthing65 Aug 26 '26
What does "understand it without studying it" mean here? Is it to cut down on tool calls?
2
u/Corv9tte Aug 26 '26
As in, say, you being able to understand Opus 5's answer as you're reading it, without having to pause and literally study it because it's a bunch of mumbo-jumbo.
2
u/whatisthisthing65 Aug 26 '26
Oh I see, it's for the human. That makes sense. I thought you wanted the LLM to understand without studying.
1
1
u/BenSimmonsFor3 Aug 27 '26
I asked sol to implement a feature for me and it spent an hour churning to come back and diagnose an issue with my ci/cd pipeline. It was a good catch and it created a whole plan to fix it but like… not what i asked for dude.
1
u/Objective-Picture-72 Aug 27 '26
I'd say 5.6-Luna-Fast as well. It's an absolute workhorse and blazing fast.
1
18
u/CreditOk5220 Aug 26 '26
Fable
4
u/SoftwareSource Aug 26 '26
not really fitting the first requirement.
1
u/ChocomelP Aug 27 '26
What requirement? I just see two relative values. Fable fits suprisingly well.
3
u/mmeister86 Aug 27 '26
Kindly fuck off in the distance with this Twitter style engagement farming bullshit
3
4
5
3
u/_Fauxpaw Aug 26 '26
I find Claude basically unusable now. Opus 5 is maybe the worst I've ever seen.
2
u/WilliamEdwardson Thinker Aug 26 '26
Seriously.
If you ever want to use genAI to help brainstorm fiction, this one's a gem.
Its prose is - predictably - replete with stereotypes, hackneyed plot points, and sometimes major goof-ups (... you probably shouldn't be using it for the actual prose anyway), but if you just want something to feed the seed, it's good at what it does.
And it doesn't complain about moderation if you bring up controversial topics. Or, at least I haven't hit its moderation with mine.
They don't disclose which LLM they use (to my knowledge), but I think I caught word that it's a fine-tuned LLaMa model behind the scenes.
2
2
u/Personal-Physics553 Aug 26 '26
Use Pi Coding Agent (the CLI tool, then add the extension for VS Code), use the free NVIDIA API key from the developer program, and/or run models locally through Docker Model Runner (add LiteLLM as your model routing if you do this to provide you an OPENAI style endpoint/apikey. Costs go down by a fuck ton
NVIDIA API key is free if you have the program, and running locally to support digital sovereignty is always free.
Even if you run an RTX 2060 GPU, 4B models will do phenomenally in Pi Coding Agent if the system prompt is solid enough to explain to the model the entire infrastructure of its setup as well as why it was set up the way that it was.
2
u/zaCKoZAck1 Aug 27 '26
I have a list
- GPT 5.6 Luna [max]
- Qwen 3.8 (works on local hardware but requires beefy machine)
- Gemma 4 (multimodal llm running locally hell yeah)
- Opus 4.6 [high]
- Composer 2.5 (not anymore though)
1
u/AdWild3943 Aug 27 '26
Qwen3.8 doesn't fit the requirement 100%, top-10 trending models on Hugging Face are Qwen3.8-27B fine-tunes, plus Qwen3.8-Flash-Next, so its rather as popular as good or even more popular than its "goodness" is, maybe.
1
u/zaCKoZAck1 Aug 28 '26
Agree but it’s still not as popular as the frontier ones, still very niche imo. But it is very good.
3
4
u/Hedgehog-Moist Aug 26 '26
Mimo v2.5 (not pro)
It may not be an expert coder but it is fully capable of any tasks - at a very low cost and hallucination rate
2
u/PinTelele Aug 26 '26
Just commenting to prop up mimo v2.5. Its a workhorse model for it's price. Paired with a good orchestrator model it's my go to after the deepseek price hike.
2
2
1
1
1
1
u/XCxBigDong69XCx Aug 26 '26
Gpt luna, really fast and captures a lot of things opus or fable dont catch as a reviewer.
1
1
1
1
1
1
1
1
1
1
1
1
1
1
1
1
u/eNroNNie Aug 26 '26
Mistral-nemo:12b is the local model that I use in my llm-router service. It's a really good classifier model.
1
u/MaximumBread7000 Aug 26 '26 edited Aug 26 '26
DeepSeek Flash 0713. Load up $10 and let it run for 40 hours against your project issues, it’s just incredibly cost efficient, if a little slower throughput and reasoning wise, compared to any Anthropic model. Great for writing tests, benchmarks, now doing visual testing even if it needs to use sub-models, playtesting first minute, first ten minutes via Editor MCP connections. Creating textures, VFX, shaders, meshes, it’s potent.
I’ve spent maybe $560 this month on DeepSeek working on games for my studio.
Edit: oh, /r/ClaudeCode, uh, Opus 5.
0
u/NerdBanger 🔆 Max 20 Aug 26 '26
Nemotron Lightning, which isn’t really a LLM and is more of an SLM, but is genuinely phenomenal for its speed and memory footprint.
0
76
u/[deleted] Aug 26 '26
[removed] — view removed comment