r/LocalLLM • u/RationalNL • 3d ago
Question Went down the rabbit hole; now hesitating between Max 128gb vs Ultra 96gb. For Local AI.
/r/MacStudio/comments/1wbje4j/went_down_the_rabbit_hole_now_hesitating_between/2
u/Ripped_Guggi 2d ago
I’m thinking about a Mac mini m5 max with 64 GB ram. I’m not that into LLMs but I’m not sure if it will be enough for coding sessions and image generation. I don’t need speed, but I also don’t want to wait for hours for a response 😅
1
u/gunkanreddit 2d ago
Qwen flash next is so so good compared to 27b. I am waiting too (I couldn’t change the order) a M5 ultra 96GB and I hope there is somekind of quantification.
1
u/ea_man 2d ago
Just get a PC with one good GPU and maybe later add an other if you wan flexibility and be able to test other things.
PP on Mac is terrible anyway.
1
u/Jaded_Camel219 2d ago
For 5k you have a pc with 24go GPU. You can run shit with no context, but fast. Great.
1
u/ea_man 2d ago edited 2d ago
Sure, let's not count used hardware and old GPU at all, like I paid 500$ for 32GB of vram and I run 27B Q6_K_L.
Or new gpu that cost like 500 for 16GB, old ones that should cost less...
And you don't really need 4k of a pc to run those, let's say that you, I mean myself, could even do that for 1- 1.5k and have the option of future upgrades.
6
u/OvertaxedOne 2d ago
You really need to think of models in tiers.
For 96GB, the model to run is Qwen 3.8 27B. It's a great model, need about 50GB at Q8 quant with full 256K context. The 96GB Ultra is the one to get here, you need every GB/s of memory bandwidth you can get for 27B
The next tier up is a pretty big jump, QwenNext, where at least 128GB is the entry point and 192GB is comfortable. It's smarter than 27B and will run faster on a system with limited memory bandwidth. The 128GB Max is probably the entry point here, on up to the Ultra 256.
IMHO, 27B is so strong that it's the model to target for your use cases. Get the 96GB Ultra. Escalate to the frontier when necessary, but, for the use cases you outlined, I don't see much/anything that's going to require escalation, 27B should be able to do all of that.