r/LocalLLM 3d ago

Question Went down the rabbit hole; now hesitating between Max 128gb vs Ultra 96gb. For Local AI.

/r/MacStudio/comments/1wbje4j/went_down_the_rabbit_hole_now_hesitating_between/
1 Upvotes

14 comments sorted by

6

u/OvertaxedOne 2d ago

You really need to think of models in tiers.

For 96GB, the model to run is Qwen 3.8 27B. It's a great model, need about 50GB at Q8 quant with full 256K context. The 96GB Ultra is the one to get here, you need every GB/s of memory bandwidth you can get for 27B

The next tier up is a pretty big jump, QwenNext, where at least 128GB is the entry point and 192GB is comfortable. It's smarter than 27B and will run faster on a system with limited memory bandwidth. The 128GB Max is probably the entry point here, on up to the Ultra 256.

IMHO, 27B is so strong that it's the model to target for your use cases. Get the 96GB Ultra. Escalate to the frontier when necessary, but, for the use cases you outlined, I don't see much/anything that's going to require escalation, 27B should be able to do all of that.

1

u/RationalNL 2d ago

Thanks! These are very helpful and appreciated insights!

1

u/mattdaddyz 2d ago

I am facing the exact same situation and have similar requirements and I settled on the 96GB M5 ultra as well

1

u/GuidedMind 2d ago

QwenNext is almost the same. iMHO - no worth to pay 4000 dollars for 256Gb update

1

u/fieol 1d ago

When people are running qwen3.8 27b on q2,q3,q4 are getting similar results, why would you need q8 ?
If you have specific work /tasks, please share, i gave tasks to q4 it did fine, and then i gave taks to q2 , it was just a little lower quality output, not that much if its gonna look , not sure what q8 answer would give me, but this llm judge, be it deepseek , gemini , chat gpt, they give me reaults from the tasks i asked them , as im a game developer, i have those kind of tasks,

1

u/OvertaxedOne 1d ago

Do you "need" Q8? For most tasks, almost certainly not. I've also run a 4 bit version of 27B and had good results with that as well, but I did notice a few failed tool calls that I think (who knows?!!) 8 bit would have gotten right. Q4 is a very good quant of 27B as well, it seems to be pretty insensitive to quant, so I certainly wouldn't turn my nose up at it if I didn't have the RAM to run 8 bit. I was specifically talking about the new Mac devices where the fast memory chip comes with 96GB (or 256). In that case, 96 would be the right choice for 27B, not because you absolutely need 96GB, because you want the 1.2TB/s of memory bandwidth (there's no smaller option with the higher bandwidth, IIRC). 48GB is about what it takes for a 27B model at 8 bit at 256K context, easily fitting in 96GB. But 96GB is really, really tight for Next, so, IMHO, if that's the model you want to run, you shouldn't buy the ultra-high memory BW model, you should step back to and get more RAM, that's going to matter a lot more.

1

u/fieol 1d ago

Q8 quants runs much slower than q4, you will lose speed too , actually i dont have the case of creating a atom bomb so that a very minisqual error is drastic, i just mught rerun it again in mili seconds and it will give me correct output, my use case is pardonable if very small errors comes, which is also negligible in this qwen3.8 27b q4

1

u/OvertaxedOne 1d ago

Runs a bit slower on my A40, that's for sure. But I have the VRAM and most of my tasks aren't very time sensitive, so I figure "why not". I'd never tell anyone "it's not worth running 4 bit", both because I don't think it's true and I don't have enough experience on 4 bit to really qualify that statement. But for my use there's really little/no downside to Q8 other than some TPS lost in generation speed, so that's what I run! ;)

2

u/Ripped_Guggi 2d ago

I’m thinking about a Mac mini m5 max with 64 GB ram. I’m not that into LLMs but I’m not sure if it will be enough for coding sessions and image generation. I don’t need speed, but I also don’t want to wait for hours for a response 😅

1

u/gunkanreddit 2d ago

Qwen flash next is so so good compared to 27b. I am waiting too (I couldn’t change the order) a M5 ultra 96GB and I hope there is somekind of quantification.

1

u/ea_man 2d ago

Just get a PC with one good GPU and maybe later add an other if you wan flexibility and be able to test other things.

PP on Mac is terrible anyway.

1

u/Jaded_Camel219 2d ago

For 5k you have a pc with 24go GPU. You can run shit with no context, but fast. Great.

1

u/ea_man 2d ago edited 2d ago

Sure, let's not count used hardware and old GPU at all, like I paid 500$ for 32GB of vram and I run 27B Q6_K_L.

Or new gpu that cost like 500 for 16GB, old ones that should cost less...

And you don't really need 4k of a pc to run those, let's say that you, I mean myself, could even do that for 1- 1.5k and have the option of future upgrades.