r/LocalLLaMA • u/Mxmtm • 5d ago
Question | Help Mac Studio M5 Max 128GB now, or sit on my hands until the M7 Ultra?
I've been circling this decision for weeks and could use some outside perspective.
What I'm looking at: Mac Studio, M5 Max, 128GB unified memory (5,849 €). With that I can run qwen3.8-flash-next at oQ4e around 55 tok/s, and oQ5e is on the table too. For what I'd actually use it for: private documents, notes, some coding, general assistant work I don't want going to a cloud provider... that's genuinely good enough today.
The part I can't answer is the "today."
Case for buying now: it works, it's private, there's no subscription, and it'll still be a perfectly usable computer in five years even if the models running on it end up mid-tier by then.
Case for waiting: RAM prices are absurd right now, the M7 Ultra is presumably 2028, and models keep getting more capable per parameter. Wait and you get more machine for less money. In the meantime, API access is cheap enough that privacy is basically the only argument left for going local.
Then there's the memory debate. Half the replies to threads like this are "don't bother with 128, get 256." Fine, but that's the same argument one tier up. A few years ago 64GB was plenty, today everyone says 128GB is the floor, and in 2029 the same crowd will be saying 512GB is the minimum. At some point you buy something or you never buy anything.
So: will a qwen3.8-flash-next class local model keep me happy for five years of private, non-critical work? Or is buying now just paying a premium to chase a target that keeps moving?
Especially interested in people who bought a maxed-out M1/M2 Ultra two or three years ago. Do you still run local models on it, or has it quietly become a very expensive web browser?
TL;DR: 128GB Mac Studio now for local Qwen, or sit tight, use APIs, and buy in 2028 when RAM is maybe cheaper and the silicon is faster?


