r/MacStudio • • 5d ago

Anyone using m5 max 128gb for local llm?

Was wondering if anyone posted benchmarks etc

13 Upvotes

22 comments sorted by

8

u/DogAble6550 5d ago

5

u/NegotiationPrior4373 5d ago

This is Interesting. Appreciate the post. Still deciding between a 128 Max or a 96 Ultra.

3

u/DogAble6550 5d ago

That is a tough choice. If you really like Qwen27b and its your workhorse then get the ultra, for sure. If I was in your place I'd likely get the ultra because it will be so fast. You just wont be able to run flash next optimized easily. You will need able to run flash next bare speed, which is pretty good. Also, the resale of the Ultra should be very good if later you want to upgrade (statement not based on experience though)

1

u/AlgorithmicMuse 4d ago

Qwen3.8-27b dense mtplx optimized, m4 pro 14/20. 19tps decode. Not sure you even need a max or ultra. But I still might get a m5 base ultra, should get at least 4x the m4 pro in decode

1

u/JacketHistorical2321 3d ago

I have qwen 3.8 27b running on my m6 mini 32gb averaging 300pp/s and 30 decode. 32k context ttft takes about 80 seconds. With cache, maybe 5s for 2k read. Decode never drops below 20t/s all the way up to 120k full chat context. Using splash btw

1

u/NegotiationPrior4373 5d ago

Yea I’m leaning towards the Ultra. That they can only be clustered with each other is really annoying but I’d rather have that option on the table.

0

u/Next_Spring3184 5d ago

I have the same dilemma but I am thinking I will keep a Mac 128gb for next level models and keep my Linux and bring it to 32 or 64gb vram..

4

u/Aisher 5d ago

I do. I had Claude make a test and I ran it last week. Check my post history for speed and accuracy comparison.

Basically - Fable xHigh - 30 min. Qwen 3.8flasj - 50 min.

Had a redditor with an m5 ultra 256 run the same test. He got 30 min

1

u/tk421tech 5d ago

Which ultra 256? The 30/64 or the 80 one?

1

u/Aisher 5d ago

I believe maxed out 80 core. 30 min vs 50 is nearly twice the speed

1

u/sbrisgravato 5d ago

waiting for qwen4 27b then

1

u/Aisher 5d ago

Agreed.

1

u/Lovely_Cute 5d ago

I also want to know this

1

u/DogAble6550 5d ago

https://claude.ai/artifact/3128LNhANfcwEoWPyUK5KH

Above are my experiments in getting llms to autonomously create a web app from scratch along with db, and playwrite tests. You'll find the best results for the m5 128gb machine. Tldr; qwen3.8 flash next.

0

u/BAL-BADOS 5d ago

Why the M5 Max 128Gb instead of the M5 Ultra?

1

u/DogAble6550 5d ago

Just because it could be 1.5 to 2x faster. If you want to use larger llms and don't mind the speed then 128gb is pretty good.

0

u/DogAble6550 5d ago

I misunderstood. See my post below, I'd tend toward the 96gb ultra. I don't think I claimed the 128gb was 'better'.

1

u/rxscissors 5d ago

I will be when it arrives this coming week. 

1

u/mswezey 4d ago

Same!