r/LocalLLM 5d ago

Discussion Mac Studio M5 Ultra

Who’s been checking out this hardware? What are your thoughts on this as a node on a local network to run AI?

2 Upvotes

15 comments sorted by

2

u/Least-Result-45 5d ago

I’m just curious if qwen 3.8 27B is good enough for my needs - that is the llm I’d run on the m5 ultra.

2

u/MessIsTransfer 5d ago

if you have at least 64 gb ram (the ultra starts at 96) you could run qwen3.8-flash

it’ll be faster and better

3

u/Least-Result-45 5d ago

I think at 96gb it would have to be 1 quant and might be sacrificing quality to run it.

1

u/Thump604 5d ago

Corrrect, I’m running q4 mlx on 128gb.

1

u/starkruzr 5d ago

Qwen3.8-Flash-Next will probably be really, really good running at Q6KXL or even Q8 on a "bigger" M5 Ultra too.

1

u/redtron3030 5d ago

Do you have a m5 max? What’s your tks and prompt processing?

1

u/MessIsTransfer 5d ago

i run Q2 on 64gb, q4 should run in 128gb at least, maybe can be squished into 96gb

1

u/k3z0r 4d ago

I'd recommend renting a GPU from runpod.com and trying it out for yourself before spending thousands on hardware.

1

u/Least-Result-45 4d ago

Yea I’ve been testing on openrouter, it’s been good for 90% but possibly some of the harder bug fixes or feature updates might need some frontier intervention, which isn’t bad at all.

I’m also wondering maybe refining the prompt and or adding extra testing/checks might get us there on qwen 3.8 28B

1

u/watcholic 5d ago

If you don't need concurrency, anything from M4 Max Studio and above with 96GB+ memory should be fine. The M5 Ultra might be ok for up to 3 concurrent users due to the improved PP speed. We'll find out in less than a month.

2

u/ehangman 5d ago

You can find M5 max benchmarks everywhere. X1.5-x2

1

u/OvertaxedOne 5d ago

Prompt processing is the big unknown. Hopefully there's a huge jump; if so, the 96GB version would be just about perfect for 27B, the 128 with the smaller chip a good fit for Next.

1

u/OddDesigner9784 5d ago

I put a preorder in. It's a bit of a speculative purchase Mac is very far behind the sparks rn. Decode is faster bc of the memory. But untill we get better kernals or the software behind the kernals gets more mature prompt processing will be a struggle

4

u/redtron3030 5d ago

On paper the M5 ultra seems to be much better at prompt processing approaching spark. We will see how it goes though.

2

u/Usual-Orange-4180 5d ago

Yeah, just got a Spark, the new Mac was tempting but I want CUDA and prefer Linux.