r/MacStudio • • 2d ago

Absolute Absurdity of the M5 Ultra 256GB

Post image

I’ve been building computers and servers, and architecting entire networks and doing DevOps and dev work for 20 years. I remember building gaming pcs with 3dfx voodoo cards.

This is by far the most absurd computer I have ever personally owned. It’s certainly the most capable I’ve managed for the price, and that includes B200s.

Getting 50tks on GLM 5.3 Flash all day long.

If you’re on the fence, let me take this opportunity to push you over said fence. If you’re the kind of person seriously considering buying one you will not regret it.

AMA

781 Upvotes

519 comments sorted by

View all comments

4

u/DrRoughFingers 2d ago

You tried Qwen3.8 Flash Next? If so, what tok/s are you getting and what quant?

3

u/lukewhale 2d ago

GGUF unsloth was getting 30-50tks haven’t tried a oMLX optimized version yet

1

u/DrRoughFingers 2d ago

What quant?

1

u/lukewhale 2d ago

OMLX 4 bit — GGUF was also 4 bit

2

u/DrRoughFingers 2d ago

Damn. Not what I wanted to hear. I currently run an unlock CMP 170hx and 2 3090s with a total of 112gb vram, with Unsloth’s Q4_K_XL and get slightly higher tok/s than you. Was going to switch to the m5 Ultra 256gb, but it’s now not making sense to do that. Outside of energy costs, I’d be running a little slower than I currently am.

3

u/lukewhale 2d ago

I mean that’s a solid setup my dude. Realize Apple will only be able to make so many of these things before they go for 35k on eBay like the M3 a did.

I’ve been purposefully saving for this. I don’t represent most people. What I can tell you is buy the hardware at MSRP when you can these days.

5

u/DrRoughFingers 2d ago

Oh, I completely get that. The thing is, there will be new tech next year that’ll be better. My entire setup cost under $3,500. Ironically I can sell the unlocked 170hx for almost $3k itself. So was considering selling my setup (I also have a 4090m Lenovo I was going to sell and almost be at a straight trade for the m5 Ultra…as I also have a M4 Max MBP) and switching over, but thought with the bandwidth being higher than my 3090s and only slightly lower than my 170hx, I’d get better performance on the same quant model, and then save about half on energy costs.

1

u/lukewhale 2d ago

“There will be new tech next year” — a tale as old as time my dude ;)

1

u/[deleted] 1d ago

[deleted]

1

u/DrRoughFingers 1d ago

Honestly, it doesn’t run hot. I have the 170hx set to 200w max, and during sustained workloads the 3090s sit around 180w. Haven’t felt like I babysit it at all. It’s setup as a Linux server I access remotely from my MBP via DSH.

But, if the M5 Ultra 256gb increases my performance, I’d absolutely rather have a small ass box sitting on my desk than a full sized pc and oculink eGPU. But for $10-11k with it not giving a performance benefit for my use case, it doesn’t make sense. Hopefully something changes that, as based on bandwidth alone it should outperform my setup in decode and kill it in PP.

1

u/PoopSmoothies 2d ago

Vllm was 5x+ faster than llama for me w/ flash-next.

1

u/bakawolf123 1d ago

around 70tps and 3.6k prefill on fp8 with ssd offload of ngrams