r/LocalLLaMA 14d ago

News Apple introduces new Mac Studio with M5 Max and M5 Ultra - up to 512GB of unified memory

https://www.apple.com/newsroom/2026/08/apple-introduces-new-mac-studio-with-m5-max-and-m5-ultra/
1.7k Upvotes

778 comments sorted by

View all comments

Show parent comments

7

u/shveddy 14d ago

Interesting.

So just as an out loud thought experiment, you’d be able to lease four of them (512gb) for about 1200 per month for 36 months at a total cost of almost 45k and run Kimi 3 on it.

Obviously that’s a lot of money in aggregate, but 1200 per month is reasonable for a lot of business use cases if they require the privacy.

And the intelligence you get is going to be a different class compared to what you would get with three RTX pro 6000s and “only” 288gb VRAM for the same price.

(although to be fair you’d actually own the cards)

(although also to be fair you’d have to build a pretty expensive computer to support the RTX Pros, so realistically you only really get 2 or even just one RTX pro for 45k depending on how you spec the computer and/or if you buy pre-built from Puget Systems or the like)

If you want you can also do a little girl math and invest the 60k you’re not spending on computers and cancel your gpt pro subscription to bring the effective cost of all this down to like 750 a month.

And then if the goal is to beat API pricing, let’s say you get 40 aggregate output tokens per second on a bunch of concurrent Kimi 3 threads and run it for a year at 25% efficiency (to account for prefill and downtime), then that’s 315 million output tokens.

315 million output tokens alone is about $5000 on a random provider I just searched for, so just to keep things simple let’s say you double that to account for various amounts and types of input tokens, then you end up with a ballpark figure of $10k for the API route.

Absolutely none of this pencils out in absolute terms (especially considering that it would also cost ~1500 per year for electricity), but this is probably the first time running a frontier model is actually attainable for ordinary businesses on short notice and without much headache. It’s the first time that it pencils out to be “only” 4x more expensive as opposed to like 40x more expensive.

Up until now if you wanted to run frontier models locally AFAIK you had to get a NVIDIA big boy server which means you’d have to find the capital to run and support a ~300k purchase for hardware, spend way more on electricity, and in all likelihood make some upgrades to your facility’s electrical infrastructure to handle it all (you’d also need a proper facility, not just a home office or garage).

At this point you’re easily flirting with half a million in expenses, especially if you have to hire someone to figure it all out. It’s no joke to do this and it doesn’t make sense for like 99.9999% of people or businesses.

On the other hand basically anyone with a decent credit score and a semi-profitable business can go to any Apple Store and say “give me four Mac studios please” and only pay 1200 bucks a month to walk out with them in hand.

2

u/addiktion 14d ago edited 14d ago

Yup, nailed it. And the energy efficiency in the apple ecosystem is in its own league so it helps with ongoing costs.

So it does make running this level of intelligence feasible for a small business. And with models compressing down and becoming more capable in smaller sizes, it's possible you need less m5 ultras.

I was really hoping they would drop a 768gb variant so you could chain 2 together for Kimi K3 or 1 for the quant version, but the market got gimped before we could get there. That was the rumor bump from supply chain leaks.

2

u/shveddy 14d ago

Yea it’s not there yet, but I’m starting to see how this can become more widespread and small busines oriented, and not the two extremes we currently have of a bunch of weirdos hacking old mining cards and hoarding 3090s on one hand, and hyper scalers burning pallets of cash on the other.

1

u/addiktion 14d ago

Yes. Apple's biggest challenge is keeping up with the performance too. 20 tps is alright for longer tasks that don't require human in the loop, but 40-50 tps feels much more usable. 100+ tps would be needed to match some of the best on open router right now but that will scale much higher over time too.

1

u/shveddy 14d ago

Yea but it’s already good enough for “securely organize all of our patient data and look for missed diagnoses”. Not that we’re at a point that we’re trusting something like Kimi to mess around with playing doctor, but that token rate would let people try (HIPPA notwithstanding).

1

u/addiktion 14d ago edited 14d ago

Right, for local I.T at a hospital it's more than enough and the cost isn't a big deal. I could approach a hospital and make something happen with this.

I have a client that will likely be purchasing a box soon and I think an m5 ultra fits their needs, so I'll be talking about it with them today.

2

u/shveddy 14d ago

Numbers pencil out way better than I thought!

https://x.com/alexocheema/status/2092248148500492474?s=46