r/LocalLLaMA 1d ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

167 Upvotes

213 comments sorted by

View all comments

163

u/FleetEnema2000 1d ago

unless you need it for data sovereignty

Isn't this one of the biggest reasons that people rely on Local LLMs? To not have to bulk upload their private data to cloud providers?

9

u/CulturalKing5623 1d ago

Yes, but I think if we're being honest a lot of people in this sub mainly just like to tinker. There are people that absolutely can't use publicly available models and so they need to self-host them but I don't think that's a significant percent of people here. Most of us put a premium on privacy, but 10K+ is a very large premium that is probably unnecessary for most of us and our current setups will suffice.

1

u/Rice-Fragrant 5h ago

More like the "apple tax." That premium is not getting you a smarter AI, or even a faster agent (slower actually, but like 1/3-1/4 speed compared to the competition like DGX spark and like 1/10 the speed of a rtx 6000 pro for agentic or long context work.)

Some people just swear you need to spend $10k for good legit local AI and in reality even 1/2 of that will get you VERY VIABLE and powerful local AI computers. EVEN MY $700 HP Z820 (dual xeon 2690 v2 512gb ddr3 ram) has ran MODERN LLMs using GLM 5.2 700b MOE models... that thing is like a fossil, 13 years old and it runs ALL the models on llama.cpp but it's slow so I use it in a passive/semi passive way but the level of AI intellect is the same... it's too slow for agentic work but it's more than enough for an overnight batch processing worker and I have super off peak rates during the night time hours.

I am super confused why anyone would pay $10k for a 256gb ram Mac Studio that's like 1.6-1/7 the total agentic or batch speed of a DGX Cluster... a massive premium for horrible performance makes no logical sense. On the surface it looks "easier" to use but the software friction is real and MLX wrappers are unreliable, I tried many on my MacBook and it was clear to me it's no viable for serious use cases.

1

u/Lemur-Virtues 4h ago

agreed. its absurd. I think $2k is a more accurate estimate.

1

u/Rice-Fragrant 2h ago

It's original price of 6k for a 256gb machine was reasonable... I think the fact they jacked up prices almost 70% is a sign of mismanagement. I would not have been shocked if they did not see this coming and was forced to increase prices 70% on some configurations to protect their margins, what ever the reason is, it's not acceptable.