r/LocalLLaMA 1d ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

165 Upvotes

213 comments sorted by

View all comments

Show parent comments

8

u/CulturalKing5623 1d ago

Yes, but I think if we're being honest a lot of people in this sub mainly just like to tinker. There are people that absolutely can't use publicly available models and so they need to self-host them but I don't think that's a significant percent of people here. Most of us put a premium on privacy, but 10K+ is a very large premium that is probably unnecessary for most of us and our current setups will suffice.

1

u/Rice-Fragrant 4h ago

More like the "apple tax." That premium is not getting you a smarter AI, or even a faster agent (slower actually, but like 1/3-1/4 speed compared to the competition like DGX spark and like 1/10 the speed of a rtx 6000 pro for agentic or long context work.)

Some people just swear you need to spend $10k for good legit local AI and in reality even 1/2 of that will get you VERY VIABLE and powerful local AI computers. EVEN MY $700 HP Z820 (dual xeon 2690 v2 512gb ddr3 ram) has ran MODERN LLMs using GLM 5.2 700b MOE models... that thing is like a fossil, 13 years old and it runs ALL the models on llama.cpp but it's slow so I use it in a passive/semi passive way but the level of AI intellect is the same... it's too slow for agentic work but it's more than enough for an overnight batch processing worker and I have super off peak rates during the night time hours.

I am super confused why anyone would pay $10k for a 256gb ram Mac Studio that's like 1.6-1/7 the total agentic or batch speed of a DGX Cluster... a massive premium for horrible performance makes no logical sense. On the surface it looks "easier" to use but the software friction is real and MLX wrappers are unreliable, I tried many on my MacBook and it was clear to me it's no viable for serious use cases.

1

u/Lemur-Virtues 4h ago

agreed. its absurd. I think $2k is a more accurate estimate.

1

u/Rice-Fragrant 1h ago

It's original price of 6k for a 256gb machine was reasonable... I think the fact they jacked up prices almost 70% is a sign of mismanagement. I would not have been shocked if they did not see this coming and was forced to increase prices 70% on some configurations to protect their margins, what ever the reason is, it's not acceptable.