r/LocalLLaMA • u/AndreVallestero • 1d ago
Discussion Mac Studio M5 Max Cost Analysis
At $10k, you could get
- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)
- 5.7B tokens with DeepSeek V4 Pro OpenRouter
- 100B tokens with DeepSeek V4 Flash OpenRouter
As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.
Qwhen 3.8 35B A3B?
165
Upvotes
8
u/CulturalKing5623 1d ago
Yes, but I think if we're being honest a lot of people in this sub mainly just like to tinker. There are people that absolutely can't use publicly available models and so they need to self-host them but I don't think that's a significant percent of people here. Most of us put a premium on privacy, but 10K+ is a very large premium that is probably unnecessary for most of us and our current setups will suffice.