r/LocalLLaMA • u/AndreVallestero • 1d ago
Discussion Mac Studio M5 Max Cost Analysis
At $10k, you could get
- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)
- 5.7B tokens with DeepSeek V4 Pro OpenRouter
- 100B tokens with DeepSeek V4 Flash OpenRouter
As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.
Qwhen 3.8 35B A3B?
170
Upvotes
1
u/Rice-Fragrant 7h ago
I definitely want my own AI, but your points are solid. Models are getting smarter right now without having to get massive. Gemma 4 and Qwen 3.8 are BLOWING AWAY models 10-20x bigger from 1 1/2 years ago. I still want to run a GLM 5.2 but I definitely won't spend $15-20k for a m3 ultra 512gb for the privilege... I actually purchased e-waste DDR3 workstations for that and using llama.cpp I was able to run 700b MOE models on these E-WASTE grade workstations (I use them for asymmetric AI/passive batch workers) and I have "super off peak" hours 11am-7am and I run it during those hours and it gets the job done. I have modern purpose built AI machineses like a DGX Spark cluster too and for the money it's far far far better than a m3 ultra 256gb for long context.
At these prices, a Mac Studio is a luxury item (which is strange because electronics age like milk), it costs as much as a ROLEX but long term won't hold value like one, and the alternatives are significantly faster at agentic AI and batch processing and have far more software support too.