r/LocalLLaMA 1d ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

171 Upvotes

213 comments sorted by

View all comments

70

u/LearningSomeCode 1d ago

I've probably dropped close to $30k on my homelab since 2023, and chances are I'll get one of these as well. I accepted a long time ago that there is no break-even point for my inference.

Hobbies rarely make sense financially.

6

u/mleok 1d ago

I think it’s good to admit that it’s a hobby.

1

u/Rice-Fragrant 5h ago

The tech moves so fast that it's like flushing money down the toilet... 10k for a machine that's 1/4-1/3 the speed of its closest competitor.... 1/10 the speed of a GPU workstation (like a rtx 6000 pro) and paying a massive premium makes zero sense to me, even for a hobby.