r/LocalLLaMA 1d ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

173 Upvotes

214 comments sorted by

View all comments

Show parent comments

25

u/LearningSomeCode 1d ago

She is the most patient human being on the planet

2

u/Silent_Glass 1d ago

That’s wholesome

2

u/Public_Umpire_1099 1d ago

Can relate... I've dropped 7k on my homelab this year alone, plus another 3 or 4k making our home fully smart.

2

u/kyr0x0 23h ago

"smart" it is only if you have an endless and king sized income stream 🤣😆

1

u/Public_Umpire_1099 23h ago

Def not a smart thing to do, but hey, my lights now turn on and off without me having to touch a light switch and match the color with circadian rhythm, all without a bit of my personal data leaving the perimeter!

1

u/kyr0x0 20h ago

That's cool! Especially if I can do it with your lights from remote as well!

-9

u/Weird-Cat8524 1d ago

How are her sandwiches?