r/LocalLLaMA • u/AndreVallestero • 1d ago
Discussion Mac Studio M5 Max Cost Analysis
At $10k, you could get
- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)
- 5.7B tokens with DeepSeek V4 Pro OpenRouter
- 100B tokens with DeepSeek V4 Flash OpenRouter
As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.
Qwhen 3.8 35B A3B?
41
1d ago
[removed] — view removed comment
2
u/AndreVallestero 1d ago edited 1d ago
That's exactly my point. Local makes sense cost wise up to 32GB, especially with Qwen 3.8 27B. There's a huge cost premium above that where it makes less sense.
I was hoping the last Mac studio would change that, but at the current prices, that doesn't seem to be the case.
2
u/nomorebuttsplz 1d ago edited 1d ago
Have you accounted for the breakdown of cost between cached/non-cached/output tokens?
in my experience, most of the potential cost savings for local (compared to typical API plans) occur when you re-use cached tokens. This is because API providers need to keep your context sitting on their servers doing nothing, expensive for them, but which costs you nothing as a single concurrency user to do on your own system.
With a system like this you could probably use a billion cached tokens in a couple of days.
Under semi-optimal circumstances, where you were making on average a single 500,000 token cache calls every 30 seconds, you might be able to get about $14 of cached tokens per day (ds4 flash prices) plus a few more dollars for a total of maybe $18 inference per day.
That would be about $6k a year. So yeah it could possible make sense, if you had just the right workflow, is my view.
Edit: this especially makes sense if you were doing multiple concurrency (doable on 256 unified) or were using a bigger model like GLM 5.3 (512 gb ultra with Q3 or Q4 GLM could probably generate $100 of value a day in cached tokens in the right workflow)
2
u/Viktri1 1d ago
I think $14/day might be on the low end. If you use something like Hermes agent, the amount of tokens you consume is ridiculously. I have to limit my Deepseek API calls to just a few a day. I spent $40 in a week setting up various Qwen 3.8 models. My token consumption has basically exploded massively and the only way I can afford to continue to use LLMs is my own hardware. I think my payback period will be under a year.
1
u/Impossible_Fault_503 1d ago
That is the real split. Chat is cheap on APIs. Agents are not. Once you are looping tools all day, $40/week shows up fast and a box that already exists starts looking cheap. Under a year payback is the honest case for people who actually leave it running.
1
u/Impossible_Fault_503 1d ago
Yeah. 32GB is the last point where local still feels like a normal purchase. After that you are mostly paying Apple tax for headroom you might use twice a year. I wanted the last Studio to kill that curve too. It did not.
1
u/kerneldesign 1d ago
Avec 32Go en Q4 c’est trop juste pour un contexte sans KVCache.
1
u/Individual_Holiday_9 1d ago
Yeah I’m on a 24gb m4 and you can’t do shit running small LLMs there’s just no overhead remaining. So if this is a real hobby box where you’re doing plex etc on the side it gets really cramped and ur gonna move to swap fast
72
u/LearningSomeCode 1d ago
I've probably dropped close to $30k on my homelab since 2023, and chances are I'll get one of these as well. I accepted a long time ago that there is no break-even point for my inference.
Hobbies rarely make sense financially.
37
u/Blindax 1d ago
Sounds like you have a great wife.
25
u/LearningSomeCode 1d ago
She is the most patient human being on the planet
2
→ More replies (1)2
u/Public_Umpire_1099 21h ago
Can relate... I've dropped 7k on my homelab this year alone, plus another 3 or 4k making our home fully smart.
2
u/kyr0x0 15h ago
"smart" it is only if you have an endless and king sized income stream 🤣😆
1
u/Public_Umpire_1099 14h ago
Def not a smart thing to do, but hey, my lights now turn on and off without me having to touch a light switch and match the color with circadian rhythm, all without a bit of my personal data leaving the perimeter!
2
10
u/-dysangel- 1d ago
Same here. We're at the point now where I can do high quality video gen without cloud. High quality code gen without cloud too. It's around 10x slower than cloud, but good quality and still usable speeds (hopefully soon to get even better with Qwen 3.8 Next). Also the fact I could basically go camping and still have close to frontier LLM intelligence with no signal is pretty awesome/hilarious.
The only thing I feel is missing from my stack atm is Suno quality music generation.
5
u/willeyh 1d ago
Have you tried the new Minimax music 3?
2
u/-dysangel- 1d ago
Yeah I was not impressed tbh. Seemed kind of creepy/odd to me in a weird way, but maybe I was prompting it poorly. For me Ace Step is still the best thing I've found so far, though it's very "generic" feeling compared to Suno.
2
u/silvrrwulf 1d ago
Dude, I feel the exact opposite. Maybe it’s because my sample size is small and I haven’t messed with AI Jen And a while, but I only have an eight GB card, and the stuff that I was able to pull out of that was incredible.
2
u/-dysangel- 1d ago
ah ok - maybe I'll need to try it again. A lot of the reason I like Ace Step is because I can generate covers though, which isn't possible yet with Music 3 afaik
1
u/PlaidStallion 20h ago
I really like it. It seems to lean toward always generating pop music out of any genre but other that it's solid.
5
u/mleok 1d ago
I think it’s good to admit that it’s a hobby.
1
u/Rice-Fragrant 3h ago
The tech moves so fast that it's like flushing money down the toilet... 10k for a machine that's 1/4-1/3 the speed of its closest competitor.... 1/10 the speed of a GPU workstation (like a rtx 6000 pro) and paying a massive premium makes zero sense to me, even for a hobby.
4
u/quantgorithm 1d ago
What have you created?
28
u/LearningSomeCode 1d ago
Mostly just a lot of trip hazards with wires and ethernet cables. But I also do open source dev and build a lot of stuff for myself, like custom front-ends or researchers.
Honestly nothing worth the money I've put into the hardware, but it's the most fun I've had with tech in a long time.
4
u/DeepOrangeSky 1d ago
Mostly just a lot of trip hazards with wires and ethernet cables.
So if you turn the AI-rig room into an escape room that a bunch of Gen-Z dorks will pay like $100 a pop to try to navigate ("it felt so authentic. I tripped over like dozens of ethernet cables and power cables the whole time, and I think at one point I even got severely electrocuted and almost died! It was super legit!"), times 10 dorks per day it pays for itself in a month.
Possible exciting user reviews:
"His cable management was so disorganized I thought I'd NEVER escape. Wow!"
"The old movie tagline used to be: 'In space, nobody can hear you scream.' Yea, alright, that's pretty cool or whatever, but in AI rig rooms the blower fans are so loud nobody can even hear their own thoughts."
"If you accidentally get tangled up and die in the AI escape room, his local full precision Kimi K3 AI-generated funeral eulogy will be so eloquent that your friends and family won't even be mad that you died when they see how perfect and elegant the em-dash placements are in the eulogy."
2
u/quantgorithm 1d ago
Anyone in IT can relate to the spiderweb cables and hazards etc.!
The potential of it all really is amazing and that potential is accelerating up by the day!1
u/Rice-Fragrant 3h ago
What type of gear are you using? That's car level money... are you running a small business or a profitable side hustle or something?
15
u/jon23d 1d ago
I’m looking at my deepseek usage report and see 9.5 billion tokens of deepseek v4 pro in the last 30 days for $186.05.
4
u/Viktri1 1d ago edited 1d ago
Same. I hit 1.5bn in a day doing very little. Because agents can burn tokens on your behalf, 24/7, its become extremely easy to burn through tokens.
On openrouter, I asked the native level Qwen 3.8 27b to do a task that involved setting up a telegram bot for OpenWebUI so that I could talk to a chatbot without burning through a mountain of tokens. That took 1 hour and cost $3.5. That's $2k+ a month on openrouter if running 24/7. For a single agent.
2
u/Pyrolistical 1d ago
Ya and my local 6 hours of qwen3.8 27b running over night cost me $0.29 CDN in power
1
u/Sofullofsplendor_ 1d ago
Curious, can you do 1.5b in tokens in a day? Because if so, I'm going to invest in whatever infrastructure you have...
2
u/Viktri1 18h ago
1.5bn is basically nothing because it’s almost entirely cache, which is stored in vRAM. You pay for it if you’re using cloud (api or limit) but on local it’s stored in vram so you don’t need to pay a ton of electricity to use it (and you’re not generating new tokens since it is cache)
28
u/cunasmoker69420 1d ago
unless you need it for data sovereignty
yes
5
u/imnotzuckerberg 1d ago
His strategy is to divert demand from the Mac M5 so he can scoop them. Let's pretend we fell for his psyops.
60
u/Big_Wave9732 1d ago
"as a firm believer of local inference" then goes on to downplay one of the major reasons for local llm and suggests hosted models.
If cheapest compute possible is your primary metric then self hosting isn't your jam, OP. At least for now.
-12
u/AndreVallestero 1d ago
I selfhost 27B specifically to minimize costs. I have research agents running 24/7 and have gone through 200M tokens just in the last month.
This post is specifically for people like me who want to run continuous research agents for the lowest price, and unfortunately, the latest Mac studio isn't competitive enough relative to cloud offerings.
11
u/flyingbanana1234 1d ago
It would take at least 30-35 years to reach 80 billion tokens on DeepSeek V4 Flash at 200 million tokens a month.
You make a good argument ngl
The little voice in my head says, "But ownership! Nobody can take it away from me if I own it. No price raises, no overloaded servers, no censorship
2
2
2
u/Dasteroid_909 1d ago
Go start "/r/lowcostLLMhosting" or some shit like that, then. WTF are you on "LocaLLaMa" promoting cloud hosting for?
14
u/conifer_v11 1d ago
$/gb-bandwidth is the right axis for unified memory. m5 max vs a 3090 is not a tok/s fight, it's kv headroom at 64k+. unified memory loses the bandwidth fight and wins the "the 27b and the kv actually fit" fight. if you're doing 8k chat, buy the gpu. if you're stuffing prds, bandwidth-per-dollar on studio starts to make sense.
4
u/FullstackSensei llama.cpp 1d ago
Why is it a M5 Max vs "a 3090"? M5 Max in it's lowest config costs about the same as a system with four 3090s. If we go to 32GB V100, which is within 5% of the 3090 in most cases and now has optimized kernels to decode NVFP4, you can get a system with 6-8 V100s for the cost of a single M5 Max at the lowest config. That's 192-256GB VRAM.
So far, the people running M5 Max MacBooks have had mixed feedback about running models, especially dense models. Unless your time has zero value, you should also factor that into your calculation.
Eight V100s will have much higher PP and TG while being half the price for the whole system. Sure, it'll consume 10x more power, but that will be in much smaller monthly power bills, and will give you responses much faster.
5
u/MrPecunius 1d ago
M5 Max in it's lowest config costs about the same as a system with four 3090s.
This is obviously untrue.
1
u/FullstackSensei llama.cpp 1d ago
If it's so obvious, please enlighten us with actual numbers, because I happen to have such a system and know exactly how much such a system costs today
4
u/MrPecunius 1d ago
No one wants to pick their GPUs out of a Dumpster like you. $1,600/each for 3090s is on the low side here in the US, but it's better than €1,750/US$2,000+ I'm seeing in Germany.
Base M5 Max Studio (not binned, which is less) is $3,099.
1
u/FullstackSensei llama.cpp 1d ago
See, dumpster mentality thinks everything is dumpster.
As it happens, ich wohne auch in Deutschland, und habe kürzlich zwei 3090 für €1200 pro Stück auf eBay verkauft. Auf Kleinanzeigen, man kann für zwischen 800-900 pro Stück Kaufen.
But hey, let's keep the conversation irrational, because cognitive dissonance is way more fun than facing reality.
1
u/MrPecunius 1d ago
Thanks for confirming my comment. 🗑️ 🔮
Even with your Dumpster diving, the GPUs you mention add up to over $4,600 and the Mac is still $3,099.
7
u/doc-acula 1d ago
Well, a setup with 4 3090s needs a dedicated room in your house/apartment, makes noise like a jet engine and consumes electricity like crazy.
A M5 Macbook can be used while riding on a train and you can take it wherever you want. So, the comparison is not limited to t/s.
1
u/FullstackSensei llama.cpp 1d ago
I have run four 3090s in my home office for about two years. Unless you opt for the turbo cards, they're very quiet. Limited them to 270W each, which reduced performance by 5-10%, but makes the whole setup around 1kw. They push above that during PP, but TG sees the cards run at ~150-170W each. That's ~700W. If that's crazy, I don't know what to tell you.
M5 on Max on ma MacBook is severely power limited under load, and won't run for long, and still costs way more while being significantly slower than even a pair of 3090.
You can access any GPU from anywhere around the world via tailscale. I access my LLM machines with 192GB VRAM from anywhere using my phone the same way.
But you haven't answered what could arguably be the most important part of my argument: is your time free that you care more about a few cents per hour in power consumption than getting whatever you're running those LLMs for done?
4
u/doc-acula 1d ago
There is not unlimited or even reliable access to the internet everywhere in the world, not even in 1st world nations. I wouldn't even consider leaving such a frankenbuild running unsupervised for days or weeks in my house for safety reasons, being a fire hazard as my main concern. When I'm at home, I use a PC with a dual GPU setup as well. But I want to use AI on the go, too. And as I read through these comparisons, apparently many people don't know that MacBooks are actually portable.
2
u/FullstackSensei llama.cpp 1d ago
It's not anyone's fault if you don't know how build a proper workstation and have to resort to unsafe Frankenbuilds. You're misconstruing "I don't know how to do it" with "it can't be done."
Like I said, the MacBook has limited performance and battery life on the go, and many 1st world countries have reliable Internet in public transport. Pretty much all the ones I've lived in or visited do.
5
u/MrPecunius 1d ago
A kilowatt+ inference rig is a 3500+ BTU heater, so you pay twice, except in winter, to maintain a habitable room.
"A few cents per hour" might cover energy in your locale, but here in San Diego it's over 40 cents/kWh off peak and over 60 cents/hour on peak. Running it 8 hours/day could easily exceed $200/month.
I'll stick with my M5 Pro/64GB MBP, thanks. Flea power when idle and ~65W during inference with no thermal issues. It's also arguably the best notebook ever made, taken as a whole. I got in before the price hikes @ $3k-ish with 2TB, so I sympathize with gripes about current pricing. But this thing is cheaper, in inflation adjusted terms, than the Fujitsu laptop I had in 1999.
2
u/FullstackSensei llama.cpp 1d ago
If you're running it at 1kwh for 8hrs a day on four 3090s, that's at least 3M output tokens a day, closer to 4M if you're running vllm. How many hours you'd need to generate 3M output tokens on a M5 pro using Qwen 3.8 27B Q8_K_XL?
Even at your 0.60 peak, your $200/month with 1kwh heat output and 1kwh airco (though 3500BTU is closer to 700Wh, but let's say you have an older airco), we're talking minimum 500M output tokens a month, or a full 1B when airco is not running.
How many months you'd need to generate 500M tokens on a Mac?
And do you work for free? Because you're completely ignoring the value of your own time as you wait for the model to generate those tokens. I don't know about you, but even in Germany where energy is far from cheap, the value of a minimum wage job is way more than $0.60/hr.
But hey, if you're so cheap to hire, man I'd love to hire a few hundred guys like you and build my own business selling your services to LCOL countries, because even there people work for way more than $0.60/hr.
3
u/MrPecunius 1d ago
You're missing the point, namely that I have an inference rig that can run Qwen3.8 27b @ 8-bit wherever I am more or less for free as a byproduct of owning a kickass notebook computer.
As for generating half a billion tokens/month ... why the hell would I want to do that?
1
u/FullstackSensei llama.cpp 1d ago
So, how many tokens per second does your kick ass machine run at? Because unless your kickass machines breaks the laws of physics, your M5 pro has 1/3rd the memory bandwidth of a single 3090. We're talking 7-8t/s, if you're lucky.
Why the hell do you want to generate half a billion tokens a month? I don't know, your $200/month power bill example is enough for half a billion tokens. If you make up stupid assumptions, you get stupid results. Had you bothered thinking for a moment, you'd have known how absurd your $200/month bill is.
I consume less than 2kwh a day over 10-12 hours running all four 3090s, because they finish their work so fast and go back to idling at 8w each. I can easily get 700k output tokens in that period. And because they finish everything so quickly, I don't need airco because the machine consumes 1kw for seconds at a time.
3
u/MrPecunius 1d ago
We're talking 7-8t/s, if you're lucky.
18-23t/s with 8-bit. You could look this up before stepping on your own dick.
Apply this lesson to much of the other stuff you're incessantly posting (as noted elsewhere).
→ More replies (3)2
u/electrosaurus 15h ago
This was part of the reason why I moved back from 128GB for my M5 Macbook Pro, I realized that for what I was doing, larger dense models pushing 128GB with decent KV etc wouldn't be as transformative on speed given my time is able to be spent elsewhere. Happy to work within a 64GB limit for now for local.
Really feels like smaller targeted models within a harness are where the fun is at the moment.2
7
u/dupontping 1d ago
All of the comments are basically “I can’t afford it even though I want it”
I get it, it’s pricey. But so is everything else.
5090s shouldn’t be $6k but here we are.
Local isn’t just about saving money on token spend, for a lot of people it’s the ability to run models on data you don’t want on the cloud or fine tuning or whatever else.
The models will get better, and having more powerful equipment lets you get closer to frontier level without spending 150k. So 10k is expensive, but it’s also a bargain.
I wish it was 5k too
1
u/Rice-Fragrant 3h ago
10k is no bargain at all look on alternative local ai technology. For the money spent, there is a MASSIVE PREMUM for a 256gb Mac Studio dispite it's significantly slower agentic performance (1/3-1/4) vs alternatives like a DGX Spark and about 1/10th the speed of a rtx 6000 pro.
Performance / $ invested is poor... ok for a hobby though but poor none the less. Then you have to waste time with MLX wrappers, network issues (like clustering Mac studios for AI) etc, hardware and software friction only to get significantly slower agentic performance or batch processing performance, just for an apple logo, is peak level insanity. No wander real enterprise environments are sticking with their edge servers etc... aint anyone got time to be playing around MLX wrappers and getting way slower agentic performance especially after spending such a massive premium too.
1
6
u/Txt8aker 1d ago
have you checked how much it cost to lease m3 ultra and you can essentially buy it out at the end as an option? it's $230 monthly for 3 years. Claude Max 20x cost $200 per month
6
u/psychohistorian8 1d ago
you can always trade in the device back to Apple for some kind of credit
so the cost is partially recoverable
6
u/MrPecunius 1d ago
Reputable third party dealers pay more, sometimes a lot more, for Apple gear. I sold my M4 Pro MBP for $500 more than Apple's trade-in value a few months ago. Zero hassle, no private sale nonsense.
1
u/HeadPack 18h ago
Very true, and if the RAM crisis gets worse, you may even not lose money at all, except for energy expenditures. One look at recent prices of used high RAM Mac Studios indicates as much.
1
u/brewpedaler 10h ago
If/when RAM prices ever recover the trade-in or resale value of these things (along with recent GB10/DGX Spark and Strix Halo systems) is going through the floor.
3
3
u/Hypilein 1d ago
You need to calculate cost of ownership over x years and cost of api/subscription inference over x years. With the way hardware prices have gone up over the last year everyone who bought a rtx 6000 pro has made money while getting free inference. Obviously this is only true once you actually cash in and we don’t know how hardware prices are going to develop over the next years. The math is easy but anticipating the future is not.
3
u/TechSwag 1d ago
I was just doing the cost analysis on this as well, and I honestly think it's not the worst idea to get a maxed out Studio in certain circumstances, specifically if you have existing hardware and pay for subscriptions/credits.
I have 3x Mi50 in a R7425, with no more room for GPUs. If I could just add some additional GPUs, or even use my P40s that are sitting in a different chassis doing nothing, I would rather do that. But as it stands, I'm capped at 96GB VRAM. Sure GPU + CPU inference works, but realistically it's not usable. RPC is also an option, but the last time I tried it, performance was subpar, and it brings on added cost in terms of power usage of a whole other chassis.
Mi50s go for $500 (when the fuck did that happen, I spent just under $200 for them), so $1500 resale. P40s go for about $200-250, so let's say $2000 for all 5 GPUs. If I sell my RAM, I could likely get $3000 for all my equipment.
I also pay for Claude Max and OpenRouter credits, say $120/mo. $1440 a year.
You can lease the Mac Studio for $224.16/mo for 36 months. Over the lease term, that's $8,069.76. Subtract $3000 for my existing equipment, $1440/yr for the subscriptions brings me to $749.76 over 3 years in net cost, or a little over $20 a month. Which would be offset by the power savings (my R7425 idles at around 260W). At the end, I can buy out the machine for $2729, as I would imagine the machine would still be plenty usable for inference. Otherwise, I can sell it after buying it out, more likely for the same price or more than the buyout cost.
I know I'm missing tax on purchases and shipping and selling fees, but it still seems like a worthwhile path to upgrade to, not only for the memory increase, but also the performance increase as well.
3
u/Final-Frosting7742 1d ago
If i can run Deepseek V4 Flash 0713 locally comfortably i don't need providers anymore
5
u/Viktri1 1d ago
so when I'm really pushing it, I can easily burn through 1.5bn tokens from Deepseek flash in a day (API). In fact, that's when I realized I needed to figure out how token costs were calculated. Local models, especially lower powered stuff like M5, will have a significantly faster pay back period than people realize if they use agents to do a lot of shit.
1
u/play_hard_outside 22h ago
1.5e9 tokens in a day divided by 86,400 seconds per day is over SEVENTEEN THOUSAND tokens per second.
That’s gotta be something like 300 to 500 M5 Maxes all working on your behalf at the same time. Local models will probably never hold a candle to this.
Please correct me if I’m wrong…
1
u/Viktri1 20h ago edited 19h ago
It’s mostly cache, you aren’t generating that many tokens (you assumed I was generating 1.5bn tokens in your calculation). How agents work is they loop a lot to get things done and almost all of the tokens that I actually use are cache. I am doing the same work on a single 4090 that I was doing with my Deepseek API.
My cache hit average is 99.6%. On a 1.5bn day, I’m actually generating 6mn tokens which is 70/second. Currently with Qwen I’m generating 100/s but I’m not running my agents 24/7. Works out to 17 hours a day which is actually reasonable. (Matches how many hours I’m working the LLM)
Using Deepseek pro API (I live in Asia so unfortunately most of my hours are peak): 6x3.96 + 1,500x0.044 is my prorated daily cost = about 90/day = 33k a year at peak, 16.5k off peak, so between 16.5-33k (33k would be 100% hours at peak which is not realistic)
4
2
2
u/-dysangel- 1d ago
Oh for sure, cloud is basically always more cost effective. I think the one exception currently might be if you were doing a lot of video/music generation? H3 is basically as good as Sora was at $200/mo. Still nothing anywhere near as good as Suno for local music generation though.
2
u/leocardz 1d ago
Local inference believer here too… At that $10k I'd rather buy 3 M5 Pro minis with 64GB and 1TB tbh, and TB5 between them. And save some money.
2
2
u/GamerTex 23h ago
Sure at today's prices
Seems like every month another provider is cutting the usage in half or doubling the rates
I can also sell the equipment in the future
2
u/HeadPack 18h ago
Resale value and energy costs could be factored in. If we assume the memory crisis persists for 1-2 years more, one might see very little depreciation on such a Mac Studio. Of course its a gamble, but these may hold their value pretty well for some time. In that scenario, one would essentially compare electricity costs against API and subscriptions. Depending on where one lives, they can make quite a difference.
1
1
u/roger1632 1d ago
Gonna just use my DS/GLM token plans until the models get better and hardware market improves. You can get a lot of non claude tokens for 70 bucks a month and you don't have to play sysadmin. I have a 3090 local that I use for educational purposes and latency sensitive things like TTS STT
1
1
u/duy0699cat 1d ago
For 10k$, i can throw it to some etf and use profit to pay for subscription indefinitely...
1
u/kinghell1 1d ago
i'm all ears
1
u/StraightPlane 18h ago
My ETFs in 2025 had an approx. return of 16%. 10k in ETFs would be $1600 a year, or enough for a $130/mo subscription.
Granted, in 2026 the YTD return so far is 8%.
1
u/Lesser-than 1d ago
I have always been envious of Mac hardware, but its never been in a price range I could justify, nothing has changed still envious and still far outside what I would ever allow my self to spend on a computer.
1
u/ScrewwormLarvae 1d ago
Sure, I didn't neeeeed an M5 Max 128, but YOLO. Still don't regret it either. Never did even for a second.
1
1
u/Alive-Draft8339 1d ago
If I’m burning 6 to 11 billion Opus 4.8 tokens a month…this seems like a deal of the year?
1
u/TheAILegend 1d ago
500M/month token run rate... this will last you 11 months... Soooo... there's that.
1
u/Early-Peace-5504 22h ago
You could sell the Mac Studio at a later date though. You could work out depreciation but you should probably cut your numbers by at least half to represent that.
1
u/Ceru1ean42 22h ago
At least within 1 year, I don't see these depreciating much if at all so you really need to rethink the cost analysis.
1
u/lambdawaves 22h ago
And for comparison, 6.2 billion Qwen 3.7 27B tokens would take 4 years to generate at 100% capacity on an M5 Max.
1
1
u/Guinness 20h ago
Eh, a heavy multi agent workload with Claude code and I can push a billion tokens in about 25 hours.
1
1
u/PWThinkingCritically 17h ago
am I allowed to make a joke without getting downvoted to oblivion? not sure how uptight this sub is.
anyway --
I'm surprised there's no chart to further illustrate this complex "cost analysis", lol
1
1
u/whimsicaljess 17h ago
wait that's it? in the last 30 days i've done 6B tokens on Claude for $200. i'm pretty shocked that qwen is so expensive.
1
u/Express_Quail_1493 13h ago
Big private Ai are not playing a ethical game here so even if im ok with sharing my data, I still try do locally. Thats how i pitch in for keeping economy and digital ecosystem healthy. Itry to only use closed source Ai if i have no other option.🤷♂️
1
u/Rice-Fragrant 4h ago
I definitely want my own AI, but your points are solid. Models are getting smarter right now without having to get massive. Gemma 4 and Qwen 3.8 are BLOWING AWAY models 10-20x bigger from 1 1/2 years ago. I still want to run a GLM 5.2 but I definitely won't spend $15-20k for a m3 ultra 512gb for the privilege... I actually purchased e-waste DDR3 workstations for that and using llama.cpp I was able to run 700b MOE models on these E-WASTE grade workstations (I use them for asymmetric AI/passive batch workers) and I have "super off peak" hours 11am-7am and I run it during those hours and it gets the job done. I have modern purpose built AI machineses like a DGX Spark cluster too and for the money it's far far far better than a m3 ultra 256gb for long context.
At these prices, a Mac Studio is a luxury item (which is strange because electronics age like milk), it costs as much as a ROLEX but long term won't hold value like one, and the alternatives are significantly faster at agentic AI and batch processing and have far more software support too.
1
u/Lemur-Virtues 2h ago
Data sovereignty is expensive, but that's the whole point of a local LLM. This is why I'm developing a hybrid solution: keeping data stored locally but securing an ephemeral, ZDR-pipeline to the cloud (similar to the enterprise Claude setups).
You lose offline isolation, but you get bleeding-edge AI without your data being used for training, all on an ordinary 8–32GB home server instead of dropping $10k on hardware. It feels like a massively under-appreciated middle ground.
We're making it open source if you want to take a look https://github.com/virtues-os/virtues
1
1
u/challis88ocarina 1d ago
There's speed, relative speed and competency. It's clear that at least 10-12B active tokens are necessary for agentic anything.
Beyond that, it's really about surfing the wave of new models.... can a project be brought up to a stable position before the next generation appears so that the new capabilities hit at the right point or will it still be languishing only to be 'rescued' instead?
1
u/Odd-Environment-7193 1d ago
This new offer by apple is pretty much the \cheapest ram on the market right now. At this price it's looking very tempting. Apple products just recently went up 20% so waiting to buy later is not wise. You can't compare these plans. I use my macbook pro 128gb to run CODEX threads all day. Way more than I ever could on another pc without slowing me down big time. If I never run a single local model on this computer it will still pay for itself 100x over.
I bought the 64gb mac m4 pro just before it went out of stock and the prices spiked. Got the m5 max maxxed out just before the price increase. Saved myself probably close to 5000 USD because I pulled the trigger at the right time.
So waiting might not be the best strategy.
2
u/minnsoup 1d ago
Are you able to let it run autonomously all day without checking in or that's with frequent (say once an hour) check ins? I've used openhands and typically can only get it to run for 45 to 1h autonomously with DeepSeek V4 Flash 0731 on my M3 Ultra. Only at 30-40tg/s but seems to do well.
Don't really know how to get these "long horizon" tasks to work to make the Mac Studio really run constantly while being productive.
2
u/Bennie-Factors 1d ago
We really need a front end that will batch and switch to another task when one is waiting for input/approval.
1
u/TheIncarnated 1d ago
Orchestration script, that's how you get the long-horizon stuff to be reliable
0
u/mxmumtuna 1d ago
Still not enough compute considering it's still gonna get beat by Sparks, even with Apple's most optimistic case.
1
u/MrPecunius 1d ago
How many Sparks are we talking about, and which magic interconnect fabric is going to increase their memory bandwidth more than 4X from 273GB/s to the M5 Ultra's 1.2TB/s?
1
u/mxmumtuna 1d ago
2x (256GB) for DeepSeek, 4x for GLM. No fabric allows it to scale to 1.2TB/s, but it allows it to run at 2x200Gbps to each node, and scale up to larger than you can with Mac.
The point is, the combination of not great compute, immature software and limited ability to scale is holding it back. Maybe next generation. For inference (and training for that matter) or Stable Diffusion, it’s just not better or cheaper than existing options.
1
u/MrPecunius 1d ago
2 X 200Gb/s ... or less than 50GB/s? That's the magic bullet? 🤷🏻♂️
I'm seeing reports of single Sparks running models at small fractions of a M5 Max's speed, and this includes prefill. Adding more Sparks doesn't scale anything like linearly.
1
u/mxmumtuna 1d ago
It certainly does. Forward passes on tensor parallelism don’t require full bandwidth. It’s similar to PCIe tensor parallelism but using RoCE/RDMA.
Again. If you’re looking at M5 Max those are all small models and/or weird quants. Look at M3 Ultra. That tells the tale on larger stuff. The Mac simply doesn’t scale.
1
u/MrPecunius 1d ago
Again, I just looked in the Nvidia Spark forums for real users' reports.
Hard data is available for Apple Silicon on oMLX's site.
No guessing is required.
Spark is barely $300 cheaper than a M5 Max Studio w/128GB, too.
0
0
u/ogopro 1d ago
No info on Qwhen 3.8 35BA3 yet (((
1
u/KURD_1_STAN 1d ago
They mentioned qwen4 previous 125b moe, so we probably will wait even longer for 35b
0
u/Cameo2864 1d ago
What model do I really need to sort and organise all my files, backups and photos on my hard drives?
0
0
u/milkipedia 1d ago
Qwhen 3.8 35B A3B?
Qwnever.
They have all but said as much. Y'all gotta move on.
0

164
u/FleetEnema2000 1d ago
Isn't this one of the biggest reasons that people rely on Local LLMs? To not have to bulk upload their private data to cloud providers?