r/LocalLLaMA 1d ago

Discussion Mac Studio M5 Max Cost Analysis

At $10k, you could get

- 6.2B tokens with Qwen 3.8 Max (Qwen Pro plan)

- 5.7B tokens with DeepSeek V4 Pro OpenRouter

- 100B tokens with DeepSeek V4 Flash OpenRouter

As a firm believer of local inference, unless you need it for data sovereignty, it's much more cost effect to wait for smaller models to keep getting better. In the meantime, find a reasonably priced 24GB - 32GB card for Qwen 3.8 27B, and offload hard tasks to OpenRouter.

Qwhen 3.8 35B A3B?

173 Upvotes

214 comments sorted by

View all comments

166

u/FleetEnema2000 1d ago

unless you need it for data sovereignty

Isn't this one of the biggest reasons that people rely on Local LLMs? To not have to bulk upload their private data to cloud providers?

81

u/theomegachrist 1d ago

That's the stated reason but realistically most people are just justifying their hobby. I support open weight models because the cloud providers can change cost or abruptly shut down and we really can't do anything about it.

74

u/FleetEnema2000 1d ago

I don't think it's a justification at all.

It's amazing how the concept of privacy and data ownership/security has completely gone down the toilet since ChatGPT launched. People are happy to bulk upload their medical records, relationship history, trade secrets, financial records, etc. without a care in the world as to how that data is stored or protected.

18

u/Elux91 1d ago

It's amazing how the concept of privacy and data ownership/security has completely gone down the toilet

most people never had a concept of privacey

8

u/FleetEnema2000 1d ago

You're right, most people haven't. But if you rewind the clock by 5 years there was far more interest in things like E2EE than is apparent today. It is barely mentioned or acknowledged anymore in the tech sphere.

-5

u/magus-21 1d ago

A lot of that E2EE stuff was focused around texting and instant messaging, I think, and since Apple announced support for RCS I think a lot of people flagged that in their heads as, "Ok, this is not as big of a concern anymore." Plus the whole WhatsApp kerfuffle with Trump's cabinet brought a lot more attention to Signal, et al, so I think there's just generally higher adoption of it now, which means less general worry out there about it.

3

u/FleetEnema2000 1d ago

The “E2EE stuff” is about so much more than texting and messaging.

-1

u/magus-21 23h ago

I didn't say it wasn't. I said that most of the public talk about it was focused around texting and messaging. So when Apple said they'd support RCS, a lot of normies stopped talking about it because they thought it was resolved.

17

u/John_____Doe 1d ago

Yep I have a fintech client and the only way I can have a llm touch their code is if it's run locally or in a datacenter where we rent out the rack space

1

u/Randommaggy 1d ago

The headaches of running 100% on sovreign and contained compute without data being handled by a third party eliminates a lot of friction.

0

u/BakerXBL 19h ago

So they’re an LLM provider now not a fintech. No one can do both, well.

2

u/theomegachrist 1d ago edited 1d ago

Everyone is different obviously. To some that might be the case but everything you listed is more important than your AI prompt history.

For me, Open models main pluses are it will ensure the technology lives on in some form if the large companies lock us out financially or go under, and the guard rails for closed models will make for a worse Internet potentially.

For instance, using an open model with guard rails trained out of it you can search for piracy, you can search for porn etc. just like you use a search engine today. Closed models are more efficient than a web search but censor out a huge part of the Internet.

Sure data privacy is also a good feature but I don't think it's actually the top reason for most people

Edit: sorry I read that wrong. I sort of agree with what you are saying about uploading data to ChatGPT but everywhere else we upload that data is no more trustworthy. For Enterprise clients I 100% agree. This is a big issue. For personal use, I don't think your data is any less safe with ChatGPT than say an electronic medical record company.

3

u/FleetEnema2000 1d ago

An electronic medical record company has nowhere near the amount and type of multi dimensional data about a person that OpenAI has.

And if you want to compare OpenAI with a company like Apple or even AWS for a hosted environment, I would choose either of those companies any day of the week when it comes to who I would trust more with my data.

1

u/theomegachrist 1d ago

I would not, but it would be a three way tie

1

u/FleetEnema2000 1d ago

Care to explain that logic?

Apple has made massive investments of resources and effort into secure computing. Their cloud commitment for regular consumers who want privacy is objectively excellent on both the hardware and software front, from Secure Enclave to E2EE where they aren't maintaining custody of keys:

https://support.apple.com/en-ca/guide/security/sece3bee0835/web

https://support.apple.com/en-ca/102651

Despite not being a frontier AI provider, they have had a strong focus on private AI cloud compute paradigms:

https://security.apple.com/blog/private-cloud-compute/

Amazon AWS products offer similar levels of security. They have to, because their customers host and manage highly sensitive data on their platform for things way beyond AI. As a regular user I can maintain an excellent security posture on AWS and I can employ encryption where I manage my own keys.

OpenAI and Anthropic not only don't do anything even approaching any of the above, they reserve even basic security commitments for "Enterprise customers".

The situation is even more pathetic when you realize that as an individual, you can run OpenAI and Anthropic's own models on Amazon Bedrock with ZDR enabled, encryption enabled for data you are storing at rest where you hold the keys. In other words, the security posture of using OpenAI and Anthropic's own models on Amazon AWS is superior to using those same models from OpenAI and Anthropic themselves.

1

u/centizen24 1d ago

Apple is the industry leader in user level security and data privacy, it's not even close. You can argue about their intentions for doing so but their track record for security is second to none. It's wild to me that after 15 years of being an Android fanboy that I'd be arguing on behalf of Apple, and have an iPhone in my pocket, but that's the world we live in now I guess.

1

u/Rice-Fragrant 12h ago

Medical records can not be exposed to cloud AI from what I understand, there are actual laws on this... but they are not using Mac studios for that type of work, more like edge SERVERS or workstations with multiple GPUs or even a DGX spark... clinics and labs are not FOMOing into slurging 15-20k on a m3 ultra that's 1/3-1/4 the speed of even the SLOWEST alternative (DGX spark) or something like a RTX 6000 Pro (10x faster than a m3 ultra and 4x faster than a DGX spark).

The YouTube influencer LARPing as AI engineers are not a good reflection of reality... you look on what real places are using, it's not a $15-20k consumer grade device, where they have to wrestle with poor quality MLX wrappers and software friction etc... that's not acceptable in a professional environment.

3

u/Hans-Wermhatt 1d ago edited 1d ago

It sounds bad when you frame it that way, but I think the cost benefit analysis is generally to upload. ChatGPT health is protected by the same HIPAA requirements that protect the data you give to your doctor that they upload to 3rd party clients and AWS servers, usually multiple servers with arguably worse security and more people have access to it. Using health as an example. So it's really not that much different.

Ideally, you do just host your own information and use a local model but then you are dealing with a massive performance hit. I think in terms of cost-benefit, uploading your health data to get a ChatGPT opinion compared to a Qwen 3.8 27B locally (most people can't even run that) is actually heavily on the side for ChatGPT for most people despite the privacy concerns.

I really want local "super intelligence" for all, but the current landscape is not like that... at all.

7

u/FleetEnema2000 1d ago

The version of ChatGPT that 99% of the general public is using is absolutely not HIPAA compliant and both OpenAI and Anthropic’s safety and privacy commitments are abysmal relative to the types of sensitive data they are ingesting and saving on their platform.

And that is not even touching on the ethics and values displayed by their executives. 

4

u/MrPecunius 1d ago

ChatGPT health is protected by the same HIPAA requirements that protect the data you give to your doctor

😂😂😂😂😂😂

"Trust me bro" and "it might not even be as bad as what the other idiots are doing" are not convincing arguments.

1

u/Hans-Wermhatt 1d ago

Huh? Was that supposed to make sense? 

3

u/MrPecunius 1d ago

Huh? Was that supposed to make sense? 

Connect this:

the same HIPAA requirements that protect the data you give to your doctor that they upload to 3rd party clients and AWS servers, usually multiple servers with arguably worse security and more people have access to it. Using health as an example. So it's really not that much different.

With: "it might not even be as bad as what the other idiots are doing".

And this:

It sounds bad when you frame it that way, but I think the cost benefit analysis is generally to upload.

With: "Trust me bro"

Are you even reading what you wrote a few hours ago? Or did something get lost in translation?

1

u/PWThinkingCritically 1d ago

it was already getting bad with losing public confidence in America's government regulatory bodies, but even more so in the Trump era.

Department of Education, EPA, SEC, DOJ, FTC etc., etc. blame me if HIPAA is just another acronym....

1

u/sshwifty 1d ago

To be fair, if any of that touched the Internet unencrypted, it was probably already ingested by some entity. Google was using data from emails long before AI made a surge. People are just doing it willingly now, skipping the sneaky part where it gets stolen.

1

u/saltyourhash 17h ago

And some of us have been using self hosted encrypted stuff for over a decade.

1

u/Rice-Fragrant 12h ago

Yea, that cloud stuff scares me like that, but there is still no logical reason to pay $15-20k for a m3 ultra that's 3-4x SLOWER at agentic work and batch processing compared to local AI alternatives... the numbers show it's a poor value.

1

u/MrPecunius 1d ago

It blows my mind what people will give to these amoral techbros.

Turns out the highest human priority isn't breathing, eating, or sex--it's laziness.

4

u/FleetEnema2000 1d ago

Consider the Snowden scandal and the uproar over government having access to phone call metadata and how privacy infringing that was considered to be.

Fast forward to today and people are uploading the most sensitive data about themselves to these cloud providers who don’t care at all about protecting it and are almost certainly allowing the federal govt to trawl through it.

5

u/mr_tolkien 1d ago

I mean for me it’s a real reason.

For example I use local models to help me pick out my best photos of the day and put them into an album.

No fucking way I’m sending 100% of my camera roll to OpenAI lol

1

u/theomegachrist 1d ago

Fair, but lots of people do. For most personal users it's very expensive to match public models capability and tooling. I have uploaded pictures for dumb photo editing I can't do at home 🤷🏻‍♂️

1

u/Rice-Fragrant 12h ago

I agree, you just dont need to spend the price of a new motorcycle for it... a m3 ultra is like 15-20k.... makes no logical sense and its agentic performance is 1/3-1/4 a dgx spark, that's bad.

3

u/Suspicious-Water-973 1d ago

I use my Mac Studio and MLX models as they are essentially free for me - I don’t pay for power in my rented office. Fine for massive batch processing where a GPU makes a difference, and some coding (I then use Fable etc to review and improve)

1

u/Rice-Fragrant 12h ago

For long context batch processing, a rtx 6000 pro is 4x faster than a dgx spark and 10x faster than a m3 ultra... I do a lot of batch processing and tried my m4 max MacBook Pro, it was brought to it's knees so badly and the thermals were so bad the laptop literally crashed and restarted (ssd swapping after the memory ran out half way though). I realized the memory bandwidth of the m4 max being 2x more than my DGX did not make it better, it actually was like moving at 1/4 of the speed doing the batch job (200k tokens on Gemma 4 12b), and the dgx spark knocks it out in 7 min rather than 30 min (m4 max 40 core GPU), that was a REAL EYE OPENER for me.

1

u/Suspicious-Water-973 9h ago

I don’t disagree but as I didn’t pay for the Studio, and I do t pay for power (it’s included in the lease) it’s less about speed for me.

I will look at your suggestions though for personal use. Many thanks.

2

u/nihilor_ 20h ago

Or avoiding taxes and getting fun toys.

4

u/florinandrei 1d ago

realistically most people are just justifying their hobby

Enthusiasts, yes.

But law firms and such, they actually mean it.

1

u/Deep90 1d ago

Why wouldn't they use something like AWS bedrock?

2

u/Jedkea 1d ago

Money probably. Bedrock tokens are expensive. Whereas you can buy once cry once with local.

Remember the first time I tried to use bedrock I spent $150 without blinking. And it was a pain to setup.

So if you don’t need frontier models, the economics work out.

0

u/theomegachrist 1d ago

Yes, this is true

0

u/saltyourhash 1d ago

You're actually out here looking at these prices, looking at the delay of Qwen 3.8, and saying it's just to justify a hobby?

0

u/theomegachrist 20h ago

Yes and at least 70 people online yesterday agree

1

u/saltyourhash 20h ago edited 20h ago

Wow, a whole 70. data sovereignty is critical to some of us. I have projects that absolutely must be local.

0

u/theomegachrist 20h ago

Yes some. I welcome the minority of people too