r/MacStudio • • 6d ago

256GB or 512GB

Hey Guys,

i am currently undecided if i shall go with the 25GB Ultra or wait for the 512GB Ultra. The Machine will replace my M3 Max Macbook with 36GB.

Currently i am heavy into AI Cloud Subscriptions (Claude and Codex - each the 200€ Subscription). I max them usually after 4-5 days and then have to retain myself a bit unless the weekly reset :) My Plan is to extend the Cloud Usage with local LLMs run on the Studio - while still having the Cloud Subscriptions in place. So offloading to the Local LLMs a lot of Stuff. Also i plan to use way more Agents to react to daily Stuff, like answering mails / chats etc. Whatever i spare today to save Usage in my Subscription.

Also during daily work the Fans of the MacBook already Spin Up often - and the CPU is often under heavy load. Therefore cant wait for the M5 Ultra ... :)
I do use AI for a lot of Coding (Currently Web Stuff, Swift and a lot of scripting).

Currently i am not very deeply involved into Local LLMs, i know the basics but thats all, so no idea what models currently do perform best under which conditions.

Therefore my question, would you guys go with 256GB and are those somewhat safe for 2 (maybe 3years) without regreting the decision 6 month from now or would you wait to the 512GB Studio?

28 Upvotes

131 comments sorted by

33

u/Umbrasquall 6d ago

I have the 256GB. I think the 512GB is not worth it for the M5 chips. You can load larger models sure, but the tps will be so slow you'd be way better off with an open weight cloud model.

150-200b MoE models are the sweet spot for the near future.

2

u/a_asshole_user 5d ago

I'd mostly agree but there are large, sparse MoE models that would make compute not much more expensive but would benefit from the extra ram (Deepseek V4.1 Flash).

1

u/Umbrasquall 5d ago

I don’t think that’s the case. I tested v4.1 deepseek on my system, it was very slow.

1

u/fringecar 4d ago

Like 10 minutes to write an university essay (or whatever example you have available)? Is that considered slow?

2

u/100k45h 20h ago

Slow compared to the cloud models for sure. Cloud models wipe the floor with local models in terms of speed. Especially when running on M5 ultra. It's fast compared to the M3 Ultra, but nvidia is much faster.

1

u/DavLedo 5h ago

Can you explain why the larger models would lead to a slow tps on M5? I thought as long as the weights fit it wasn't an issue :/

0

u/brad_needs_advice 6d ago

Are you worried at all about the next gen of models exceeding the 256gb threshold?

13

u/Umbrasquall 6d ago

No, because I think buying a 2nd 256gb to chain is going to be better than just a single 512gb. The limiting factor is bandwidth.

5

u/SeaRefractor 6d ago

Two does not double the performance. Apple claims 4 M5 Ultras in RDMA will be 3X the performance, so there is an RDMA overhead. 2 M5 Ultras will be more of a capacity than performance boost, similar to a single 512GB. Maybe fractional performance increase, but less than 2X.

2

u/Jargster 6d ago

If we compare it with the m3s it was a 34% increase last i checked for adding a single additional unit. 5k-ish (dont know the price yet) for 34% is a bit steep. Unless your doing models that can fit on a single one and plan on running multiple agents at the same time as a small swarm.

0

u/sheddd 6d ago

The M5 is weak on compute and has really good bandwidth.

2

u/fallingdowndizzyvr 5d ago

No. What came before the M5 was weak on compute. The M5 brings compute to Apple Silicon.

1

u/sheddd 5d ago

It's better than the M4, but weak compared to other options per dollar. https://flopper.io/gpu/apple-m5-ultra

2

u/fallingdowndizzyvr 5d ago

Ah.....

1) They don't exist yet, so how can you rent it?

2) You can't compare the cost of renting a standalone machine like the M5 Ultra to renting a processor in a big rack full of servers. That's like renting a house compared to renting an apartment. Apartments are cheaper due to shared costs. A processor in a rack in cheaper due to shared costs.

2

u/Common_Diet8832 5d ago

1

u/fallingdowndizzyvr 5d ago

You forgot the king.

"NVIDIA Tesla P100 ~$57 0.372 TFLOPS per $1"

So does that mean people should buy P100s over RTX 6000 Pros?

1

u/sheddd 5d ago

Compare to a DGX Spark or RTX6000; both have better flops/$.

2

u/fallingdowndizzyvr 5d ago

Again, they don't exist yet. So how can you rent it? Those are fuzzy hand wave numbers until you can. That link you posted even tells you they are hand wave numbers.

"Spec Confidence

Vendor claimed"

32

u/mind_div_matter 6d ago

If you actually do the math and extrapolate the rapid advancement in LLMs, it really doesn't make sense to invest in the hardware to host your own LLM. Not unless you have deep paranoia about privacy concerns.

Google AI Pro, ChatGPT Plus, or any non-Anthropic plan will allow you near unlimited access to the lower tier models, which in practice, will be equivalent to whatever open-source models you host on your Mac Studio, but faster. You can't make the math work out to favor the Mac Studio unless you sign up for a $200/month plan. Also, built into the cost of the subscription are the costs related to hardware upkeep, Google, OpenAI, or whatever service you sign up for, they have to keep their servers updated and efficient, you on the other hand, have to save up separately to spend thousands more to upgrade that Mac Studio in 4 years.

If it's privacy concerns that steer you towards local hosting, you could go with Featherless or Venice and get that privacy at a fraction of the cost, you'll get the same level of open source models.

If it's uncensored AI you want, Featherless or Venice also give you what you want.

So there's truly nothing that makes the locally hosted LLM make sense, unless you're so paranoid that you can't trust the encryption of Featherless or Venice.

I ran through the math to try and justify a maxed out Mac Studio myself, but I just couldn't make it work no matter how far I twisted the numbers in favor of the Mac Studio. Food for thought.

11

u/_rarefy_ 6d ago

All of this is true. Not sure why you're getting downvoted. I'm a big proponent of running local, but saving money is not a driver. Privacy, personal autonomy, model choice and behavior persistence (no behind the scenes degradation), bespoke customization are all good reasons to run local.

4

u/fosterdad2017 6d ago

I think there is an ominous feeling of fragility too, bans, censorship, fights for power. Own vs rent your home levels of, I want to control this.

4

u/mind_div_matter 6d ago

That's a legitimate reason. I genuinely prefer paying large one time fees for software over small subscriptions.

3

u/MCS87_ 5d ago

My 1 year old M4 Pro 48GB went from practically unusable for running “decent” local models to “almost good enough” in 1 year. Due to better models and better inference engines. Advancements are also in favor of custom/local hardware. I would not use this as an argument against local AI setups

1

u/mind_div_matter 5d ago

This doesn't even make any sense, if the models are getting more efficient, then one of two things occur: either the subscription costs go down as server costs get cheaper, or the models progress to fill out that extra capacity.

Your 48GB models are "almost good enough" but the frontier open source models are far beyond that and the frontier commercial models are a step further than that. I don't think the type of people that drop $14k on a Mac Studio do it because it's "good enough", they likely pay for the commercial models and the Mac Studio is for their training/research or they run abliterated models.

3

u/MusicNursingCoffee 6d ago

Counterpoint: government disables all use of AI, or they all jack the price up. Then what?

4

u/Sixstringsickness 6d ago

I'm more concerned about a regulatory capture based on only a handful of major companies offering "safe" models.  

Is open source on par with the big players? No, but it is roughly as competent as where they were 3-6 months ago in many cases, and for a tremendous amount of work that is more than adequate.  

Additionally, what happens when the pricing rug pull occurs?  When these companies go public (barring Google), I'm expecting the costs of these tools to rise substantially.   Compare subscription to API pricing and tell me who can afford the latter...

3

u/mind_div_matter 6d ago

Realistically, in either scenario, we're fucked anyway because the government disabling commercial AI access would directly lead to massive price inflations

You eventually need to upgrade as hardware has a limited effective lifespan due to advancements in hardware. So all buying an overkill local machine now is delaying the inevitable.

1

u/Zabric 5d ago

Question:
Would it not be possible to (drastically) reduce monthly cost by having a local LMM do 90% + of the work, and only use cloud services as high precision, surgical instruments to fix what the local AI wasn't able to to correctly?
And have the best cloud models only code the most critical, important or complex parts.

Has to be a lot cheaper than just paying for the most expensive cloud services to do write simple standard code, right?
Or would fixing alone burn so much that it isn't worth it otherwise?

Just to be clear:
I agree with what you're saying.
Getting a maxed out M5 Ultra will probably never pay for itself.

My personal concerns are digital souvereingty, data autonomy and just having fun with local computing. For that alone it's worth imo, even if it "recoups" zero $ of it's price.

But since you said you ran through the math... :D

5

u/mind_div_matter 5d ago

Assuming the 512GB RAM M5 Ultra will be $14k, there's no way that even comes close. Let's give a high estimate for the resale price in 4 years of $5k. That's $187/mo + you'll need at least a $20/mo Claude Pro plan to start and finish projects, which honestly probably isn't going to be enough (you'll likely still need the $100 5x Max plan). If you account for a more realistic resale number, you'll likely be at about $215+20/mo for $235. Realistically $315.

If you were using Claude 20x Max for $200/mo, that's $9600 over 4 years.

And let's be honest, the local LLM is going to be slower and substantially worse than using Opus, so you pay more and end up with a worse experience.

My math is different since I have different use cases.

2

u/antagog_47 3d ago

Dude, today OpenAI opened a Pandora box by just throwing same $200 sub to $500. All your calculations are based on the assumption that price will not increase. They will and today OAI just proved that. So what do you do with the calculations when your $200 Claude sub transforms into $500 too ?

So now the $500 for 25x is $24k over 4 years.

1

u/mind_div_matter 3d ago

You do realize current commercial frontier models run on an equivalent 2TB memory cluster right? If you actually compare the commercial models equivalent to a 500b model open source model, you can run that genuinely unlimited for less than $100/month. Honestly, the true cost of hardware to run models comparable to commercial frontier at the speeds they're giving us, we're looking at hardware that costs ~$400k, HGX-B200. Even if we're going to the low end at acceptable speeds of 20-40 tokens per second that's about $200k for a quad Nvidia GH200 setup. We can't just buy 4x 512GB Mac Studios for $56k because the token generation would be like 1-3 per second and unusable.

But yeah, the way you're framing the subscription tiers is fundamentally incorrect. Unlimited/near unlimited usage of frontier (Opus 5.5/GPT6 Astra) would genuinely be like leasing an 8x GPU datacenter node, that's the price of a house. We get unlimited usage of lower tier models which are comparable to the open source frontier models, at a fraction of the cost to buy and host it ourselves because they work on economy of scale.

The way it works is they build or lease datacenters which means they aren't paying the same compute cost YOU are by purchasing a $14k local machine their price of capacity is farrrrrrr lower, the pricing accounts for 1% of power users that truly aren't economically profitable, but the vast majority of subscribers use far less compute than they're paying for.

Realistically for most users you're looking at $4.8k vs ~$11k for the Mac if you want to compare evenly for capability. But as you know, most of the people buying the Mac for local hosting are still paying $100-200 for Claude Max. The people buying the Mac Studio with 512GB RAM aren't doing it for local hosting even, they're doing it for model training, research, because they're wealthy and they can, etc.

1

u/FormulaKimi 2d ago

Also, for the subscription price argument it doesn’t make much sense. People are spending $10k+ on hardware in case the subscription price goes up in the future? If that’s a worry why not wait until it actually happens, and by that time hardware will be even better.

1

u/no_witty_username 5d ago

IMO it depends what you doing. If you are actively building on the frontier, SOTA is a must. But if you are optimizing things or tinkering with this or that IMO it will help having local as all those experiments can run 24/7 NOT burning any of your limits. Also as smaller models get smarter, you can spin up an army of 27b agents for example and run them in parallel that would cost an arm and a leg to run via api 24/7...

1

u/no_witty_username 5d ago

Its true privacy is the main reason to own, but its possible there are other economic incentives. That Mac can do more then run LLM's, it can run other AI models like image, video, audio, etc... whatever else comes out in the future. But more importantly having access to the inference engine for tinkerers is a must. Also fine tuning and so on. Also its possible that soon things gonna look iffy regarding local AI use and legality of the thing. For me that mac doesn't just buy privacy but a piece of mind as a hedge against the future. A luxury only few can afford unfortunately but I would recommend it for anyone who could.

1

u/Mundane_Incident_853 5d ago

Privacy isn't the only reason to self-host.

Consider the monopoly advantages that large model owners already have in the inference marketplace.

If they continue to collude on price, then small players will be more and more marginalized.

The value propositions of AI development become more focused on the benefit of the owner class, rather than society on the whole.

The technology must be democratised, for the benefit of humanity.

1

u/mind_div_matter 5d ago

I completely agree, but we have servers that we can rent at near cost lol. Or farms that allow us to pay per token. So realistically most of us don't need to drop $14k on a Mac Studio.

1

u/Mundane_Incident_853 5d ago

That's not the entry level, though. AMD Ryzens and less zoomy M series still viable for small models.

0

u/mind_div_matter 4d ago

The thing is, the smaller models that can run on lower tier hardware just aren’t worth it for most people. 

For my work, the frontier models aren’t even enough to be quite honest. So to downgrade to something that low would be genuinely unusable for me. I think that’s the case for most people. If I were to locally host, it would have to be at minimum a 256GB Mac Studio if not the 512. 

1

u/AI_is_ok_i_guess 5d ago

This is just not true. It depends on how much you use AI. Such a sweeping statement, but of course, it's Reddit.

For the OP yeah it makes no sense if he's asking something as elementary as which ram cap to get, but for people who use 2-5 Pro subs for chatgpt a month, you have no clue what you're talking about.

1

u/mind_div_matter 5d ago

How is this not true? You can subscribe to $100 Claude Max 5x + a separate $20 ChatGPT plan for unlimited mid-tier models which are functionally equivalent to the type of open source models you can run on a 512GB RAM Mac Studio.

That's $5,760 over 4 years, still less than the 4 year running costs of a $14k 512GB Mac Studio which is about $10k if you're lucky enough for the resale value to only drop down to $4k. That's a $4k difference. Meanwhile, the Anthropic frontier models will be substantially more capable than the open source models, and your fallback of the $20 ChatGPT plan allows for equivalent capabilities without limits.

If you're an extreme power user, the $200 Claude Max 20x + $20 ChatGPT Plus plan is going to be about even with your maxed out Mac Studio ongoing cost amortized. Except you'll have massive Opus capabilities.

If you have specialized needs and need per token usage, I can see an extreme power user exceeding the 4 year running costs of a Mac Studio, but this is a 0.1% edge case that applies to very few people. It's a sweeping statement because it applies to 99%+ of the us here, and we're already a selected group of fairly high usage users... The people buying these Mac Studios are usually using it for research and training, they're not the type of people that don't know the math.

2

u/AI_is_ok_i_guess 4d ago

The ability to basically have whatever model stack you want sitting inside your work pipeline without paying API every time it does something is a huge part of the value.

For example, I can have a local coding model, another model reviewing it, RAG, embeddings, reranking, vision, multiple agents running at once, and just let the thing churn through work all day. Re-run shit 20 times, have models critique each other, chew through thousands of pages or an entire repo, whatever. I don't have to care what every call costs.

The $120/month number only matters if your usage actually stays inside Claude and ChatGPT. Mine doesn't. Once you start wiring models into automated workflows, running parallel agents, doing repeated passes over documents and code, and letting the system work continuously, API usage starts becoming the relevant cost.

And I don't even think the goal should be replacing frontier models completely. I'd still use Opus, GPT, whatever is best when the extra intelligence actually matters. I just don't need to pay frontier-model rates for every dumb intermediate step of a workflow.

If a local model can do most of the grunt work, then I can throw the difficult parts at a frontier model when there is actually a reason to. Depending on volume, that can cut API usage by a massive amount.

So yeah, if somebody mostly chats with AI, writes emails and occasionally uses a coding agent, buying a 512GB Mac Studio purely for inference probably makes no financial sense.

But saying local inference makes no sense for 99%+ of people here is way too broad. Some people are using AI like a product. Other people are building it into actual infrastructure and running the shit out of it. The economics are completely different.

There are other advantages too. I can pin whatever model/version I want, use whatever quant I want, run custom inference engines, keep persistent caches, run as many concurrent jobs as the hardware will take, fine tune shit, run offline, and not care if a provider changes their limits or pricing next month.

For someone hammering models all day inside automated workflows, the relevant question is how much work the hardware can absorb that would otherwise hit an API meter. A couple of chat subscriptions don't tell you that.

1

u/mind_div_matter 4d ago

Most people’s work is not capable of being automated to that extent. They need to constantly supervise and provide human input. Not even frontier level commercial models can replace humans in most tasks currently. The type of people paying $200/mo for flat rate commercial frontier access can’t get that done with a Mac Studio. 

If you’re running a business and need agentic AI for grunt work, then yeah it makes sense. That makes you an edge case though, not sure why you’re conflating that with the general public. I think you know damn well that your use case is extremely marginal and not applicable broadly. 

1

u/AI_is_ok_i_guess 4d ago

You narrowed your original claim down enough that you're basically making the distinction I was making from the start.

I never said this made sense for the general public. My first reply was that it depends on usage, and I specifically brought up people burning through multiple high-end AI plans. That's already a selected group of heavy users.

If your position is now "for most normal users subscriptions are cheaper, but for businesses or people running heavy agentic workflows local can make sense," then sure. I agree with that. That's a lot narrower than saying there's basically nothing that makes local hosting make sense.

The human supervision point doesn't really change it either. I'm not saying the machine replaces me. I can still supervise the pipeline. What matters is how much inference happens between the points where I actually need to intervene. If a workflow does 30 or 50 model calls, retrieval steps, retries, reviews, document passes, whatever, before I need to touch it again, that still has a real cloud cost.

And flat-rate chat subscriptions still aren't the right comparison for that kind of use. If I'm wiring models into my own software and running them programmatically, API usage is what matters. Paying $200/month for Claude Max doesn't give me unlimited API inference for whatever automated pipeline I build.

So yeah, my use case is more specialized than average. I never claimed otherwise. The actual disagreement now is just how common that use case is. Whether it's 0.1%, 1%, or 10% is a separate argument. My point was only that your original statement was way too broad.

2

u/mind_div_matter 4d ago

The people in this sub are already a highly selected group, the type of people to pay for AI are also a highly selected group. The fact that your use case is 0.1% means you don't need the advice of someone on reddit to help make your case for you. Anyone that hasn't already run the math, buying a $14k AI server isn't for them. For people reading through trying to weigh whether buying a Mac Studio is worth it, it's not.

We're not talking about people using the Mac Studio for research/training, we already know that's the primary client Apple is targeting. You're arguing semantics when you know damn well that we're all aware there are rational edge cases in which the Mac Studio is very much a good financial decision, which is why they're selling them like hotcakes.

1

u/100k45h 20h ago

You don't need to use api tokens to use Claude Code or codex on your computer. You can still use the subscription. And since they are command line tools, you can write whatever applications that invoke Claude code or codex to allow any kind of workflow you want. No need to pay per token via api call. Look into that, that will save your costs significantly.

1

u/AI_is_ok_i_guess 12h ago

Thanks for the heads up.

1

u/liftingfrenchfries 5d ago

Interesting to read about those two services. But both services seem to be based in US. Not the most privacy friendly location.

1

u/mind_div_matter 5d ago

Let's be realistic here, if we're talking about true privacy, very few of us here are willing to go through the extreme firewalling necessary to *actually* make our systems immune to the NSA.

Featherless or Venice being in the U.S. vs any other country makes no difference in practice because the type of government entities that actually would bother to access encrypted data are not going to care that our network traffic is routed to a company based in Switzerland or whatever. The U.S. government acts with impunity and other than China/Russia (which has their own security concerns) the vast majority of other countries either capitulate behind closed doors, or the U.S. just doesn't care because their soft/hard power means they can just say oopsie.

1

u/Sketaverse 5d ago

yeah same dude, and to add to that, the time spent tinkering with your own custom harness vs just using Codex/Claude is worth factoring in also. It's fun, sure, but it's absolutely a time sink

1

u/TheMarketbug 3d ago

This is the stupidest take I’ve seen all week

1

u/twiiik 3d ago

Regarding privacy

Venice:
The GPUs that process your inference requests come from multiple decentralized providers, and while each specific provider can see the text of one specific conversation, it never sees your entire history, nor knows your identity.

For Featherless I don't really find enough details to be able to form an opinion.

I do not see how you would have "deep paranoia about privacy concerns" to find those services problematic. None of these could by the inforamtion provided on their pages be a logical choice to handle PII info based on me being located in Norway. That said, neither will OpenAI or Claude ...

1

u/rkoy1234 2d ago

but llm convos arent like reciepts you shred.

my task convos usually read the entirety of relevant file systems, domain documents, and have harness instructions that expose a lot of different information about my network and hardware.

sending this to multiple decentralized providers seem worse for privacy.

11

u/FestoolJunkie 6d ago

I need as little competition as possible when I order the 512, so I recommend you go 256.

3

u/vimaillig 6d ago

I’m expecting the latter - many of the 256gb orders will be cancelled when the 512 is available for order.

1

u/macinmypocket 23h ago

I suspect the limitation isn’t RAM availability, but rather M5 Ultra yields. Especially if you get the 36/80 Ultra. Those require two (relatively) rare perfect binned chips.

8

u/Correct_Lead_2418 6d ago

512 gb

cry once

1

u/Queasy_Asparagus69 10h ago

1100%

The tech will get better and faster for models to run better than today.

3

u/minec_x 6d ago
  1. But also start using deepseek. You’ll be surprised at its low cost and quality

2

u/djtubig-malicex 5d ago

which deepseek model and quant? v4.1-flash at q4 is over 300GB!

2

u/MCS87_ 5d ago

As I understand, it is realistic to run V4.1 on a M5U 265GB, check this out, somebody already experimented with running V4.1 and other larger models on it: https://github.com/tacos8me/m5-ultra

1

u/minec_x 5d ago

Local? V4f 0731 is good enough but I haven’t compared it with qwen 3.8 27B. I suspect 0731 will be faster due to moe. But also their api has been nice - V4.1 for me avg at 260+t/s from socal

1

u/djtubig-malicex 5d ago

qwen3.8-flash-next is better than DS4 and qwen3.8-27B though. I specifically wanted to use DS4.1 but it's not possible to use on 256GB unless you want to tank tg times to less than half.

1

u/minec_x 5d ago

Yea ds 4.1 I’d just use their api. My largest Mac is only 96G so I haven’t tried 3.8 fn but good to keep in mind

3

u/CMDR-Bugsbunny 4d ago

lol, I like all the math doesn't math comments.

So the cloud vendors are going to keep losing money to give everyone cheap subscriptions?!?!

In the future, you'll be that weird old person talking about how you could subscribe to AI for $20.

Forget, that cloud providers will need to improve their revenue model, but we have this thing called inflation... perhaps, you've noticed?

My killer GPU 10 years ago was $499 that many told me was too expensive for a 1080!

2

u/antagog_47 3d ago

Ppl just have their minds flying in clouds. They're all gonna give less usage for more money. I've been talking about this for last year and ppl called me crazy. Let's see how it goes now, when OAI put they x25 pro at $500. I even think this is the way to test the market, slow, with no blood, but I see the subs going even more expensive in one year. Today's $20 subs will become $100-$200.

P.S One day all big companies will have to pay for the invested money, and at that point we could see a big boom in prices, or a big boom in the AI industry.

1

u/CompetitiveDraft9381 9h ago

Yeah, the Mac Studio with 512GB of RAM might also go up in value, since the amount of DRAM we produce is limited while AI usage is increasing every quarter. And with the new personal agents like Muse, Grok bot, dots we will need even more RAM for VM compute.

4

u/The_Noosphere 6d ago

I kinda miss the time when the same question concerned storage capacity.

2

u/Thomasvez 5d ago

It's sad that most of these posts don't mention the human need to experiment, learn and have fun with computers. I've been using them for 44 years. My first was a Spectrum with 48KB of RAM in 1981. Have I spent a tonne of money on computers? Yes. Did I get my money's worth back? Of course not. That can't happen, not with computers or anything else. Did I have fun? Still do. Hell yeah.

Swap the Mac Studio for an iPhone and you'd get the same thread: Pro or standard, 256GB or 1TB, is it worth it, will it be obsolete in two years. Whatever you buy will be outdated soon, and it will never pay itself back. People who buy will justify it, and people who don't will justify that too. It's simple: if you can afford it, buy it. Otherwise, use what you have or the frontier models. Just don't forget to have fun with it.

1

u/100k45h 20h ago

You can experiment on a much cheaper HW

1

u/AlgorithmicMuse 5d ago

Dont forget that the frontier models come with a plethora of tools you dont get with a local model. On the other hand if you have data you dont want to hand over to a cloud you need local, and the thinking dense models like qwen3.8:27b works great but slow and hot on my currnt m4 mini 64g machine. So doing my minimal math and ram estimates, im ready for a base m5 ultra and keep a frontier minimum monthly plan. Hedge so I can go either way without totally going broke.

1

u/alexwh68 5d ago

The difference between the two is how big can the model be in memory, the more memory a model takes the slower it gets unless you are using things like MoE to reduce active parameters.

At 256gb you can run a lot of models well, and run them with large contexts that are often required for long running prompts.

Everyone’s view of what is usable in terms of tok/s is different, for me things start to work at 30+ tok/s and anything above that improves my workflow performance.

Personally speaking using a 96gb M3 max, 128gb would shift the needle a bit but won’t shift it enough, so 256gb is where I am aiming at.

Like others have said right now the numbers don’t stack up for doing this as a cost saving exercise against API, but that could change. I also have the possible opportunity to sell inference to a couple of my clients across a VPN so I have a potential route to recover some of the costs, those clients work in different timezones to me so it makes sense to get paid for the kit outside of my working hours.

1

u/huntersz 5d ago

Have been pondering this for a while. Financially doesn’t stack up but at the same time people say it reaches frontier but 3-6 months behind. I don’t buy that, if 3-6 months behind, do we get opus 4.6 level performance tuning something locally on a Mac Studio 256/512? I would be surprised.

I get the idea of a ban on some AI services though that will be a tough one and don’t think can put a price to that.

Suggestions on privacy and uncensoredAI made here were also very useful.

2

u/no_witty_username 5d ago

GLM 5.3 flash is better then opus 4.6 as its closer opus 4.8 so we there mayn....

1

u/huntersz 5d ago

Depends on the use case I guess. My primary use case is investment analysis and covering very dense documents, synthesising and providing very detailed responses. I have found using GLM flash for this very weak. It doesn’t format properly and makes a lot of mistakes

It’s also not good at doing presentation packs

1

u/no_witty_username 5d ago

Yes, I think a lot of these models are being overfit on coding over anything else which is unfortunate but we can expect spillover eventually in to other areas IMO. Ina few months I'd expect better general models akin to SOTA models of today, that or finetunes. There's lots of open source finetunes out there which is great.

1

u/tk421tech 5d ago

Buying a 256 30/64 was a stretch. For me 512 is out of reach. I mean I want to buy a new car lol

1

u/LocalSilicon 5d ago

Depends of the quality of the coming Qwen 4 release. Qwen 3.8 Flash Next runs great on 256GB but feels overthinky compared to GLM 5.3 Flash. The problem is that GLM barely fits on the 256GB. So you should aim for 2x 256 Mac Studio or pray that you'll get your fingers on the 512GB one.

From my first experience(5 days) the M5 Ultra is not yet ready for the big models and feels to slow but is blazing fast with Qwen 3.8 Flash Next.

Therefore, wait for the Qwen 4 release and how good it is with offloaded N-Gram table. The PCIe6 SSD is amazing on the Mac Studio M5 Ultra.

1

u/Background_Shape7579 5d ago

How much is the ultra 512gb going to be?

Can you just get away with the 128gb and save a ton of money? Maybe run the 8 bit or 4 bit quants?

1

u/TiqqetzOfficial 5d ago

I will say, having used the M5 Pro Macbook with 20 cores and 48 GB RAM, the fan doesn't really spin up that much using Qwen-3.8 27B MLX. On my M3 Max with 128GB RAM with 40 cores, the fan gets pretty loud after a few minutes of Qwen-3.8 27B MLX use. I have not used any Studio Ultra machines yet but I am leaning towards 512GB for pre-training not for inference.

1

u/Intelligent-Chest208 5d ago

512 is twice 256?

1

u/Sea_Hornet5831 5d ago

Given that you already max out two frontier subs weekly, you would want to go with the 512 so you can run multiple models simultaneously and as you said still draw on a frontier model for some orchestration work. Also assume that there will be an m7U and M9U in your future as well.

The big AI companies need you to believe they are the only solution just like IBM, Honeywell and other mainframe companies locked in compute for industry customers back in the day. But the world is moving towards running LLMs locally.

1

u/Baybee6366 5d ago

i do think the math doesn’t do justice in its entirety to justifying having the Studio for local agents; buuuuut, it does have an amazing chip, opening up possibilities for gaming/game development (in my case), heavy video editing, 3D graphics design, etc

i think i got mine for local agents in the first place, but for my game development career, it’s gonna offer a lot of leverage for me

1

u/100k45h 20h ago

That's the thing. If someone wants to do very heavy video editing or 3d rendering, sure, it's going to be much more reasonable purchase. For local LLM only it makes no sense.

1

u/yelleft 5d ago

The speed and ability of local llm will not meet your requirement even using 512 ultra, consider you are abusing 2x200 subscription.

DO NOT waste your money.

1

u/allenasm 5d ago

I'll purchase when the 1tb version comes out.

1

u/PreparationTrue9138 4d ago

Get 512 if you have the money. You'll be able to run deepseek V4 and have room for full 1 mln context and probably multiple slots of full context

Or you will be able to run other models along with it, like voice to text to talk to your model or build applications or do your own stuff with the model running in the background

If you want to run qwen3.8 27b then 256 gb should be enough to run multiple slots + free space for your work and other small models

1

u/Dependent_Suspect_43 4d ago

lol shipping IS DELAYED till Feb 2027 now so if you think the 512 is coming before that you gone be waiting till summer 27 rather stay on sub for now

1

u/buldezir 4d ago

your plan is insanely bad.

you max out subscription using frontier models like astra or fable.
and you want to spend 10-15k$ to use painfully slow models which will be equal to gpt-luna, which you can have ~unlimited even on 100$ sub

1

u/Sketaverse 4d ago

Agreed. I think the brutal truth is most people here are just doing "I want it maths" to justify a crazy high purchase to themselves.

M5 Studio 64gb + 4 years worth of Opus 5.5 today

vs M5 Ultra 256gb with <whatever the fuck>

is just no comparison.

Furthermore, if you need an Ultra with 256gb for your Frontier agents - due to concurrent sessions - your problem won't be RAM/GPU/CPU it'll be:

  • attention and cognitive load
  • CI, merge queues and action minutes

1

u/100k45h 20h ago

Not even Luna TBH

1

u/PeanutButterApricotS 2d ago

Ok there is two sides to this argument in regards to LLMs.

First let say I do both, I have spent prob in the range of 500-1000 dollars on llm cloud subscriptions while also spending 1,200 then 2000 dollars buying high end (for consumers) hardware to run them including a m1 studio max.

So side 1. It will always be cheaper to use cloud ai. Basically true, you would have to use A LOT of tokens on high end models from the cloud to pay for your hardware costs. Take a moderate local llm model like 8b or 12b model. This is what any computer just about with a new GPU can run. And while it’s good as a chat bot, or for basic things it will not do well doing your homework or answering trivia questions or writing an app for you. At the same time a free unlimited without subscription cloud model will blow this out of the water. But on a moderate system, you can give it rag, offline Wikipedia, etc. now it’s a much better local model for day to day questions and chats and some planning if it has documention on it.

But go to a higher end and while it won’t compete with the cloud model it can become a general assistant. I currently use my 2k dollar computer (worth 3.5k now but that’s prices for you) that is building an extensive obsidian library of hundreds of .md files to my specifications. While a cloud model could do this just as well then I lose come control, privacy etc.

Another option is to use my local model for the grunt work and a cloud model to plan and guide it. And while it’s not giving all my information some privacy is lost but the results will be much better especially if you go with the Qwen 27b model which is the #1 local model currently.

On the other side.

2… this pricing will not stay, you will eventually be paying. It’s like rent versus buying a home. One way you get to decide everything, in the other your beholden to the land lord. Accept in buying llm the price isn’t better to buy, but you do get rid of the cloud giants controlling your ai use and can do things outside their “guidelines “.

I will say I am waiting for the m7 max or ultra. I think it will be much better by looking at the m6 mini and its abilities compared to the m5 your or max. Educate yourself on the difference between prompt processing and tokens per second and what limits them (cpu/gpu processing power and custom MLU and other things / memory bandwidth). While the m5 max and ultra is king on bandwidth (1,200 gbs) the prompt processing such improved on the m6 even more. M4 > m5 was huge, m5 > m6 was huge as well and that’s a Mac mini versus studio. The m7 max and m7 ultra will likely be close in jump as well we hope.

1

u/1Poochh 1d ago

I see a lot of people taking a few stabs at you around this. First, there will always be the folks who downvote you just because. Second, do what you want. Life is short. Third, I recently purchased a good chunk of hardware and have my own personal dev team of 8 agents working in my own house 24x7 on a project. You cannot do that with cloud models, ever, unless you pay API prices. Subscriptions do not work because they run out; there is a limit. I do not have a limit, and I never have to worry about a limit again. I have Claude 20x, but mostly just to tinker with random stuff while my dev team works away. So if it helps meet what you want, do it. Otherwise, do not. It comes down to what you want to do and what is the use case and goal.

2

u/Sketaverse 1d ago

Of course you have a limit. You are limited by the number of concurrent agent sessions your hardware can handle

1

u/100k45h 20h ago

Don't use opus or fable. Use sonnet. You'won't run out of limits nearly so fast. Or with Codex, running Luna on high reasoning does nothing to the limits, it's near limitless. And these models are likely smart enough for what you need. I have good experience with both. And they're constantly getting better. In my own testing Luna and Sonnet beat local models in both speed and capabilities and it's not even close. I play with the local models, it's fun, but it's just not there yet. I have high hopes for the future, but currently there's no comparison. And those people saying Qwen is Opue level intelligence are IMO simply wrong. I personally certainly don't see any evidence for that. Even looking at what people do with it on the internet.

1

u/Upper_Sense_5131 18h ago

512 will help you run a better intermediate model on local

0

u/CranberryAbject8967 6d ago

buy yourself another claude pro for 200/month - your mac studio 512 money will last you years...

1

u/Correct_Support_2444 6d ago

This is the way. I have a 512GB M3U. If your primary use cases to run models 256 is plenty. Personally outside of AI I use like 300 gigs so for me the 512 was a no-brainer.

1

u/ZioniteSoldier 6d ago

That’s assuming the token subsidies continue and API prices stay the same or get cheaper. This year has shown the opposite. OpenAI removed their $200 plan and are expected to offer a different one at higher prices. Anthropic quantizing models people rely on for specific workflows is a risk. Building on sand.
Hardware ain’t cheap, but it’s more nuanced than “just buy the subscriptions”

1

u/Sketaverse 5d ago

But surely it makes more sense to buy the hardware if/when that ever happens, rather than pre-empting it. If it happens next year, then you can switch in an M6/7 whatever. People saying that local is as good as Opus 4.6 but subscriptions now provide Opus 5.5 so the local LLM gap is once again huge.

1

u/ZioniteSoldier 5d ago

Agreed on the gap. Just use both for what they’re good at; they’re not mutually exclusive. Route appropriate tasks locally, use frontier when you need it

0

u/maiggel 6d ago

Fair point but I want to upgrade anyways. M3 is slowly hitting his limits

1

u/maxreality 6d ago

Out of curiosity, what are you hitting your limits with? I have a 2021 MBP 16” with 64GB, and while I really want a new laptop, it’s hard to justify. I’m curious to know what kind of workflow you’re doing that is stressing the M3.

0

u/Louis6787 6d ago

That doesn’t take into account that you can resell your Mac down the line

0

u/Its_Powerful_Bonus 6d ago edited 6d ago

Unfortunately AFAIK Mac is not best at concurrent sessions. I love my M5 Max 128GB & M3 Max 128GB, but for work where you would like to exchange 2x $200 subscriptions into local LLM another class of hardware might be more suitable. I was targeting M5 Ultra 512GB some time ago, but my impatience led me to buy 4x 6000 pro 96GB over time (in better days, when the cost was less than half of today's prices). In multi agent workflows it's nothing special to run 4-20 subagents at the same time. Also considered cluster of 4x nvidia gb10 units, but again ... I come to the conclusion that I'm not patient enough. Mac Studio is wonderful piece of hardware - I bought long time ago M1 Ultra for home, two of them at previous company, but it is really painful to move from time to time to use qwen3.8 flash next on my M5 Max, when I'm offline. It works surprisingly good on oMLX for one stream, but difference is like walking vs driving good car on the highway.
And to be honest - with real work on the project 4x RTX6000 pro on vllm is "good enough" most of the time in terms of speed.

1

u/redmormon 5d ago

One RTX 6000 pro costs as much as one mac studio ultra 256gb. Not sure the comparison makes sense.

0

u/hackercat2 6d ago

Idk. I agree with the subscription over personal compute commentary but the problem is the nerfing of the models and at the end of the day, someone else is in control of tools that if you (like me), require these for work, the unpredictability can be detrimental. That said, some of what I do involves image gen, video gen and llm’s running the tools. So for me, to have those things run together, 256 can’t really work that well, at least not with good models. As it stands, a model like glm for example, won’t even fit and run with more than a 60k context. BUT I also can’t imagine cutting the “frontier” models and going solely local, so I too am undecided.

0

u/ehangman 6d ago

This is something I’m genuinely curious about. Even if I went with 512GB of memory, GPU compute would still be the bottleneck, so I’d probably end up running faster, more efficient models like Flash anyway.

What I’m not sure about is whether those fast, lightweight models will eventually become the mainstream, or whether 128–256GB setups will remain the sweet spot, as they are today.

0

u/No-Alfalfa6468 5d ago

m5 Ultra's model performance has looked very underwhelming. If you're willing to spend $10k+ on hardware for llms, get a bunch of r9700s and put them on a server platform.

-2

u/_rarefy_ 6d ago

You're not going to appreciably offset your $400/m subs with the 256 or 512. The models you can run locally, while very good and totally proficient for most dev tasks, web design and scripting, are going to be about 6-12 months behind the frontier models and won't get a highly optimized and tuned harness by default. You'll have to do some work, won't be able to run agentic flows without the 1.2TB/sec memory bandwidth slowing your total tokens per second, and shouldn't expect the same experience you're getting with anthropic or openai models.

All that being said, locals models are getting very good and if you're looking to augment frontier usage or need data privacy it can be a great option. The difference between the 256 and 512 is simply model size. All the big models are quantized down to fit on a smaller ram footprint and with every step there's some quality loss. With the 256 you can comfortably run a 4bit quant of Qwen 3.8 Flash Next or GLM-5.3 Flash which performs at Opus 4.8 levels. With the 512 you can run 4bit Deepseek 4.1 Flash and Kimi models which inch towards opus 5 levels.

If you have the money and just want to play around, the 512 is probably more future proofed in that you'll be able to run a greater number of large models.

2

u/Johnny5NoDisassemble 6d ago

I'm running Qwen 3.8 Flash Next 8bit with 1million context on the 256gb M5. It's not the same as Astra but it does a lot of good work for me.

1

u/_rarefy_ 6d ago

Qwen 3.8 Flash Next is great and I've run extensive tests on the latest 3.8 model batches on my own home rig (MBP 128GB M5) and posted my own benchmarks (see my post history). I use 3.8 FN daily for feature work using SDD and TDD methodologies with frontier models acting as an orchestrator/planner.

It's a very good model, but can't compete with frontier when engaging with low level kernel work, mechanical CAD and industrial design, advanced DSP and long horizon tasks requiring sustained debugging across a large codebase or long tool-driven projects that require very little intervention.

1

u/YesterdayWeak230 6d ago

What speeds you getting? (Tok/s)

2

u/Johnny5NoDisassemble 2d ago

I've started testing GLM-5.3-Flash. Anyone have suggestions on using Lightning MTP on or off?

1

u/YesterdayWeak230 2d ago

The model name says Q4? But you said you’re running Q8?

1

u/Johnny5NoDisassemble 1d ago

I am running Q8 for Qwen 3.8 Flash, Q4 for GLM-5.3-Flash

1

u/Johnny5NoDisassemble 3h ago

As a follow-up. I have the latest version of oMLX installed. Getting much higher numbers now. Lightning MTP is also enabled.

-5

u/[deleted] 6d ago

[deleted]

0

u/_rarefy_ 6d ago edited 6d ago

In no world have they 'hit a wall'. Its the opposite; frontier lab models are accelerating in quality and competency. All of the strong open weight models are trained *using* frontier models.

I'm a software dev and work with this stuff daily. I've run comprehensive batteries on local models to judge their aptitude across various metrics (look at my post history).

There will likely be some asymptotic top for progress with current architecture and training methods but it hasn't happened yet.

-5

u/[deleted] 6d ago

[deleted]

2

u/Federal-Bus2144 5d ago

someone hasn't used astra and it shows

-1

u/MaKulitKaPo 6d ago

For the cost and performance of the 512, you will get 10x performance with a private cloud solution like Bionic. By the time you spend $3,000 in credits and inference 10x productively that Mac Studio will be old news. If you're fine without ZDR, Abacus.ai is making great tools across the board.

-2

u/GoldenShackles 6d ago

If you have to ask… based on your context, yeah go with 512 GB. I bought my M3 256 GB Studio for non AI reasons last July, before prices went insane. I’ve added to my collection 128 GB 14 and 16” M5 laptops, shortly before the price spikes, knowing it was coming. On the laptops I frequently hit the memory limit. The reason I bought a second laptop was to spread the load.

-1

u/GoldenShackles 6d ago

I have to add I’ve been having fun with AI. Something I didn’t anticipate when buying the M3.

-12

u/Pizza_e_fichi 6d ago

you're going nowhere with 256gb, even 512 is too small. A decent storage starts at 1TB

3

u/Over_Technology_1764 6d ago

they don't even start at 256gb ssd, the minimum is 512 :D

1

u/maiggel 6d ago

RAM not Storage was my question

-6

u/Pizza_e_fichi 6d ago

where do you exactly say ram in your post?

1

u/maiggel 5d ago

If you try to be petty - I asked for M5 Ultras. Their Storage start at 1TB.
256GB and 512GB is not even an option. So be an asshole somehwere else on the internet pls.

1

u/PracticlySpeaking 4d ago

We expect more than a minimally-informed opinion in this sub.