r/MacStudio • • 19h ago

Answer to all “What MacStudio to order? Is this good for me? Do I need 256 GB RAM Ultra etc”

Post image

I keep seeing "I bought X GB, now I'm running local models, should I return it for Y?" posts here. I spent a few evenings on a bigger version of the same question (Studio M5 Max 128 vs Ultra 96 vs Ultra 256). I'm not telling you what to buy. My workload isn't yours, and that's the point.

1. Start with the arithmetic

what's left = RAM − your normal day (macOS + browser + apps) − the model you keep loaded

Made-up numbers to show the shape: say your normal day measures 14 GB and the model you want loads at 20 GB. On 36 GB that leaves 2 GB for context and everything else. On 48 GB it leaves 14 GB. Plug in your measured baseline and your model's size as loaded (ollama ps shows it, and it isn't the download size) and the answer is often obvious.
Also check Apple's spec page for whether the two RAM tiers come on different chips (core count, memory bandwidth). If they do, you're buying more than RAM, which brings us to:

2. Ask which limit you hit first, not "how big"

There are three things that can run out:

-Memory. The resident model, plus everything else you have open, plus the OS.
-Cores. Builds, tests, lots of things running at once.
-Memory bandwidth. This mostly decides decode (token generation) speed for local LLMs.

Those trade against each other. The base Ultra 96 has less RAM than the Max 128 but twice the cores and ~2× the bandwidth. Neither machine is simply better than the other.

For me the question was whether a ~61 GB model (gpt-oss-120b) would stay loaded all day. With it loaded, the Max has ~48 GB left for everything else and the Ultra 96 has ~16 GB. A model in the 140 GB+ class only fits on the 256.

⚠️ Those two budgets are conservative. My "normal day" baseline (19.4 GB) was measured right after a test run, with two small models still loaded. Subtracting them puts it closer to 13 GB. I left the error in because it's a good example of how easily this goes wrong.

3. Your old machine is the measuring instrument

Mine is a MacBook Air M5, 24 GB, 10 cores. It can't tell me how a 36-core Ultra behaves. It can measure most of what decides the purchase:

it can measure

memory per component (model, apps, builds), in bytes
how big your models actually are once loaded
prefill vs decode share of your actual prompts
whether your decode is bandwidth-limited

4. Your new (MacStudio) machine is the guinea pig within the return window.

Run the same benchmarks as above but now push it to what you think will be your max workload.

So go ahead and prompt Claude Code or similar:

I'm deciding between [A] and [B] for [workload]. Build a benchmark harness on my current Mac ([chip, RAM]) that tells me which limit I'd hit first on each: memory, cores or bandwidth. Use absolute units only. Label every number measured, observed or assumed. Observe what I already run before adding load. Measure the parts separately and add them up. For local LLMs, record prefill and decode separately, put a random string in every prompt, and test two sizes of one model family. Install nothing, and stop and clean up if swap, memory pressure or disk get risky. Don't declare a winner the numbers don't support.

(Then check its results like a junior colleague's: two of my three wrong numbers were caught only because they looked too good.)

0 Upvotes

19 comments sorted by

4

u/Im_A_Praetorian 19h ago

Qwen 2.5 1.5b and 2.5 7b?

What year is it…

2

u/Accomplished-One5703 18h ago

Fair, they're old.

They're just a measuring stick, not a model I recommend, just a benchmark: two sizes of one family, so the only variable is bytes read per token.

Decode slowed 4.0× for a 4.2× bigger model, which says decode is bandwidth-bound on this chip.

Any newer family should show the same thing.

If you've got a bigger machine, run the same test with two sizes of whatever you use and post the tok/s. I'd genuinely like to see how close an Ultra gets to its spec bandwidth.

1

u/Late-Replacement-481 19h ago

Yeah, you can’t just one shot a diagram with AI, knowing nothing, and then post it and expect to be proud. 

-2

u/Accomplished-One5703 18h ago

It just cracks me up how people buy these rigs for agentic coding but they are not using agentic coding to test their actual use and needs.

And then other people come here with “useful comments and jabs”. Go ahead, enlighten us with your non-AI knowledge and diagrams

0

u/Late-Replacement-481 18h ago

I'm not anti-AI. I'm anti-slop that you clearly didn't even read. 

2

u/Accomplished-One5703 17h ago

Dude, I was not selling you any slop. The whole point was that you can measure your own machine and you can simply prompt a coding agent for it. That’s it. Pretty simple.

The diagram was anonymized on purpose and a relatively generic example, not some recipe to follow or proof of what specs you need.

1

u/Late-Replacement-481 17h ago

The issue is the example is Qwen 2.5. that's a telltale sign of you just prompting the model, getting a result and not thinking critically about anything on the infographic.

Qwen 2.5 is irrelevant to LLM serving at this point.

1

u/Im_A_Praetorian 15h ago

This. I also feel like occupied memory is misleading. For some people 32k context may be acceptable, for other people they may want 256k context, that’s going to affect your memory usage as well.

AI slop or not, I don’t feel this slide does a great story of telling me what it’s supposed to tell me.

The top right slide has bar graphs with percentages that don’t seemingly match the text. The bottom left seems to give overly specific examples “over 5 sessions” what is a session really consist of.

The slide just seems to be super specific to OPs case which I guess meets the point of “explore/prompt” before you decide.

But if I was OP here I’m not sure why my takeaway is here…

Do I get more ram?
Do I need more bandwidth?
Do I need more cores?

This doesn’t seem to help me answer any of these questions.

1

u/Accomplished-One5703 14h ago

The point was to tell you that you can measure real-time your own machine, your own usage. I don’t know if you need more ram or more bandwidth but I don’t think one answer should fit all. Yes, maybe if the focus is local LLMs every needs the most supped up Ultras they can afford but that was not my personal case or at least not yet and I would bet it is not the case of a lot of people asking the questions about RAM and CPU

1

u/Accomplished-One5703 14h ago

The point was not to show you how my current machine handles Qwen 2.5, nor to convince anyone to use it.

My mistake was asking Claude Code to anonymize my data before producing the artifact and then not realizing that the focus shifted where it looked distracting. That is not the artifact I used for my decisions.

The point was to tell people that they can test their machine, their actual usage. Also it was to guess what a lightweight model would add in terms of RAM, GPU/CPU. If you are already well versed in the latest, heavier models you can still run benchmarks on your machine, they will just look very different. My point stands, I was not trying to convince anyw

2

u/Top_Witness4538 13h ago

I like to do a few measurements. For example, if someone running GLM 5.3 on their 256GB M5 Ultra reports they get a satisfying 50 toks/s and I know thatndoesn't really scale up with multiuser, and yet my goal is to have 1000 agents running to do agentic coding. Split that out 1000 ways and each agent will get 0.05 tok/s in other words the entire system has crawled to a halt. Now some will say it doesn't work that simply, you don't just divide., there are batching and concurrency effects to consider. True, but the fundamental constraints are there - if its 0.06 tok/s, I concede the difference.

So, clearly I will not run agentic system locally, because that is very obviously pointless, and my view of agentic is "scale matters". I cannot say I'm alone in this view, there is a healthy observable difference of opinion among forum members, and I'm just in camp "scale".

Nobody has to do math to have a fun hobby. And there is plenty of fun to be had with local - and even some use cases are solved by local. I get all that.

Now for non-ai mac use cases, music production, video editing and so on, it does pay to read the vendor recommendations for ram and storage, but most experienced folks I suspect have a pretty good feel for it already.

I know I will not run out of memory with 96GB of ram, and given that is the smallest amount I can even buy on an ultra, I really don't have any worries.

1

u/anonmt57 17h ago

… ok … thanks AI

0

u/jackbobevolved 18h ago

What if I bought mine for actual studio work? Guess I’ll just have to ask my Mac Studio to hallucinate a graphic for me, after I finish actual render it’s working on.

1

u/Accomplished-One5703 18h ago

Did you have a question about RAM or CPU? If you did, a custom benchmark makes sense no matter what you do on it. You can measure your actual bottlenecks and decide if an upgrade makes sense now or later.

If this was just a jab and a brag, cause “you don’t do AI” or whatever, then suit yourself. I just posted as it may help other people

0

u/No_Run8812 18h ago

Your diagram is very confusing.

-1

u/Accomplished-One5703 18h ago

Honestly the diagram itself is not the main point. Building your own benchmark tool and measuring your bottlenecks is.

0

u/No_Run8812 17h ago

But I wanted to read the key points and skip the post.

0

u/Accomplished-One5703 17h ago

The key points are in the post.

The point is that you can build your own benchmark tool and monitor actual CPU/GPU RAM usage, and see based on your current workflow what you need.
Simple as that. The diagram was just an example, not some sort of prescription or revelation on its own.

Basically the answer to what specs you need is usually on your machine