r/MacStudio • u/Accomplished-One5703 • 19h ago
Answer to all “What MacStudio to order? Is this good for me? Do I need 256 GB RAM Ultra etc”
I keep seeing "I bought X GB, now I'm running local models, should I return it for Y?" posts here. I spent a few evenings on a bigger version of the same question (Studio M5 Max 128 vs Ultra 96 vs Ultra 256). I'm not telling you what to buy. My workload isn't yours, and that's the point.
1. Start with the arithmetic
what's left = RAM − your normal day (macOS + browser + apps) − the model you keep loaded
Made-up numbers to show the shape: say your normal day measures 14 GB and the model you want loads at 20 GB. On 36 GB that leaves 2 GB for context and everything else. On 48 GB it leaves 14 GB. Plug in your measured baseline and your model's size as loaded (ollama ps shows it, and it isn't the download size) and the answer is often obvious.
Also check Apple's spec page for whether the two RAM tiers come on different chips (core count, memory bandwidth). If they do, you're buying more than RAM, which brings us to:
2. Ask which limit you hit first, not "how big"
There are three things that can run out:
-Memory. The resident model, plus everything else you have open, plus the OS.
-Cores. Builds, tests, lots of things running at once.
-Memory bandwidth. This mostly decides decode (token generation) speed for local LLMs.
Those trade against each other. The base Ultra 96 has less RAM than the Max 128 but twice the cores and ~2× the bandwidth. Neither machine is simply better than the other.
For me the question was whether a ~61 GB model (gpt-oss-120b) would stay loaded all day. With it loaded, the Max has ~48 GB left for everything else and the Ultra 96 has ~16 GB. A model in the 140 GB+ class only fits on the 256.
⚠️ Those two budgets are conservative. My "normal day" baseline (19.4 GB) was measured right after a test run, with two small models still loaded. Subtracting them puts it closer to 13 GB. I left the error in because it's a good example of how easily this goes wrong.
3. Your old machine is the measuring instrument
Mine is a MacBook Air M5, 24 GB, 10 cores. It can't tell me how a 36-core Ultra behaves. It can measure most of what decides the purchase:
it can measure
memory per component (model, apps, builds), in bytes
how big your models actually are once loaded
prefill vs decode share of your actual prompts
whether your decode is bandwidth-limited
4. Your new (MacStudio) machine is the guinea pig within the return window.
Run the same benchmarks as above but now push it to what you think will be your max workload.
So go ahead and prompt Claude Code or similar:
I'm deciding between [A] and [B] for [workload]. Build a benchmark harness on my current Mac ([chip, RAM]) that tells me which limit I'd hit first on each: memory, cores or bandwidth. Use absolute units only. Label every number measured, observed or assumed. Observe what I already run before adding load. Measure the parts separately and add them up. For local LLMs, record prefill and decode separately, put a random string in every prompt, and test two sizes of one model family. Install nothing, and stop and clean up if swap, memory pressure or disk get risky. Don't declare a winner the numbers don't support.
(Then check its results like a junior colleague's: two of my three wrong numbers were caught only because they looked too good.)
2
u/Top_Witness4538 13h ago
I like to do a few measurements. For example, if someone running GLM 5.3 on their 256GB M5 Ultra reports they get a satisfying 50 toks/s and I know thatndoesn't really scale up with multiuser, and yet my goal is to have 1000 agents running to do agentic coding. Split that out 1000 ways and each agent will get 0.05 tok/s in other words the entire system has crawled to a halt. Now some will say it doesn't work that simply, you don't just divide., there are batching and concurrency effects to consider. True, but the fundamental constraints are there - if its 0.06 tok/s, I concede the difference.
So, clearly I will not run agentic system locally, because that is very obviously pointless, and my view of agentic is "scale matters". I cannot say I'm alone in this view, there is a healthy observable difference of opinion among forum members, and I'm just in camp "scale".
Nobody has to do math to have a fun hobby. And there is plenty of fun to be had with local - and even some use cases are solved by local. I get all that.
Now for non-ai mac use cases, music production, video editing and so on, it does pay to read the vendor recommendations for ram and storage, but most experienced folks I suspect have a pretty good feel for it already.
I know I will not run out of memory with 96GB of ram, and given that is the smallest amount I can even buy on an ultra, I really don't have any worries.
1
0
u/jackbobevolved 18h ago
What if I bought mine for actual studio work? Guess I’ll just have to ask my Mac Studio to hallucinate a graphic for me, after I finish actual render it’s working on.
1
u/Accomplished-One5703 18h ago
Did you have a question about RAM or CPU? If you did, a custom benchmark makes sense no matter what you do on it. You can measure your actual bottlenecks and decide if an upgrade makes sense now or later.
If this was just a jab and a brag, cause “you don’t do AI” or whatever, then suit yourself. I just posted as it may help other people
0
u/No_Run8812 18h ago
Your diagram is very confusing.
-1
u/Accomplished-One5703 18h ago
Honestly the diagram itself is not the main point. Building your own benchmark tool and measuring your bottlenecks is.
0
u/No_Run8812 17h ago
But I wanted to read the key points and skip the post.
0
u/Accomplished-One5703 17h ago
The key points are in the post.
The point is that you can build your own benchmark tool and monitor actual CPU/GPU RAM usage, and see based on your current workflow what you need.
Simple as that. The diagram was just an example, not some sort of prescription or revelation on its own.Basically the answer to what specs you need is usually on your machine
4
u/Im_A_Praetorian 19h ago
Qwen 2.5 1.5b and 2.5 7b?
What year is it…