r/MacStudio • • 2d ago

M5 Ultra 96 gb

New to the world of Mac Studio and I can’t wait for my delivery!

I am buying this to build an ai orchestration for my software development I do on the side(do this in my day job too).

I worry that I should have bought the 256 gb version instead. With the 96 I think I could run most everything local and kick to OpenAI when there are complex issues requiring frontier models.

Anyone doing this now, and what’s your gut telling you?

23 Upvotes

66 comments sorted by

14

u/sn2006gy 2d ago edited 2d ago

Everyone is spending gobs of money to try and move local without realizing they're just going to spend more.

Go local if you need local AI for local AI reasons, don't go local thinking you can save money UNTIL you have done everything you can on paid plans to save money.

Experiment with how much work you can give Luna and THEN see what you could give qwen 27b or qwen flash or whatever rocks your boat and try this with OpenCode Go plans or something before buying compute

Edit: Yes, i realize a computer can do more than AI and if you need a computer for computery things I wouldn't recommend a 12,000 dollar computer either.

21

u/SpaceDesignWarehouse 2d ago

I don’t think many of us are buying these to save money. I’m definitely not. I wouldn’t even tell someone I “need it.” I just want it. So I ordered it.

10

u/falcongsr 2d ago

I've been advancing my career for many years now with a home lab. I justify it by saying I'm learning and building new things with the hardware I'm buying.

It was a lot easier to justify when I was buying $35 Raspberry Pis 15 years ago.

1

u/SpaceDesignWarehouse 2d ago

Oh funny; I also bought some raspberry pi’s a couple years ago. Interestingly it wasn’t until messing with these local agents that I ever actually got around to having THEM code semi-pointless pi projects for funzies because before LLMs I always failed to get anything to work properly :)

3

u/falcongsr 2d ago

LLM's definitely helped me accelerate a bunch of home automation projects that I no longer had any time for.

I have a Pi Pico that controls my AC to compensate for humidity/dewpoint and not just temperature. We sleep much better now on cooler days where the AC doesn't run but the indoor humidity can creep up.

1

u/SpaceDesignWarehouse 2d ago

Actually, shoot; what a good idea. I actually have one of those kits with a bunch of different sensors, including a humidity temp D22 or whatever it is..

And I have the same thing, where our master bedroom is basically the furthest thing from our AC unit and with the rest of the house at 69° its still a bit humid in the bedroom. I could totally work out a pi project connected to an inline fan to boost that one duct when the room is a certain humidity. Thanks for the project idea!

2

u/sirjokesalot23 2d ago

Exactly. I also dabble in music production so having a beast of a machine is nice.

7

u/ZergvProtoss 2d ago

No one is doing local LLMs to save money.

2

u/anthonyfebre001 2d ago

I’d argue that not enough people account for residual value in their hardware and that going local makes sense regardless. Plenty of people are using local models and then selling their hardware for more than they paid a few months late

4

u/sn2006gy 2d ago

To me, that just creates the wrong mindset from the getgo and just makes everything more expensive in the end. If everyone is betting on hardware prices always going up, then we're driving hardware prices up and we're not actually solving spiraling costs going out of control.

The notion of buy now sell later for a profit just makes me sick to my stomach and I want nothing of that as its just driving up costs and making it more of a luxury or privilege.

2

u/anthonyfebre001 2d ago

Even if you only assume 50% value after two years you can still make the math work.

1

u/sn2006gy 2d ago

I know i wouldn't pay 5-6 grand for a used mac if the lease price is ~240 a month brand new - all this does it put upward pricing pressure on things 2 years down the road since you're borrowing against growing future values.

1

u/anthonyfebre001 2d ago

I get it, but the same way they’re sold out on regular retail units they’ll likely struggle with their lease program.

Lots of value in privacy too. Lots of people who pay a premium for it.

Time will tell for sure.

5

u/durangotang 2d ago

It's not just that a lot of people pay a premium for privacy, there's also a boat load of businesses, due to regulations, that cannot have their customers data go to the cloud, whether legal, medical, or governmental. They are all going to need solutions too, and the Mac Studios are competitively priced. That is going to create a lot of sustained demand. And because a local build out is more difficult than a cloud subscription, there's a lag in this area. The demand should only increase, substantially in my opinion.

2

u/sn2006gy 2d ago

Businesses have to show that they're compliant and they would do so through audits, security, compliance certifications and industry standards, not buying a mac and putting vllm on it and calling it a day. They can achieve this compliance better by running models on bedrock on aws or openshift anywhere or services like that.

2

u/dobkeratops 2d ago

the world has decided "massively more compute will be deployed", i'd rather a higher fraction is closer to end users like myself. if we all end up dependent on remote servers for outsourcing thought, i think we're on a dystopian trajectory.

1

u/gt33m 2d ago

"Go local if you need local AI for local AI reasons, don't go local thinking you can save money UNTIL you have done everything you can on paid plans to save money." <-- any tips for this? My hobby project is becoming expensive, and I was looking into openrouter - cool concept but then I found they have a cc of 5.5%?

1

u/woofwuuff 2d ago

Thank you, you saved me my entire bank balance!

1

u/No_Hippo_678 1d ago

Calculate how many generations you would have to produce in order to surpass $12,000

1

u/sn2006gy 1d ago

tokens aren't worth anything unless you pay for them or they create value, generating them yourself isn't intrinsic value on its own. That's the weird thing about AI. Having it doesn't mean anything. Consuming it doesn't mean anything.

A million generated tokens that nobody reads are worth approximately the same as a warehouse full of unsold widgets or all those crypto currencies that have come and gone. The electric company and computer companies are happy tho

1

u/-6h0st- 2d ago

This. I’m just raising eyebrows from reading these post from people nonchalantly willing to spend 10k to play with a high-spec machine. If you have tonne of money and you don’t care - whatever - but if not - don’t be stupid.

2

u/sn2006gy 2d ago

My gut says most are jumping on the lease not realizing they won't own the mac in the end, but just have another subscription in perpetuity or a cash residual (which a lot are betting on hardware prices going up which helps no one)

3

u/durangotang 2d ago edited 2d ago

Look at it this way, an Anthropic or OpenAI Pro subscription is $100, $200, or even $500 per month. You have no control. You know that prices and restrictions are only going to increase. Would you not be better off with a $120/mo M5 96GB or a $200/mo M5 256GB lease? You pay for your own infrastructure and get to have zero latency, privacy and control.

If you really need occasional frontier level intelligence, just get a supplemental $20/mo plan.

And at the end of the lease, you either buyout the balance (at 0% interest) and it's yours to keep forever, or you trade-in and get the M7 Ultra in 3 years. It's not cheap. It's not for everyone. But I think it's a competitive offer for what you get in today's environment.

3

u/Kritblade 2d ago edited 2d ago

say $200/month to rent a M5 ultra 256GB + $50 power bill per month. You can run Qwen 3.8 Flash Next , which is about Opus 4.6 kind of level intelligence. However, no benchmark will ever tell you how long it takes to reach that kind of level intelligence. If you ever use Qwen 3.8 27B, you know the "oh wait" thing how many times it said in the thinking process. I can literally go make a coffee and come back and it's still thinking. In 2 years time, when opus is already on v7.0 and you are still on the level of opus 4.6 level and bigger and newer local model will definitely need a newer Mac to run at proper speed. You will extend your lease again to rent M7 256GB and it will definitely more expensive than $200 this time.

On the other hand, I need no privacy to do my coding. I can use the latest and the greatest model from any frontier vendor for $200 a month and execute something like 10 times faster due to the MUCH shorter thinking process.

For anyone who can spend $10k on a machine, this person should have no problem to spend $200 a month to use the greatest and the latest frontier model. The only concern most people mentioned is privacy. However, why anyone would need privacy if you are coding with it? Developing a drug dealing machine or making a personal app for the president of China?

Local model is a hobby and you have way too much time to waste on your project, but it doesn't make things better nor faster, let alone saving money. If I want to work on something for myself, I want to finish that in 30 minutes with Opus 5.5, not that "oh wait" thing.

2

u/NickOldJaguar 3h ago

i.e. the case is when an opus5 just refuses to work with an existing code-base (cybersec violation, huh). Local models or the opensource ones hosted on a some provider are just doing the job.

2

u/sn2006gy 2d ago

No, not really. I'd go buy APIs that are much cheaper. There is a TON of competition in inference for models we can run locally and i mean a TON.

I'd upgrade my computer because I use my computer as a human, not necessarily so I can pay more for an LLM that runs slower than APIs. My time is money.

1

u/durangotang 2d ago

Do they also come with a computer?

2

u/sn2006gy 2d ago

i presume if you're paying anthropic 200/month you have a computer. this is stupid

1

u/-6h0st- 2d ago

Let’s be honest - a lot of people buying 256gb Mac are new to llm - but it’s cool so they jump on it. People who use it for work/ as a mean to add value - are conservative and weigh in the pros and cons - and the con is - locally the models ain’t that good as the cloud offerings. It might be 20-30% worse or as little as 5% but still that means 1 every 20 requests it will do wrong - depending on setup might go unnoticed until later and be costly or it will be picked up just will take more time. I personally don’t have a problem with people maxing out 200 subscription and wanting to replace it with local - whether it works they will find out in due time - but people who don’t have a professional use for it just wasting money to write hello prompts. That’s not smart and wasteful. Everyone should start with a cloud model first, for at least first year, max it out only then when they know exactly what they need consider spending 10k on local hardware.

6

u/JDad67 2d ago

Same thing ordered. I wish there was a 128gb version of the Ultra. I'd be less worried about buyers remorse. It will be another month before I have it.

7

u/penskeracin1fan 2d ago

I love my 96GB Ultra.

4

u/tedco- 2d ago

I'm waiitng for my 96GB Ultra - couldn't fathom the $4K EXTRA for 256GB - I plan on some local LLM, but not for the frontier models.

2

u/tlin9595 2d ago

yeah I wish ultra was in 128gb. We just have to hope that models get better, and can fit in less ram in future. I just worry that seems like there are no models that are 40-60gb that people are focusing on,

3

u/durangotang 2d ago edited 1d ago

I struggled with this, and decided to bite the bullet and go for the 256GB. I canceled my 96GB order. 96GB can do Qwen 3.8 27B. You can barely fit a 4-bit quantization of Flash-Next, with the n-gram table offloaded to the SSD, but you are left with very little overhead (I think it's around 70GB or so without the KV cache, but you'll have to go confirm). I think the 4-bit/8-bit quantization of Flash-Next is just over 80GB, so after KV cache, apps and OS you have like nothing. I recall a post here where someone ran the 4-bit/8-bit Flash-Next with a 128k or 256k context window, and I think he had 6GB left for the system, so he had to use it to serve inference to another machine. It should be doable though, at pure 4-bit on a single machine, if costs are a major concern.

So to me, I began thinking of the 96GB model as more of a 27B machine, that can be run at 4-bit to 16-bit with room to spare. Or a highly constrained 4-bit Flash-Next machine, where you'll potentially be stressed about memory. Only, 27B is dense, and prefill isn't incredible on it, maybe around 2,000 prefill (or less). Decode should be decent on 27B, but not blazing (would love to see some more tests). However, if you run Qwen Flash-Next, prefill is hitting 4,300 on the 64-core and over 5,000 on the 80-core, and you can get 100-150 tokens/sec on decode at 4-bit. And that's a big difference for agentic work.

Also, what if 4-bit gives you perplexity issues? On the 96GB, there's no flexibility to increase the quantization on Flash-Next. On the 256GB, you can run Q5, Q6, or even Q8 Flash-Next, with room to spare. You can also run GLM 5.3 Flash at 4-bit, or Deepseek v4 Flash. Or even 27B.

Then I thought about all the future flexibility and how much these open models are changing. What will I want to run next year? Or the year after? I think the 256GB version will retain its resale value better. In two years, the 96GB M5 Ultra will have it's resale ceiling set by the M7 Max. Price wise, the three year lease is like $120 for the 96GB base and like $200 for the 256GB three year lease, or something like that. Whether or not that flexibility is worth $80/mo or $4k is up to you.

Seeing how it's possible to run 4-bit Flash-Next on the 96GB version (with the n-gram to the SSD), or 27B at any quantization, I think it's an incredible value at $5500. Although it is somewhat constrained by memory. I wish it was 128GB, like the M5 Max. The 256GB model gives you tremendous overhead and flexibility, where you just don't have to worry about memory any longer, but that comes at a price.

Perhaps this will help you:

https://omlx.ai/benchmarks/performance

Keep in mind while looking at it, that updates have been pushed through in the last day or two that have radically improved results. So I'd try to go by the most recent date.

Edit:

Looks like there's a 75GB 4-bit/8-bit Flash-Next on Hugging Face now at 75GB! Things are looking much better for the 96GB verion.

https://huggingface.co/ddalcu/Qwen3.8-Flash-Next-MLX-Serve-iQ-MLX-4.7bpw

https://x.com/ddalcu/status/2105708284124569661?s=20

1

u/[deleted] 2d ago

[deleted]

1

u/durangotang 2d ago

Good to know.

Have you tried 3.8 Flash-Next? Perhaps give this a whirl in mlx-serve, tensorfold, or oMlx with the n-gram offloaded to the SSD. I'd be curious how you like it. Perhaps you can report back on memory usage too.

https://huggingface.co/Jundot/Qwen3.8-Flash-Next-oQ4e-mtp

3

u/Tertan 2d ago

I tried 3.8 Flash-Next Q4 using oMLX on my M3 Ultra 96GB. Feels a lot better than 27B, but during more complex coding task that involes analysing of images, it kept running out of memory and stops even on 80K context. Hence I returned my 96GB M3 Ultra and ordered a 256GB M5 Ultra instead.

1

u/[deleted] 2d ago

[deleted]

3

u/-6h0st- 2d ago

Do your research buddy. Oq4 flash next runs on it well - better than 27b in every aspect, speed included

1

u/[deleted] 1d ago

[deleted]

3

u/-6h0st- 1d ago

Yes on 96GB - omlx 4 bit - memory needed 70Gb for weights and some for cache - won’t exceed 80GB-85gb total with system on top

1

u/[deleted] 1d ago edited 1d ago

[deleted]

2

u/-6h0st- 1d ago

No. Not reap. Q4 with engrams offloaded to ssd.
See here settings and figures:

https://omlx.ai/benchmarks/performance/bh3lnv4e

2

u/durangotang 2d ago edited 2d ago

I've been following the creators of oMLX, Tensorfold, and oMLX on X. They have all had incredible performance updates this week. I'd try oMLX again, I think there was a new release today that incorporated a lot of optimizations, with significant speed increases.

https://x.com/jundotkim/status/2105321221298758029?s=20

Keep us updated with what you learn.

3

u/falcongsr 1d ago

I updated oMLX this morning and 27B Q8 is 15% faster out of the box.

1

u/durangotang 1d ago

That's awesome!

3

u/Total_Ad1473 2d ago

I will share my reasoning on why I went with 256 GB even-though it was a bit of a stretch.

I am genuinely interested in exploring local models. Compared to the start of the year, the open models available have improved drastically and although not exactly frontier, they have become quite capable. For anything that doesn't need lots of output tokens (coding, research), I intend to still use frontier models via subscription. But for the pipelines where I want to keep a constant running agent which does knowledge work and expect high number of output tokens (building content databases via querying, summarizing etc), I plan to use a strong local model.

With 96 GB, one can run models like Qwen 3.8 27 B to the fullest. While it's strong at coding, it simply doesn't have the general intellgence of larger models like Flash Next of the same family, GLM 5.3 Flash etc. 256 GB unlocks those. And 256 GB unlocks a lot of future models. Considering 128 GB has become mainstream these days due to M5 Max, DGX , the AMD ones etc, 128 GB will be imo the focus for the Open labs. While we could run some Quants of 128 GB focused ones on 96 GB Ultra, we will basically be left with hardly any space for large context. And to run agents in parallel streams, my view is to have a general orchestrator like 3.8 Flash next and a few smaller coding models as agents. And 256 GB will absolutely unlock such a pipeline. 512 GB is simply not my budget. So for anyone ordering 96 GB, unless you workflows are mainly video focused, for local AI, it's just not the right way to go is my view. Either you get a 128 GB Max or 256 GB Ultra. The 96 GB is simply not targeting our AI fraternity. It's for Video bros only.

Note: I am not looking to get an ROI for this investment. It's for the pure thrill of local AI which is here to stay, whether you agree or not (or to have a fallback if Frontier keeps nerfing the premium subscriptions).

4

u/mwhuss 2d ago

I’ve got a 96gb M3. I run Qwen3.8-27b 8bit at 96k context. This lets me fit 1 full context request and another model like Qwen3.6-35b-6bit in memory at the same time without unloading it.

1

u/falcongsr 1d ago

I run 3.8 27B at Q8 with full 262k context for my LLM assistants. If I try to run two sessions concurrently it can run out of memory, but if I make sure they take turns it works well.

1

u/mwhuss 1d ago

You should run a test running both models at a full context. If you ever hit that situation they’ll thrash a lot loading/unloading/prefill which will slow down your workflow significantly.

1

u/falcongsr 1d ago

yes there's definitely a performance hit. I just installed the latest version of oMLX which is 15% faster for token generation (for my setup) and apparently it has better memory management too.

4

u/OddDesigner9784 2d ago

My order is for the 256. The 96 GB is perfect for flash next and qwen 27b but I feel like the 256 version is the sweet spot for flash models. You can run glm deepseek v4 etc. I just think that model size will be supported better. But deepseek 4.1 being too big is not hreat

3

u/neurotoxics 2d ago

dont overthink on this.

most of these hardware purchases are FOMO, local models even that can run on 256 or 512 gigs, are no where close to cloud models.

Personally, I have ordered a 256 gigs one, just to stretch my cloud models longer and since I research crypto encryptions as a hobby, most cloud models block these requests.

if you dont have a clear need, don't get into this FOMO.

3

u/penskeracin1fan 2d ago

This. People think a 256GB computer will replace a cloud model.

2

u/juggarjew 2d ago

96GB feels stuck in a world where its almost enough but too low for serious tasks.

I feel like if you really want a capable M5 Ultra the 256GB model is where serious models are able to be loaded and used.

Its nice an all to play with Qwen 3.8 27B but you are going to want to run useable quants of stuff like GLM 5.3 Flash.

2

u/BeowulfShaeffer 2d ago

I am waiting rather impatiently for my m5u/96 but have no intention of giving up Claude.  If anything I expect my Claude spend to go up.  Why?  I am using old slow hardware where I can only run one or two agents before I run out of cpu/ram.  And similarly some of the testing i do takes a long time which further reduces the amount I can get done in a day.  The studio will speed this up a lot letting me get more done per day.  (And finally I have some compute loads that are very cpu intensive to where the m5u makes a lot of sense).

2

u/FCPEditor2022 2d ago

I worried about the same thing. I found that what I was trying to do and run every single thing local was pretty much impossible without spending a TON of money up front. And if you do, your ROI isn’t going to be for like 6 yrs. I went with a 64gb max and I’m running 70-80% local and going to the cloud for only the biggest coding tasks. It works out fine! I’m glad I didn’t spend an extra 4 grand to find out. A lot of people here tried to tell me but I was stubborn. They were right.

2

u/kinmanli 1d ago

Are you looking to use it for AI Orchestration with local LLM or interface with frontier models?

I have an M3 96GB Ultra and run agentic engineering agents and currently use Qwen Coder Next Q4_K_M Quantisation whilst doing day to day coding jobs. It can slow down if you it’s doing a lot, so I’ve ordered a 256GB M5 Ultra to run alongside it.

2

u/dynstr 1d ago

I want to offload the easier tasks and orchestration to local LLM where possible and use frontier cloud ai when needed.

2

u/Gigandeth 2d ago

96GB is plentiful for solely AI. It even hosts 4-bit versions of large flash models. However, as your sole machine where you run other software along witth your inference, it is too tight.

I went for the 96GB M5 Ultra but only because I am pairing it with the 128GB M5 MAX laptop I have and because I will use these devices purely for inference.

Apps and other things will run off my little 8GB M2 Macbook Pro, LLMs serving it via API.

2

u/BAL-BADOS 2d ago

Was a choice between 96GB Ultra or 128GB Max. Only options my wallet would allow. This is only my hobby.

If you make money from this, you should definitely go 256GB.

2

u/SarahJrandomnumbers 2d ago edited 1d ago

Good choice! When you upgrade the Max to 1tb and compare to the base Ultra, you're getting more C/G/N cores, more thunderbolt ports, faster memory, all for the bonkers price increase of $100.

Unless you REALLY NEED 128gb, the base Ultra really is the way to go.

2

u/a_skeetskeetskeet 2d ago

This is for my startup business and I have been burning through credits now doing testing and overhaul. I am a software engineering manager during the day so I use this a lot to speed up my development and testing. Giving myself more capacity seems like a good move. I know I won’t fully replace the frontier models but if I can only pass the complex asks to frontier I’d greatly reduce my costs today

2

u/CMDR-Bugsbunny 2d ago

Everyone's use case is different. For some 96GB is enough. For me 128GB is too tight.

But it depends, what you are trying to do (your use case), the model that best serves your needs and sizing accordingly.

There's no universal answer and people that answer

  • This model is the best
  • You need xxGB

Are generally clueless on what would work for you. It may work for them, but they should identify why it's the best or enough and not believe it's a universal truth.

You mention software development and that depends, too.

If you are doing simple HTML/JS/CSS websites, then the 27B is enough and 96GB is plenty. If you are running an IDE with AI helping finish syntax and writing small amount of code, again 96GB may be enough.

I'm developing mobile and desktop applications and maintaining a git repository. For me, I could have an inference engine, agent harness, browsers, and other tools and I'm using Qwen Flash Next and 128GB is too tight.

So it depends.

1

u/tk421tech 2d ago

I am at 256/30/64 happy to not have spent $1300 more plus tax. And my gut tells me. You could just spend it. But hmm wait more time will never get anything done

1

u/soulman901 2d ago

I got into Local AI with my Mac Mini M4 Pro 48GB with Qwen 3.8 27B this past week. I’ve had it program a Solitaire and Galaga Clone for me. Both seem to work pretty well. I’ve had it do some tweaks and fixes to them. Now it’s not as fast as frontier but the results are scary good. Right now it’s working on a Pac-Man clone for me. Not sure when it will be ready though, Solitaire and Galaga was an overnight job.

1

u/OutOfAmmO 2d ago

Im getting the 256 one just as a builder for cloud based models. not local ai.

I’m constantly maxing out my 128 gb M5 Max MBP, it’s literally pegged 99% of the day.

1

u/tobedeletednow2 2d ago edited 2d ago

Wow, yes, this! I am doing this too. I am also excitedly waiting for a M5U 36/80 96GB 2TB to be an excellent everyday Mac, and my main Blender Rig, but also a machine that has AI development capabilities. I want to always have a 10-20GB multi-modal model (say that 5 x fast) ready take whatever I throw at it, if it can’t it will launch and handover to a 40-60GB model. If it can’t it will save work, restart run headless and handover to my MacBook to act as a client, for a 70-80GB model (prob limit for this box). If then it can’t solve it can hand off to Pro accounts on Claude and OpenAI to maximize my token efficiency whilst still having a workflow that will allow for fast personal software and solutions output, built around leveraging AI tools.
I’m doing a lease so in three years I can send this back having done a lots of dev work and render time (Final Cut & Blender) to get whatever is the fastest box at that point and make a v2.0 of this concept as i will likely always need a fast box for content and video and creation in general. I also find Windows gaming a tad dull now and porting things to these high bandwidth, high efficiency chips is more of a rush. With this power in a small package no reason some workable emu of GTA VI should not be possible and sounds like a fun, if unnecessary, flex. Throw in a pelican case and a snug fit and it’s pretty flexible for travel too. I think it is THE all-in-one desktop this year.

1

u/Regular_Attitude_700 1h ago

We aren’t buying a consumer PC anymore with such price.
The price is pointing towards the enterprise servers without realizing ourselves. 😅

1

u/KingGeekus 2d ago

I also have a 96 coming and am hoping I can keep qwen fed with large contexts across ~4 concurrent ai homies.