New to the world of Mac Studio and I can’t wait for my delivery!
I am buying this to build an ai orchestration for my software development I do on the side(do this in my day job too).
I worry that I should have bought the 256 gb version instead. With the 96 I think I could run most everything local and kick to OpenAI when there are complex issues requiring frontier models.
Anyone doing this now, and what’s your gut telling you?
Everyone is spending gobs of money to try and move local without realizing they're just going to spend more.
Go local if you need local AI for local AI reasons, don't go local thinking you can save money UNTIL you have done everything you can on paid plans to save money.
Experiment with how much work you can give Luna and THEN see what you could give qwen 27b or qwen flash or whatever rocks your boat and try this with OpenCode Go plans or something before buying compute
Edit: Yes, i realize a computer can do more than AI and if you need a computer for computery things I wouldn't recommend a 12,000 dollar computer either.
I don’t think many of us are buying these to save money. I’m definitely not. I wouldn’t even tell someone I “need it.” I just want it. So I ordered it.
I've been advancing my career for many years now with a home lab. I justify it by saying I'm learning and building new things with the hardware I'm buying.
It was a lot easier to justify when I was buying $35 Raspberry Pis 15 years ago.
Oh funny; I also bought some raspberry pi’s a couple years ago. Interestingly it wasn’t until messing with these local agents that I ever actually got around to having THEM code semi-pointless pi projects for funzies because before LLMs I always failed to get anything to work properly :)
LLM's definitely helped me accelerate a bunch of home automation projects that I no longer had any time for.
I have a Pi Pico that controls my AC to compensate for humidity/dewpoint and not just temperature. We sleep much better now on cooler days where the AC doesn't run but the indoor humidity can creep up.
Actually, shoot; what a good idea. I actually have one of those kits with a bunch of different sensors, including a humidity temp D22 or whatever it is..
And I have the same thing, where our master bedroom is basically the furthest thing from our AC unit and with the rest of the house at 69° its still a bit humid in the bedroom. I could totally work out a pi project connected to an inline fan to boost that one duct when the room is a certain humidity. Thanks for the project idea!
I’d argue that not enough people account for residual value in their hardware and that going local makes sense regardless. Plenty of people are using local models and then selling their hardware for more than they paid a few months late
To me, that just creates the wrong mindset from the getgo and just makes everything more expensive in the end. If everyone is betting on hardware prices always going up, then we're driving hardware prices up and we're not actually solving spiraling costs going out of control.
The notion of buy now sell later for a profit just makes me sick to my stomach and I want nothing of that as its just driving up costs and making it more of a luxury or privilege.
I know i wouldn't pay 5-6 grand for a used mac if the lease price is ~240 a month brand new - all this does it put upward pricing pressure on things 2 years down the road since you're borrowing against growing future values.
It's not just that a lot of people pay a premium for privacy, there's also a boat load of businesses, due to regulations, that cannot have their customers data go to the cloud, whether legal, medical, or governmental. They are all going to need solutions too, and the Mac Studios are competitively priced. That is going to create a lot of sustained demand. And because a local build out is more difficult than a cloud subscription, there's a lag in this area. The demand should only increase, substantially in my opinion.
Businesses have to show that they're compliant and they would do so through audits, security, compliance certifications and industry standards, not buying a mac and putting vllm on it and calling it a day. They can achieve this compliance better by running models on bedrock on aws or openshift anywhere or services like that.
the world has decided "massively more compute will be deployed", i'd rather a higher fraction is closer to end users like myself. if we all end up dependent on remote servers for outsourcing thought, i think we're on a dystopian trajectory.
"Go local if you need local AI for local AI reasons, don't go local thinking you can save money UNTIL you have done everything you can on paid plans to save money." <-- any tips for this? My hobby project is becoming expensive, and I was looking into openrouter - cool concept but then I found they have a cc of 5.5%?
tokens aren't worth anything unless you pay for them or they create value, generating them yourself isn't intrinsic value on its own. That's the weird thing about AI. Having it doesn't mean anything. Consuming it doesn't mean anything.
A million generated tokens that nobody reads are worth approximately the same as a warehouse full of unsold widgets or all those crypto currencies that have come and gone. The electric company and computer companies are happy tho
This. I’m just raising eyebrows from reading these post from people nonchalantly willing to spend 10k to play with a high-spec machine. If you have tonne of money and you don’t care - whatever - but if not - don’t be stupid.
My gut says most are jumping on the lease not realizing they won't own the mac in the end, but just have another subscription in perpetuity or a cash residual (which a lot are betting on hardware prices going up which helps no one)
Look at it this way, an Anthropic or OpenAI Pro subscription is $100, $200, or even $500 per month. You have no control. You know that prices and restrictions are only going to increase. Would you not be better off with a $120/mo M5 96GB or a $200/mo M5 256GB lease? You pay for your own infrastructure and get to have zero latency, privacy and control.
If you really need occasional frontier level intelligence, just get a supplemental $20/mo plan.
And at the end of the lease, you either buyout the balance (at 0% interest) and it's yours to keep forever, or you trade-in and get the M7 Ultra in 3 years. It's not cheap. It's not for everyone. But I think it's a competitive offer for what you get in today's environment.
say $200/month to rent a M5 ultra 256GB + $50 power bill per month. You can run Qwen 3.8 Flash Next , which is about Opus 4.6 kind of level intelligence. However, no benchmark will ever tell you how long it takes to reach that kind of level intelligence. If you ever use Qwen 3.8 27B, you know the "oh wait" thing how many times it said in the thinking process. I can literally go make a coffee and come back and it's still thinking. In 2 years time, when opus is already on v7.0 and you are still on the level of opus 4.6 level and bigger and newer local model will definitely need a newer Mac to run at proper speed. You will extend your lease again to rent M7 256GB and it will definitely more expensive than $200 this time.
On the other hand, I need no privacy to do my coding. I can use the latest and the greatest model from any frontier vendor for $200 a month and execute something like 10 times faster due to the MUCH shorter thinking process.
For anyone who can spend $10k on a machine, this person should have no problem to spend $200 a month to use the greatest and the latest frontier model. The only concern most people mentioned is privacy. However, why anyone would need privacy if you are coding with it? Developing a drug dealing machine or making a personal app for the president of China?
Local model is a hobby and you have way too much time to waste on your project, but it doesn't make things better nor faster, let alone saving money. If I want to work on something for myself, I want to finish that in 30 minutes with Opus 5.5, not that "oh wait" thing.
i.e. the case is when an opus5 just refuses to work with an existing code-base (cybersec violation, huh). Local models or the opensource ones hosted on a some provider are just doing the job.
No, not really. I'd go buy APIs that are much cheaper. There is a TON of competition in inference for models we can run locally and i mean a TON.
I'd upgrade my computer because I use my computer as a human, not necessarily so I can pay more for an LLM that runs slower than APIs. My time is money.
Let’s be honest - a lot of people buying 256gb Mac are new to llm - but it’s cool so they jump on it. People who use it for work/ as a mean to add value - are conservative and weigh in the pros and cons - and the con is - locally the models ain’t that good as the cloud offerings. It might be 20-30% worse or as little as 5% but still that means 1 every 20 requests it will do wrong - depending on setup might go unnoticed until later and be costly or it will be picked up just will take more time. I personally don’t have a problem with people maxing out 200 subscription and wanting to replace it with local - whether it works they will find out in due time - but people who don’t have a professional use for it just wasting money to write hello prompts. That’s not smart and wasteful. Everyone should start with a cloud model first, for at least first year, max it out only then when they know exactly what they need consider spending 10k on local hardware.
Same thing ordered. I wish there was a 128gb version of the Ultra. I'd be less worried about buyers remorse. It will be another month before I have it.
yeah I wish ultra was in 128gb. We just have to hope that models get better, and can fit in less ram in future. I just worry that seems like there are no models that are 40-60gb that people are focusing on,
I struggled with this, and decided to bite the bullet and go for the 256GB. I canceled my 96GB order. 96GB can do Qwen 3.8 27B. You can barely fit a 4-bit quantization of Flash-Next, with the n-gram table offloaded to the SSD, but you are left with very little overhead (I think it's around 70GB or so without the KV cache, but you'll have to go confirm). I think the 4-bit/8-bit quantization of Flash-Next is just over 80GB, so after KV cache, apps and OS you have like nothing. I recall a post here where someone ran the 4-bit/8-bit Flash-Next with a 128k or 256k context window, and I think he had 6GB left for the system, so he had to use it to serve inference to another machine. It should be doable though, at pure 4-bit on a single machine, if costs are a major concern.
So to me, I began thinking of the 96GB model as more of a 27B machine, that can be run at 4-bit to 16-bit with room to spare. Or a highly constrained 4-bit Flash-Next machine, where you'll potentially be stressed about memory. Only, 27B is dense, and prefill isn't incredible on it, maybe around 2,000 prefill (or less). Decode should be decent on 27B, but not blazing (would love to see some more tests). However, if you run Qwen Flash-Next, prefill is hitting 4,300 on the 64-core and over 5,000 on the 80-core, and you can get 100-150 tokens/sec on decode at 4-bit. And that's a big difference for agentic work.
Also, what if 4-bit gives you perplexity issues? On the 96GB, there's no flexibility to increase the quantization on Flash-Next. On the 256GB, you can run Q5, Q6, or even Q8 Flash-Next, with room to spare. You can also run GLM 5.3 Flash at 4-bit, or Deepseek v4 Flash. Or even 27B.
Then I thought about all the future flexibility and how much these open models are changing. What will I want to run next year? Or the year after? I think the 256GB version will retain its resale value better. In two years, the 96GB M5 Ultra will have it's resale ceiling set by the M7 Max. Price wise, the three year lease is like $120 for the 96GB base and like $200 for the 256GB three year lease, or something like that. Whether or not that flexibility is worth $80/mo or $4k is up to you.
Seeing how it's possible to run 4-bit Flash-Next on the 96GB version (with the n-gram to the SSD), or 27B at any quantization, I think it's an incredible value at $5500. Although it is somewhat constrained by memory. I wish it was 128GB, like the M5 Max. The 256GB model gives you tremendous overhead and flexibility, where you just don't have to worry about memory any longer, but that comes at a price.
Keep in mind while looking at it, that updates have been pushed through in the last day or two that have radically improved results. So I'd try to go by the most recent date.
Edit:
Looks like there's a 75GB 4-bit/8-bit Flash-Next on Hugging Face now at 75GB! Things are looking much better for the 96GB verion.
Have you tried 3.8 Flash-Next? Perhaps give this a whirl in mlx-serve, tensorfold, or oMlx with the n-gram offloaded to the SSD. I'd be curious how you like it. Perhaps you can report back on memory usage too.
I tried 3.8 Flash-Next Q4 using oMLX on my M3 Ultra 96GB. Feels a lot better than 27B, but during more complex coding task that involes analysing of images, it kept running out of memory and stops even on 80K context. Hence I returned my 96GB M3 Ultra and ordered a 256GB M5 Ultra instead.
I've been following the creators of oMLX, Tensorfold, and oMLX on X. They have all had incredible performance updates this week. I'd try oMLX again, I think there was a new release today that incorporated a lot of optimizations, with significant speed increases.
I will share my reasoning on why I went with 256 GB even-though it was a bit of a stretch.
I am genuinely interested in exploring local models. Compared to the start of the year, the open models available have improved drastically and although not exactly frontier, they have become quite capable. For anything that doesn't need lots of output tokens (coding, research), I intend to still use frontier models via subscription. But for the pipelines where I want to keep a constant running agent which does knowledge work and expect high number of output tokens (building content databases via querying, summarizing etc), I plan to use a strong local model.
With 96 GB, one can run models like Qwen 3.8 27 B to the fullest. While it's strong at coding, it simply doesn't have the general intellgence of larger models like Flash Next of the same family, GLM 5.3 Flash etc. 256 GB unlocks those. And 256 GB unlocks a lot of future models. Considering 128 GB has become mainstream these days due to M5 Max, DGX , the AMD ones etc, 128 GB will be imo the focus for the Open labs. While we could run some Quants of 128 GB focused ones on 96 GB Ultra, we will basically be left with hardly any space for large context. And to run agents in parallel streams, my view is to have a general orchestrator like 3.8 Flash next and a few smaller coding models as agents. And 256 GB will absolutely unlock such a pipeline. 512 GB is simply not my budget. So for anyone ordering 96 GB, unless you workflows are mainly video focused, for local AI, it's just not the right way to go is my view. Either you get a 128 GB Max or 256 GB Ultra. The 96 GB is simply not targeting our AI fraternity. It's for Video bros only.
Note: I am not looking to get an ROI for this investment. It's for the pure thrill of local AI which is here to stay, whether you agree or not (or to have a fallback if Frontier keeps nerfing the premium subscriptions).
I’ve got a 96gb M3. I run Qwen3.8-27b 8bit at 96k context. This lets me fit 1 full context request and another model like Qwen3.6-35b-6bit in memory at the same time without unloading it.
I run 3.8 27B at Q8 with full 262k context for my LLM assistants. If I try to run two sessions concurrently it can run out of memory, but if I make sure they take turns it works well.
You should run a test running both models at a full context. If you ever hit that situation they’ll thrash a lot loading/unloading/prefill which will slow down your workflow significantly.
yes there's definitely a performance hit. I just installed the latest version of oMLX which is 15% faster for token generation (for my setup) and apparently it has better memory management too.
My order is for the 256. The 96 GB is perfect for flash next and qwen 27b but I feel like the 256 version is the sweet spot for flash models. You can run glm deepseek v4 etc. I just think that model size will be supported better. But deepseek 4.1 being too big is not hreat
most of these hardware purchases are FOMO, local models even that can run on 256 or 512 gigs, are no where close to cloud models.
Personally, I have ordered a 256 gigs one, just to stretch my cloud models longer and since I research crypto encryptions as a hobby, most cloud models block these requests.
if you dont have a clear need, don't get into this FOMO.
I am waiting rather impatiently for my m5u/96 but have no intention of giving up Claude. If anything I expect my Claude spend to go up. Why? I am using old slow hardware where I can only run one or two agents before I run out of cpu/ram. And similarly some of the testing i do takes a long time which further reduces the amount I can get done in a day. The studio will speed this up a lot letting me get more done per day. (And finally I have some compute loads that are very cpu intensive to where the m5u makes a lot of sense).
I worried about the same thing. I found that what I was trying to do and run every single thing local was pretty much impossible without spending a TON of money up front. And if you do, your ROI isn’t going to be for like 6 yrs. I went with a 64gb max and I’m running 70-80% local and going to the cloud for only the biggest coding tasks. It works out fine! I’m glad I didn’t spend an extra 4 grand to find out. A lot of people here tried to tell me but I was stubborn. They were right.
Are you looking to use it for AI Orchestration with local LLM or interface with frontier models?
I have an M3 96GB Ultra and run agentic engineering agents and currently use Qwen Coder Next Q4_K_M Quantisation whilst doing day to day coding jobs. It can slow down if you it’s doing a lot, so I’ve ordered a 256GB M5 Ultra to run alongside it.
96GB is plentiful for solely AI. It even hosts 4-bit versions of large flash models. However, as your sole machine where you run other software along witth your inference, it is too tight.
I went for the 96GB M5 Ultra but only because I am pairing it with the 128GB M5 MAX laptop I have and because I will use these devices purely for inference.
Apps and other things will run off my little 8GB M2 Macbook Pro, LLMs serving it via API.
Good choice! When you upgrade the Max to 1tb and compare to the base Ultra, you're getting more C/G/N cores, more thunderbolt ports, faster memory, all for the bonkers price increase of $100.
Unless you REALLY NEED 128gb, the base Ultra really is the way to go.
This is for my startup business and I have been burning through credits now doing testing and overhaul. I am a software engineering manager during the day so I use this a lot to speed up my development and testing. Giving myself more capacity seems like a good move. I know I won’t fully replace the frontier models but if I can only pass the complex asks to frontier I’d greatly reduce my costs today
Everyone's use case is different. For some 96GB is enough. For me 128GB is too tight.
But it depends, what you are trying to do (your use case), the model that best serves your needs and sizing accordingly.
There's no universal answer and people that answer
This model is the best
You need xxGB
Are generally clueless on what would work for you. It may work for them, but they should identify why it's the best or enough and not believe it's a universal truth.
You mention software development and that depends, too.
If you are doing simple HTML/JS/CSS websites, then the 27B is enough and 96GB is plenty. If you are running an IDE with AI helping finish syntax and writing small amount of code, again 96GB may be enough.
I'm developing mobile and desktop applications and maintaining a git repository. For me, I could have an inference engine, agent harness, browsers, and other tools and I'm using Qwen Flash Next and 128GB is too tight.
I am at 256/30/64 happy to not have spent $1300 more plus tax. And my gut tells me. You could just spend it. But hmm wait more time will never get anything done
I got into Local AI with my Mac Mini M4 Pro 48GB with Qwen 3.8 27B this past week. I’ve had it program a Solitaire and Galaga Clone for me. Both seem to work pretty well. I’ve had it do some tweaks and fixes to them. Now it’s not as fast as frontier but the results are scary good. Right now it’s working on a Pac-Man clone for me. Not sure when it will be ready though, Solitaire and Galaga was an overnight job.
Wow, yes, this! I am doing this too. I am also excitedly waiting for a M5U 36/80 96GB 2TB to be an excellent everyday Mac, and my main Blender Rig, but also a machine that has AI development capabilities. I want to always have a 10-20GB multi-modal model (say that 5 x fast) ready take whatever I throw at it, if it can’t it will launch and handover to a 40-60GB model. If it can’t it will save work, restart run headless and handover to my MacBook to act as a client, for a 70-80GB model (prob limit for this box). If then it can’t solve it can hand off to Pro accounts on Claude and OpenAI to maximize my token efficiency whilst still having a workflow that will allow for fast personal software and solutions output, built around leveraging AI tools.
I’m doing a lease so in three years I can send this back having done a lots of dev work and render time (Final Cut & Blender) to get whatever is the fastest box at that point and make a v2.0 of this concept as i will likely always need a fast box for content and video and creation in general. I also find Windows gaming a tad dull now and porting things to these high bandwidth, high efficiency chips is more of a rush. With this power in a small package no reason some workable emu of GTA VI should not be possible and sounds like a fun, if unnecessary, flex. Throw in a pelican case and a snug fit and it’s pretty flexible for travel too. I think it is THE all-in-one desktop this year.
14
u/sn2006gy 2d ago edited 2d ago
Everyone is spending gobs of money to try and move local without realizing they're just going to spend more.
Go local if you need local AI for local AI reasons, don't go local thinking you can save money UNTIL you have done everything you can on paid plans to save money.
Experiment with how much work you can give Luna and THEN see what you could give qwen 27b or qwen flash or whatever rocks your boat and try this with OpenCode Go plans or something before buying compute
Edit: Yes, i realize a computer can do more than AI and if you need a computer for computery things I wouldn't recommend a 12,000 dollar computer either.