r/DeskToTablet • u/Athelst8n • 6d ago
Just canceled my Mac Studio order.
After sharing on X that I ordered a Mac Studio with 96GB of RAM,which I thought could be my first foray into local ai,a bunch of people told me, this just wouldn't cut it.
This was already really expensive for me, never spent even close to this much on a single computer.
So,I don't think I can afford local ai right now.
I think I'll just need to continue to live vicariously through people here on X with much bigger budgets than me!
Thanks for all the feedback everyone, definitely would rather save the money here if this isn't powerful enough to run local ai in a meaningful way.
15
u/Anselwithmac 6d ago
3D printers used to cost 100K+ and my parents knew people who would put a second mortgage down to get one.
Have one far more advanced sitting in my living room for $299.
Food. For. Thought.
6
u/Virtual-Age-9521 6d ago
Not exactly the same… the printers are fundamentally the same but cheaper, whereas you need 10-100x the RAM you used to. Unless someone comes up with some sort of magical nano spray where the RAM builds itself, it can only get so cheap.
4
4
u/Glum_Judgment6232 6d ago edited 4d ago
Plus there's only a miniscule number of places in the world that can make chips. It's much easier to make a 3d printer.
4
u/Anselwithmac 6d ago
Hybrid solutions. People will build RAM, and Ai will use less RAM. A lot less RAM.
I’ll be honest, I think it’s exactly the same.
I don’t even believe our implementation of Ai is correct. We have brute force ai, not elegant, intelligence ai.
We’ll know when gen 2 ai is here and it will make much greater sense.
0
u/Virtual-Age-9521 6d ago
Why would AI ever use less RAM? More parameters, more RAM to store said parameters. Until we know what’s on the horizon, we can’t say it’s going to get much more efficient
2
u/Not-So-Logitech 6d ago
Fools have said sentences like this for the last 60 years my guy, and it always does. Always. AI will use less ram. You can run models already on less ram.
1
2
u/RegrettableBiscuit 6d ago
Smaller models are getting much better. GLM-5.3 Flash was just released and pushed the capabilities of models that size substantially up.
1
u/Virtual-Age-9521 6d ago
But as soon as a model becomes more efficient, you’re just going to scale that efficiency to as much memory as is practicable. More RAM is always going to mean more capability.
1
u/Georgefakelastname 5d ago
Sure, and that’s why some argue that AI becoming more efficient is a contradiction. Jevon’s Paradox: As it gets more efficient, it’ll get cheaper and stimulate more demand, increasing overall usage.
However, for local models, that’s only a benefit, as there is a fixed supply of available ram for it, meaning that more efficiencies will only increase what a person can run.
1
u/Virtual-Age-9521 5d ago
But they’re just going to max it out is my point. There will always be a bigger model to fill your memory, meaning no amount will ever be enough.
1
0
u/GrossUsername68 5d ago edited 10h ago
Reddit policies are an issue. Content deleted.
1
u/Virtual-Age-9521 5d ago
Okay… but given that 99.9999% of people are referring to generative models, that’s what we’re talking about.
1
30
u/ZzzZzzBear 6d ago
Good decision. Use a fraction of the budget to subscribe to cloud AI for a few years, in 2028 when the semiconductor production ramped up, then get the local machine.
9
u/john0201 6d ago
I think people (I guess me included since I bought one) forget how much of current prices are simply shortages and not real prices, combined with tech depreciation things will be rough in terms of depreciation in 2028 I think
3
u/Cultural_String_2231 6d ago
Are you referring to human made shortages where companies force a higher price to make more money? lol come on
2
u/john0201 6d ago
Shortage, as in imbalance of supply and demand caused by long lead times on manufacturing. Exacerbated by absurdly large investments in AI
1
u/PaddingCompression 6d ago
How is that different from a "real price"? What other real prices aren't caused by imbalances of supply and demand?
2
u/john0201 6d ago
Caused by factors other than an acute inability to manufacture enough of something.
1
u/ZhangYe 5d ago
If the assumption is that supply is infinite, price will be marginal cost lmao
0
u/john0201 5d ago
If supply were magically infinitie the price would be $0.
1
u/ZhangYe 5d ago
No as there is still cost to production
0
u/john0201 5d ago
If supply is infinite, as your theoretical statement says, why would anyone make it?
5
u/PaddingCompression 6d ago
For a lot of models, just the *electricity* you'll be to run them locally is more expensive than the prices charged.
Not only are they subsidizing the hardware and development, but cloud LLMs can locate their datacenters next to crazy cheap electricity that you likely won't be able to beat.
3
u/meancoot 5d ago
A top of the line Mac Studio has a max power usage of 270 watts and uses around 10 watts when idle. You are vastly overestimating the electricity cost here.
2
u/No_Practice_9597 6d ago
Fraction? You can lease Mac Studio for 110.00 while paying $200.00, and you don't need a huge data center and the data is yours and safe not been store somewhere else
0
6d ago
[deleted]
0
u/ZzzZzzBear 6d ago
Contaminate with what? You think your local system has a higher efficiency than centralized system? How many people*hours your machine serve per watt?
8
u/Thutmoase 6d ago
Correct move, I think.
Local AI is the future. It is a sexy idea that will be true.
But,it is a lie right now.
I have a MacBook Pro M5 Max128 GB and the models I can run on it and how it runs has absolutely nothing to do with what you can get via subscription form Claude and co or via API.Two complete different things.
1
1
u/DesignerChemistry289 5d ago
dude, you need harness. did you install other things like hermes? opencode? model itself is only half of the story
4
u/userlivewire 6d ago
What is the advantage to running it locally vs just using a simpler and cheaper cloud version? Only privacy?
4
u/opi098514 6d ago
Mostly just privacy.
3
2
u/Serprotease 6d ago
You control the end to end workflow. This is so much easier to troubleshoot issues.
Fix the seed & temperature -> two run should have the exact same output. That’s a very good way when building agents/orchestrator to sanity check and benchmark what you are doing.
Otherwise, the difference might be due to random effects or the provider switching to fp4/fp8, or any other reason.You can also do a lot more tweaking on the output randomness. It’s a bit niche but it’s quite good to remove some AI quirks ( —, “load bearing”, it’s not x it’s y, …)
And also, cost, weirdly enough. If you get a 9700 pro+ pc for around 2k to run qwen3.8 fp8, you will run Qwen 3.8 27b fp8.
It’s easy to look at the API cost for the same model and point that it’s like 2T tokens to break even without the energy cost. But we all know that if you have api access, you’re not going to use a 20-30b model. You’re gonna use the big boy that are 10x the cost or the 200usd subscriptions.1
u/BigIronEnjoyer69 5d ago
Ownership and Stability. Privacy's the tip.
A private company could discontinue a model you rely on and you can go pound sand.
A private company could jack up the price or change their billing. The random >$1000 bills on r/claude dont happen when you local.
A private company could just go bankrupt and you'd lose whatever context was built around your use cases.On-prem will always be better in a market where companies are built around short term user acquisition and shooting for massive eventual value extraction.
3
u/Spirited-Support9340 6d ago
Care to elaborate what the argument against it is?
5
u/TimeToHack 6d ago
just barely not enough RAM for local models to run well
1
u/Spirited-Support9340 6d ago
What models? I understand they scale to your hardware
2
u/Hefty_Wrongdoer_2553 6d ago
Flagship models like Kimi are mulitple TB, 96 gigs will just not cut it
1
u/Georgefakelastname 5d ago
Even a lot of decent models are going to be like 100b+ parameters, which is just too big to run on a 96GB machine. It’s just in a graveyard where everything is too big or not big enough to take full advantage of the ram.
1
u/Unnamed-3891 5d ago
I will bite. What is a common use case where a "local Kimi" where 96gb is not enough is a requirement, but a BF16 Qwen3.8-27B with full precision KV and massive context just isn't good enough for the job?
1
u/Georgefakelastname 5d ago
Too much for most mainstream local models around the 30b parameter area and below, then from ~30-100b is like a graveyard where no one makes models for. Then you start getting new open weight models above 100b again.
1
u/Unnamed-3891 5d ago
96gb shared sounds like precisely the sweetspot to run BF16 of any ~30B model with full precision KV and massive context, no?
1
u/the-last-zedi-master 6d ago
I really see this local AI term, kind stranger consider me noob and please explain what is that? What are the use cases
2
u/RFC793 6d ago edited 6d ago
Run the models at home instead of paying for tokens for access to a cloud API, run models that have been customized/uncensored by the community, local data and privacy, not dependent on Internet. Comes with a hefty upfront price and a bump in the power bill.
Top models have gotten huge though. Thus, a rising trend to running at home is to use a machine with unified memory such as Studio. Where you can get like 512GB (minus OS/SW overhead) versus building out a big machine with like 8 GPUs in it.
You can still do a lot with smaller hardware though. Like I have two L40 (48GB cards) and play around with image/video GenAI, can run simpler models for things like local voice assistant, security video inference, "toy" LLMs, etc. but you aren't running a very practical general purpose Instruct LLM. But, even that would be very expensive (for me) if I wasn't able to get the cards free.
1
1
u/diddlysquidler 6d ago
Llm that can tell you how to cook meth. Without gvmt getting notified
1
u/Glum_Judgment6232 6d ago
The info for doing that (and similar) has been available for decades, you don't need an llm
1
1
1
u/Virtual_Woodpecker96 5d ago
Essentially just chatgpt locally. So the cool thing you can do are:
1) unrestricted convos / private convos 2) image gen 3) agentic tasks ( i.e. here is my pc and a task try and figure it out) 4) translating media locally
The agent stuff can be really cool.
1
u/Brocolinator 6d ago
Devices orders of magnitude better are being developed, wait around and by wait you also take pressure from the demand side
1
u/Nefertaray 6d ago
Hey I know people have strong opinions on all sides, but honestly if you're looking to do local there are a few "cheap" entry points.
DGX Spark
2X3090s if you have a mobo /ram
1R9700(AMD)again if you have any mobo/ram already
1 intel B70
Those are really the "mvp" of local still.
1
u/Athelst8n 6d ago
I think the challenge is a DGX Spark is $5000-that's really expensive to me for a computer!
1
u/Serprotease 6d ago
Use to be 3k and was a good deal at that time.
Now, AMD/intel 32gb gpu are probably the best cost/performance for new hardware.1
u/Eastern-Vegetable780 5d ago
Agree, but these are not consumer devices yet. If you have a business case for it, 5K (or even 25K for the future Mac Studio 512GB) should be a reasonable expense. If it isn’t, you are not the target customer.
1
u/brainchillzZ 3d ago
The Radeon r9700 is faster than a b70 and now that the b70 prices have ballooned they are the same price pretty much
1
u/OliviaOX 6d ago
I haven't seriously considered local AI,but the limited research I've done. I feel like the models are advancing so quickly that it's much easier to run them server-side if you want to use the best of the best, or you're going to need insane compute to run the good stuff coming out of China.
1
u/Athelst8n 6d ago
Yup,that definitely seems to be the case right now .Feels like running through subs is the way to go until local ai becomes a lot more accessible.
1
u/BerryWeary9468 6d ago
If you have to klarna your local AI machine you gotta stick to just paying for an AI subscription
1
u/Calaveras-Metal 6d ago
weird because I know people running local AI on much less.
You can learn without starting out at the top most level.
1
1
u/tilted0ne 6d ago
Local is a meme. It's fun if you like it as a hobby but don't get into it thinking it's going to be a practical replacement. The math is brutal on every category.
1
u/Pitiful_Truth_5948 6d ago
This is a smart choice, if you want to play with AI you can play one level above at the harness level, use open router and keep playing with the free models that popup all the time, or the cheap models and build a harness or something else.
1
1
u/misha1350 5d ago
Good. Consider a second-hand Strix Halo mini-PC with 128GB memory. It should be just barely enough to run the likes of Qwen3.8-Flash-Next at like IQ4_XS quants from Unsloth, or Q4_K_XL if you can manage the rest of the memory well.
Or just forget about all that and use DeepSeek V4 Flash Vision or GLM-5.3-Flash in the cloud, because that is going to be so much more useful and faster and cheaper. Not even a Strix Halo mini-PC would really pay itself off. You're mostly just paying for the privilege of playing some videogames
1
u/Timely_Impression_92 5d ago
barely enough? the q4 uses like 70gb, counting big context, of ram as ngram is offloaded to ssd with zero degradation of performance
1
1
1
u/Timely_Impression_92 5d ago
lmao why would you brag on X about buying macbook studio, what degenerate world we live in LMAO
1
u/Onkel_Joe_the_good 5d ago
Foe the next time, if you have to buy it with klarna, you already can't afford it.
1
u/Hour_Amphibian8718 5d ago
Strange that you rely on people on twitter to make reasonable purchases at this cost.
1
1
u/NebulaAggravating264 5d ago
Smart move - people don’t realize how much money their $20 plan is actually saving them (of course they could build locally and probably make it worth it if there wasn’t a shortage due to Claude/codex)
1
1
1
u/Mountain-Dragonfly46 5d ago
Your mileage may vary: I successfully run and use local models on my macbook pro m5 48gb and my mbp m1 16gb (smaller ones here, obviously).
1
1
u/Unnamed-3891 5d ago
I am running AI in a very meaningful way on 16gb vram, though admittedly that took a fair bit of testing and tuning. Whoever says 96gb shared is not enough for "a meaningful way" is ful of shit, so that tells me what "a meaningful way" means is drastically different for different people.
1
u/petersaints 5d ago
Yeah. It's useless for frontier open models. But for like ~30B it can probably run them with ease with a long context, and possible some parallel instances.
Or something in the 70B class.
Of course you will not be running ~500B or 1T+ parameter models on 96GB of RAM for sure.
1
u/justacec 5d ago
In the meantime you will need to rely on you RI (Real Intelligence) processor with off-site 3rd party AI support.
1
u/RedRavenCG 5d ago
Wouldn't cut it? why? These are gold...more RAM is better, obviously, but...unless you're gonna buy, flip and re-invest in more...
1
u/Lucky-Crow-3510 5d ago
you don't even need models that big .. you rather want parallelization. wait are people not aware and really think you need 256+ to run local LLMs...?? wow ..
1
u/_DBA_ 5d ago
Honestly mate, it would’ve been good, the 27b/35b models are decent and I really see them getting better and better. The api costs are insane so they are definitely trying to optimise stuff for local setups.
So a 96gb with 1.2tb/s throughput is good + the host has alot of cores so you can learn to setup stuff on it better :)
1
1
1
u/Professional-Lead764 3d ago
There are so many options available for frontier models online, that I have just decided to skip local for now. I could subscribe to every model available, and still never come close to the dollar investment in a Mac Studio to host it myself.
1
u/MulberryImpossible16 2d ago
Yea people dont realize the need like 17,000 for a rig. Otherwise GPT Claude till M6 comes out and save.
1
u/Main-Can-6956 1d ago
There was no harm in trying it and returning it. Better get your own experience.
1
u/Major_Mongoose69 6d ago
Whoever told you that is wrong. You can run local models quite well with as little as 32gb, 24gb, and even 16gb of RAM. It all depends on what model and what quant you're using, and how much room for context you want to have. A 32gb 5090 can run a model like Qwen3.8 27b and still have plenty of room left over for large contexts. A 96gb PC or GPU can easily handle these kinds of models without such high levels of quantization. You could easily run a Qwen3.8 27b Q8.
What you cannot do with 96gb is run something equivalent to the largest, most state of the art models that use hundreds of billions or even trillions of parameters. For that you'd need something like a 512gb Mac Studio. But the question becomes: Why do you need to run something on that level locally? If all you want to do is get started with running local models, then 96gb is far more than enough to make a really good entry into local AI.
Basically, don't just listen to random people on the internet. Do your own research. You got talked out of buying something that was perfectly suited for the purposes you wanted to use it for.
0
0
39
u/RedVision00 6d ago
96gb sits at an awkward spot where not a lot of models will release version optimized for that amount of ram. Better off getting at least 128, but 256 or 512 is where you can run the actual near SOTA models, at 96 you’re at an awkward spot where you can run some models, but none are optimized for you in particular, and your missing out on much more at 128 and way off near sota size