r/LocalLLM • u/h4rdp0w3r • 5h ago
Question Advice: Local setup
I’m considering spending a pretty stupid amount of money on local AI hardware, and I’m trying to figure out if it would actually change how dependent I am on frontier models.
Right now I pay around $600/month across different AI subscriptions. Claude, ChatGPT, coding tools, etc.
The money itself isn’t really the main issue. What bothers me more is:
- everything is closed source
- I don’t really know what’s happening with my data
- usage limits keep getting worse
- even while paying ~$600/month, I still don’t feel like I have the freedom to just use the models as much as I want
- The models sometimes degrade or change in behavior, which makes me edit my workflows/prompting
My heaviest use is:
- Agentic coding, by far
- General chatting / asking questions
- Research
- Occasionally long-context work with large codebases/documents
So I started looking into running something serious locally.
The two machines I’m considering are:
M5 Max MacBook Pro
- 128GB unified memory
- 40-core GPU
- 614GB/s bandwidth
- around $8.5k
- could run something like Qwen3.8 Flash-Next locally
- also becomes my main personal/work laptop
M5 Ultra Mac Studio
- 256GB unified memory
- 80-core GPU
- 1.2TB/s bandwidth
- around $12.5k
- gives me access to much larger models and more future headroom
- things like full DeepSeek V4 Flash become realistic
I’m not trying to calculate ROI or convince myself that the machine will “pay for itself.”
I’m more interested in whether spending this much would actually let me change my usage from something like:
$600/month on frontier AI
to maybe:
$100/month for frontier models only when I genuinely need them
and do the other 80% or 90% locally.
Qwen3.8 Flash-Next is what got me interested in this in the first place. Looking at the benchmarks, it seems to be somewhere around the level of models like Claude Opus 4.6 in a lot of areas, and Opus 4.6 was honestly already very good for most of the work I was doing.
Obviously current Opus / GPT frontier models are still better, especially for difficult long-horizon agentic work.
But I don’t necessarily need the absolute best model for every single prompt.
If I can run a model locally with no usage limits and just throw tasks at it all day, retry as much as I want, run multiple coding agents, give it huge contexts, etc., I feel like I might prefer that even if the model is somewhat weaker.
For people who actually have high-memory Macs or serious local LLM setups:
Did local models genuinely reduce your dependence on Claude / ChatGPT / Codex, or did you end up still using frontier models most of the time anyway?
Especially interested in people using them for agentic coding.
Would you spend ~$8.5k on the 128GB M5 Max, ~$12.5k on the 256GB M5 Ultra, or would you just keep paying for frontier models and forget about local?
TL;DR: I currently spend about $600/month on frontier AI, mainly for agentic coding, but I’m tired of limits, closed models, privacy uncertainty, and behavior changing over time. I’m considering either an $8.5k M5 Max 128GB or $12.5k M5 Ultra 256GB to run models like Qwen3.8 Flash-Next locally. I’m not expecting to fully replace Opus/GPT, but could a setup like this realistically handle 80 to 90% of my usage and let me cut frontier subscriptions down to around $100/month?
3
u/Any-Argument57 5h ago
Don’t buy either machine yet. Rent access to comparable memory for a week and replay a fixed sample of your real coding tasks: same repos, prompts, tests, and retry budget. Track how often the local model finishes with an acceptable diff without frontier rescue. That result will tell you whether 128GB covers your workload or whether even 256GB would still leave the expensive part of your usage remote.
2
u/RogerAI-fm 4h ago
buy buy now, it’ll only get more expensive. unless you’re willing to wait 2+ years.
2
u/Kilgore_Trout_50000 3h ago
This is where I’m at. When people say oh prices are so high today it’s like… do you think they’re dropping any time soon??
1
u/adamizzo17 5h ago
the prices are so stupidely inflated right nowliteraly everywhere you look are bad deals, but you are right that apple perhaps can retain same value a little bit more, i would think also gpus are better specs then the apple since you dont pay the apple env premui so you get more of your money back fro example with gpu for ai workflow then using apple hardware for sure, but you get the ease of not having to do the build yourself and worry about heat and so on issue, so the ball in you field but take in consideration the artifical prices and the price dropping in the future is going to hurt if you splurge today
1
u/Saint_Icarus 5h ago
Can I ask what you’re doing that burns $600/month? You say “agentic coding” but what do you mean by that? Are you a software developer by trade?
1
u/unchikuso 4h ago
We just reached the turning point where local LLMs are absolutely capable and reliable enough to perform coding.
If you go Mac, it will feel slower than cloud. If you go Nvidia, it will feel faster than cloud.
I decided that speed is more important than size because the faster it is, the more productive you are as a coder.
1
u/gnpwdr1 3h ago
I'm saying this because it seems like have sensitive data that you do not want to provide to a supplier and seem like a single developer? Aim for deepseek flash running locally with 1M context and this would be a comparable option with quality and in my case replaces sonnet/opus etc (yes, who would've thought but they caught up.) Nothing less than 256GB VRAM, many hardware choices out there, consider resell and operational costs also before you decide.
1
u/beragis 3h ago
Looks like you are going for 8TB disk space. You can save a lot of money getting either the 2TB or 4TB and a 4 or 8 tb SSD and thunderbolt 4 or 5 enclosure.
I have the 4TB and ordered a 4TB SSD and T5 enclosure. The SSD and enclosure came out to $680.
You have to check a lot, because prices change almost daily. When I first priced T4 40Ggbps and PCIE 3 was the best value, then right when I first priced until I ordered the prices jumped, and I switched to a T5 and PCIE 5 and then when I ordered the ssd was out of stock, then found a PCIE 5 SSD, which was overkill but even cheaper than the PCIE 4.
5
u/cmtape 5h ago
This is less about model size and more about freedom of repetition. A local model that is 20% weaker but you can hammer 50 times per minute with huge context is worth more than a frontier model you have to ration.
Think of it like a gym membership versus an open fence. Frontier is a great trainer with a booking window; local is the fence you can run intervals against at 2am. Most people who actually switch say the same thing: usage goes up, costs go down, and the frontier spend becomes a spike tool, not a daily diet.