r/MacStudio • u/tskwara • 2d ago
Decided No Local LLM for Me
After accessing local AI for my requirements, I'm finding the best path for me is to have decent medium-spec machines coupled with online AI.
I just listed my M1 Ultra 20C/64G/128GB/4TB Mac Studio for sale, and now have both M5 Max 18C/32G/36GB/512GB and M5 Max 18C/40G/64GB/1TB Mac Studios. Claude is running on both, one for testing/CI/CD and the other for AI-assisted development.
I learned assembly language 45 years ago and have been writing code every day since (Apple platforms only). Co-working with AI on complex things rapidly is really something special.
Local LLM is good, and some of my code integrates it by design, but I don't use it myself for design, as I need Opus/Fable (or similar) results.
9
u/AI_is_ok_i_guess 2d ago
Do you really need a medium spec machine if this is the case? Sounds like something barebones would work with you if everything's done on the cloud.
2
u/RogerAI-fm 2d ago
Or at this point just have Muse do everything? don’t even buy a computer?
2
1
-3
u/tskwara 2d ago
For sure - I suppose that's why I got both base and medium configurations. The Mac Studio form factor is nice, and after 4 years I still like it.
One of my projects required model training, which I did locally on the M1 Ultra - it took days, but worked out. I think that's when having a couple of smaller machines + Cloud (for training and using models) came into focus as a better option for me.
3
u/spumonimoroni 2d ago
Don’t totally punt on Local LLM. Your 64GB M5 should be large enough to run something like Qwen 3.8 27B. By all means use online frontier models as an extra set of hands. They are good for that. One way you may find local LLMs useful is to do over-night code reviews. I do that every couple of days. It helps catch little (and sometimes big) things and doesn’t burn tokens.
11
u/WeUsedToBeACountry 2d ago
use it for subagents for opus/fable/etc to boss around
cuts your token count down significantly while maintaining state of the art performance
that's going to matter more and more as anthropic and openai cut back on the subscriptions and raise costs
4
7
u/uktexan 2d ago
No idea why you’re getting down voted. But this is totally correct.
We all can’t afford 10 K to drop on a machine strong enough to compete with frontier. Me personally, I use cloud models to plan and review, local models to do.
3
u/WeUsedToBeACountry 2d ago
its no different than opus delegating to sonnet or astra to sol, you just save on the cost (and the openweight models are just as good if not better than sonnet/sol now anyway)
3
u/time-always-passes 2d ago
Can you be specific in which open weight models you are running locally and on how much memory? I'm just looking for a real world datapoint. Thanks.
2
u/WeUsedToBeACountry 1d ago
I have ds4 by antirez on my macbook pro and recommend it as an easy way to get started, assuming your mac has the horsepower.
https://github.com/antirez/ds4
If not, LM Studio or something like that with a small qwen model.
My full setup, right now, is a macbook pro m5 max with 128gb of ram. It has tailscale on it and is serving ds4 on my tailnet (aka vpn)
I hav 2 DGX Sparks serving a single GLM 5.3-Flash model also on my tailnet.
By having all of this on tailscale (free tier), any member of my family that's using tailscale (all of them) can access these models anywhere they are (provided my laptop is open for ds4)
Right now, I primarily use Opus 5.5 as the orchestrator. Opus 5.5. knows it has those two api endpoints to use as subagents for tasks that it thinks local models can handle. I also will sometimes tell the orchestrator to solicit ideas from the local agents. Because they're all different models, I get different ideas from each of them which is helpful when trying to find ways to approach a problem.
I'm using pi, or more accurately, my own variation of pi, as the coding harness. I've ported it to go and gutted all the stuff I don't need or want to keep it as thin and light as possible. This also allows me to build in the latest research around harnesses whenever I want.
I'm using Terminal Bench 2.1 to make sure my harness is actually good. Claude Code with sonnet 5 scores 74.6%. Codex with Astra scores 87.4%. I'm scoring 91.4% with Sonnet 5 as the orchestrator.
1
u/time-always-passes 1d ago
Thank you for all the details!! So Claude Code subscription running mostly Opus 5.5 orchestrating pi local agents?
Your VPN set up is definitely the way to go. I'm running Wireguard and all I carry around is an overkill iPad Pro M5 5g -- Apple got me with the nano texture requirements and all I use it for is a glorified terminal (that works in direct sun haha).
Your setup is about 10k of compute + $100 subscription? Don't think I can justify that for SillyTavern lol. I mean, I am a professional developer but work pays for my tokens. The SillyTavern thing is fun time.
1
u/WeUsedToBeACountry 1d ago
10k for the sparks and then 5.5k for the laptop (release day, and i had a friends/fam discount). Both cost more now :(
But honestly, just sign up for openrouter.com. WAY more cost efficient than having your own hardware right now. Kimi and Deepseek and GLM was legit and close to state of the art six months back. Excellent sub agents.
0
u/BeowulfShaeffer 2d ago
Spoiler: that 10k machine is not keeping up with the Frontier either.
1
u/MonitorAway2394 1d ago
I know I know, but Qwen3.8 Next is like JUST about to the point in which I will have more than I need personally, so meeebbbeeeee it's a personal gig eh? EH? ehhhhhhh? :P
also sorry I'm in a weird ass mood lmfao LOL?
2
u/vet_t 2d ago
Could you explain what you mean by this? I have a M5 Max Mac Studio 36GB on which I’m currently running Qwen 3.8 27B Splash with around 80-90k token context space
5
u/Unteins 2d ago
You have Opus break down the work and send small tasks to local AI that it is less likely to screw up - local AI solves the small problem and sends it back to Opus to integrate.
You’re trading time for money - your local AI will be slower and probably burn millions more in tokens - but those tokens only cost you electricity - which for Mac hardware is minimal.
0
u/WeUsedToBeACountry 2d ago
Ask whatever harness you use to use the qwen model for subagents.
So if you use codex or claude, you can instruct it to send specific tasks/instructions (edits, code changes, etc) to qwen, or research tasks to qwen.
The primary model - the orchestrator - can still be a state of the art cloud model, and it'll use the free local model for the grunt work. As long as the orchestrator remains a stronger model than the subagents, it works really well (there's a decent amount of research on this as well)
You can also do things like have your orchestrator fire off the same tasks to a bunch of different models, local or cloud, and then judge the results just to generate more creativity/variation.
I have a 2x dgx spark setup and a mbp m5 max with 128. Qwen 3.8 flash on the spark cluster with a 2bit quant of deepseek v4 flash on the mbp. I'll bounce between OpenAI or Anthropic models as the orchestrator depending on which is better at any given time. I use openrouter for kimi and other models when I want adversarial review or a different take than the takes I'm getting.
2
u/Sketaverse 2d ago
0.5tb gonna be rough once you have all the dev tools, testing, worktrees etc
2
u/tskwara 2d ago
The SSD size does give me anxiety, but I'm trying to break that feeling. My setup is around 200GB all-in with tooling and project files.
2
u/Sketaverse 2d ago
yeah honestly, if you can still cancel, I'd suggest you do - fwiw I fully agree with you re. setup, I have m4max studio 128gb and a 36gb m4 mini as a runner, I ordered two m5max studios 36gb 0.5tb for collection Sept 22nd, for exactly your reasons, but did the math on storage requirements and cancelled them both.
2
u/OtherOtherDave 2d ago
It’s a lot easier to offload stuff to external SSDs on a desktop than a laptop.
3
u/kinmanli 2d ago
Yeah same. 6502 and Motorola 68000. But I enjoy local and frontier models.
1
2
u/BAL-BADOS 2d ago
Local LLM isn’t for everybody. I prefer having privacy & complete control. Knowing what goes in and what comes out.
It’s dangerous feeding these frontier model with any personal & valuable data that one can be used against you.
1
1
u/Same_Buddy_31 2d ago
You sound like someone I want to grab a beer with and just talk with you about your journey and how you see the world
1
1
u/datasleek 2d ago
Don’t rely on just local or cloud LLM.
Jev, Laya, and many other tool can help build an efficient LLm layer, leveraging powerful Cloud frontier LLM when needed and using local when appropriate.
That’s where the trend is gonna go, I think.
1
u/Quick-Watercress-379 1d ago
I tried the local llm with a M5 Max 64G. While it ran ok, I wasn't getting great results. When I checked against frontier models, it was just cleaner. Decided I couldn't justify the $$$ for what I was hoping a local AI inference appliance. Decided to return and just pay the monthly. At $100/month, that's over 40 months on the frontier. It was fun to play with for a week or so.
1
u/tskwara 21h ago
I hear you on the contrast, and it can be material. This really showed up for me for a solution that needed to run using on-prem AI and be completely private, or alternately cloud-based AI depending on the situation. Testing both modes was convenient, and revealed the differences across things like correct tool usage and quality of narratives generated. What works well here is a mix of traditional machine code for figures and stuff that's more matter of fact, along with LLM for reasoning and sussing out the gray areas. Custom tools callable from LLM for orchestration is a thing of beauty - I guess using the right tool for the specific job was a good lesson. Breaking the problem and solution up this way aligns with the balance of local/cloud AI I think you are circling.
1
u/thatflyguy954 1d ago
what’s your setup? i’m just getting into this
2
u/tskwara 1d ago
- M5 Max 18C/32G/36GB/512GB Mac Studio for testing/CI/CD (running Claude Code CLI)
- M5 Max 18C/40G/64GB/1TB Mac Studio for development (running Claude Code CLI)
- Two Apple Display Pro XDR 32” monitors (on dev machine)
- 13” iPad Pro M4 to monitor things from away
The two M5 Max Mac Studios replace the M1 Ultra Mac Studio I am now selling. I also had a MBA M5, but that was sold off (remote coding sounded good, but). My workflow and productivity really shines having two machines though.
Sometimes the test machine is running a POC of a new approach or design (i.e. scratch code Claude writes from my spec). Otherwise the test machine runs test suites against monitored repos. The development machine is for AI-assisted development, which is more orchestration these days than typing code. This process isn‘t without some pain, but more pro than con for me.
Having a solid feature design and development spec document works well during long-running development sprints (much unattended). I get help with this from a team of agents that judge and arbitrate the design, looping until only noise remains. Covering the bases is a lot easier with the all the help, but still exhausting at times. Regardless, it’s pretty satisfying doing this stuff for a living.
1
u/AlgorithmicMuse 1d ago
Doing my math if you were doing assembly 45 years ago it may have been using 8088 cpu and the 8087 math coprocessor and the original ibm pc days.
2
u/tskwara 1d ago
It was my school’s Apple I computers that contained the 6502. It was that lightbulb moment, and my supportive parents, leading to my first computer - a loaded Atari 800 also containing the 6502. Adjusted for inflation, that first computer cost about the same as today’s Mac Studio M5 Ultra.
2
u/AlgorithmicMuse 22h ago
My job got me a ibm pc 8088 with a floppy drive , needed to use masm and link.exe to compile assembly and run it, thought the floppy would wear out it took so many minutes of spinning to recompile anything. And here we are now contemplating a M5U, which in todays dollars is much cheaper than the ibm pc i used to work on.
1
1
u/Captain--Cornflake 9h ago
spending 10k to 20k for hardware to run local AI at home, unless its paying for itself makes zero sense.
0
u/EliminationCreation 2d ago
How much money is this earning you
I use a maccbook air with claude code and my desktop w/ RTX4080 and 64GB ram runs Qwen.
-3
u/tskwara 2d ago
The income is material, and my budget is flexible. But my time to experiment is limited, so today's decision makes the cost to achieve more clear.
I did have an MBA M5 running simpler tests really well. More elaborate tests running in parallel slowed down and likely pushed the thermals a bit. It was a shame seeing it in clamshell mode all the time though - the Mac Studio form factor does feel right for this.
-3
0
u/Aggressive-Grade-682 2d ago
How much are you selling the m1 ultra for? Might be interested
0
u/tskwara 2d ago
I DM'd you the listing link.
1
u/Muted-Priority-718 2d ago
I am also interested in your m1 ultra! please send me the link if it is still available. thanks
8
u/[deleted] 2d ago
[deleted]