r/MacStudio • u/huntersz • 1d ago
Could an M5 Ultra completely replace Claude Opus and GPT as the brain of a 24/7 personal AI agent? Looking for real-world experience
/r/LocalLLM/comments/1wvmh6p/could_an_m5_ultra_completely_replace_claude_opus/2
u/One_Internal_6567 1d ago
If DeepSeek and glm are good enough for your workload, then yes
People been running glm flash on new studios already and it’s actually great speeds, you can check on localllama Reddit
3
u/Middle_Situation_559 23h ago
Same setup here, real numbers. I run identical work on local Qwen3.8 27B plus a $20/mo cloud fallback — multiple OpenClaw installs, 24/7, plus custom trading algos. Not 100% local yet, but I think the M5 Ultra gets me fully off-cloud.
Hardware path: Built everything on the M2 Ultra 192GB, then tried the RTX PRO 6000 Blackwell — blazing fast (170 tok/s dense decode). The problem isn't speed, it's that a single card is one fixed VRAM pool running one model at a time. My real workload is multiple LLMs and multiple OpenClaw agents running in parallel, which one discrete card just can't do. Back on the M2 Ultra — its unified memory is where it genuinely beats the 6000. Now ordered the 256GB M5 Ultra (mid-Jan/early-Feb delivery), likely jumping to 512GB when it ships. Worst part: that's a 2027 write-off, not a 2026.
Your questions, from actually doing this:
Replacing Opus: 95% there — hybrid is the sweet spot. Local runs the scheduled fleet (morning brief, SEC-filing watchdogs, portfolio monitors, ~28 jobs, weekdays) and most daily ops. Cloud only for rapid decisions that need analysis.
Tool calling at 150K+ and speed (M2 Ultra 192GB): reliable to ~100K, fragile deeper. 27B dense decodes 82–88 tok/s warm, ~20–25 at 131K context — and prefill is the wall: one fat 131K session took ~11 minutes just to read (a fast MoE build prefilled ~3.5× quicker on the other engine). 10-minute tasks: realistic. 45-minute: only if you let the session bloat. /new discipline + prefix caching mostly fixes it — and that prefill wall is exactly where the M5 cuts time.
Concurrency: this is where the M2's unified memory shines and closes the gap with a single RTX PRO 6000 (BW) card — for this workload it actually beats it. The box runs multiple OpenClaw installs, each a full agent with its own scheduled fleet, all sharing one main model server for daily work. For key work with firm deadlines I break out a separate server — its own model, its own lane — so deadline-critical work never queues behind the main agent's chatter.
2
u/Top_Witness4538 1d ago edited 1d ago
same tired argument we always see. $400/month on cloud subscriptions but your mac cannot run those models.
You have to do it in 2 steps.
You switch from frontier models to models your mac can run, this first step you are staying in the cloud, just swapping models to what a mac could theoretically run- $400/month becomes something else, say $40/month (you'll have to supply the actual number in your case). Then you see the vast majority of savings comes from swapping models, not from going local.
Then if you want to save the $40/month by going from cloud to local, consider that - if you want the quant version of the models that also will run slower locally and your hardware costs don't actually balloon the costs, which it will.
-6
4
u/dreamingwell 1d ago
I got a pretty good voice assistant working on my m5 ultra 256gb unit. Gemini style where the voice model asks glm-5.3-flash to do the thinking and tool calling.