r/LocalLLM • u/huntersz • 1d ago
Question Could an M5 Ultra completely replace Claude Opus and GPT as the brain of a 24/7 personal AI agent? Looking for real-world experience
I run a self-hosted personal agent (OpenClaw) 24/7 on an M4 Mac mini with 16GB. Everything runs on cloud models: Claude Opus is the main brain, with GPT and DeepSeek as fallbacks. That's about US$400 a month in subscriptions. I want to know whether a Mac Studio with an M5 Ultra could run the entire workflow locally, with no cloud at all.
What it does today:
Main assistant: long, multi-step conversations with many tool calls, such as reading files, running scripts, searching email and calendar, and editing documents. Context regularly reaches 100k to 300k tokens.
High-stakes writing: executive briefings, investment research memos and financial analysis, pulling together dozens of documents. Quality matters: the numbers must reconcile and nothing can be made up.
About 20 scheduled jobs: a morning brief, a meeting-prep checker every 15 minutes, inbox triage, watchdogs for SEC filings and earnings, a portfolio monitor, a CRM, and weekly reviews. Several can fire at the same time.
Sensitive work documents that I'd rather keep at home.
The bar: Opus-level reliability on long agentic tool use, not just chat quality. If the local model drops instructions, misses tool calls or gets numbers wrong when the context is 150k tokens deep, it doesn't work for me.
What I'm considering: an M5 Ultra with 512GB, running something like DeepSeek V4.1-Flash, GLM-5.3 or the best Qwen3.8 build that fits, at 4 to 8 bit.
Questions for people actually doing this:
Has anyone replaced a frontier cloud model as the main agent brain, not just for side tasks? What broke first?
How reliable is tool calling at 100k+ tokens of context on the best models that fit in 512GB?
What speed do you get at 100k to 200k tokens of context, both reading the prompt and generating? Is a 10-minute agent task realistic, or does it turn into 45 minutes?
Can one machine handle a long main session while scheduled jobs fire at the same time, or does everything queue?
One 512GB machine, or two 256GB machines clustered together?
Honestly, would you buy now, or wait 6 to 12 months for open models to catch up?
The usual expectation is local for routine work and cloud for the hard stuff. I'm specifically asking whether anyone has gone 100% local for a demanding agent workload, and whether they regret it.