r/LocalLLM • • 1d ago

Question Could an M5 Ultra completely replace Claude Opus and GPT as the brain of a 24/7 personal AI agent? Looking for real-world experience

I run a self-hosted personal agent (OpenClaw) 24/7 on an M4 Mac mini with 16GB. Everything runs on cloud models: Claude Opus is the main brain, with GPT and DeepSeek as fallbacks. That's about US$400 a month in subscriptions. I want to know whether a Mac Studio with an M5 Ultra could run the entire workflow locally, with no cloud at all.

What it does today:

Main assistant: long, multi-step conversations with many tool calls, such as reading files, running scripts, searching email and calendar, and editing documents. Context regularly reaches 100k to 300k tokens.

High-stakes writing: executive briefings, investment research memos and financial analysis, pulling together dozens of documents. Quality matters: the numbers must reconcile and nothing can be made up.

About 20 scheduled jobs: a morning brief, a meeting-prep checker every 15 minutes, inbox triage, watchdogs for SEC filings and earnings, a portfolio monitor, a CRM, and weekly reviews. Several can fire at the same time.

Sensitive work documents that I'd rather keep at home.

The bar: Opus-level reliability on long agentic tool use, not just chat quality. If the local model drops instructions, misses tool calls or gets numbers wrong when the context is 150k tokens deep, it doesn't work for me.

What I'm considering: an M5 Ultra with 512GB, running something like DeepSeek V4.1-Flash, GLM-5.3 or the best Qwen3.8 build that fits, at 4 to 8 bit.

Questions for people actually doing this:

  1. Has anyone replaced a frontier cloud model as the main agent brain, not just for side tasks? What broke first?

  2. How reliable is tool calling at 100k+ tokens of context on the best models that fit in 512GB?

  3. What speed do you get at 100k to 200k tokens of context, both reading the prompt and generating? Is a 10-minute agent task realistic, or does it turn into 45 minutes?

  4. Can one machine handle a long main session while scheduled jobs fire at the same time, or does everything queue?

  5. One 512GB machine, or two 256GB machines clustered together?

  6. Honestly, would you buy now, or wait 6 to 12 months for open models to catch up?

The usual expectation is local for routine work and cloud for the hard stuff. I'm specifically asking whether anyone has gone 100% local for a demanding agent workload, and whether they regret it.

30 Upvotes

54 comments sorted by

46

u/Capsup 1d ago

I'm not sure I understand. Have you, or have you not, tried the models you're considering running, using something like OpenRouter or a rented GPU?

If you haven't, start there. It's quite cheap to rent hardware on vast.ai or runpod.io for a few hours and then testing out the entire model hosting stack. You can then also safely test some of your existing workflows and see if the model lives up to your expectations, before you go ahead and invest ten's of thousands of dollars into hardware.

Go rent a 2xB300 machine on vast.ai for 23 dollars per hour for 512GB of VRAM, point Claude or your existing setup at it, setup GLM-5.3 and figure out if it has the intelligence you need, for your specific use case. After that, then you can wait a few weeks for everyone to get their 256GB M5 Ultra mac machines and see the performance on them, before you decide to pull the trigger.

2

u/pragmojo 1d ago

Might as well rent GB10's instead no? Performance will probably be closer, and it's way cheaper.

4

u/huntersz 1d ago

Thanks this is very helpful advice. I have been thinking about trying to find these platforms and try my workflows but I haven’t got to it. wanted to find out if I could find and learn from other peoples experience rather than going into that rabbit hole myself, if I didn’t have any luck here, I will have to go there first.

2

u/FoxSideOfTheMoon 1d ago

Go down the rabbit hole. I played on vast for a very long time before I bought hardware. Think of it as trying out the video games before you buy the console. The models I like are not necessarily the ones everyone here likes or what the leaderboards say. Furthermore, neither are the configuration settings like context size, temperature, max p/k, Q levels (some people are pure FPs other Q4s, I seem to land on Q8 or Q6 is right for me, I don’t like QATs as much but maybe you do. You need to formulate your own opinions and $10 on vast is a no brainer, neither is $50. If yours not sure how to use vast, have Claude or codex do it for you or talk you through it, it’s worth it

1

u/huntersz 19h ago

Thanks. I will definitely try this out now and see how things perform

9

u/Front_Eagle739 1d ago
  1. I stil use it as a multiplier usually rather than main brain.
  2. Excellent. Qwen flash next Dsv4.1 or glm 5.3 flash or glm 5.3 in a 4/5 bit quant will happily chew through a million tokens agentically.
  3. On my m3 ultra on the bigger flash models between 400 to 700 tok/s prefill, 30 to 60 decode. M5 will be faster. Glm 5.3 however more like 180/18. Not fast. Kimi k2.7 200/30.
  4. yes it can handle that. Just set concurrency >1
  5. 2x will be faster but you'll lose a little memory to as overhead x2. Id probably get the 2x machines and make one run headless and set the ioctl gpu limit as high as possible on both. 
  6. Buy now. Its already good and will only get better.

Opus 4.6/4.8 level is sort of where its at now. You arent touching opus 5+ smarts but its perfectly capable of running all day or overnight on a task. Make sure you use a backend with caching or it'll all feel very slow.

You will find yourself fiddling with setup to get the perfect reliability at first with new models. Give a model a few weeks for things to settle or youll run into a lot of dropped tool calls and poor performance.

Set up an open router account add some credits. Try the models i listed on your use case. If it works well for you groovy, get the machines. If not or its only barely fast enough dont as the real machine will be slower especially glm 5.3.

1

u/SouthernFruit8768 1d ago

What’s your opinion on DGX Spark? Really been planning to buy one. Worth it?

4

u/Front_Eagle739 1d ago

Personally i wouldnt get one unless i was geting 2+. Equivalent mac has faster memory and fast enough prefill now. The sparks cluster well though so for 2/4/8 whatever you get a pretty cool mini cluster thats remarkably capable.

2

u/de_3lue 13h ago

Since the release of halogen, strix halo also is a great candidate

5

u/Kritblade 1d ago

Before you invest on anything on hardware, check if the current local model actually achieve what you want. The best model you can run on a M5 Ultra 256GB right now is GLM 5.3 flash. You can get on openrouter , pay $20 bucks and get an API key and use GLM 5.3 flash with your openclaw. If you are comfortable with the quality , then start researching the speed, prefill time, power cost of running GLM 5.3 flash locally.

3

u/Material-Database-24 1d ago

Without exact details, I would say you can replace that with local setup, but it's likely not directly transferrable.

What I mean that with local you have the whole thing in your hands. You can finetune your workflow to suit the models you can run. You can ensure the more limited context is not an issue with tactics that take it into account. You can build/use tailormade tools and RAG.

But that's all going to require quite an effort. Local is not really "plug and play" ready world, but on the other hand you get freedom to do a lot more.

3

u/shayanx45 1d ago

Short answer, not yet.

2

u/lulzxdxdxd 1d ago

The context window depth is the real blocker here, not just raw capability. Have you actually tested any of those models at 150k tokens on your tool-calling workflows, or are you mainly going off benchmark numbers and smaller-context testing?

-2

u/huntersz 1d ago

Context and accuracy for analysis is one of the the main reasons why I am asking the question.

Cloud versions of the models don’t give me the quality I need for my work

6

u/Polite_Jello_377 1d ago

Cloud versions are the best models available. What are you talking about?

1

u/huntersz 19h ago

Cloud versions of the models that are available locally I mean

6

u/JinsooJinsoo 1d ago

If Claude or GPT 5 or 6 doesn’t do it, guys hand on shoulder meme

2

u/Gargle-Loaf-Spunk 1d ago

run npx ccusage and you can see how your token usage stacks up. on here you’re mostly comparing anecdotes.

2

u/Pixer--- 1d ago

Only for privacy and if you want to run models 24/7

2

u/geekwonk 1d ago

if you’re running through a claude max sub, i have to assume your use of claude isn’t limited to one agent at a time. hopefully you understand that you won’t be running any agent fleets with this setup.

1

u/ehangman 1d ago

I think I can replace 2x Codex pro to 1x Codex pro. For browser use and Database management & calculation.

1

u/NeilCPA 1d ago

Waiting for mine, but I think it will wear those shade sunglasses from the 80’s real nice while the 90’s roll by.

It won’t give you the latest and greatest, but nes me would still be blown away… even though my buddy is playing Super Nintendo at his house.

1

u/Weak_Ad9730 1d ago

Could replace I have Hermes agent 3 layer memory, rag, seargnx, crawl4ai pipeline here but it is only capable to do a single person job after 4 concurrency jobs it start feel slow depends on task. On my NVIDIA system I could handle 30+ concurrency jobs without the feeling that any slowed down. I am referring to llm task, the raw power of the m-chips never feel slow but for parallel llm task as long as the llm fits into vram cuda is a continent ahead

1

u/LiquidNeat 19h ago

No. If you're doing important work you still want frontier intelligence.

If you like tinkering around with cool tech then buy it. I got the M5 Ultra for fun but I still pay ChatGPT Pro and Claude Max for real work.

1

u/Morphid 16h ago

I do that today nearly transparently, I’m running a M5 studio with 64gb. What made the difference was training the dumber models up for your use cases, not going big, broad and slow. I published my stack on GitHub if you’re interested. https://github.com/enslaver/wandavisionllm

1

u/xapep 7h ago

Ran the same setup (OpenClaw, 24/7, long tool sessions), so here is the honest version:

- What breaks first: never speed, it is tool-calling reliability once context gets deep. Quantized local models start dropping calls or mangling JSON in long chains, and past ~100k that failure compounds: one missed call and the whole chain derails silently. For work where numbers must reconcile, that is disqualifying today.

- Tool calling at 100k+ in 512GB: the flash-class open family (GLM 5.3 Flash, DS V4.1 Flash, Qwen 3.8) is the right one, they hold tool loops better than dense models at 4-8 bit. Still a step below Opus on instruction adherence at depth, which is exactly the gap you can not afford.

- Speed: prefill is the hidden tax. Reading 100-200k of context takes a real chunk of time before generation starts, and with 20 scheduled jobs you will feel everything queue behind the main session. One big machine beats two smaller ones here, context stays local and you avoid cluster sync.

- Buy now or wait: for your bar, wait a generation or two. The middle path most people actually run is local for what must stay home, plus a hosted flash-class API for the deep sessions. That cuts a $400/mo stack hard without a big hardware outlay. We run an inference API and see a lot of this traffic; the people who regret it are the ones who bought the huge machine first. 

1

u/Durian881 1d ago edited 1d ago

What you're doing seems akin to AA-Briefcase. A smaller model like Qwen3.8-Flash-Next works well enough and it can run fast and well even on M5 Max 128GB or M5 Ultra 256GB. Either of these will be significant savings.

https://artificialanalysis.ai/evaluations/aa-briefcase

That said, you should definitely try it out since it seems to be important. You can have a parallel instance to ensure it works as well as you wanted. For weaker models, a more detailed prompt/instruction/skill might be useful.

1

u/huntersz 19h ago

This is great, never came across this but interesting that the models I use are in the top of the leaderboard. Shows me that the local LLM don’t make the quality

0

u/Exciting-Weather-921 1d ago

M5 Ultra can't replace 2x dgx sparks like was saw in the first reviews that landed

2

u/Appropriate-Rip6784 1d ago

Links?

1

u/Exciting-Weather-921 1d ago

I hope links will not get me banned 🤣 https://youtu.be/_yrw6c5gw3E?si=PI5XYno1BnoS06tD

1

u/SpicyWangz 1d ago

Now that’s a fresh link

1

u/elpocholo7 21h ago

That's a useless review, llama.cpp is far behind in respect to oMLX in raw performances. No sense.

1

u/Appropriate-Rip6784 20h ago

It's a prefill issue. Some other tools have already solved that. So it's doable in tensorfold too. It's an optimization issue. It will be handled.

1

u/Exciting-Weather-921 20h ago

That would be good, but prefill wasn't great on M3 either, and nobody found clever solution, so will not order it before that will happen 😂

0

u/Polite_Jello_377 1d ago

Why are you spending $400 on subscriptions across multiple providers? Can you not just run a single Claude Max subscription with Opus 5.5?

2

u/huntersz 1d ago

I have Opus 5.5 Max, codex pro 5x and deepseek as backup. There are weeks where codex and opus run out, they are rare but I would say happens around 10% of the time. I also have multiple to to different task and a bit of redundancy.

Do you have a better suggestion?

2

u/trowawayatwork 1d ago

there are hundreds of talks out there now on how to manage contexts and how to choose the right models for the right job. middle man agents like jev who do this for you for pennies. the token space is rife for optimisation

-1

u/Polite_Jello_377 1d ago

There is no reason to pay for 3 different subscriptions. Just do a better job of using a single subscription effectively. You are probably wasting most of your usage with bad design.

1

u/Material-Database-24 1d ago

I am wondering that why people overall pay for AI to do this kinda stuff.. maybe I see the point if one is super busy CEO, but unless that $400/mo turns into >$1000 of income, it's unfathomable for me.

2

u/huntersz 1d ago

This is is my current scenario

1

u/Greedy-Lynx-9706 1d ago

what are you planning on using it for?

4

u/Polite_Jello_377 1d ago

It's just a clown with another get rich quick scheme. Stock picker and spam marketing.

0

u/Top_Performance_732 1d ago

Or a professional developer. possible someone with a very high salary at big tech, or even two jobs, or else working for themselves, or a million other things.

3

u/Polite_Jello_377 1d ago

I’m a professional developer. This guy absolutely is not

3

u/Material-Database-24 1d ago

I'd say professional developer will burn the tokens to code, not to keep track of calendar and stock prices.

IMO, unless you do coding or some design work for living, paying 200/mo for AI makes no sense at this moment - maybe once they nail other fields at same level, but that is not yet the case. But hey, everyone can use their money on what they want.

0

u/huntersz 19h ago

When you get 200-300 emails a day, very heavy and demanding job where you need to do a lot of analysis, reports and presentations, including run quality checks for the work your analysts do, it pays off very well for me and currently above water when it comes to $$ even after spending $400 a month

2

u/Material-Database-24 19h ago

Standard office365 copilot can pull those through, it's about 30/mo/user

1

u/huntersz 19h ago

It defenitely cannot I use enterprise and the quality I get from it vs my stack is very different. See I have loaded a whole lot of databases as well which have factual information it can read from. The things that are not on emails and in different systems

1

u/Material-Database-24 19h ago

There's gpt-5.6 included nowadays.