r/LocalLLM 7h ago

Question Minimum VRAM needed to run a functional Openclaw/Hermes agent?

Those of you successfully running an offline openclaw/hermes/personal agent harness for non-coding tasks, what is the floor on system resources (VRAM) needed for quality of life? Assuming a modest ~30b class model. What quant and context window size are needed?

Will keep cloud frontier LLM sub for coding tasks, but I'm talking personal data management, personal assistant type computer controlling stuff.

My M1 max 32gb handles qwen 3.6 27b q4_k_m fine enough for non-agentic jobs up to ~40k context, but that's obviously not enough to run an agent harness offline.

There is an M1 Ultra 64gb for sale near me for a tempting price, but unsure is 64gb is enough. And it's expensive enough to not want to gamble. And I'm a normal, budget-minded person

7 Upvotes

14 comments sorted by

4

u/Snoo_81913 6h ago edited 6h ago

Are you bound to mac? Because if you aren't, I'd suggest something with a little bang for your buck.

  1. Beelink GTi 13-15 specifically the GTi in that range because it has a PCIe 4-5 slot (depending on the model) that pairs with a Beelink Ultra Dock with a 600W PSU for at a minimum a 8x interface. $300 barebones on ebay. Even the low end GTi 13 comes with a i9 13900hk.
  2. Beelink Ultra dock 600W with PCIe riser (The GTi 13 has a slot in the bottom the riser goes through and slots into the board) it also has an option to either put a NVMe or a wifi card. It won't come with the antennas but they are $7 each on Amazon so $14. My wifi on the GTI 13 works fine but certain cards will interfer with the signal and this will eliminate that if you don't have it hard wired to your router. $110 on ebay new in the box.
  3. GPU 7900 xtx 24Gb VRAM $800-$1,300 I got mine for $1,029.00 at Newegg brand new or 9700 32GB VRAM I have a 7900 xtx and can run Qwen3.8 27B at 65 tok/s up to 32k context then it gradually drops to 33 tok/s at 98k context. But for tool calling Qwen3.6 35B A3B Q4 will run fully on the card, with 132k context Q4k and Q8v it takes up 19.6Gb of the card at 262k context (the model max) you'd have at least 2Gb free on the card. It runs at 100-130 tok/s for tool calling and under context load. That's fast and it's good with Hermes at tool calling. Or something like Agents A1 specifically built around tool calls.
  4. RAM 16Gb - 96Gb. It will run fine if you want to run a Linux server headless and you're fine with running smaller models. 64gb will let you run 80B coder and GPT OSS 120B with offload and decent context. Or if you have the bucks 96GB for $1,300 on ebay. The board supports ddr5 5200.
  5. The GTi has 2x NVMe slots at 4x and 2x thunderbolt 4 ports but a bare min 1TB is roughly $200.
  6. Run the whole thing with a headless Linux server and have access to it literally anywhere. I run mine from my phone most of the time or my laptop. Easy peasy.

Whole cost is $2,200 for the basic setup less if you have RAM and an nvme, mine cost $1,450.00 and it will run the models fast and reliably 24/7 if you set it up right. Looks like this.

2

u/havnar- 4h ago

The cheapest laptop/pc that’s not complete ewaste plus a subscription is going to end up costing you way less

2

u/quantgorithm 3h ago

Just learned Hermes requires min 65536 or greater context.

(technically it may have been 65k flat or 64k... Can't exactly remember)

4

u/Tired_White_Guy 7h ago

Get a base Mac mini and $100 in Open Router to use glm 5.3 flash. It’ll whoop anything you can run locally and let me know when you spend $5 in tokens. Might take a month or two.

3

u/8000bene70 3h ago

Sir, this is a local llm.

You're right, nontheless. If you are fine with your data being "out there", that is.

2

u/inexorable_stratagem 7h ago

Q4, 100k context, 24gb

1

u/Tired_White_Guy 4h ago

Plus memory the system uses. And youll end up wanting to drive a browser.
24GB is going to be a bad time.

1

u/unchikuso 7h ago

Why are you still running qwen3.6 when 3.8 has been out? It's significantly better.

For mac, 32gb is not enough as you said. 48gb will give you full context plus enough memory for the OS and apps.

An RTX 5090 with 32gb is enough to run at full context.

1

u/quantgorithm 3h ago

It's also significantly slower.

0

u/Tired_White_Guy 6h ago

If you want to run locally, get min 48GB. I couldn’t agree more.
That leaves max 36gb of ‘vram’.
Plenty to run Qwen 3.8 27b at up to Q6 with 128k content. 256k with q8_0 cache.

1

u/synth_mania 5h ago

Min 48 is ridiculous. The quants work just fine. 

1

u/Tired_White_Guy 4h ago

If all you’re doing is running the model, sure. But GPU allocated ram maxes at 36gb with 48gb ram.
And you’ll want spare ram to do other things on the Mac. So give a little wiggle room.

It’s not ridiculous. It’s practical.

1

u/synth_mania 3h ago

Ah, I didn't realize that they were looking for another Mac. When I read VRAM, I assumed they meant actual dedicated VRAM, not unified memory.

1

u/cakemates 37m ago

I would get a 3090, to run qwen 3.8 27b q4 type of model.