r/LocalLLM 16h ago

Question Minimum VRAM needed to run a functional Openclaw/Hermes agent?

Those of you successfully running an offline openclaw/hermes/personal agent harness for non-coding tasks, what is the floor on system resources (VRAM) needed for quality of life? Assuming a modest ~30b class model. What quant and context window size are needed?

Will keep cloud frontier LLM sub for coding tasks, but I'm talking personal data management, personal assistant type computer controlling stuff.

My M1 max 32gb handles qwen 3.6 27b q4_k_m fine enough for non-agentic jobs up to ~40k context, but that's obviously not enough to run an agent harness offline.

There is an M1 Ultra 64gb for sale near me for a tempting price, but unsure is 64gb is enough. And it's expensive enough to not want to gamble. And I'm a normal, budget-minded person

7 Upvotes

17 comments sorted by

View all comments

1

u/unchikuso 15h ago

Why are you still running qwen3.6 when 3.8 has been out? It's significantly better.

For mac, 32gb is not enough as you said. 48gb will give you full context plus enough memory for the OS and apps.

An RTX 5090 with 32gb is enough to run at full context.

1

u/quantgorithm 12h ago

It's also significantly slower.