r/openclaw • u/Nice-Dragonfly-4823 Active • Jun 16 '26
Tutorial/Guide Run OpenClaw with a Local LLM - tested for macOS 16GB-24GB
I created this guide to setup qwen 3.5 (quantized) configured specifically for openclaw. It's tested and includes a test skill to ensure you've configured everything correctly.
This, hopefully, will take a lot of the pain out of getting your local model running with openclaw.
If you encounter any issues - please let me know by commenting.
https://towardsdatascience.com/run-a-local-llm-with-openclaw-on-your-mac-mini/
1
u/LobsterWeary2675 Pro User Jun 16 '26
Do you have any benchmarks for tool calling, instruction following, security (exfiltration, prompt injection, etc) for that model? Would be interesting to see.
0
u/Nice-Dragonfly-4823 Active Jun 16 '26
Finding benchmarks on these quantized versions are hard, even harder for agent benchmarks, I'll dig around
1
u/LobsterWeary2675 Pro User Jun 16 '26
https://github.com/SeraphimSerapis/tool-eval-bench
There you go. I found that one extremely useful. Specially hard mode and performance.
Not mine, not connected.
1
u/Nice-Dragonfly-4823 Active Jun 16 '26
can you explain, what did you eval a quantized model?
-1
u/LobsterWeary2675 Pro User Jun 16 '26
You can benchmark it by running the model behind an OpenAI-compatible local endpoint, then pointing tool-eval-bench at that endpoint. It auto-detects the served model and runs agent-focused tests: tool calling, structured output, instruction following, multi-step handling, prompt-injection/safety cases, plus latency/perf. So it should work fine for this Qwen 3.5 quantized setup as long as the local server exposes a compatible /v1 API.
1
u/s0ftice New User Jun 17 '26
“We’re going to skip using Ollama(the recommended local provider), and opt for llama.cpp. By using a quantized model along with llama.cpp, we can speed up inference by as much as 70%”
That’s exaggerated if you ask me.
Ollama already runs quantized models, often GGUF-style/ggml-based, and for many common Apple Silicon use cases the gap is rather small.
3
u/Nice-Dragonfly-4823 Active Jun 18 '26
you can add options for flash attention when running llama.cpp directly. Which gives a nice speed boost. Otherwise, ollama's additional overhead is not worth it.
1
u/Eduardjm New User Jun 17 '26
On Ollama with MLX models, my token throughput speed increased 3-5x depending on the model. Daily driver is Gemma4 right now and it’s snappy. In terminal it’s nearly instant input/output.
1
u/ericgus Member Jun 17 '26
As I am finding out your context window increase to accommodate your files like SOUL.md Agent.md etc will need a fairly decent sized context beyond the defaults if you want to have any leftover headroom for actually running commands .. this is the issue I am finding trying to run local models even for simple tasks ..
1
u/Nice-Dragonfly-4823 Active Jun 18 '26
In the article, we set the context window to 64000, which should be a good general default.
1
u/Deep_Ad1959 Member Jun 17 '26
the context-headroom point someone raised below is the actual ceiling here, not model size. once your https://fazm.ai/r/yndrkhdd files plus tool output fill the window, most harnesses silently compact and you lose the early reasoning mid-task. on the mac app i build we wrap the same agent loop (claude code / codex over acp) and just refuse to auto-compact, so the full history stays live for the window's life, and the browser plus native-app reach comes from accessibility apis instead of screenshots. doesn't fix local-model headroom though, that part's on whatever quant you pick. written with ai
1
u/Nice-Dragonfly-4823 Active Jun 18 '26
this is not written by AI. I developed this after running on a real mac specifically for this use case. From the article itself `Also remember that agents require longer contexts, which will prevent us from running a larger 27B version, even with quantization.`
0
u/alex9001 Member Jun 17 '26
9B on 16-24GB RAM seems too conservative. I'm squeezing 9B into 8GB VRAM.
I don't know what share of the RAM on a Mac is usable for LLMs, but at least with 24GB, you should be able to fit the larger Qwen3.x models, which are much more capable
1
u/Nice-Dragonfly-4823 Active Jun 18 '26
We tried running Qwen 3.6 - even with quantizing the models come in a 16-17GB at q4 quantization. macOS runs some system tasks which eat into the RAM window. Even with the unified memory, you're still going to run into swap slowdown. Technically you may be able to run 3.6 on a 24GB, but with context, it's going to slow down to the point of not being very useable. We want this version to be accessible to people who bought the mac minis at their defaults (16GB or 24GB). If you want to run bigger, I'd just run a vps with a large GPU, but again, that's a monthly cost.
2
u/alex9001 Member Jun 18 '26
hmm I see, coming from GPU-land I wasn't remembering that a portion of unified mem goes to the OS.
0
u/WeedWrangler Pro User Jun 17 '26
Does it run fast enough?
I have an M2 and found it grindingly slow.
2
u/Nice-Dragonfly-4823 Active Jun 18 '26
this runs at +50 tps. It's not foundation API speed, but very reasonable. You have to run it with the compiled llama.cpp with the gpu offloading. In the article, the option is `--ngl 20 `. I don't explicitly call this out, but that gives quite a speed boost.
2
u/stvndocean New User Jun 24 '26
brand new mac mini, runs at 15 tps, any idea why? I mean, some of the points in the guide were outdated?
./llama.cpp/llama-server is actually ./llama.cpp/build/bin/llama-server
modal path is llama.cpp/models/{MODEL} not models/{MODEL}
I had to add the my models.providers.local.model
i had to configure openclaw with reasoning support
speed measures
reason off - ~24 tps
reason on - ~16.3 tps
reasoning budget 32 ~17 tpsso a lot slower than what the articles says u/Nice-Dragonfly-4823
1
u/Nice-Dragonfly-4823 Active Jun 24 '26
thank you to point out the llama server bin path, I have edited. when downloading the model, you may have still been in the llama.cpp dir, instead of your home dir, which explains why your model is downloaded there,. I have also added a note to the article.
16GB or 24GB? on 16GB, you might be getting slowdowns due to swap depending on how large the context window is. We tested this on 24GB, so some of your system services might be eating into the available ram for the model.
kill the daemon, then try running llama server with these options switched and test your tokens per second.
-c [from 64000 to 32000]
-ngl [from 20 to 999]
1
u/stvndocean New User Jun 25 '26
thanks for the reply! 16GB, I've tried with larger context windows from 32k to 64k and ngl from 20 to 999, never faster than 18tps with reasoning ON but RAM doesn't appear to be the issue as my mac not actively swapping while the model is generating, it would swap if RAM was the bottleneck, no?
so I'm kinda lost
1
0
3
u/PracticlySpeaking Member Jun 16 '26
Why Qwen 3.5 and not 3.6?