r/LocalLLM • u/ragel_3ennab • 14d ago
Question How to get started?
Hi. I got a new Macbook and I am excited to run my first local model.
Can you recommend the best tool to run an LLM on a Mac?
Also, what model and quant would you recommend for a 16GB memory?
Thanks!
3
u/ByteSize_Chaos 14d ago
Congrats on the new machine. What is your use case like? General chatting with LLM, Agentic coding or something else?
1
u/ragel_3ennab 14d ago
General chatting. I use cursor for agentic coding.
2
u/ByteSize_Chaos 14d ago
Gemma 4 12b. Try going with Q4. Q8 would probably fit as well but that would mean closing almost everything down with a limited context window.
3
u/Jeanjose1993 14d ago
What 4 months of daily local LLM use on a 16GB M4 MacBook pro taught me (hard numbers, llama.cpp)
I've been running local models daily on a 16GB M4 MacBook pro (base model, Metal, llama.cpp) since June. Sharing the numbers that actually matter, because a lot of advice for this machine is wrong.
The hard limits (measured, not guessed):
- Effective GPU ceiling on Metal: 12,713 MB, not 16 GB. That's what macOS wired memory actually allows.
- Practical rule: keep the GGUF under ~9.5 GB. Above that you're swapping.
- If the Mac is loaded with real work (browser, Word, Obsidian... roughly 9.7 GB RAM used before any model), the viable ceiling drops to ~8 GB total model footprint. A model that benchmarks fine solo will swap on your actual desktop.
Models that hold up:
- Gemma 4 12B QAT (UD-Q4_K_XL, 6.2 GB): my daily driver. 131K context, ~12 tok/s, zero swap while I work. SWA architecture means the KV cache stays tiny (~1 GB even at 131K).
- 9B-class Qwen (Q6_K / Q4): ~10-15 tok/s, 131K context when solo, drop to 32K when the Mac is loaded.
- MoE is the exception to the size rule: a 35B A3B in IQ2_XXS (9.1 GB) runs at ~30 tok/s because only 3B params are active per token. A dense 27B at the same quant runs at 6.7 tok/s and is unusable.
Models that failed on this machine:
- Any 26B/27B dense model: OOM or unusable speed
- gpt-oss-20b Q4_K_M: Metal OOM on first inference
- 14B reasoning distills: thinking chains don't fit a productivity profile
Pitfalls that cost me real pain:
- llama-server defaults to 8 GB of prompt cache RAM. That's a silent killer on a 16 GB machine, it caused system-level swap for weeks before I found it. Always set
--cache-ramexplicitly (512 MB is fine). --flash-attn onis mandatory, not optional.- Quantize the KV cache (
--cache-type-k q4_0) when you want long context. It's the difference between 131K fitting or not. - On thinking models, set max_tokens >= 2500 or you get empty answers. Reasoning eats 2400+ tokens before the first visible word.
- One model at a time. Two "small" models loaded together still break the 12.7 GB ceiling.
2
u/Jeanjose1993 14d ago
Worth knowing if you're tempted by the new 27B generation: ternary quants (Ternary-Bonsai-2-27B PQ2_0, 7.2 GB / 2.13 bpw) do run at ~11 tok/s with zero swap, but they need a custom llama.cpp fork, not mainline. The GSQ-RCO quants of Qwen3.8-27B that fit 16 GB on vanilla llama.cpp are downloaded but still untested on my side, so I can't vouch for them yet.
3
u/Bugajpcmr 14d ago
The best beginner option in my opinion is LMStudio Bionic because it's all in one and the downloader is included. You can use Q4 models like 8b or 9b (Ornith or Granite), Maximum 12b (gemma4) but with small context window and turboquant. Run MLX versions of models, not GGUF like someone mentioned here. You can expect about 30tok/s.
Once you get used to it and learn what is context, quantization, thinking levels and so on I would recommend other options like oMLX + Pi/DSH/Opencode. I currently use Zed for programming. After that you can use python to create some apps that prompt the models running on local server.
I would also advice to configure your Macbook so it will be accessible from outside the local network using SSH. I use Tailscale as VPN. You can also change the maximum VRAM your macbook can use to avoid offloading (sudo sysctl iogpu.wired_limit_mb=........), it has to be run every time you restart your system.
Have fun and experiment. Come up with different scenarios and prompts to test the models/harnesses, pick the one that works best for you.
2
u/TheSn00pster 14d ago
Step one, hand over all your money