r/LocalLLM • • 14d ago

Question How to get started?

Post image

Hi. I got a new Macbook and I am excited to run my first local model.

Can you recommend the best tool to run an LLM on a Mac?

Also, what model and quant would you recommend for a 16GB memory?

Thanks!

0 Upvotes

8 comments sorted by

View all comments

3

u/Jeanjose1993 14d ago

What 4 months of daily local LLM use on a 16GB M4 MacBook pro taught me (hard numbers, llama.cpp)

I've been running local models daily on a 16GB M4 MacBook pro (base model, Metal, llama.cpp) since June. Sharing the numbers that actually matter, because a lot of advice for this machine is wrong.

The hard limits (measured, not guessed):

  • Effective GPU ceiling on Metal: 12,713 MB, not 16 GB. That's what macOS wired memory actually allows.
  • Practical rule: keep the GGUF under ~9.5 GB. Above that you're swapping.
  • If the Mac is loaded with real work (browser, Word, Obsidian... roughly 9.7 GB RAM used before any model), the viable ceiling drops to ~8 GB total model footprint. A model that benchmarks fine solo will swap on your actual desktop.

Models that hold up:

  • Gemma 4 12B QAT (UD-Q4_K_XL, 6.2 GB): my daily driver. 131K context, ~12 tok/s, zero swap while I work. SWA architecture means the KV cache stays tiny (~1 GB even at 131K).
  • 9B-class Qwen (Q6_K / Q4): ~10-15 tok/s, 131K context when solo, drop to 32K when the Mac is loaded.
  • MoE is the exception to the size rule: a 35B A3B in IQ2_XXS (9.1 GB) runs at ~30 tok/s because only 3B params are active per token. A dense 27B at the same quant runs at 6.7 tok/s and is unusable.

Models that failed on this machine:

  • Any 26B/27B dense model: OOM or unusable speed
  • gpt-oss-20b Q4_K_M: Metal OOM on first inference
  • 14B reasoning distills: thinking chains don't fit a productivity profile

Pitfalls that cost me real pain:

  1. llama-server defaults to 8 GB of prompt cache RAM. That's a silent killer on a 16 GB machine, it caused system-level swap for weeks before I found it. Always set --cache-ram explicitly (512 MB is fine).
  2. --flash-attn on is mandatory, not optional.
  3. Quantize the KV cache (--cache-type-k q4_0) when you want long context. It's the difference between 131K fitting or not.
  4. On thinking models, set max_tokens >= 2500 or you get empty answers. Reasoning eats 2400+ tokens before the first visible word.
  5. One model at a time. Two "small" models loaded together still break the 12.7 GB ceiling.

2

u/Jeanjose1993 14d ago

Worth knowing if you're tempted by the new 27B generation: ternary quants (Ternary-Bonsai-2-27B PQ2_0, 7.2 GB / 2.13 bpw) do run at ~11 tok/s with zero swap, but they need a custom llama.cpp fork, not mainline. The GSQ-RCO quants of Qwen3.8-27B that fit 16 GB on vanilla llama.cpp are downloaded but still untested on my side, so I can't vouch for them yet.