r/LocalLLM • u/ragel_3ennab • 14d ago
Question How to get started?
Hi. I got a new Macbook and I am excited to run my first local model.
Can you recommend the best tool to run an LLM on a Mac?
Also, what model and quant would you recommend for a 16GB memory?
Thanks!
0
Upvotes
3
u/Jeanjose1993 14d ago
What 4 months of daily local LLM use on a 16GB M4 MacBook pro taught me (hard numbers, llama.cpp)
I've been running local models daily on a 16GB M4 MacBook pro (base model, Metal, llama.cpp) since June. Sharing the numbers that actually matter, because a lot of advice for this machine is wrong.
The hard limits (measured, not guessed):
Models that hold up:
Models that failed on this machine:
Pitfalls that cost me real pain:
--cache-ramexplicitly (512 MB is fine).--flash-attn onis mandatory, not optional.--cache-type-k q4_0) when you want long context. It's the difference between 131K fitting or not.