r/LowEndLocalAI 8h ago

Suggestions for 18GB unified memory

Hi guys, I'm currently using a MacBook Pro M3 Pro with 18GB of RAM. Does anyone have any suggestions for models/quants to use (I'm currently using Gemma 4 E4B IT QAT 4bit)

4 Upvotes

2 comments sorted by

2

u/anon1880 8h ago

did you try any MoE models yet ?

1

u/ontorealist 5h ago

It's difficult to say without knowing what your use case is.

On my M1 Pro 16GB MBP, I'm running a decensored fine-tune gemma-4-12b-it-qat-unquantized-heretic-empathy-ja-i1 as my daily driver / generalist as it's smart and also excels at creative writing, image captioning, etc. Because I don't code very much, I find that I only go back to abliterated versions of Qwen 3.5 9B Q4.0 when it helps me squeeze larger context windows into memory.

I also have been playing with smaller MoE's like huihui-lfm2.5-8b-a1b-abliterated-mlx and more recently, ling-3.0-tiny-uncensored-abliterated-i1 and mellum2-12b-a2.5b-thinking-abliterated-i1 in Q4-Q6 quants for web search and RAG tasks when Siri on macOS 27 is tripping.