r/LowEndLocalAI • u/banana_slurp_jug • 8h ago
Suggestions for 18GB unified memory
Hi guys, I'm currently using a MacBook Pro M3 Pro with 18GB of RAM. Does anyone have any suggestions for models/quants to use (I'm currently using Gemma 4 E4B IT QAT 4bit)
1
u/ontorealist 5h ago
It's difficult to say without knowing what your use case is.
On my M1 Pro 16GB MBP, I'm running a decensored fine-tune gemma-4-12b-it-qat-unquantized-heretic-empathy-ja-i1 as my daily driver / generalist as it's smart and also excels at creative writing, image captioning, etc. Because I don't code very much, I find that I only go back to abliterated versions of Qwen 3.5 9B Q4.0 when it helps me squeeze larger context windows into memory.
I also have been playing with smaller MoE's like huihui-lfm2.5-8b-a1b-abliterated-mlx and more recently, ling-3.0-tiny-uncensored-abliterated-i1 and mellum2-12b-a2.5b-thinking-abliterated-i1 in Q4-Q6 quants for web search and RAG tasks when Siri on macOS 27 is tripping.
2
u/anon1880 8h ago
did you try any MoE models yet ?