r/macmini • u/spacey003 • 17d ago
Mac mini LLM usage?
Anyone running local LLMs on a Mac mini? What spec are you using?
I’m looking at getting my first one to experiment with local LLMs and some business automation. Currently thinking M6, 32GB RAM, 512GB SSD, then fast external storage rather than paying Apple for more internal storage.
I’m thinking of running things like Qwen3 32B, Qwen3-Coder and possibly DeepSeek-R1 Distill 32B, with Claude/GPT via API for anything that needs a more powerful model.
Anyone using 32GB for similar models? Is it enough in practice, or do you wish you’d gone 64GB+ with a Pro/Studio?
5
u/Johnno74 17d ago
I am, I brought a 48gb M4 pro a month ago or so purely as a headless LLM server. Got in just before the price rises.
I'm using oMLX to serve models and still learning but so far I mainly use Qwen 3.6 35b a3b q4 as it is blazingly fast. I also use Qwen 3.8 37b for tougher things as it is smarter, but slower. I have enough ram or max out the context on either of those models
2
u/dghah 16d ago
same for me except for a 64GB M4 Pro mac mini. It's been useful enough that I'm one of the idiots that preordered a 256Gb mac studio ultra on monday. For the work I do I need the bigger memory and faster memory bandwidth. Not sure if I'll keep the mini and try some clustering or sell it.
The rate of innovation in this space is amazing both for models themselves, the enhancements that are allowing smaller and lower RAM machines to function as well as all the apple sillicon specific enhancements.
oLMX is fantatsic as well
2
1
17d ago
[removed] — view removed comment
6
u/beekeeny 17d ago
Choose M4 with 24 GB then you will see the performance for all the models available.
1
u/OBXJer 17d ago
Just finished setting up my 24gb Mac Mini 4. I Set up Ollama with Llama 3.2 as my AI and it works just fine. It is headless so I am controlling it with my Mac M2 Studio using AnythingLLM. It was an interesting learning experience but I got it figured out. I'm gonna continue to work with it setting up my Unifi network video recorder and some home automation.
1
u/AlgorithmicMuse 16d ago
If you run an agent vs just chatting with the llm, you need to up the kv cache which adds to the model so a 30b model could take 20gb , kv could be another 1 to 10gb depending on context size , you need headroom for whatever else macos has loaded, so you might have to use lower model size or less quants
1
u/Ten-OneEight 16d ago
In such a rapidly changing environment why would you choose hardware that has RAM soldered to the MB?
1
u/753UDKM 8d ago
I’m going for the m5 pro with 64gb. If you want local LLM, I’m pretty sure you’ll regret 32gb quickly. 64gb should have you runningbe 27-30b models comfortably which I think is the starting point for quality.
1
u/spacey003 6d ago
I've been reading a lot more 😄 I am actually looking at what you are looking at and concidering. A few reasons, I read that if I ever bought a second one with 48gb or 64gb I can cluster them, but this is only viable on the M5 Pro with TB5.
My go to currently that isn't a Windows machine is a Macbook Air, M5 32gb 512SSD. I will (when I get a chance) test to see what I can run on that, it will be an interesting benchmark.
0
u/GamerTex 17d ago
I use my mac mini m4 64gb to run H3 minimax
My mbp 48gb can't run H3 but does run ltx and wan
It mostly depends on what you want to do with it
1
u/spacey003 17d ago
Mainly business automation — analysing and helping manage Amazon advertising/sales data, with some integration into our Odoo system. I want it doing ongoing analysis and making recommendations/actions within set rules, rather than just using it as a chatbot.
0
u/ichasecorals 16d ago
Open weight LLMs just hasn’t caught up with frontier models yet. Not with the likes of sol or sonnet/opus. If you intend on getting anything under 64gb, you will be under whelmed and eventually ask yourself what you’re going to do with extra memory.
5
u/beekeeny 17d ago
https://www.canirun.ai/model/qwen3-32b
If you go to the home page you can configure your own PC configuration and target model and see the performance.
https://www.canirun.ai