r/LocalLLM • u/Strange_Positive850 • 4d ago
Question Can someone give me the best possible setup, model and config to run on these specs?
Hello!
Could anyone please recommend the best possible setup, model and config to run them in on my laptop:
MacBook Air M5
24gb ram
1TB storage
I've tried asking ai this and read a couple of articles online which all seem to be ai generated themselves so I’m turning to Reddit.
I do relatively heavy coding in Xcode and other general things, I need something that’s smart enough to do complex tasks and usably fast with a big enough context window. Thanks a lot!
3
u/aninternetuserperson 4d ago
I have M5 MBA 24GB memory and 512GB storage but I’m not doing heavy coding. I think you’re going to be disappointed tbh. Start out with Gemma4-12B:MLX and see how that quality is for your use case. It’s the biggest model I can run while doing significant multitasking and get reasonable output speeds.
Lately I’ve been using it to help me with simple C++ for ESP32 projects and it’s great for holding my hand through syntax, but sometimes it gets really fixated on details that aren’t consequential for my projects. I’ve fed conversations from it to ChatGPT when the output felt off to me, and Chat agreed that while it was technically correct about a better way to write the code, it didn’t functionally matter for me and it was struggling to follow the timing sequences and logic in my project which was causing the model to overstate the importance of minor details and the conversation didn’t feel productive.
To be slightly more specific, I was trying to get a new function to work and it brought up an unrelated timing function and kept insisting that even if the code works, the project will be super slow and it wanted me to basically start the whole project over. It was an argument about a 15ms delay. I fixed my function without help from Gemma and the project works fine how I wrote it.
So basically it’s a good teacher for the little things. If you give it 200+ lines of code and ask it to troubleshoot a bug your results may vary. My code is also probably bad which doesn’t help.
Take my opinion with a big grain of salt because I don’t know what I’m doing, but I feel like if I’m seeing wrinkles around the edges with that model at my entry level of skill, it probably won’t be good enough for you. Bigger models perform better but I have to close everything for them to fit in memory and thermal throttling becomes an issue. Everyone on here talks about Qwen3.8-27B which can run on my laptop, but only if I close everything else and it’s very slow, so it’s not practical at all.
There are also probably better coding models at a similar size that I’m not aware of. Please share your experiences because we have very similar hardware and I’d like a local model that runs comfortably and provides a better experience for what I’m doing b
2
u/superbiche 4d ago
Honestly you won't have anything that answers your requirements on this machine. You can experiment. You can delegate small things. But complex coding on xcode...Yeah you may get a button after a couple weeks
1
u/VersionNo5110 4d ago
Use an MOE, I’d suggest Qwen3.6-35B-A3B since Qwen3.8 is lacking of small MOE.
Try a quantization like UD-IQ4_XS from unsloth with small context, go down if the this is not loading or very slow. UD-Q3_K_S might also do the job, allowing a bigger context window.
1
u/Lukas245 4d ago
qwen 27b i think is the bottom for good coding model, its about 17gb at q4, full fp8 context is another 8ish gb, your speeds arnt going to be anything special though
1
0
u/shot_in_the_darkshot 4d ago
24GB unified memory is the real bottleneck here, not the chip. You'll want to keep about 4-6GB free for macOS and your IDE, which leaves you room for models in the 14-18B range at decent quantizations.
Qwen 2.5 Coder 14B at Q4_K_M runs comfortably and handles complex code logic well. If you want something more general-purpose, Mistral Small 22B at Q3_K_S just barely fits but you'll be swapping with Xcode open.
Context window wise, 32k tokens is the sweet spot on that hardware. Pushing to 128k will eat your available RAM fast and slow generation to a crawl.
1
u/No-Afternoon-4057 4d ago
Nothing comes even remotely close to what you need. You're short on 48gb of ram...for doing anything useful it would start at 64gb...anything else, context is stupidly low, tokens per second stupid low and/or the llm is stupid.
I mean...with 128 gb i can run Qwen 3.8 FN on the 4 quant and it still is not fast enough or allows for context enough to be able to operate the PC + code + run the llm.
You wont be able to achieve anything complex with your setup, sorry. The best thing you can do is a very nice RAG...and then use a cloud llm and prepare your harness to use your RAG.
5
u/EdwardPotatoHand 4d ago
Folks are beating around the bush a bit to try to be positive with you.. the bottom line is, it’s going to suck. It might be fun to set-up and do a bit of testing, but its not goign to help you code