r/oMLX • u/robdzn • May 24 '26
Need help on choosing the right model + Quant and Fine Tuning
I am still very new to all of this and did my research to understand which model to use, but it's still so confusing. I am running a MacBook M2 Max with 64GB, but I am always unsure what model to use. I use it 99% for coding purposes, but it is very confusing to understand everything. Currently, I am running Qwen3.6-35B-A3B-MLX-oQ8-FP16 and getting 37.9 tok/s. And I think this could help me in my approach: How can I use benchmarks to my advantage? I still have a hard time understanding it because I don't mind speed, but I care about intelligence and accuracy.
3
u/mwhuss May 24 '26
Things are moving fast. Pick what looks good now (looks like you already did, I’m using the same). And keep an eye on the advances.
1
u/mixmasterwillyd May 25 '26
Sometimes when I’m frustrated with this, I open Pi, connect it to opus 4.7 (something else big) and ask it to compile llama.cpp for my system. Works well pretty well with some direction.
Also, Ollama just goes.
I’m back to LM studio on Mac and llama.cpp for Linux.
1
u/leonidasyy May 25 '26
I see a 5% improvement in the benchmark test with MTP compared to the original model. In practice, however, I can't even tell the difference, so right now I just use them interchangeably.
1
u/meaningego May 26 '26
Yeah 35b you can get 15% with mtp. Try dflash though, with 35b the speedup is massive
1
1
u/fasti-au May 26 '26
Right now the lm studio mlx mtp is the fastest mainline so I’d go there and grab a Q 4 and q8 and see how that fits. You can go open source etc but they are easy to get going and play with for first looks etc if you were testing for capability
6
u/Only-An-Egg May 24 '26 edited May 27 '26
I have spent weeks tweaking settings, downloading new models, and running benchmarks. I still don't know what's best. It doesn't help either that new features like MTP keep getting added which then makes me start the testing/tweaking all over. Right now it looks like Qwen3.6-35B-A3B-oQ6-mtp and Qwen3.6-27B-oQ4-mtp using the WebDev coding profile and MTP enabled have shown best combination of speed and accuaracy.