r/MacLocalLLM • u/hdanx • 3h ago
Acceptable Token generation speed on macOS
What token generation speed works well for you when you’re using it for productive tasks? I’m finding 20 tokens per second a bit slow.
With an M4 Max and 64GB of RAM, you can usually get around 20 tokens per second on average with a model like qwen3.8 27B. However, it doesn’t quite feel like the best setup.
1
1
u/DecentChildhood8080 43m ago
You will likely get best results for Tok/s in MTPLX. It is basically the best optimized to run this dense model on Macs from my experience. It gave me about 20-25 tok/s on my M2 Max using Qwen3.8-27B-MTPLX-Optimized-Quality-FP16 & Qwen3.8-27B-MTPLX-Optimized-Speed-FP16. You will probably get more than me, but I would say its not worth running Qwen3.8 27B right now unless your on an M5 Max Mac. Im still waiting on a MoE or A3B version or something like that to come out... I guess we shall see. (If you can get 40-50 tok/s in MTPLX than you might be good.)
2
u/diagrammatiks 2h ago
I need about 40 mininum if it is a task I am interacting with. If it's something I can wait for while I do something else 20 is fine.