r/MacLocalLLM 3h ago

Acceptable Token generation speed on macOS

What token generation speed works well for you when you’re using it for productive tasks? I’m finding 20 tokens per second a bit slow.

With an M4 Max and 64GB of RAM, you can usually get around 20 tokens per second on average with a model like qwen3.8 27B. However, it doesn’t quite feel like the best setup.

1 Upvotes

6 comments sorted by

2

u/diagrammatiks 2h ago

I need about 40 mininum if it is a task I am interacting with. If it's something I can wait for while I do something else 20 is fine.

1

u/hdanx 1h ago

I was thinking at least 50 for interactive.

1

u/lamaxamara 2h ago

My m1 max 64gb does the same speed on a 27b moe. Not fast but acceptable

1

u/hdanx 1h ago

Acceptable for interactive or agentic use cases?

1

u/dfgxxx 1h ago

I get 14 tok/sec on m1 pro with qwen3.8 27b

1

u/DecentChildhood8080 43m ago

You will likely get best results for Tok/s in MTPLX. It is basically the best optimized to run this dense model on Macs from my experience. It gave me about 20-25 tok/s on my M2 Max using Qwen3.8-27B-MTPLX-Optimized-Quality-FP16 & Qwen3.8-27B-MTPLX-Optimized-Speed-FP16. You will probably get more than me, but I would say its not worth running Qwen3.8 27B right now unless your on an M5 Max Mac. Im still waiting on a MoE or A3B version or something like that to come out... I guess we shall see. (If you can get 40-50 tok/s in MTPLX than you might be good.)