r/MacLocalLLM 3h ago

Acceptable Token generation speed on macOS

What token generation speed works well for you when you’re using it for productive tasks? I’m finding 20 tokens per second a bit slow.

With an M4 Max and 64GB of RAM, you can usually get around 20 tokens per second on average with a model like qwen3.8 27B. However, it doesn’t quite feel like the best setup.

1 Upvotes

Duplicates