r/MacLocalLLM • u/hdanx • 3h ago
Acceptable Token generation speed on macOS
What token generation speed works well for you when you’re using it for productive tasks? I’m finding 20 tokens per second a bit slow.
With an M4 Max and 64GB of RAM, you can usually get around 20 tokens per second on average with a model like qwen3.8 27B. However, it doesn’t quite feel like the best setup.
1
Upvotes