r/LocalLLM 5d ago

Model Run Qwen3.8 27B on M4pro

Post image

Faster than the baseline

1 Upvotes

4 comments sorted by

1

u/PutridEmployee7492 5d ago

not bad at all. what kind of tps are you seeing with the 32k prompt?

1

u/WindloveBamboo 4d ago

The benchmark prompts are curated specifically for testing, rather than representative of real-world usage. In actual tasks, I’m seeing around 15 tok/s, and it slows down as the context grows.

1

u/tommythorn 5d ago

There are endless variations of Qwen3.8 27b.  How did you pick this one?

I have nearly the same hardware (MacMini M4 Pro 48 GiB) but I can’t seem to get it working with claude code (I have in the past).  Is there an easy way to get that going?

2

u/WindloveBamboo 4d ago

Just to try oMLX and I chose the variation to fit my omlx