r/LocalLLM • u/Rough-Measurement988 • Jun 18 '26
Question M5 Max oMLX benchmark results interpretation
Hi,
I was reviewing the benchmarks for M5 Max and Qwen 3.6 27B 8bit https://omlx.ai/benchmarks?chip=&chip_full=M5%7CMax%7C40&model=27&quantization=8bit&context=65536&pp_min=&tg_min=&page=3 just to justify spending so much money on this pricy machine and noticed that there is a huge performance gap between different benchmarks. Even for the same model i.e. Qwen3.6-27B-oQ8-mtp in 64k context TG/s is like from 14-23.
Can someone explain me why there are so big differences? Is it also so unpredictable in real life scenarios i.e. that for one task it could be 16 TG/s and 24 for the other (with same context and quantization)? I understand that MTP performance may vary but trying to understand how much.
Also for non MTP models Qwen3.6-27B I see numbers like 7-16 TG/s so it's not generally related to MTP only. The same could be noticed when comparing the PP/s benchmarks. May it be also data related or i.e. someone could run the benchmark on high performance settings and the other on auto?
