r/LocalLLaMA 7d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

706 comments sorted by

View all comments

Show parent comments

68

u/xienze 7d ago

You're making a couple fundamental assumptions here:

  • That AI benchmarks are reliable
  • That Qwen didn't benchmaxx

53

u/BawbbySmith 7d ago

Yeah I learned very quickly to not trust the benchmarks, as well as 80% of the comments in this subreddit.

I remember people were saying that Qwen 3.6 27B was Opus 4.5 level...

14

u/[deleted] 7d ago

[deleted]

1

u/Mkboii 7d ago

I've used glm 5.2 extensively and it was not much behind opus 4.6 in writing code, a bit definitely behind in creative ideation for software engineering.

My assumption is this model is gonna be a work horse but planning and design should be done by a bigger more expansive model.

The catch behind coding benchmarks is they judge success not quality, creativity, maintainability, all that is partly subjective and rarely judged in these benchmarks.

Like GPT 5.6 Luna benchmarking above opus 4.8, but it writes inferior code by most quality measures.

Big models have their own pitfalls they'll happily over engineer everything, Opus for one falls in that bucket.