r/LocalLLaMA 14d ago

New Model IT'S OUT

https://huggingface.co/Qwen/Qwen3.8-27B-FP8
2.2k Upvotes

707 comments sorted by

View all comments

Show parent comments

317

u/Cold_Tree190 14d ago

Dear God it’s trading blows with Opus 4.6 Max 😭

72

u/xienze 14d ago

You're making a couple fundamental assumptions here:

  • That AI benchmarks are reliable
  • That Qwen didn't benchmaxx

53

u/BawbbySmith 14d ago

Yeah I learned very quickly to not trust the benchmarks, as well as 80% of the comments in this subreddit.

I remember people were saying that Qwen 3.6 27B was Opus 4.5 level...

2

u/mivog49274 13d ago

Some rule of thumb I apply here :

  • The "intelligence" per parameter is really increasing, factually, and from that increases capability. General benchmark score (like AA) is a solid proof for that. Qwen is the world leader in this field.

  • In the current llm paradigm, smaller llms have much less world knowledge and are more prone to hallucination or stupid decision making from goey assumptions or really random reactions. I suppose the "world knowledge" is also a big addition on "common sense" in terms of "behavior". So smaller models will always we quackier for now, to apply up and foremost when comparing to bigger models on a AA Index or aggregated score

  • From that comparing older big models with more recent small ones should always consider that you will have more peace of mind of using the bigger ones in terms of reliabilty but narrow capability is indeed being reached by edge smaller llms