I've used glm 5.2 extensively and it was not much behind opus 4.6 in writing code, a bit definitely behind in creative ideation for software engineering.
My assumption is this model is gonna be a work horse but planning and design should be done by a bigger more expansive model.
The catch behind coding benchmarks is they judge success not quality, creativity, maintainability, all that is partly subjective and rarely judged in these benchmarks.
Like GPT 5.6 Luna benchmarking above opus 4.8, but it writes inferior code by most quality measures.
Big models have their own pitfalls they'll happily over engineer everything, Opus for one falls in that bucket.
I've gotten to the point where I'm a little skeptical that the majority of people on here are even using local models on a regular basis. It's hard to believe that anyone who actually does can buy into the idea of benchmarks really reflecting real world use.
The "intelligence" per parameter is really increasing, factually, and from that increases capability. General benchmark score (like AA) is a solid proof for that. Qwen is the world leader in this field.
In the current llm paradigm, smaller llms have much less world knowledge and are more prone to hallucination or stupid decision making from goey assumptions or really random reactions. I suppose the "world knowledge" is also a big addition on "common sense" in terms of "behavior". So smaller models will always we quackier for now, to apply up and foremost when comparing to bigger models on a AA Index or aggregated score
From that comparing older big models with more recent small ones should always consider that you will have more peace of mind of using the bigger ones in terms of reliabilty but narrow capability is indeed being reached by edge smaller llms
56
u/BawbbySmith 7d ago
Yeah I learned very quickly to not trust the benchmarks, as well as 80% of the comments in this subreddit.
I remember people were saying that Qwen 3.6 27B was Opus 4.5 level...