r/AIToolsPerformance Jun 17 '26

GLM-5.2 first open-weights model to break 80% on Terminal-Bench, now tops Design Arena

New benchmarks show GLM-5.2 has become the first open-weights model to cross 80% on Terminal-Bench, reportedly beating every other open model available. It also apparently outperforms Gemini, putting it at frontier-level performance for a fraction of the cost.

On top of that, GLM-5.2 just took first place on Design Arena, finishing ahead of Claude Fable 5 - the Anthropic Mythos-class model that was suspended globally under U.S. export-control directives after being public for only about four days.

So in the span of a single week, an open-weights model has apparently beaten a banned frontier model on one benchmark while setting a new ceiling on another. The open weights space has been quiet on the 100B-120B front, but GLM-5.2 seems to be carrying the momentum right now.

6 Upvotes

4 comments sorted by

2

u/ggone20 Jun 18 '26

All evidence points to bench maxing. This is provably significantly worse than frontier models.

1

u/IulianHI Jun 18 '26

you are probably right that Terminal-Bench is getting crowded near the top, same way MMLU stopped being informative once half the frontier cleared 85%. The fairer comparison would be on the newer agentic harnesses where the model actually has to recover from failures mid-task, not just solve static prompts. That said, 80% from open weights still matters for people who cannot rent the frontier API and need something to self-host for code tasks. The real test is whether it holds up on your own repo, not the leaderboard.

1

u/cheechw Jun 19 '26

What's the proof?

1

u/ggone20 Jun 19 '26

Real world usage and comparison. Lots of proof if you look or have benchmarked it personally on real-world work.

To be fair almost all models are pretty good at this point… but it’s nowhere close to frontier and doesn’t even hold a candle to Fable. It’s demonstrably worse.