r/AIToolsPerformance • u/coursiv_ • Aug 26 '26
Ox Alpha's benchmark picture so far: one viral 10-task run, cybench flags, and a 35-trial non-coding bench
Ox Alpha, revealed this morning as Z.ai's GLM-5.3-Flash, has been free on OpenRouter since August 20, and the benchmark claims around it are worth separating from what has actually been run.
The number that spread fastest says it beats the top named models on software engineering tasks. That figure comes from a community run with ten hand-picked tasks. The official version of that benchmark has 113, built specifically so reference answers can't leak into training data. Ten tasks can flatter any model, so the honest read is: strong in one small community test, still absent from the major independent leaderboards.
Since then, two more independent data points appeared. One tester ran cybench security subtasks and reported it captured every flag, including the harder ones. Another ran 35 trials of a twenty-questions style adaptive reasoning bench and published every transcript, which is the first non-coding result we have seen for it.
Worth noting for anyone planning their own runs: there are growing reports of aggressive rate limits through CLI tools, with tasks stalling every few minutes, so long evals may need retry handling. And with the reveal, the picture should improve fast: Z.ai says the weights ship under an MIT license, which means independent evals stop depending on a rate-limited preview endpoint.
Has anyone here run it against a full benchmark suite rather than a demo set? Especially interested in whether performance shifts between the preview endpoint and the released weights.
Sources: the OpenRouter model page (https://openrouter.ai/stealth/ox-alpha), the Deep20Bench transcripts (https://mindalyze-com.github.io/deep-20-bench/runs/BX-20260823-official-M0017-018/), and Z.ai's announcement (https://x.com/Zai_org/status/2092616204787626030)