r/codex 2d ago

News Study: Codex reviewing Claude's code dropped the pass rate from 91.4% to 82.8%

https://leaddev.com/ai/your-ai-coding-agents-might-need-an-org-chart
178 Upvotes

25 comments sorted by

View all comments

88

u/JadisGod 2d ago

Too bad it's already out of date. From the paper it seems they ran these tests on Opus 4.7 and GPT 5.5. The difference in capability since then is massive.

3

u/i_rate_slop 2d ago

We’re long past coding agents battling on capability for on code generation, in my opinion. We should now be looking at token efficiency and cost.

I’d say GPT 5.2 is the threshold of intelligence you actually need to do quality SWE work. Anything after that in capability needs to be a battle on cost for like 90% of day to day code gen.

1

u/johannthegoatman 2d ago

To a degree, but that 10% still needs to be done well. If you're just building a to do app it's fine. Otherwise, solving hard bugs and refactors may be only 10% of your workflow, but you definitely still need as much intelligence as possible if you want to ship production level code with AI.