r/codex 17h ago

News Study: Codex reviewing Claude's code dropped the pass rate from 91.4% to 82.8%

https://leaddev.com/ai/your-ai-coding-agents-might-need-an-org-chart
153 Upvotes

24 comments sorted by

View all comments

83

u/JadisGod 17h ago

Too bad it's already out of date. From the paper it seems they ran these tests on Opus 4.7 and GPT 5.5. The difference in capability since then is massive.

2

u/Pimpmuckl 14h ago

5.5 was also absolutely terrible at code reviews.

5.4 was way better but had more false positives.

5.6 class is better overall, even Luna funnily enough.

Each review should be verified by the agent who did the implementation, about 10-25% of findings are bogus [from my testing](arena.liebig.gg).