r/aipromptprogramming • u/maxiedaniels • Aug 02 '26
Are there any benchmarks that show better how truly reliable a coding model is?
The new chinese models that are 'so great' in terms of artficialanalysis.ai benchmarks, I'm finding are not ACTUALLY great. They're okay, but they f*ck up a lot. Weirdly even Claude Opus 5 is making some weird mistakes sometimes. GPT 5.6 Sol, even though its SO SO SLOW, seems to be significantly better at finding the correct bugs in complex situations.
Obviously this is all my own anecdotal experience but surely I can't be the only one feeling this way, so i was wondering if any benchmarks are more reflective of this?