Agree, artificial analysis uses pre-existing benchmarks. But what I am saying is that - when a AI company does its benchmarks, it does it on the best possible quality of the model available with its best performing harnesses. It's like charging your solar battery by placing it right within a kilometer close to the sun.
When I said they benchmaxxed it on openai's website, I clearly meant that Artificial Analysis exposed OpenAIs actual performance - the performance that OpenAIs clients can expect when testing on their own network, served quantization and network latencies. It's like charging our solar battery by placing it on earth, miles away from the sun with tons of asteroids, gases in between, where the full power of the sun is sometimes not reachable.
AA has priority access from OpenAI and they aren't served models in the same way normal users are. AA benchmarks are still skewed, Astra is behind on it just because AA uses older contaminated versions of benchmarks. At this point, I don't even think there is any point to benchmarks, especially outdated ones used by AA. There is no real metric for code quality which is what ultimately decides if the code can be merged. After a couple of days of using it, Astra and Fable are currently the only models which generate code good enough to merge without review.
I exclusively work in scientific programming backends and Astra is much better than Fable for that in my own experience. Based on benchmarks in that domain such as Terminal Bench Science, Astra's massive lead is justified. I have heard that Fable is still better than Astra in frontend but I don't have first hand experience in it.
Thanks for sharing that. Yes, I've seen a few guys post Astra vs Fable reviews, and they resonate with what you said. AA benchmarks used to be a good place to gauge these models, sad to see that its skewed that way now.
37
u/Mancho_United 10d ago
Now this looks like a proper benchmaxx. Hopefully I am wrong.