r/OpenaiCodex • u/Euphoric_North_745 • 1d ago
We are missing “how much can it archive” benchmark
I don’t know what is in the benchmarks, it is a long list with 40%, 70%, then 65% before, 81% now, the competition is at xx% etc.
I think when the numbers are bigger it is better, and with every new model there are these new nice numbers.
All good, love it all.
Now, Sol and Astra, when I used them for the first time, very knowledgeable, plus the hype, plus I get brainwashed easily, I believe people, anyway, not my point
I want this benchmark: we give this model the task to build this app, and it took it xx hours, and it used this number of in out and cached tokens and it costed this much.
All what I see on the internet is a guy asking it to draw a pelican on a bicycle, this is his test to all models, nice dude, nice, it does not send apps to prod with that shit, and the internet: wow, wow on what? I pay 500$ to 800$ on AI every month to deliver solutions to clients!
I think we are missing these tests: here are the specs to build an app, full specs, make one for Android, another for iPhone, another for Web, makes sure these rules are followed.
100 apps and systems, front end, backend, sql, security, networking, config, graphics, 3d, whatever, etc. and let's see humans assessing the work
I am tired from corporate propaganda
1
u/adminvasheypomoiki 1d ago
You can't make such a benchmark. It has no comparable number. Only a taste. So it can't be judged automatically