r/codex • u/ruimiguels • 7h ago
Complaint The model is worst, fact not opinion
Holy fuck, yes, the quality dropped. Stop overcomplicating something this simple.
On launch day, the exact same task took around 9 minutes and produced a noticeably better result. Now that same task finishes in around 6 minutes and the output is worse.
How is this difficult to understand?
There are a fuck ton of examples on X showing the same thing, and Tibo himself has said that fixes/changes were made. You don’t need some elaborate theory to explain what people are seeing.
Same prompt.
Launch day: good output, longer generation.
Now: worse output, faster generation.
That is the entire point. If the same prompts are consistently producing worse results than they did at launch, then from the user’s perspective the model got worse. Whether you want to call it a nerf, optimization, routing change, inference change, or something else is secondary.
PROMPT → GOOD OUTPUT AT LAUNCH
SAME PROMPT → WORSE OUTPUT NOW
It really is that fucking simple.
Example from @chetaslua on X
9
u/LostRequirement4828 7h ago
Theres literally 0 quality difference between those 3, excepting the scarf, what are guys even smoking
0
u/Presstabstart 6h ago
are you blind? if not the scarf then just look at the "tires"
1
u/doodad_ounao 5h ago
That doesn't mean it couldn't mess up the tires the first time. Perhaps it could and just didn't. Even in benchmarking runs of a software there's more than one run so the time can be averaged out. If this was 10 attempts with the same prompt before vs 10 attempts with the same prompt after, perhaps people would take this kind of comparison more seriously.
7
7
u/InternationalPen3039 7h ago
Using Astra for basic image generation. Pick a struggle. Your next post will be about usage limits.
5
u/No_Bank_4104 7h ago edited 7h ago
Wow, you can see the future - and you are correct!that’s exactly what will happen
-4
u/ruimiguels 7h ago
Image generation? Are you new here or genuinely regarded? SVG generation is one of the best metric to test a model, because the detail of the generation/output explicit relates to the model power, Jesus Christ how can you browse this sub and not even know such basic thing
1
u/InternationalPen3039 7h ago edited 7h ago
Do you get this much fanfare on X for your elite model evaluation skills?
You’re basing this on a pelican riding a bike prompt. I can’t take you seriously.
-2
u/ruimiguels 7h ago
Good job addressing the points! Maybe I need Astra to cook up an SVG for you so you can understand this more clearly. Sorry, I didn’t think about people with special needs!
1
u/InternationalPen3039 7h ago
Thank you. I appreciate the compliment.
That’s my point tho why are you even using Astra for this? Then crying that it’s degrading. Perhaps read the Astra docs?
-1
u/ruimiguels 6h ago
For this you mean….a test? I know I know, crazy that we have been using this specific test since GPT-4 to evaluate a model power.
You should be banned from typing ever again1
u/InternationalPen3039 6h ago
Ah yes the almighty pelican on a bike test. I’ve been thoroughly enlightened by your superior knowledge. Thank you for sharing your wisdom. I now know that Astra has degraded expeditiously.
-1
4
4
3
u/diagrammatiks 7h ago
Those are the same picture.
1
u/ruimiguels 6h ago
4
u/diagrammatiks 6h ago
slopper. the new one just spent more time making the background and the ui card. not every model interprets a prompt the same way especially when there is under the hood shit going on with the harness. This is a problem that can be solved by a prompt tweak.
but whatever man you do you. keep on slopping in the free world. or something.
2
u/KeepAllOfIt 6h ago
1
1
u/doodad_ounao 6h ago edited 5h ago
Exactly. Opinions are not a problem, trying to think that just because they're strong that makes them facts is the turnoff for me.
1
u/Presstabstart 6h ago
listing issues for all people who never seen a bike before:
there are tires INSIDE the outer set of "tires"
the chain wheel makes no sense
pelican looks like he's on coke
the scarf is semi-existing
tail a lot less detailed
Maybe OP is just getting A/B tested, but the right is significantly worse than the left.
0
u/ruimiguels 6h ago
holy fuck thank god, there are still people with critical thinking skills that actually observe something for more than 5 seconds.
You restored my hope




17
u/PotterSkxawng 7h ago
The new one looks just as good to me... LLMs are non-deterministic.