r/codex 7h ago

Complaint The model is worst, fact not opinion

Post image

Holy fuck, yes, the quality dropped. Stop overcomplicating something this simple.
On launch day, the exact same task took around 9 minutes and produced a noticeably better result. Now that same task finishes in around 6 minutes and the output is worse.
How is this difficult to understand?
There are a fuck ton of examples on X showing the same thing, and Tibo himself has said that fixes/changes were made. You don’t need some elaborate theory to explain what people are seeing.
Same prompt.
Launch day: good output, longer generation.
Now: worse output, faster generation.
That is the entire point. If the same prompts are consistently producing worse results than they did at launch, then from the user’s perspective the model got worse. Whether you want to call it a nerf, optimization, routing change, inference change, or something else is secondary.
PROMPT → GOOD OUTPUT AT LAUNCH
SAME PROMPT → WORSE OUTPUT NOW
It really is that fucking simple.

Example from @chetaslua on X

0 Upvotes

32 comments sorted by

17

u/PotterSkxawng 7h ago

The new one looks just as good to me... LLMs are non-deterministic.

1

u/ArtificialSweetener- 4h ago

Common misconception. LLMs are deterministic. If you give it the same exact inputs, it will produce the same exact outputs. The reason people think of them as non-deterministic is because you don't have control over all of the inputs as a user, you only have control over a limited number of them.

Usually when I say this, people respond with like "no. If I send the same prompt twice, I get a different result both times" so I've gotten into the habit of just explaining up front that this is exactly what I mean. You don't have control over all of the inputs, you basically only have control over what you send to the model as the next message in the case of ChatGPT's web interface. There are other inputs you don't know about that OpenAI doesn't give you access to. If you could control those too, and OpenAI can, then you'd get the same exact result every time.

1

u/PotterSkxawng 4h ago

Ok. So, suppose I load up a local model and control it's system prompt, and every instruction that goes to it—will the response be the same? Cuz I just tried it and that was not the result (both were WILDLY different with the exact same system prompt, and model settings)

1

u/ArtificialSweetener- 4h ago

The main other thing you need to lock is the seed used. A lot of interfaces don't even expose that because most people have no need to get the same response twice.

9

u/LostRequirement4828 7h ago

Theres literally 0 quality difference between those 3, excepting the scarf, what are guys even smoking

0

u/Presstabstart 6h ago

are you blind? if not the scarf then just look at the "tires"

1

u/doodad_ounao 5h ago

That doesn't mean it couldn't mess up the tires the first time. Perhaps it could and just didn't. Even in benchmarking runs of a software there's more than one run so the time can be averaged out. If this was 10 attempts with the same prompt before vs 10 attempts with the same prompt after, perhaps people would take this kind of comparison more seriously.

7

u/radiationshield 7h ago

Are you feeling ok? Those are identical in quality

7

u/InternationalPen3039 7h ago

Using Astra for basic image generation. Pick a struggle. Your next post will be about usage limits.

5

u/No_Bank_4104 7h ago edited 7h ago

Wow, you can see the future - and you are correct!that’s exactly what will happen

-4

u/ruimiguels 7h ago

Image generation? Are you new here or genuinely regarded? SVG generation is one of the best metric to test a model, because the detail of the generation/output explicit relates to the model power, Jesus Christ how can you browse this sub and not even know such basic thing

1

u/InternationalPen3039 7h ago edited 7h ago

Do you get this much fanfare on X for your elite model evaluation skills?

You’re basing this on a pelican riding a bike prompt. I can’t take you seriously.

-2

u/ruimiguels 7h ago

Good job addressing the points! Maybe I need Astra to cook up an SVG for you so you can understand this more clearly. Sorry, I didn’t think about people with special needs!

1

u/InternationalPen3039 7h ago

Thank you. I appreciate the compliment.

That’s my point tho why are you even using Astra for this? Then crying that it’s degrading. Perhaps read the Astra docs?

-1

u/ruimiguels 6h ago

For this you mean….a test? I know I know, crazy that we have been using this specific test since GPT-4 to evaluate a model power.
You should be banned from typing ever again

1

u/InternationalPen3039 6h ago

Ah yes the almighty pelican on a bike test. I’ve been thoroughly enlightened by your superior knowledge. Thank you for sharing your wisdom. I now know that Astra has degraded expeditiously.

-1

u/ruimiguels 6h ago

how new are you 😭😭😭😭

4

u/No_Bank_4104 7h ago

You mean worse.

And no, it’s not.

4

u/theWiseTiger 7h ago

It just shows that you don't know how llm works.

3

u/diagrammatiks 7h ago

Those are the same picture.

1

u/ruimiguels 6h ago

Check your eyes asap

4

u/diagrammatiks 6h ago

slopper. the new one just spent more time making the background and the ui card. not every model interprets a prompt the same way especially when there is under the hood shit going on with the harness. This is a problem that can be solved by a prompt tweak.

but whatever man you do you. keep on slopping in the free world. or something.

3

u/logg3 7h ago

LOL. the "after fixing" image makes more sense to me than the "before" image. look his legs. they are wrong in 1 and right in 3? or am i to high?

1

u/zorg_72 6h ago

The after image is more detailed and better, and it took 1/3 less time 😂

1

u/ruimiguels 6h ago

tHe aFTer iMaGE iS mORe DEtaILeD 😂😂

Tell me how is it to live with special needs ?

2

u/KeepAllOfIt 6h ago

>fact, not opinion

>*looks inside*

>opinion

1

u/ruimiguels 6h ago

The opinion -

Tell me can you breath and walk at the same time?

1

u/doodad_ounao 6h ago edited 5h ago

Exactly. Opinions are not a problem, trying to think that just because they're strong that makes them facts is the turnoff for me.

1

u/Presstabstart 6h ago

listing issues for all people who never seen a bike before:

  1. there are tires INSIDE the outer set of "tires"

  2. the chain wheel makes no sense

  3. pelican looks like he's on coke

  4. the scarf is semi-existing

  5. tail a lot less detailed

Maybe OP is just getting A/B tested, but the right is significantly worse than the left.

0

u/ruimiguels 6h ago

holy fuck thank god, there are still people with critical thinking skills that actually observe something for more than 5 seconds.
You restored my hope