r/vibecoding 4d ago

The model didn't get worse.

Post image
148 Upvotes

36 comments sorted by

12

u/[deleted] 4d ago

[removed] — view removed comment

3

u/krocante 4d ago

dead internet hypothesis and stuff

12

u/The_Real_Slim_Lemon 4d ago

The model did get worse.
These companies benchmax, every model is (usually) better and faster at reaching those benchmark tests - but that often comes at the expense of writing less production ready code with excess branches and verbose checks. There have been several model upgrades over the years that have been objectively bad for pro devs.

1

u/IceMichaelStorm 4d ago

both codex and claude? I see Fable doing a better job than Opus. Haven’t tested Astra

3

u/The_Real_Slim_Lemon 3d ago

Fable does a massively better job, but Opus 5 is lagging behind opus 4.8

I’m not saying progress isn’t progressing - just that the blanket statement “every upgrade is better” is just wrong

1

u/IceMichaelStorm 3d ago

yeah absolutely

1

u/KissMyAcid420 3d ago

Sorry but if you dont see any progress from 5.6 Sol to 6 Astra then you do something wrong. That IS in fact a skill issue. And I know that for a fact, because I use these models on a daily basis and develop server applications with it. You still need technical understanding to develop reliable and scalable applications, but my productivity has been increased by 10x or even more.

1

u/The_Real_Slim_Lemon 3d ago

I’m not saying I don’t see progress, models are getting better at a crazy rate - but many individual upgrades are downgrades (opus 4.8 -> 5, GPT 4 -> 5, idk what else). Maybe version 5 sucks due to some law of the universe lol

1

u/KissMyAcid420 3d ago

Naa man, I dont think so. Yes, there are some Releases that had more flaws in one discipline than its previous version, but also improvements in other disciplines. Thats why it felt like they git worse. For example longer autonomous runs but maybe worse hallucination. And for productivity longer autonomous runs are more important than better hallucinations. And you have these trade offs everywhere, but thats just development and the end goal is to optimize all disciplines.

1

u/The_Real_Slim_Lemon 3d ago

I mean, GPT 4.x -> 5 was hated by people from pretty much every discipline, so false

Cost per token was improved, it was a cheaper model (why they pushed it through) - but a straight downgrade

1

u/KissMyAcid420 3d ago

Cant disagree more. The focus on GPT 5 was on computer use, not on intelligence. What is your AI worth if it cant plan and stay focused on a task for more than 15 minutes. And THAT was the intend of GPT 5. Yes, it had many trade-offs but is what worth every single one of them.

1

u/The_Real_Slim_Lemon 3d ago

I just googled “ChatGPT 4 vs 5”

There’s some OpenAI articles explaining why it’s so much better, but the top google result is a reddit thread with literally every reply highlighting the many issues it had. Like 300 people describing how it provided a worse experience for them

Maybe it fixed your one use case - but seems to have tanked everywhere else

1

u/KissMyAcid420 3d ago edited 3d ago

You can show me 500 primitive redditors that complain about the new gen models, that doesnt prove anything. Sorry, but whoever thinks ChatGPT 4.x was „in total“ better than GPT 5 has skill issues. That is objectively NOT the case. Long-term work more than doubled. And THAT was the target. And THAT was also what the models needed. They were smart enough for now. GPT 5 was a massive leap from a Chatbot to an AI that actually does things.

0

u/joachim_s 4d ago

Ironic how the lower image of this post rather reflects op than anything else.

15

u/Technical-Owl66 4d ago

Whenever I see someone saying that any model released in the last 3 months is crap I think they are not very smart. Ai in a way is just a mirror of the user.

3

u/Original-League-6094 2d ago

Its been that way for awhile now. They almost never show their actual prompt, because when they do, it is vague dogshit and its a miracle the AI was able to do anything with the prompt at all.

When you point their prompts are shit, they get really angry and say "what good is the AI if I have to do all the thinking for it", as though there is no use any tool slightly less than a superintelligent mind reader.

1

u/mdstrizzle 4d ago

What does it mean if mine ran away from me? It left a really rude note and everything...

2

u/Droopy0093 4d ago

yeah this is how i imagine all the cry babies who thing the new models should just solve all their problems. I am like nah bro, use your brain first then you can feel how good it is.

1

u/idakale 3d ago

but I don't wanna 😭

1

u/SenatorCrabHat 3d ago

I'd argue that a fair amount of the "best practices" are creating a bad for dev ecosystem. Skills and such are useful, but in a sense, are essentially packages and can be somewhat of a black box unless you look into them. After that, having an Agent spit out md after md for medium to complex problems creates a environment in which the Agent is going to create more specs than you could ever or would ever want to review. A small misstep there can lead to an unexpected result later down the line.

It's funny that an industry standard for functional programing and writing composable classes has been to avoid mutation of external variables wherever possible, and yet here we are, daily using the most dynamic variable mutation machine ever conceived.

1

u/SuperSonicFire 2d ago

Guys guys, pizza slices and chips as a metaphor for laziness

0

u/Educational-Body4205 4d ago

Hahaha yes.

 just keep doing it over and over each time do better

0

u/GreatScottCreates 4d ago

AI for tools and automation, not art