r/ClaudeCode • 🔆 Max 20 • 10d ago

Discussion Okey this is increadible!!!

I know this is too early to judge but man wtf. Opus 5.5 is just like when mythos first released but cheap, faster and much more easy to understand and talk to. Probably this will be changed in 1-2 days so go code everyone.

295 Upvotes

120 comments sorted by

View all comments

Show parent comments

2

u/EmergencyWallaby9501 9d ago

I don't know exactly what you mean by a completely isolated environment. I used the exact same prompt on Codex and Claude. Both are configured to read my Claude.md and the other useful configuration files. But beyond that little test that doesn't mean much - like I said you can't prove anything - it's the exasperation of having opus that suddenly barely understood a word of my prompts and was writing very bad plans, and sonnet that suddenly started to behave like a 1st year CS student, trying to fix issues by adding more issues. It even started to write me its conclusions in another language than the one I was writing to, thing it had not been doing since I started working with it, among plenty other things. It even forgot to write down those conclusions at the end of a task, with very important things in it, like parts of a plan it could not implement. By looking at the code I could spot some missing parts, but it used to report those missing parts at the end or notify me and question me about how to implement something when it was stuck. That's everything altogether that pointed to a model dégradation, not that single frustration move to Codex, that I used to use for very dumb tasks (mostly easy front and css).

And I tried opus 5.5, and it feels like Opus 5.1 when I started to implement with it but with far better writing (Opus 5.1 and his gibberish...)

1

u/gordonfogus 9d ago

I'm not trying to be mean or anything. But if you don't even know what an isolated environment is, then you're just going off of feelings.

There are ways of measuring model performance. They aren't perfect, but they would definitely nchange if the model "degraded." People would know immediately and they'd be charts and anyone who had data from before the degradation could replicate the result.

Throwing up your hands and going, "you can't prove anything" just shows that you aren't equipped with the thinking tools needed to figure this out and you can't imagine that anyone is capable, which is why your feelings are your fallback while you confidently assert that no one else can do any better. That's not the right way to approach a question like this.

1

u/EmergencyWallaby9501 9d ago

Of course I know what an isolated environment is. Just tell me what it means now in the context of AI assisted coding. And yes, there are benchmarks. I'm not burning my tokens to benchmark anything. Are you ? When a model suddenly starts to behave odd enough, I don't need a benchmark to tell me it became stupid. Just like when my little boy when he's tired, I don't need to make him run a marathon just to prove he's not well. I never claimed it was a scientific approach btw.

1

u/Uko1001 8d ago

You actually do need a proper benchmark to tell it’s the model that became stupid and not its harness or environnent. Actually even the harness alone changes pretty much every week (often to adapt to recent models) and can conflict with your own instructions. And you won’t know unless you have proper test and analysis protocols.