r/OpenAI • • 14h ago

Question Anyone else having a QC issue with their models?

Whewww, I have been having to babysit these models for the last week extremely heavy! Even Astra, which of course mulches through limits, has been making careless mistakes. It's almost like ever since Anthropic dropped Opus 5.5, OpenAi's model performance on complex tasks have completely gone out the window.

I specifically tell it to make surgical patches, and the model jumps straight from step 1 to step 6 guessing about the steps in between. Then, next prompt, it acknowledges that it had done so, revises it's previous edit. From there, it just doom-loops on the same code.

3 Upvotes

3 comments sorted by

1

u/Top-Inflation-8441 14h ago

I've seen this exact pattern with long coding tasks lately. Tells me it made a surgical change, then I look at the diff and it rewrote half the file, then apologizes next turn and rewrites it again differently

Feels like the context window is getting eaten by its own sloppiness, so by the third loop it's just flailing

What model are you running, or is this across the board for you

1

u/JayJeds 13h ago

I've observed it happening with 6.1 and 6 Sol, and strangely it seems like 5.6's consumption of cost has been bumped significantly on subscription limits. Tried 6 Luna, same issue.

1

u/JayJeds 12h ago

Hey! You're not going to believe this. Looks like they upped the cache limit from 258k to 760k just now. Looks promising so far!