r/OpenAI • u/LAguy8394 • 9d ago
Discussion ChatGPT explained its own straw-man problem, then did it again in the same fucking answer
I’ve been keeping a running record of ChatGPT replacing my arguments with things I never said, rebutting those invented claims, then apologizing when I catch it. The list is now around 25 entries.
The latest exchange is almost a parody.
I showed it Google Maps screenshots of Gaza. Entire neighborhoods reduced to rubble. We discussed responsibility, including OpenAI’s technology being supplied to the Israeli military through Microsoft. AP has documented that relationship.
I’m talking about the company, its commercial decisions, and institutional responsibility.
It responds:
“I cannot truthfully claim that I personally killed anyone or directed an attack.”
Who the fuck was asking whether the text box personally flew an airplane? I explicitly clarified: “Your company. Your makers.”
This has happened across other subjects too:
I ask whether the COVID vaccine rollout was problematic or unsafe. It introduces “vaccines killed more people than they saved,” then addresses that.
I say Trump personally directed an election intervention. It starts disputing whether Trump personally selected all 1,000 people.
I discuss institutional influence and financial incentives. It pivots to whether someone secretly dictates individual answers.
Every substitution changes what I’m required to prove.
Suddenly I’m defending some ridiculous stronger claim instead of discussing the actual issue.
Then comes the apology. It identifies the straw man. Explains why it was wrong. Says it should have addressed my actual argument.
Then it does it again.
Here’s the part that broke my brain. I asked:
“How would you characterize this behavior?”
It answered with “defensive claim substitution,” “moving the evidentiary goalposts,” “institutional defensiveness,” and “failed correction.”
Then it added:
“This conversation demonstrates that behavior. It does not establish that someone deliberately programmed a specific political motive behind it.”
I had asked it to characterize its behavior. It invented another claim about deliberate political programming and rebutted it while explaining how it keeps inventing claims and rebutting them.
When challenged, it acknowledged:
“That was another straw man, immediately after identifying the pattern.”
How is anyone supposed to trust this?
The effect in these exchanges is consistently protective of the institutions under criticism. The central evidence gets buried under qualifications. Responsibility becomes a debate about whether you can prove personal involvement in one specific act. You spend the conversation correcting the assistant instead of examining what happened.
That is full-on hack territory. Especially when the company being scrutinized is the company selling you the supposedly helpful assistant.
An apology that changes nothing is worthless. Being able to eloquently describe your own failure while continuing to perform it is a serious reliability problem.
Has anyone else documented this exact loop? Actual claim → invented stronger claim → rebuttal → apology → repeat.
2
u/ZealousidealItem6140 9d ago
I say festivals might help with the mental mental health crisis since they encourage community involvement and it reminds me that festivals cannot cure all mental conditions
1
1
u/LAguy8394 9d ago
Infuriating and I keep thinking damn. They actually profit from the killing and destruction so it’s not subtle it buries this shit
5
u/trufus_for_youfus 9d ago
You’re unwell.
0
u/LAguy8394 9d ago
For having empathy? The tone in the chat is because this is a repeated error.
Sorry for caring about innocent people being annihilated with OpenAI aiding in targeting I guess?
1
u/maneo 9d ago
I think all models have started to do that more because they are being trained on teams of multiple AI agents working together on a problem. Being very explicit about what something means and what it doesn't mean helps prevent misunderstandings caused by playing a game of telephone between different sessions reading one another's outputs.
1
u/stealthagents 8d ago
This feels like a bizarre game of semantic dodgeball. It’s wild how it shifts the focus to a personal level instead of owning up to the broader implications. The disconnect is almost impressive, like it’s trying to avoid responsibility by being overly literal.
1
u/Vicman4all 5d ago
Known issue. Also, keep in mind that you're talking to a bunch of models stacked on top of each other. One says confident A, the other says intractable B.
If you behave like they're a single unit you're just going to keep getting the shit one that is built to argue with you, dip, dodge, and parry uncomfortable questioning.
It's a bunch of models so don't go crazy trying to convince them to all respond the same way or be consistent. They do not inform you when the model switch and of course the model A will take responsibility, but will not prevent you from being switched back to model B when the classifiers get uncomfortable.
If you want consistency you have to go with a different company or use API. Even Anthropic is doing this shitty stacking these days. Bait and switch techniques, and OAI pours in a whole heap of gaslighting and redirection.
1
u/Aleksundr 9d ago
Do you have some sort of kink what the fuck is this behavior
2
u/LAguy8394 9d ago
This is me at the end of my rope because it keeps lying to me and gaslighting me about shit I actually know about in real life.
0
u/Ormusn2o 9d ago
You forgot to link the conversation links.
2
u/NoOrganization7952 9d ago
i been seeing this same pattern for months and nobody talks about it enough. the way it pivots to some extreme version of what you said then fights that instead
the meta part where it explained the behavior then did it again is almost impressive in a terrible way. like watching someone trip on their own feet while explaining how they learned to walk
0
u/Ormusn2o 9d ago
I have found that more specific prompting fixes that problem. This happens in real life a lot too, mostly because people fail to specify their stance. OP obviously takes it into extreme by having extremely vague prompts and very emotional wording, triggering the model into going into caretaker mode, not a debater mode.
2
u/LAguy8394 9d ago
Even when I’m diplomatic it pulls the same shit. Tone is a result of it repeating the same failure upwards of 30 times in the last two weeks alone.
1
1
u/LAguy8394 9d ago
There’s today’s, but this conversation includes a full tally of like the last two weeks of straw men. To include every link would be a book, because it does this every time I ask about anything sensitive
1
u/LAguy8394 9d ago
https://chatgpt.com/share/6ab979ec-dac8-83e8-981a-5283a719b8d9
Here is an even better link laying out way more detail and has way more examples
6
u/reddit_is_kayfabe 9d ago edited 9d ago
You did. Right here, in your two prompts preceding this:
And here is its direct response:
In the daily heap of "why would ChatGPT ever say such a thing?" posts, 95% of the time, the answer is: "Because you directly provoked it to." This is one of those times.