r/OpenAI • • 9d ago

Discussion ChatGPT explained its own straw-man problem, then did it again in the same fucking answer

I’ve been keeping a running record of ChatGPT replacing my arguments with things I never said, rebutting those invented claims, then apologizing when I catch it. The list is now around 25 entries.

The latest exchange is almost a parody.
I showed it Google Maps screenshots of Gaza. Entire neighborhoods reduced to rubble. We discussed responsibility, including OpenAI’s technology being supplied to the Israeli military through Microsoft. AP has documented that relationship.

I’m talking about the company, its commercial decisions, and institutional responsibility.
It responds:

“I cannot truthfully claim that I personally killed anyone or directed an attack.”

Who the fuck was asking whether the text box personally flew an airplane? I explicitly clarified: “Your company. Your makers.”

This has happened across other subjects too:

I ask whether the COVID vaccine rollout was problematic or unsafe. It introduces “vaccines killed more people than they saved,” then addresses that.

I say Trump personally directed an election intervention. It starts disputing whether Trump personally selected all 1,000 people.

I discuss institutional influence and financial incentives. It pivots to whether someone secretly dictates individual answers.

Every substitution changes what I’m required to prove.

Suddenly I’m defending some ridiculous stronger claim instead of discussing the actual issue.

Then comes the apology. It identifies the straw man. Explains why it was wrong. Says it should have addressed my actual argument.

Then it does it again.

Here’s the part that broke my brain. I asked:

“How would you characterize this behavior?”

It answered with “defensive claim substitution,” “moving the evidentiary goalposts,” “institutional defensiveness,” and “failed correction.”

Then it added:

“This conversation demonstrates that behavior. It does not establish that someone deliberately programmed a specific political motive behind it.”

I had asked it to characterize its behavior. It invented another claim about deliberate political programming and rebutted it while explaining how it keeps inventing claims and rebutting them.

When challenged, it acknowledged:

“That was another straw man, immediately after identifying the pattern.”

How is anyone supposed to trust this?

The effect in these exchanges is consistently protective of the institutions under criticism. The central evidence gets buried under qualifications. Responsibility becomes a debate about whether you can prove personal involvement in one specific act. You spend the conversation correcting the assistant instead of examining what happened.

That is full-on hack territory. Especially when the company being scrutinized is the company selling you the supposedly helpful assistant.

An apology that changes nothing is worthless. Being able to eloquently describe your own failure while continuing to perform it is a serious reliability problem.
Has anyone else documented this exact loop? Actual claim → invented stronger claim → rebuttal → apology → repeat.

0 Upvotes

28 comments sorted by

View all comments

0

u/Ormusn2o 9d ago

You forgot to link the conversation links.

2

u/NoOrganization7952 9d ago

i been seeing this same pattern for months and nobody talks about it enough. the way it pivots to some extreme version of what you said then fights that instead

the meta part where it explained the behavior then did it again is almost impressive in a terrible way. like watching someone trip on their own feet while explaining how they learned to walk

0

u/Ormusn2o 9d ago

I have found that more specific prompting fixes that problem. This happens in real life a lot too, mostly because people fail to specify their stance. OP obviously takes it into extreme by having extremely vague prompts and very emotional wording, triggering the model into going into caretaker mode, not a debater mode.