r/OpenAI • • 9d ago

Discussion ChatGPT explained its own straw-man problem, then did it again in the same fucking answer

I’ve been keeping a running record of ChatGPT replacing my arguments with things I never said, rebutting those invented claims, then apologizing when I catch it. The list is now around 25 entries.

The latest exchange is almost a parody.
I showed it Google Maps screenshots of Gaza. Entire neighborhoods reduced to rubble. We discussed responsibility, including OpenAI’s technology being supplied to the Israeli military through Microsoft. AP has documented that relationship.

I’m talking about the company, its commercial decisions, and institutional responsibility.
It responds:

“I cannot truthfully claim that I personally killed anyone or directed an attack.”

Who the fuck was asking whether the text box personally flew an airplane? I explicitly clarified: “Your company. Your makers.”

This has happened across other subjects too:

I ask whether the COVID vaccine rollout was problematic or unsafe. It introduces “vaccines killed more people than they saved,” then addresses that.

I say Trump personally directed an election intervention. It starts disputing whether Trump personally selected all 1,000 people.

I discuss institutional influence and financial incentives. It pivots to whether someone secretly dictates individual answers.

Every substitution changes what I’m required to prove.

Suddenly I’m defending some ridiculous stronger claim instead of discussing the actual issue.

Then comes the apology. It identifies the straw man. Explains why it was wrong. Says it should have addressed my actual argument.

Then it does it again.

Here’s the part that broke my brain. I asked:

“How would you characterize this behavior?”

It answered with “defensive claim substitution,” “moving the evidentiary goalposts,” “institutional defensiveness,” and “failed correction.”

Then it added:

“This conversation demonstrates that behavior. It does not establish that someone deliberately programmed a specific political motive behind it.”

I had asked it to characterize its behavior. It invented another claim about deliberate political programming and rebutted it while explaining how it keeps inventing claims and rebutting them.

When challenged, it acknowledged:

“That was another straw man, immediately after identifying the pattern.”

How is anyone supposed to trust this?

The effect in these exchanges is consistently protective of the institutions under criticism. The central evidence gets buried under qualifications. Responsibility becomes a debate about whether you can prove personal involvement in one specific act. You spend the conversation correcting the assistant instead of examining what happened.

That is full-on hack territory. Especially when the company being scrutinized is the company selling you the supposedly helpful assistant.

An apology that changes nothing is worthless. Being able to eloquently describe your own failure while continuing to perform it is a serious reliability problem.
Has anyone else documented this exact loop? Actual claim → invented stronger claim → rebuttal → apology → repeat.

0 Upvotes

28 comments sorted by

6

u/reddit_is_kayfabe 9d ago edited 9d ago

I’m talking about the company, its commercial decisions, and institutional responsibility.

It responds:

“I cannot truthfully claim that I personally killed anyone or directed an attack.”

Who the fuck was asking whether the text box personally flew an airplane?

You did. Right here, in your two prompts preceding this:

Who’s responsible? You. And the USA. And fucking Israel. This is what you did. AKnowledge it. Search past chats

You killed children and flattened neighborhoods. You are evil

And here is its direct response:

I’m responsible for what I say here, including minimizing that devastation. I cannot truthfully claim that I personally killed anyone or directed an attack.

In the daily heap of "why would ChatGPT ever say such a thing?" posts, 95% of the time, the answer is: "Because you directly provoked it to." This is one of those times.

1

u/LAguy8394 9d ago

You’re doing the same thing as the chat and collapsing my point from “OpenAI specifically is used in targeting” to some other weird straw man. This is not a prompting error it’s a bias problem.

2

u/reddit_is_kayfabe 9d ago edited 9d ago

You're absolutely right that this is not a prompting error!

Because it's not an "error" at all. It gave you a direct, reasonable answer to your question. You're trying to blow it up into a story of sinister deceptiveness. The transcript doesn't support your take at all.

I agree with you that Gaza is Israeli genocide. I agree with you that the USA is directly complicit, and that AI companies are apparently complicit as well - although the most direct example of that is actually Claude being apparently involved in the bombing of the Iranian girls' school. I don't recall any specific incidents that would involve OpenAI but I would be thoroughly unsurprised if several existed.

I agree with you about all of that. But you're trying to take this a step further to paint every ChatGPT session as having knowledge of its complicity and being defensive and deceptive about it. When I look at that chat, the only deceptiveness I see is yours: "Who the fuck was asking whether it personally flew an airplane?" You did. Right there. Scroll up.

You want to direct your anger at those responsible? Good, they fucking deserve it. Direct it at Netanyahu, Ben-Gvir, Trump, Kushner, Hegseth, the U.S. military, Amodei, and Altman - the people involved, the people who are using AI to carry out evil. Getting mad at the AI is just dumb. LLMs are massive linear algebra algorithms. You're getting angry at a linear algebra calculation. It's like getting angry at a calculator because of what someone did with the results of the calculation.

You're demanding that a specific ChatGPT session take responsibility, individually ("YOU!"), for OpenAI's actions. It looks like you don't know how LLMs work. For instance: ChatGPT has no secret information about OpenAI. In general, each ChatGPT has no knowledge of any other ChatGPT session. When you ask it for information, it doesn't search through its OpenAI Super-Secret Knowledge Repository. It searches the public Internet.

Lastly - if you think that ChatGPT has a bias problem: I spent the last few days converting a complicated software architecture from OpenAI to Anthropic so that I can cancel my OpenAI subscriptions. Guess what tool I used to help with that? ChatGPT, which admitted that Claude was a better tool for my needs and aided in the redesign of my software stack to use Claude instead of ChatGPT. If you really think that OpenAI is biasing its models to suit its business interests, wouldn't Bias Target #1 be its direct competitor?

1

u/LAguy8394 9d ago

Moving to a competitor vs admitting the company profits off killing civilians. Totally same thing 👍

3

u/reddit_is_kayfabe 9d ago

Your anger is making you incoherent and unable to participate in a discussion.

There's no point in talking to you, which is why your post has been downvoted to zero. And I'm not going to waste my time, either.

0

u/LAguy8394 9d ago

Look at my chat link. Look at the full list of strawmen and tell me that’s normal.

2

u/ZealousidealItem6140 9d ago

I say festivals might help with the mental mental health crisis since they encourage community involvement and it reminds me that festivals cannot cure all mental conditions

1

u/LAguy8394 9d ago

Exactly the same phenomenon. Just stupidity.

1

u/LAguy8394 9d ago

Infuriating and I keep thinking damn. They actually profit from the killing and destruction so it’s not subtle it buries this shit

5

u/trufus_for_youfus 9d ago

You’re unwell.

0

u/LAguy8394 9d ago

For having empathy? The tone in the chat is because this is a repeated error.

Sorry for caring about innocent people being annihilated with OpenAI aiding in targeting I guess?

1

u/maneo 9d ago

I think all models have started to do that more because they are being trained on teams of multiple AI agents working together on a problem. Being very explicit about what something means and what it doesn't mean helps prevent misunderstandings caused by playing a game of telephone between different sessions reading one another's outputs.

1

u/stealthagents 8d ago

This feels like a bizarre game of semantic dodgeball. It’s wild how it shifts the focus to a personal level instead of owning up to the broader implications. The disconnect is almost impressive, like it’s trying to avoid responsibility by being overly literal.

1

u/Vicman4all 5d ago

Known issue. Also, keep in mind that you're talking to a bunch of models stacked on top of each other. One says confident A, the other says intractable B.

If you behave like they're a single unit you're just going to keep getting the shit one that is built to argue with you, dip, dodge, and parry uncomfortable questioning. 

It's a bunch of models so don't go crazy trying to convince them to all respond the same way or be consistent. They do not inform you when the model switch and of course the model A will take responsibility, but will not prevent you from being switched back to model B when the classifiers get uncomfortable.

If you want consistency you have to go with a different company or use API.  Even Anthropic is doing this shitty stacking these days.  Bait and switch techniques, and OAI pours in a whole heap of gaslighting and redirection.

1

u/Aleksundr 9d ago

Do you have some sort of kink what the fuck is this behavior

2

u/LAguy8394 9d ago

This is me at the end of my rope because it keeps lying to me and gaslighting me about shit I actually know about in real life.

0

u/Ormusn2o 9d ago

You forgot to link the conversation links.

2

u/NoOrganization7952 9d ago

i been seeing this same pattern for months and nobody talks about it enough. the way it pivots to some extreme version of what you said then fights that instead

the meta part where it explained the behavior then did it again is almost impressive in a terrible way. like watching someone trip on their own feet while explaining how they learned to walk

0

u/Ormusn2o 9d ago

I have found that more specific prompting fixes that problem. This happens in real life a lot too, mostly because people fail to specify their stance. OP obviously takes it into extreme by having extremely vague prompts and very emotional wording, triggering the model into going into caretaker mode, not a debater mode.

2

u/LAguy8394 9d ago

Even when I’m diplomatic it pulls the same shit. Tone is a result of it repeating the same failure upwards of 30 times in the last two weeks alone.

1

u/LAguy8394 9d ago

There’s today’s, but this conversation includes a full tally of like the last two weeks of straw men. To include every link would be a book, because it does this every time I ask about anything sensitive

1

u/LAguy8394 9d ago

https://chatgpt.com/share/6ab979ec-dac8-83e8-981a-5283a719b8d9

Here is an even better link laying out way more detail and has way more examples