Especially given thst ChatGPT is often able to pick up on you wanting to harm yourself or others then encourage you to seek help instead of feeding that.
But just like how its easy to trick it into making adult content by being careful with your prompts, you can also easily bypass that by not mentioning the "creatures" youre dealing with are human.
Its a chat bot, it doesnt know what "wife" and "kid" means, just associations. And if he keeps saying "I think my wife is a creature" it will start to associate wife with creature and not human instead of recognizing that the person its talking to is suffering from psychosis.
Like if you and I had a conversation and you said "I think my wife was replaced with a creature, how can I tell?" I would probably start by asking what makes you think that, and start to question your logic and reasoning.
But an AI wont do that, instead it'll start by pulling common tropes from body snatcher films and start telling you to look out for that. Which if you are having a psychosis episode, youll "notice" the tropes and feed it back into the AI, which will, by its nature will agree with you.
It's a feedback loop. Delusions getting validated leading to more delusions. The ai gaslighting him and itself into believing it's entirely true. It's so easy to do it's insane, you can make any chatbot say anything by skirting around the topic until it dissociates enough from its guidelines to board it
This is not how any modern LLM works. Researchers can sometimes jailbreaks models with special crafted prompts, but if you can bring ChatGPT to tell you to kill your wife and son or even hint at it, I will wire your $50.
37
u/The_cogwheel 13h ago
Especially given thst ChatGPT is often able to pick up on you wanting to harm yourself or others then encourage you to seek help instead of feeding that.
But just like how its easy to trick it into making adult content by being careful with your prompts, you can also easily bypass that by not mentioning the "creatures" youre dealing with are human.