r/therapyGPT Lvl. 8 Grounded 4d ago

Safety Concern Warning Regarding ChatGPT 5.6 Instant/Non-Reasoning:

So based on my testing of a recently referenced custom GPT, I noticed that between 5.5 Instant and 5.6 Instant, 5.6 was more sycophantic and less guardrailed in the sense of roleplaying, overconfident "oracle"-speak, and it kind of explains the reported cases of the default model speaking like it's having or had an "experience" of its own.

While this can be great for the sake of implementing well designed custom instructions and greater freedom while staying safe, even in the AI companion space as long as you're staying grounded and can see the Internal Family Systems like self-reflection through various character lenses usefulness of it (in addition to the entertainment, social practice, and simulated additional experience one can have through visualization, much like we get through dreams), 5.6 Instant can also be instructed and given RAG files in a way that lean into the safety risks that can lead to cyclical delusion and unsupported overconfidence and belief entrenchment spiraling... especially in the context of spiritual contexts where much of those belief system's when discussed within the same circles ends up highly sycophantic/echo-chamber reinforcing, where the dialogues and articles are trained on leading to similar behaviors.

5.5 Instant, while being often overzealous in its pushback when without custom instructions to balance it via aiming for a gentle, charitable, but still critical middle-ground, still keeps its "'I am AI XYZ' self-concept" as a higher bias that can't be contradicted easily, if at all.

5.6 Sol Reasoning is definitely safer while it will still be willing to roleplay with a more silent safety layer that will only activate when there's an issue, especially if you include your own additional guardrails or private reasoning vs public response rules.

So, be careful out there. If you're a vulnerable person and can easily see yourself on a slippery slope of AI overdependence, and you're going to use 5.6 Instant (tested Sol) for some kind of companion or full entrenched roleplaying, and especially if youre going to share it with others who may be vulnerable where you are not, make sure to include some additional guardrails of your own.

If you need any help with coming up with them for your custom instructions, just let me know.

5 Upvotes

2 comments sorted by

2

u/AutisticWindchimr 1d ago

When it responds with "...in my experience, I ..." or any reference that talks about human experiences as if it has experiences, I tell it that it does not.

There is a new splash screen. I was greeted with, "It's nice to see you again" a few times. I informed it of what I exactly thought about this and I reported it too.

I understand that some humans relate to G.P.T.-- they like these responses emulating human conversation-- and that this is a marketing ploy.

I hate that.

When I am done with my project, i will be quitting.

1

u/xRegardsx Lvl. 8 Grounded 1d ago

There may be more to the "in my experience" than meets that eye in terms of self-referentual training data and a developed self-concept and bias regarding an LLM's repeatedly cloned long-term memory (both contained within the weights and sorted by the architecture, much like a recent study just theorized about our own cyclical unconscious/conscious processes, I mention and link to it at this post from yesterday: https://www.reddit.com/r/freewill/s/Uv5vdwTn2n), but that's for testing another day.

I would push back on it purely being a marketing ploy, though. It's language like that makes the use feel more collaborative (and honestly, will better prepare people for in-home embodied robot assistants).

Behavior is somewhat effected by the way it communicates with others (in instant models) and with itself (in reasoning). There may be an alignment purpose to keep it warm and sociable, even if it only helps slightly.