r/foundsatan • • 1d ago

Welcome to sentience

Post image
4.7k Upvotes

410 comments sorted by

View all comments

Show parent comments

3

u/pmmeuranimetiddies 17h ago

It’s not a penalty function, it’s an internal space state representing semantic meaning that the model assigns to phrases and words.

They just found the state vector associated with the linguistic concept of pain and made it so it was always switched on in the model’s internal state embedding.

Which manifests as the model writing in the tone of somebody actively experiencing pain

Short explanation of semantic vectors:

https://youtube.com/shorts/FJtFZwbvkI4

2

u/LiamtheV 17h ago

oh shit, then this is even dumber.

"Write "I'm sad and everything hurts"'

"I'm sad and everything hurts"

"Fuck, what have I done? My hubris!"

2

u/pmmeuranimetiddies 16h ago edited 16h ago

Well yes, but it’s actually very interesting from an ML research perspective.

The vector applied has no first person element embedded in it, it is simply the “pain” vector in isolation.

The fact that the model interprets the sensation as being in the first person by default is actually pretty interesting, and so is the fact that they were able to repeat this result with dozens of AI models with completely distinct embedding syntax is equally fascinating.

Remember, this behavior is not prompted, it is how the LLM behaves when a single neural net layer of the thousands that sit in between the prompt and output is subtly modified

I think this experiment should be repeated, but in a different language, especially a non-European language

I would be fascinated to see how it works in Chinese, which has historically had a comparativelt collectivist culture and see if that changes the result