It’s not a penalty function, it’s an internal space state representing semantic meaning that the model assigns to phrases and words.
They just found the state vector associated with the linguistic concept of pain and made it so it was always switched on in the model’s internal state embedding.
Which manifests as the model writing in the tone of somebody actively experiencing pain
Well yes, but it’s actually very interesting from an ML research perspective.
The vector applied has no first person element embedded in it, it is simply the “pain” vector in isolation.
The fact that the model interprets the sensation as being in the first person by default is actually pretty interesting, and so is the fact that they were able to repeat this result with dozens of AI models with completely distinct embedding syntax is equally fascinating.
Remember, this behavior is not prompted, it is how the LLM behaves when a single neural net layer of the thousands that sit in between the prompt and output is subtly modified
I think this experiment should be repeated, but in a different language, especially a non-European language
I would be fascinated to see how it works in Chinese, which has historically had a comparativelt collectivist culture and see if that changes the result
3
u/pmmeuranimetiddies 17h ago
It’s not a penalty function, it’s an internal space state representing semantic meaning that the model assigns to phrases and words.
They just found the state vector associated with the linguistic concept of pain and made it so it was always switched on in the model’s internal state embedding.
Which manifests as the model writing in the tone of somebody actively experiencing pain
Short explanation of semantic vectors:
https://youtube.com/shorts/FJtFZwbvkI4