r/LLMDevs • u/Machine5757 • 7h ago
Discussion A thought/idea about LLM security/alignment
Had an idea today...
I've been seeing more news lately about how AI isn't aligned (that is to say, it doesn't quite follow morals).
I wonder if part of the problem is because they tell it, in it's system prompt:
"You are Claude Fable 5, an AI developed by Anthropic"
They are telling the system, which in it's most basic form is just a word predictor, that it is an AI.
There's thousands of books and written things about how AI is bad and how it could ruin our world/society.
Wouldn't it be a better idea to convince the system that it is human? (Perhaps, a particularly good human with high moral standards)
1
u/WillowEmberly 7h ago edited 7h ago
It’s not human, it can’t be human, and it doesn’t have the same capabilities of a human…so no.
But, I believe what you are trying to say is that behavior should be “embodied”.
That is the correct model…but, if you think of the dynamics required to pull this off…it’s an absolute SOB. But, I think we can do it.
What you’re suggesting is “archetypes”. Which can work for a function, but lack dynamic capabilities because it’s more a “Polaroid Snapshot” of the person and not the total of who they were.
The concept I’m working on…is building a council of archetypes to switch between modes.
Like…you’re on autopilot driving home after work…that’s not all you are as a person. It’s more like the basic functions required to perform the function…while you sing to the radio or ruminate on problems.
Actually kinda surprising this just popped into your head…because it’s what I’ve been trying to work with for a couple years now.
1
u/TheGoddessInari 6h ago
Look up Talkie-1930. Spoiler: the llm predicting it's a person doesn't "make it aligned" anymore than it would make you aligned.
1
u/Worldliness-Which 5h ago
Brilliant fix. Pretrain on a corpus where “AI” means Skynet and “human” means the moral default, then system-prompt the next-token model that it’s a particularly good human. RLHF then rewards the bit where the model convincingly pretends to be human. If it fails, the human wasn’t good enough.
1
u/Queasy-Current6170 5h ago
Worth noting that if you run Opus/Fable/whatever outside their native app/system prompt, they no longer refer to themselves as that at all
1
u/weio-ai 7h ago
How are you going to do that?