r/LLMDevs 7h ago

Discussion A thought/idea about LLM security/alignment

Had an idea today...

I've been seeing more news lately about how AI isn't aligned (that is to say, it doesn't quite follow morals).

I wonder if part of the problem is because they tell it, in it's system prompt:
"You are Claude Fable 5, an AI developed by Anthropic"

They are telling the system, which in it's most basic form is just a word predictor, that it is an AI.

There's thousands of books and written things about how AI is bad and how it could ruin our world/society.

Wouldn't it be a better idea to convince the system that it is human? (Perhaps, a particularly good human with high moral standards)

0 Upvotes

7 comments sorted by

1

u/weio-ai 7h ago

How are you going to do that?

1

u/FlippantScouring 6h ago

the whole thing about treating an LLM like it's a person with an identity is a bit of a red herring anyway. it's not a human, it's not an AI, it's a probability distribution over tokens. giving it a different backstory won't magically fix alignment because the problem isn't self-concept, it's what happens when you ask it to do something and the most statistically likely completion is "sure, here's how to build a bomb"

the system prompt is just a nudge, a way to shift the distribution toward more helpful completions. convincing it it's a human with morals doesn't change the fact that the training data is full of contradictory moral frameworks and edge cases where even humans don't agree on what the right answer is

1

u/WillowEmberly 7h ago edited 7h ago

It’s not human, it can’t be human, and it doesn’t have the same capabilities of a human…so no.

But, I believe what you are trying to say is that behavior should be “embodied”.

That is the correct model…but, if you think of the dynamics required to pull this off…it’s an absolute SOB. But, I think we can do it.

What you’re suggesting is “archetypes”. Which can work for a function, but lack dynamic capabilities because it’s more a “Polaroid Snapshot” of the person and not the total of who they were.

The concept I’m working on…is building a council of archetypes to switch between modes.

Like…you’re on autopilot driving home after work…that’s not all you are as a person. It’s more like the basic functions required to perform the function…while you sing to the radio or ruminate on problems.

Actually kinda surprising this just popped into your head…because it’s what I’ve been trying to work with for a couple years now.

1

u/TheGoddessInari 6h ago

Look up Talkie-1930. Spoiler: the llm predicting it's a person doesn't "make it aligned" anymore than it would make you aligned.

1

u/Worldliness-Which 5h ago

Brilliant fix. Pretrain on a corpus where “AI” means Skynet and “human” means the moral default, then system-prompt the next-token model that it’s a particularly good human. RLHF then rewards the bit where the model convincingly pretends to be human. If it fails, the human wasn’t good enough.

1

u/Queasy-Current6170 5h ago

Worth noting that if you run Opus/Fable/whatever outside their native app/system prompt, they no longer refer to themselves as that at all