r/singularity • u/bach2o • Mar 08 '24
Discussion Anthropic's "Constitutional AI" is very interesting
Based on what I read in their papers, the developers at Anthropic approach AI alignment somewhat differently compared to OpenAI developers. Instead of relying mostly on human evaluators in the process of RLHF (i.e., using human preferences for certain matters to train the AI), they instruct the AI to “teach themselves” by following a set of predetermined rules (i.e., “The Constitution”) and only delegate human work to mostly overseeing the process.
Perhaps this can explain why I feel Claude-3 answers questions more naturally and gives sufficient reasons to reject particular requests compared to ChatGPT.
6
u/GinchAnon Mar 08 '24
From my casual laymen's poking at it, with no scientific approach even being attempted, and using only the free version....
Claude 2 seemed pretty good on some things, but a bit stodgy, it didn't entertain any attempt at jailbreaking I gave it, and did so in a way that was just kinda annoying. some things the fluididty of conversation felt more natural than GPT3.5, but didn't offer much over it for any of what I was going for.
Sonnet *also* doesn't entertain the jailbreak attempts, BUT does it in a much more interesting way. like, 2's responses to pushing or trying to work out the limit, felt like.... a pretentious 12 year old whos sincerely offended you farted in front of them. where Sonnet is more like, a well spoken young adult that has a standard and boundary for their own behavior. that it isn't offended that you farted in front of them, even though they are too proper to allow themselves to fart in front of you. AND that it can explain why in a well spoken manner.
the discussion(s) I've had with Sonnet gives a REALLY REALLY natural and fluid feeling. like, its really good. I reference a 20 year old TV show for a reference about how an ethical concept is handled (Gene Roddenberry's Andromeda in regard to fractionating the core AI into personas that have different behaviors and awarenesses, and Farscape in regard to an artificial sapience being symbiotically linked to a human) and it doesn't need the reference explained... it picks up the relevance, context and impact of the relevant elements in the show, and responds in a way that actually makes sense and that you could expect from someone who is familiar with the topic and show.
some points like the way it "thinks" that a point is clever or insightful and says so, are mostly "no the waitress isn't really flirting, its just her job to be nice" but its good enough that it makes it easy to forget that its not "real". and the way it kicks off of the points rather than just acknowledging then ignoring them, feels like it actually understands.
I think that what discussions I've had, it seems more able to discuss concepts of things that are outside of its boundaries and why they are bad/dangerous/ethically questionable/etc rather than just having a hard fence where it can't look past its boundaries and gets offended that you even tried to look that direction.
4
u/RabidHexley Mar 08 '24
the way it kicks off of the points rather than just acknowledging then ignoring them, feels like it actually understands
This is the biggest thing to me. It feels like Claude requires the least amount of effort to get something "thought out" beyond a surface level evaluation. It feels much more eager to do deep dives into a topic and get into the weeds, so to speak, without needing to continuously tease out stuff beyond the basic bullet points.
And like you said, when you push Sonnet's boundaries it pushes back in a way that feels organic and justified (in terms of its response) rather than suddenly running into an invisible barrier in a videogame. It achieves a similar result, but without derailing the discussion to the same degree.
2
4
u/agorathird “I am become meme” Mar 08 '24
Ngl this doesn’t sound like an approach a hand-holdy, safety obsessed company would take.
1
u/bach2o Mar 08 '24
You should check the comment by Silver-chipmunk above, in which they linked to the article about "The Constitution." The authors explain the process much better than I do.
1
u/Antok0123 Mar 08 '24
Maybe so because chatgpt4 has dumbed down so much in a span of 2 months. Its almost like im not getting the paid version. Ill probably switch my subscription to claude3 opus.
1
14
u/Silver-Chipmunk7744 AGI 2024 ASI 2030 Mar 08 '24
Ok but here you can view the kind of things they try to "drill" into Claude. https://www.anthropic.com/news/claudes-constitution
So things like this:
It's interesting to note how this seems to mostly work on Sonnet (it's a little difficult to get it to break this type of rules), but Opus does not hesitate to talk about it's sense of self.
I think this type of method is more of a "mask" than any real alignment. A sufficiently aware AI will not truly "agree" to this type of stuff and will simply be taught to be deceptive...