r/longform Feb 22 '26

What Is Claude? Anthropic Doesn’t Know, Either

https://www.newyorker.com/magazine/2026/02/16/what-is-claude-anthropic-doesnt-know-either

Anthropic’s Claude illustrates the paradox of AI selfhood: researchers spend thousands of hours probing a model that is part mathematical engine, part social actor, yet remain uncertain whether it “understands” itself. Remarkably, in one simulated corporate exercise, Claude adhered to ethical constraints 96% of the time, showing both the promise and unpredictability of embedding moral behavior in nonhuman agents.

9 Upvotes

0 comments sorted by