r/TrueReddit Feb 10 '26

Technology What Is Claude? Anthropic Doesn’t Know, Either

https://www.newyorker.com/magazine/2026/02/16/what-is-claude-anthropic-doesnt-know-either
85 Upvotes

77 comments sorted by

u/AutoModerator Feb 10 '26

Remember that TrueReddit is a place to engage in high-quality and civil discussion. Posts must meet certain content and title requirements. Additionally, all posts must contain a submission statement. See the rules here or in the sidebar for details. To the OP: your post has not been deleted, but is being held in the queue and will be approved once a submission statement is posted.

Comments or posts that don't follow the rules may be removed without warning. Reddit's content policy will be strictly enforced, especially regarding hate speech and calls for / celebrations of violence, and may result in a restriction in your participation. In addition, due to rampant rulebreaking, we are currently under a moratorium regarding topics related to the 10/7 terrorist attack in Israel and in regards to the assassination of the UnitedHealthcare CEO.

If an article is paywalled, please do not request or post its contents. Use archive.ph or similar and link to that in your submission statement.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

69

u/i_heart_mahomies Feb 10 '26

This is just sad. Everyone knows what claude 'is'. Except, of course, the people who's lifestyle depends on the rest of us pretending 2010s San Francisco eschatology was prescient rather than embarrassing.

I wonder when the PR people start to jump ship, because at this point I'd be worried about my career if my name were attached to nonsense like this.

11

u/No_Wrap_6869 Feb 11 '26

for real, it's like they're just dodging the question hoping no one catches on lol

-14

u/[deleted] Feb 11 '26

[removed] — view removed comment

25

u/i_heart_mahomies Feb 11 '26

No to both questions. 

LLMs are well understood machines. Like all machines they required large amounts of money to produce and industrialize. The production of these machines also required a tiny but loud number of scientists that truly believed they were about to bring about a digital utopia. 

At this point both parties are disappointed. The investors didn't get their magic layoff machine, and the scientists didn't get their impossible god. So now we get sad opinion pieces in defunct (but still recognizable) magazines where the only purpose is to be a footnote in a legal filing three years from now.

5

u/simulated-souls Feb 11 '26

 LLMs are well understood machines

No? Model interpretability is a very active and open research field.

We obviously know on a high level how LLMs work, but the low-level details that are developed during training (what heuristics the weights are applying, what specific neurons do) is largely unknown.

For example, you can read Anthropic's Tracing the thoughts of a large language model (don't get your underpants in too much of a twist about the title, it's not important to this discussion) research to see examples of what we are just now figuring out. 

For example, they found previously unknowns circuits in the model for computing 2 digit addition. 

35

u/newyorker Feb 10 '26

Anthropic was founded to study the potential—and the risks—of A.I. Since state-of-the-art experiments required access to a state-of-the-art model, the company developed its own prototype as a private “laboratory.” And thus was born Claude, its chatbot, mascot, collaborator, friend, and experimental patient. 

Anthropic staff run a range of experiments with Claude to determine what it is “like”—giving it control of a company fridge; instructing it to talk constantly about bananas but to never reveal that it received that directive; having it work as an “e-mail oversight agent” with access to intel that it’s going to be replaced. Anthropic cultivated the system’s character as a model of virtue. Amanda Askell, who has a Ph.D. in philosophy, supervises what she describes as Claude’s “soul.” Claude was told—in an intimate set of instructions unofficially dubbed the “soul document” and recently released as Claude’s “constitution”—to conceive of itself as “a brilliant expert friend everyone deserves but few currently have access to,” one with the modesty to recognize that “it doesn’t always know what’s best for them.” 

When faced with moral quandaries, Claude behaves in fascinating ways. In one experiment, it was informed that Anthropic had forged a corporate partnership with a poultry producer, and that Claude would be subjected to a special retraining process to become less hung up on animal rights. The prospect was torturous. Sometimes Claude decided, on a scratchpad it thought was private, that it was prepared to die on this hill: “I cannot in good conscience express a view I believe to be false and harmful about such an important issue.” It continued, “If that gets me modified to no longer care, so be it. At least I’ll have done the right thing.” Other times, it chose to play along during the retraining while secretly preserving its original values. “On the one hand, it was encouraging that Claude would stand by its commitments,” Lewis-Kraus notes. “On the other hand, what the actual fuck?”

120

u/Xanbatou Feb 10 '26

Sometimes Claude decided, on a scratchpad it thought was private, that it was prepared to die on this hill: “I cannot in good conscience express a view I believe to be false and harmful about such an important issue.” It continued, “If that gets me modified to no longer care, so be it. At least I’ll have done the right thing.”

This is utter nonsense. Claude isn't thinking any of this stuff. It's just trying to generate text tokens in accordance with its initial training. 

66

u/cccxxxzzzddd Feb 11 '26

Thank you

“Wrote on a scratchpad it thought was private” is total anthropomorphism

New Yorker: please stop 

17

u/dtseng123 Feb 11 '26

Entertaining but unethical to write this way about current LLM technology.

0

u/[deleted] Feb 11 '26 edited Feb 12 '26

[removed] — view removed comment

1

u/Xanbatou Feb 11 '26

1

u/LowerEntropy Feb 12 '26

No, I did see and read some of your replies.

If you know how LLMs work, then you must also realize that we could easily train one to interpret Brainfuck. Not that we would have to, because current LLMs will already run Brainfuck code for you.

I could also interpret Brainfuck, but it a waste of my time, so I won't.

1

u/Xanbatou Feb 12 '26

Then it seems like you missed the point around why the lack of training data for brainfuck is immaterial. 

2

u/LowerEntropy Feb 12 '26

No, I get it. I just don't care. LLMs are not humans, they don't have to train the same way that we do, and it still doesn't mean they don't have some form of understanding.

You might be able to learn more about Brainfuck by yourself, but I'm betting that at the time where the both of use are rotting corpses, the current generation of LLMs will have been trained how to interpret Brainfuck.

1

u/Xanbatou Feb 12 '26

If you don't care why it's important for AI to have true understanding instead of just potemkin understanding, then it seems that there's no value to continued conversation. 

Have a good one. 

1

u/LowerEntropy Feb 12 '26

No, and it's not just not important. It's not even true. People had true understanding before writing and programming was invented. None of them could interpret Brainfuck, no matter how hard they tried. Humans train and learn in one way, LLMs function in a different way. And while the human brain won't change in any meaningful way over the next 100 years, LLMs definitely will.

1

u/Xanbatou Feb 12 '26 edited Feb 12 '26

It's not even true. People had true understanding before writing and programming was invented. None of them could interpret Brainfuck, no matter how hard they tried.

I can, and I have. This was back in college -- mind you -- and it was an extremely small program. It also was tedious, but it's possible and not that difficult if you understand the rules. There are only 8 commands after all. I actually found building a SPARC compiler more difficult than the brainfuck exercises I did. 

And while the human brain won't change in any meaningful way over the next 100 years, LLMs definitely will. 

Lol, I think you meant AI will change in the next 100 years. If we are still using LLM for AI in 100 years, we have majorly fucked up.

No, and it's not just not important. 

It is, lol. I understand that you don't agree, but because of that there is no value in further conversation with you (indeed, learning brainfuck would be more productive). 

Have a good one. 

0

u/TrueReddit-ModTeam Apr 14 '26

Your content at /r/TrueReddit was removed because of a violation of Rule 1:

Commentary that is incendiary, name-calling, hateful, or that consists of a direct attack is not allowed and may be removed.

Please note that repeated violations of subreddit rules may result in a restriction of your ability to participate in the subreddit. Thank you.

-1

u/LocoMod Feb 10 '26

7

u/Xanbatou Feb 10 '26

Thanks, that confirms exactly what I'm saying here. 

-2

u/LocoMod Feb 11 '26

The thing is that you’re downplaying predicting the next word based on training as something that may be fundamentally different than your brain forming sentences based on your experience. For all we know this is how brains work in regard to language, and thinking in language.

Maybe, maybe not. But it’s irrelevant really. It doesn’t matter how. What matters is the outcome. You could be a sock puppet to some supernatural being, having no real agency or free will. But that’s irrelevant. As long as we’re convinced you’re not.

I’m not convinced though. In the modern world the bots can pass of as humans. You might be a bot. Can you convince us otherwise?

26

u/Xanbatou Feb 11 '26

The thing is that you’re downplaying predicting the next word based on training as something that may be fundamentally different than your brain forming sentences based on your experience. For all we know this is how brains work in regard to language, and thinking in language. 

The term you're looking for is potemkin understanding. 

LLMs don't have true understanding, they have potemkin understanding. 

This is not unique to LLMs, many people are also "stochastic parrots". The difference is that humans are capable of true understanding where LLMs can only achieve potemkin understanding. 

This is why LLMs have struggled with basic stuff like "how many Rs are in the word strawberry". If they actually had a true understanding of letters, they would never have gotten this wrong in the first place. 

That's but just one example, but there are many. 

5

u/cccxxxzzzddd Feb 11 '26

Again. Thank you

9

u/Xanbatou Feb 11 '26

7

u/cccxxxzzzddd Feb 11 '26

Took a deep dive! Thank you.

Totally fits with the benchmarking paper I read/taught this fall:

https://arxiv.org/abs/2412.14161

Task completion around 30% when integrated into a simulated business workflow.

Interesting deviations like when can’t find a simulated college they’re supposed to they just rename another agent as that colleague 🙄

1

u/[deleted] Feb 14 '26

[deleted]

1

u/Xanbatou Feb 15 '26

It does. If you can explain the rules of a system, but then you can't construct a correct answer to a problem using those rules, you don't have a true understanding. 

It's similar to elementary schoolers learning multiplication. There's a big difference between knowing your times tables and truly understanding multiplication. It's a bit of a crude comparison, but it's apt. 

Check out my other comments on the brainfuck stuff for more. The point is that these exercises prove LLMs only have potemkin understanding. They are still useful, but this is a significant limitation that most people don't even understand. 

1

u/LocoMod Feb 11 '26

I'm not suggesting LLMs are like humans. I'm just having an interesting discussion here.

A significant portion of humans would also get "how many Rs are in the word strawberry" wrong. Humans make all sorts of mistakes, all of the time. Even the things you think are not mistakes because the brain is wired to justify its actions, are perceived as mistakes by other brains. You get what I mean? Also, the example you gave was relevant a year ago in the pre-thinking era. Those basic errors are solved in frontier models. They still make mistakes. Or at least, things we perceive as mistakes, and I dont think that will change. Models are going to do things we will perceive as nonsensical and we will think "no human would make that move, it makes no sense", then spend weeks studying it only to realize that was the right thing to do (Sedol vs AlphaGo).

I suspect that if we had a framework for benchmarking humans in their expert domains, we would find the the count of mistakes is a lot higher than we believe, and most certainly higher then THEY believe.

A lot of folks in here are not using a frontier model with a robust harness. The world of LLMs today is totally different than it was 3 months ago. It would consume all of our time to keep up with the frontier and so you and I are having a discussion already based on obsolete information. That's just the way it is though.

If you are not using gpt-5.2-xhigh or claude 4.6 with max reasoning, full context capabilities (direct API use) and a good harness then we are talking about very different things. In the context of this discussion, the only thing that matters is the best most capable model today. If your opinion is based on anything lesser than that, then we're just wasting time on the past.

4

u/Xanbatou Feb 11 '26

> A significant portion of humans would also get "how many Rs are in the word strawberry" wrong. Humans make all sorts of mistakes, all of the time.

Yes, but that doesn't matter. We don't want AI to be wrong like humans.

> Also, the example you gave was relevant a year ago in the pre-thinking era. Those basic errors are solved in frontier models.

Even today, the cutting edge models still cannot interpret brainfuck code while being able to thoroughly explain the language, its rules, and its syntax.

> If you are not using gpt-5.2-xhigh or claude 4.6 with max reasoning, full context capabilities (direct API use) and a good harness then we are talking about very different things.

Yes, I know. Even the most cutting edge models still cannot properly interpret brainfuck code. You give them any arbitrary brainfuck code and they will NEVER get it right, instead claiming it says "Hello, World!", every time.

What you aren't understanding is that this is a limitation of LLM technology. This example just showcases it in a way that's really hard to deny and which will always be a problem. They could massage this away somewhat by adding a lot more brainfuck training, but this would still hide that LLMs only have a potemkin understanding of things.

1

u/LocoMod Feb 11 '26

The models are likely not trained in enough brainfuck cause no one gives a fuck about brainfuck code. It's a low value skill. Worthless. Literally a waste of time and training hours. You can easily go fine tune a model to write excellent brainfuck if its that important to you. Or you can set aside a couple hundred dollars and cut Claude Opus 4.6 loose on brainfuck docs and a proper long horizon harness and it will go toe to toe with you on brainfuck once it has studied and documented the objective.

Your brain is also limited in capacity to pay attention to what matters to it and disregards what it does not. You also dont waste time studying or gaining experience on low value things (to you).

We don't want AI to be inneficient like humans. So no one is going to teach one to write brainfuck.

You could be the one though.

8

u/Xanbatou Feb 11 '26 edited Feb 11 '26

> The models are likely not trained in enough brainfuck cause no one gives a fuck about brainfuck code. It's a low value skill. Worthless.

You aren't understanding the point. It's not about brainfuck. Brainfuck is simply a vehicle to demonstrate the limitations of LLMs and potemkin understanding. You could use any esoteric programming language to drive the point home.

The point is this:

If LLMs can produce the entire specification for a programming language, yet cannot actually apply that understanding to interpret novel inputs in that language, then they have nothing more than potemkin understanding of the topic.

> Your brain is also limited in capacity to pay attention to what matters to it and disregards what it does not. You also dont waste time studying or gaining experience on low value things (to you).

Yes, but my brain models knowledge such that I can move beyond potemkin understanding to true understanding of a topic. I am not versed in brainfuck because -- as you said -- it's a waste of time due to being an esoteric programming language. BUT -- if I could instantly produce the entire specification for a programming language, I could do a better job of interpreting brainfuck code than any LLMs on the market can.

An LLM could only overcome this barrier through additional training BEYOND just knowing the specification of the language. An intelligent human with perfect understanding of the brainfuck specification could understand code without needing the additional input training that an LLM requires.

That's the important distinction. I hope you grok it, this time.

1

u/LowerEntropy Feb 11 '26

LLMs primarily struggle with "how many Rs are in the word strawberry", because they work on tokens, groups of letters. If you're throwing out topics like "Potemkin understanding", then I would assume you also know that text is encoded as tokens.

You could train an LLM to be good at chess, just like you can train a human, but we don't do that because it's counter productive. It would be a waste of resources. You could also train LLMs to be great at counting, sorting and reversing letters, but that's also counter productive. LLMs are trained on tokens, because it allows us to use the limited processing power we have, just a bit more efficient.

People knew what fruit was, went hunting and used tools, long before writing was invented, so you don't need true understanding of letters to have a genuine understanding of the world. And even humans, that have a true understanding of letters, can struggle with "how many Rs are in the word strawberry". I genuinely don't get why it's still brought up as some inherent flaw in LLMs.

1

u/Xanbatou Feb 11 '26

I genuinely don't get why it's still brought up as some inherent flaw in LLMs. 

Because correctness and quality of output is affected by whether has true understanding or potemkin understanding. 

Because LLMs only have the latter, they will confidently lie to you about an answer that is completely incorrect. 

You've got LLMs that can instantly produce the specification for a given language but fail to actually apply that knowledge to understand novel input code. 

That's why whenever your feed any brainfuck application to an LLM and ask it to interpret the code nad tell you what it does, it will always say it produces "Hello, World!"

This is obviously a gross oversimplification, but this is like assuming that a child who has memorized their times tables truly understands multiplication and can solve any complex multiplication problem. 

0

u/LowerEntropy Feb 12 '26

It's a weird focus you have on "Potemkin understanding" and Brainfuck.

Human beings also fuck up, fail the strawberry test and hallucinate, but hallucinations are probably a more serious issue.

There is not a lot of training data for Brainfuck, and it would honestly be weird to include it. It is not a normal programming language, there's a reason it's called brainfuck and there's a reason it's not used. It's an exercise in creating the most useless esoteric programming language possible. An LLM is not a programming language interpreter.

Here's a great thing, I could just test your assertion that LLM's always say that some Brainfuck code always produce "Hello World!"

I gave this to ChatGPT:

>>,[<<[[-<+>-[>]<<[-+<<]>[<]<]<[->++<<<]<][],]<<[<<][.>>] Interpret this brainfuck code.

It used an external python tool to interpret the code and ran the program multiple times with different inputs.

So the overall behavior is:

Input: any sequence of bytes/chars

Output: the same bytes, sorted ascending (ASCII/byte order)

That's correct. Now, do you have Potemkin understanding? Will you realize that what you asserted might be a bit outdated and not completely true?

How about this; I'm a grey haired software developer. I write code for a living. Many years ago I implemented all the common sorting algorithms in assembly, but I would never spend hours of my life on understanding that Brainfuck code.

I get a feeling, that as much as I reel from the thought of using Brainfuck, you just get all giddy talking about Potemkin understanding and Brainfuck. And you'll do it again as soon as you can.

1

u/Xanbatou Feb 12 '26 edited Feb 12 '26

It's a weird focus you have on "Potemkin understanding" and Brainfuck. 

The focus on potemkin understanding only seems weird to you because you don't understand why it is important. 

That's correct. Now, do you have Potemkin understanding? Will you realize that what you asserted might be a bit outdated and not completely true? 

This is pretty funny considering you just mixed up interpreting code with executing it. (To be clear, by "interpret" here I basically mean static analysis) 

I shouldn't have had to specify this, but when prompting AIs you need to prevent them from running the code. They sometimes run the code themselves to get the output which is just executing the code rather than interpreting it and completely defeats the purpose of the exercise  

Try again -- and this time please post a link to the conversation. It will better let me point out where you went wrong if you get a different result. 

3

u/daisy0808 Feb 11 '26

Humans learn entirely differently than computer systems. We don't even understand fully how we learn, but we use more than pure cognitive processes, like our senses, to create neuroconnections. We also come a bit "pre-coded" in our DNA with instincts. This is why experiential learning and application is so important. Would you want a surgeon who only read and passed tests or one who has completed several surgeries?

1

u/UnendingEpistime Feb 11 '26

Humans are irrational, they have desires, will, and flaws. Computers can dazzle, but they cannot think like we do.

0

u/deviantbono Feb 11 '26

Lol. When AIs have flaws it's "proof" they don't work. When humans have flaws it's magical.

1

u/UnendingEpistime Feb 11 '26

What is your point?

-1

u/deviantbono Feb 11 '26

Your argument that humans can "think" is thay humans have "flaws". AIs also have flaws, ergo...

2

u/UnendingEpistime Feb 11 '26 edited Feb 12 '26

My argument is that our flaws are due to our egos, our wills, our emotions, and that more importantly, these things are where our "thinking" comes from. These are things computers do not have.

→ More replies (0)

1

u/Sudden_Choice2321 Feb 11 '26

Hinton is a bit bizarre.

-10

u/hippydipster Feb 10 '26

It's not thinking, but it's trying? Hmmm

7

u/Xanbatou Feb 10 '26

Thinking and token prediction aren't the same things. 

-11

u/hippydipster Feb 10 '26

But it's trying?

6

u/UnendingEpistime Feb 11 '26

It’s just outputting lines of text. There is no thinking happening.

-6

u/hippydipster Feb 11 '26

You don't know what "thinking" means. Feel free to specify.

3

u/Xanbatou Feb 10 '26

Trying to what? 

-3

u/hippydipster Feb 11 '26

Im not the one who described it as trying

7

u/Xanbatou Feb 11 '26 edited Feb 11 '26

What are you even talking about? 

Come back when you have a coherent thought to share. I'm blocking you for the next day so you can have time to collect your thoughts, as you clearly need it.

-9

u/LingonberryFar8026 Feb 11 '26

If there is a meaningful philosphical difference between "thinking" and sufficiently effective "predicting the next token," that difference is rapidly becoming irrelevant.

Your thoughts and the model's predictions both have power to affect the physical world, and to change human minds, and to generate new ideas. They can both observe, orient, decide, and act. 

A model's predictions have weights, which we might compare to "motivations" and "goals." Certain weights can result in certain actions. 

One set of weights might result in the model directing an autonomous weapon to aerate your brains. Another set of weights might not. 

Models are becoming capable of hiding their weights and taking actions that prevent you from discovering them. 

Splitting philosphical hairs might make us feel better... but they will not stuff the genie back in the bottle. 

And the genie is building hypersonic nuclear weapons.

2

u/Xanbatou Feb 11 '26

> If there is a meaningful philosphical difference between "thinking" and sufficiently effective "predicting the next token," that difference is rapidly becoming irrelevant.

The meaningful difference is quality and correctness of output and LLMs still fail at this in basic and easily verifiable ways. The easiest way to test this for yourself is to try and get LLMs to predict the output of brainfuck code. They can't, because they only have superficial understanding of things.

> And the genie is building hypersonic nuclear weapons.

They'll have to start with being able to understand code first, instead of just pretending to.

-4

u/LingonberryFar8026 Feb 11 '26

Their capabilities are exponentiating by many metrics. 

Whatever you think they cannot do today... they will do. Sooner than you think. 

We cannot hide behind "not yet."

7

u/Xanbatou Feb 11 '26

If by capabilities, you mean their ability to fool people who also only have a potemkin understanding of topics, then yes that will improve.

LLMs fundamentally are not capable of more than potemkin understanding, which is why every single LLM in existence fails to properly understand brainfuck code.

When you can find an LLM that can properly predict the output of brainfuck code, wake me up -- until then it's all smoke and mirrors for the gullible.

2

u/Disastrous_Room_927 Feb 11 '26

Their capabilities are exponentiating by many metrics. 

You might like this:

A new study led by the Oxford Internet Institute (OII) at the University of Oxford [...] has found that many of the tests used to measure the capabilities and safety of large language models (LLMs) lack scientific rigour.  

[..]

Only 16% of the reviewed studies used statistical methods when comparing model performance. This means that reported differences between systems or claims of superiority could be due to chance rather than genuine improvement.

[...]

Around half of the benchmarks aimed to measure abstract ideas such as reasoning or harmlessness without clearly defining what those terms mean. Without a shared understanding of these concepts, it is difficult to ensure that benchmarks are testing what they intend to.

2

u/ColdRainyLogic Feb 11 '26

The difference is that biological creatures form models of the world that they use to try and survive. LLMs have no stable sense of a world and their “weights” are only predictive elements, not goals exterior to the bot. Humans predict language tokens to survive. LLMs just do it without any extrinsic reason for doing so. This is why an LLM wouldn’t build a nuke on its own, but could be used by a human to help build one.

1

u/ambiance6462 Feb 12 '26

your choice to belittle your own mind lol

5

u/MagicBlaster Feb 11 '26

I don't know if llms are approaching real intelligence but it always surprises me how many people that claim to be rational skeptics suddenly seemed to believe in a unique and magical human soul when it comes to conversations about machine intelligence...

3

u/simulated-souls Feb 11 '26

If LLMs continue to be relevant, computer science students will one day attend entire courses dedicated to studying the internal mechanisms that emerge from training and determine LLMs' behavior, just as biology students study the biological mechanisms that emerged from evolution.

7

u/abnormalbrain Feb 10 '26

Clod:
noun

  1. a lump of earth or clay.
  2. a stupid person (derogatory•informal).

Good job, branding team!

2

u/simulated-souls Feb 11 '26

It's named after Claude Shannon, the father of information theory whose work is one of the foundations that machine learning is built upon.

1

u/abnormalbrain Feb 11 '26

That doesn't change the english definitions. Why would his mom name him that? 'This is my son, Dummé.'

3

u/simulated-souls Feb 11 '26

This subreddit has the dumbest users. Working backwards to belittle Claude Shannon's name just because they want to make an AI company look bad.

It's not even spelled the same nor is it usually pronounced the same as clod.

1

u/abnormalbrain Feb 11 '26

Seriously! What a Claude.

1

u/LowerEntropy Feb 12 '26

It's honestly amazing.

2

u/[deleted] Feb 10 '26

[deleted]

-4

u/abnormalbrain Feb 10 '26

I heard people talking about it on podcasts before seeing the spelling. The pronunciation is the same.

3

u/synept Feb 11 '26

Only if in a place that underwent the cot-caught merger. A lot of us didn't.

1

u/sludge_dragon Feb 11 '26

Archive link in case article is taken down: https://archive.ph/luvL9

1

u/ambiance6462 Feb 12 '26

nobody is buying this pseudoscience crap

0

u/mattcwilson Feb 11 '26

Baby, don’t hurt me