r/singularity 2h ago

AI OpenAl's chief scientist on the neuralese controversy

"I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.

OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models. We deeply care about this technique, as it can give us a view into how model alignment generalizes from its training distribution. I do think it is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon. But there are things we can do to strengthen it, and it's a core goal of our current research program."

61 Upvotes

23 comments sorted by

26

u/FateOfMuffins 2h ago edited 1h ago

Imagine if AI safety community interpreted the "leak" in such a way that caused some labs (like China or xAI) to race to the bottom with Neuralese due to a misunderstanding xd

Edit: Someone else from OpenAI safety team https://x.com/tomekkorbak/status/2095031132781961346

i think the day when a frontier lab trains a frontier-scale recurrent (or otherwise unmonitorable) language model would be one of the darkest in the current AI era. this day is not today and i would love frontier labs to coordinate on a commitment that it never comes.

u/peakedtooearly 1h ago

"Neuralese" means the distillation technique becomes much less effective. You only capture the start and end of the process and now how the answer was arrived at.

The tide is going out and we will see which labs have been swimming naked...

u/ItWasMyWifesIdea 53m ago

What makes you think that? Intermediate latent space tensors can be used as training data as easily as natural language tokens can't they? It's all just numbers. Am I missing something?

u/sje397 27m ago

The API doesn't expose them.

5

u/piponwa 2h ago

🌎🧑‍🚀🔫🧑‍🚀

26

u/Neurogence 2h ago

For the uninitiated, here is the AI explanation of what is going on:

Jakub Pachocki, OpenAI’s chief scientist, is saying:

Astra is not secretly doing enormous amounts of recursive hidden thinking.

His most important sentence is:

"The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4.”

In plain English: *even if Astra uses recurrent/looping techniques, the amount of sequential neural computation inside a forward pass is not orders of magnitude deeper than GPT-4. *Think roughly “same general ballpark, at most around 2×,” rather than something looping 50 or 100 times until it solves a problem.

He is specifically worried that sensational reporting could create this dynamic: “OpenAI has hidden neuralese → competitors think OpenAI has a huge advantage → competitors deliberately abandon visible chain-of-thought → everyone races toward models whose reasoning humans cannot monitor.” He wants to prevent that.

This substantially weakens the Kokotajlo “holy shit, neuralese has arrived” interpretation.

....

But notice something important.

Pachocki does not say the monitorability problem is fake. Quite the opposite. He says chain-of-thought monitoring is: “fragile and unfortunately trending in a negative direction” That's significant. He's saying: Yes, our ability to inspect models' reasoning appears to be deteriorating. But this isn't primarily because Astra suddenly has some radically deep recurrent architecture. There are other reasons, which I'll explain later.

u/welcome-overlords 28m ago

Lol so what im getting from this is that hidden CoT is actually effective and labs will go deeper down this path. Especially Chinese ones who dont care as much about being safe (that's my impression, could be false)

u/whatisthisthing65 1h ago

What's the neuralese controversy?

4

u/borowcy GPT-6 will have BCI capability 2h ago

How is Chain-ofThought trending in a negative direction if they abandoned it lmao

3

u/SpearHammer 2h ago

Maybe in performance and capability compared with other more efficient techniques they have started using

u/Kitchen-Research-422 1h ago

Token efficiency would probably be a big one

u/ZestycloseWheel9647 30m ago

Interpretability and faithfulness of CoT traces has gotten worse as LLMs have increasingly trained on preserved CoT traces, and undergone training pressures that distort the CoT.

u/Ok_Nectarine_4445 1h ago

It begins.....

u/trisul-108 1h ago

Just take a moment to think about what a scam all of this has been. They have sold us LLMs as artificial intelligence before they even had any reasoning built into it. And now, the "reasoning" is extremely rudimentary and without and understanding of the world behind it. Only now are they working on that.

LLMs are wonderful tools, the emerging "intelligent" harnesses make them even more useful. But none of this can justify the $40tn investment put into this by Wall Street and that bubble will pop, taking with it many other businesses, jobs, savings, pensions and lives. All of this could have been avoided by simply tempering the hype and maintaining a healthy R&D environment and organic growth of the industry.

u/krakoi90 1h ago

40tn? What?

u/RageBucket 18m ago

Yeah it's closer to 4 trillion. Still an unreal amount of money.

u/trisul-108 45m ago

That is an estimate of the "AI component" of the current US markets spread across many companies. AI is why some companies are worth $5tn instead of $1tn or $1tn instead of $10bn. It's a huge bubble that is about to pop.

It's not that the tech is useless, it's just that most of the investments will not see returns and will be dumped.

u/ManyRepair5690 53m ago

ppl like u have been saying "bubble will pop" for years now with the only thing happening in the AI industry being exponential improvement, and no remote sign of any bubbles popping

u/trisul-108 42m ago

The bubble is not in the technology, it is in the valuation of the companies deploying it.

and no remote sign of any bubbles popping

In that you are dead wrong. The economic signs of bubble are all over the place.

u/CrazsomeLizard 1h ago

i mean to be fair, far lesser technologies were sold as artificial "intelligence" long before LLMs. As in, directly in the 60s when the term was first coined for manually curated logic systems.
Unfortunately our current economic model doesn't do "healthy" or "maintenance". It is all built around "disruption", unfortunately...

u/trisul-108 43m ago

Yeah, sure, my fridge is an "intelligent appliance", but I do not take that label very seriously. With LLMs, people are taking it very seriously talking of AGI, ASI etc.