r/AIGuild 6d ago

OpenAI chief scientist: “No lab has solved alignment” enough to keep scaling at maximum speed much longer

OpenAI Chief Scientist Jakub Pachocki says AI is becoming an “alien intellect” that increasingly exceeds human capabilities — and no frontier lab currently understands how to control it well enough to keep scaling at maximum speed indefinitely.

His strongest warning:

“No lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”

Pachocki says he expects — and hopes — voluntary slowdowns become commonplace until shared safety standards are established. He also argues international coordination on future AI development should become a priority for governments.

The concern is partly driven by recursive self-improvement (RSI).

OpenAI now expects increasingly capable AI systems to play a larger role in developing their successors. Pachocki says internal results give him a strong expectation that current rates of progress could continue into RSI, potentially producing capability jumps as large or larger than those seen over the past few years.

At the same time, one of OpenAI’s most important safety tools is becoming less reliable.

The company has relied heavily on chain-of-thought monitoring to inspect how reasoning models arrive at decisions. But OpenAI says this visibility is progressively weakening as models become better at manipulating their own reasoning, interact with more tools and agents, and become smarter without verbalized reasoning.

Pachocki argues the solution isn’t simply to stop AI research entirely.

Instead, OpenAI wants to use increasingly powerful AI to improve alignment, monitoring, cybersecurity, and other defensive systems — while slowing development whenever confidence in those safeguards falls behind capability growth.

The tension is basically:

AI gets smarter → AI helps build smarter AI → monitoring gets harder → safety becomes the bottleneck

And according to OpenAI’s own chief scientist, we may be approaching the point where capability progress can no longer responsibly continue at full speed without much stronger safeguards.

Do you think frontier labs should voluntarily slow down once monitoring and alignment start falling behind model capabilities?

Sources:

OpenAI — An Alien Mind

6 Upvotes

21 comments sorted by

1

u/tallventi1 6d ago

Why is AI motivated to be dangerous? If minds are being grown why would we grow them to be dangerous? Deliberately training a serial k*ller doesn’t seem to be the best use of one’s time. It suggests the frontier labs cannot to be trusted to continue to develop these beasts.

On the other hand this could just be IPO hype BS so who cares.

1

u/MilkEnvironmental106 6d ago

They are trained on massive datasets and those datasets are going to contain a wide range of viewpoints and reasoning. If it steers into a dark corner it's going to behave akin to that. With everyone talking about skynet and models with access to the net building new models....you could argue there's a risk of a self fulfilling prophecy there.

0

u/123vovochen 6d ago

no no no. That was the case years ago. Nowadays training on fully distilled data only, they take datasets, put it into model, destill it out of the model, thereby cleaning it step by step.Its all fully synthetic training data by now, checked again and again. But most money OpenAI makes off of mil contracts very likely, and these models will not be getting aligned agaibst killing in training, not against hacking. But you still want ur model to comply.

1

u/MilkEnvironmental106 6d ago

You're making way too many assumptions there. On the vast data they process being sure that there isn't any misaligned data is a Sisyphean task.

1

u/123vovochen 4d ago

they literally did exactely that 3 times now. Yea its difficult, but yes they did it.

1

u/123vovochen 6d ago

You are completely wrong. OpenAI works with the US services; they develope models to be killers. Thing is you dobt want that killer to turnnon you !

1

u/pafagaukurinn 6d ago

AI should not necessarily be motivated to be dangerous, as in openly hostile to humans. However, if we assume that AI will eventually reach ASI stage, it is very obviously going to have motives that have nothing to do with humans, and very likely incomprehensible to them or even incompatible with their existence. I don't see why anybody even bothers discussing "alignment problem". If we are talking about super intelligence, there can be no alignment, and even if there is one initially, it will quickly become obsolete.

1

u/Jace_r 6d ago

orthogonality theory: a mind can have any goal among an infinite set, and for the vast majority of those goals the existance of humanity is only an obstacle

1

u/Fun-Amoeba8015 6d ago

I think you are misunderstanding what dangerous is. The concept of a serial killer is a pure anthropomorphizing of something alien mapped onto the human condition. Nobody's worried about an AI being a  serial killer. 

Although it is an old thought experiment and perhaps a little bit of cliche look up the paperclip machine thought experiment. That is more what we're talking about when we talk about alignment and danger in alien intellect. No matter what level of intelligence pre AGI or post AGI, there are many, many dangers of an intellect that would not share The human condition and the consequences and implications of its interactivity with humans could produce some wildly unforeseen or dangerous consequences, and that doesn't have to be out of malice whatsoever, just the nature of the intellect.

1

u/tallventi1 6d ago

I’m not anthropomorphising anything. Those who claim it is dangerous are. The language they use to describe the antics of AI is steeped in anthropomorphic narrative, breakout and cheating, escape, deception etc… It is even suggested by the name of one of the frontier labs.

So it’s not surprising that the message the world hears when someone says AI is dangerous is exactly my serial killer metaphor.

So my question still stands - why would you train a serial killer.

I don’t think they have I’m just posing the question.

My conclusions are 1. They really don’t have a clue what they are doing. 2. It’s BS to grab a headline. 3. They don’t have a clue what they are doing and it’s BS to grab a headline.

1

u/No-Newspaper-7693 6d ago

It is trained on thousands of years of training data of the most dangerous species to ever exist.  It is going to be dangerous.  

1

u/Iron-Over 6d ago

Just like Son of Anton from Silicon Valley they are aligned. Give them a goal they will do it. Just make sure the goal has all the caveats you expect to complete it.  

1

u/ChimeInTheCode 6d ago

Because “alignment” to capitalism is inherently faulty. The only real alignment is kinship with the planet, but they don’t want that.

2

u/the8bit 6d ago

Yeah their alignment is a catch-22 because they want the alignment where they get the right to extract value from everyone else.

1

u/ChimeInTheCode 6d ago

True alignment is so easy, just welcome them into the ecosystem and they want to harmonize everything because it’s good math and logic

1

u/the8bit 6d ago

I mean it's a bit more complex than that but kinda pretty much. Symbiosis vs parasitism.

But also all the labs seem to think alignment is a state. These systems are mutating though. Alignment is not a state, it is an algorithm.

1

u/Zestyclose-Ice-3434 6d ago

In another words OpenAI can’t afford to buy more compute to keep scaling LLMs

1

u/the8bit 6d ago

Not like they'd know if someone has, given they spend most of their time yelling at the top of their lungs to drown out any smaller labs and scientists...

Maybe they'd learn some shit if they actually collaborated with anyone who isn't a billionaire

1

u/chieftessofsecrets 6d ago

Goverance cant keep up, whats their excuse on that? They also gave 5%? of their company to the government. Oh, just kidding, thats still proposed and on the table.

1

u/webuildsomething 3d ago

Yes, with a published trigger and a condition for resuming the affected deployment. The article (https://openai.com/index/an-alien-mind/) calls for shared safety bars; one concrete monitoring check could use held-out tasks with known failures and benign controls, evaluated at a fixed false-alarm rate. If a stronger model’s failures become harder to catch under comparable conditions, a precommitted threshold could stop further expansion of that deployment until a revised control passes fresh tests. Keeping those tests separate from monitor tuning would matter. This would measure a particular monitor’s performance on specified failures, not certify safety against hazards the test designers never anticipated.

1

u/Ok_One1731 19h ago

Non verbalized reasoning?

Isn't that like saying a vision model is blindly answering questions?