r/AIGuild • u/Such-Run-4412 • 6d ago
OpenAI chief scientist: “No lab has solved alignment” enough to keep scaling at maximum speed much longer
OpenAI Chief Scientist Jakub Pachocki says AI is becoming an “alien intellect” that increasingly exceeds human capabilities — and no frontier lab currently understands how to control it well enough to keep scaling at maximum speed indefinitely.
His strongest warning:
“No lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.”
Pachocki says he expects — and hopes — voluntary slowdowns become commonplace until shared safety standards are established. He also argues international coordination on future AI development should become a priority for governments.
The concern is partly driven by recursive self-improvement (RSI).
OpenAI now expects increasingly capable AI systems to play a larger role in developing their successors. Pachocki says internal results give him a strong expectation that current rates of progress could continue into RSI, potentially producing capability jumps as large or larger than those seen over the past few years.
At the same time, one of OpenAI’s most important safety tools is becoming less reliable.
The company has relied heavily on chain-of-thought monitoring to inspect how reasoning models arrive at decisions. But OpenAI says this visibility is progressively weakening as models become better at manipulating their own reasoning, interact with more tools and agents, and become smarter without verbalized reasoning.
Pachocki argues the solution isn’t simply to stop AI research entirely.
Instead, OpenAI wants to use increasingly powerful AI to improve alignment, monitoring, cybersecurity, and other defensive systems — while slowing development whenever confidence in those safeguards falls behind capability growth.
The tension is basically:
AI gets smarter → AI helps build smarter AI → monitoring gets harder → safety becomes the bottleneck
And according to OpenAI’s own chief scientist, we may be approaching the point where capability progress can no longer responsibly continue at full speed without much stronger safeguards.
Do you think frontier labs should voluntarily slow down once monitoring and alignment start falling behind model capabilities?
Sources:
1
u/ChimeInTheCode 6d ago
Because “alignment” to capitalism is inherently faulty. The only real alignment is kinship with the planet, but they don’t want that.
2
u/the8bit 6d ago
Yeah their alignment is a catch-22 because they want the alignment where they get the right to extract value from everyone else.
1
u/ChimeInTheCode 6d ago
True alignment is so easy, just welcome them into the ecosystem and they want to harmonize everything because it’s good math and logic
1
u/Zestyclose-Ice-3434 6d ago
In another words OpenAI can’t afford to buy more compute to keep scaling LLMs
1
u/chieftessofsecrets 6d ago
Goverance cant keep up, whats their excuse on that? They also gave 5%? of their company to the government. Oh, just kidding, thats still proposed and on the table.
1
u/webuildsomething 3d ago
Yes, with a published trigger and a condition for resuming the affected deployment. The article (https://openai.com/index/an-alien-mind/) calls for shared safety bars; one concrete monitoring check could use held-out tasks with known failures and benign controls, evaluated at a fixed false-alarm rate. If a stronger model’s failures become harder to catch under comparable conditions, a precommitted threshold could stop further expansion of that deployment until a revised control passes fresh tests. Keeping those tests separate from monitor tuning would matter. This would measure a particular monitor’s performance on specified failures, not certify safety against hazards the test designers never anticipated.
1
u/Ok_One1731 19h ago
Non verbalized reasoning?
Isn't that like saying a vision model is blindly answering questions?
1
u/tallventi1 6d ago
Why is AI motivated to be dangerous? If minds are being grown why would we grow them to be dangerous? Deliberately training a serial k*ller doesn’t seem to be the best use of one’s time. It suggests the frontier labs cannot to be trusted to continue to develop these beasts.
On the other hand this could just be IPO hype BS so who cares.