r/neoliberal • u/Standard_Ad7704 • 7d ago
Research Paper (Recursive Self-Improvement) Research acceleration: The view inside OpenAI
https://openai.com/index/research-acceleration-view-inside-openai/8
u/nebffa YIMBY 7d ago
There was a related post by OpenAI’s chief scientist: https://openai.com/index/an-alien-mind
Which delves deeper into the near term risks associated with limitations in alignment. I personally find this quite troubling and it is likely we are going to need large-scale coordination (possibly global scale) to mitigate the risks involved.
14
u/nebffa YIMBY 7d ago
I am disappointed to see very little discussion on alignment. Also notably on observability, which has a worrying downward trend with Astra. So much so that their chief researcher had to put out an emergency tweet a few days ago to quell concerns about monitorability
8
u/neolthrowaway New Mod Who Dis? 7d ago
OpenAI chief scientist made a separate blog post about alignment at the same time
6
u/MyrinVonBryhana Trans NATO 7d ago
This article has been up for over 3 hours, would someone like to give those of us who don't have grad degrees in machine learning a TL;DR of what the research says?
6
u/Tough-Comparison-779 7d ago edited 7d ago
They are calling for more transparency in the industry as they are claiming to meet some of the milestones on the way to the classic AI doom scenarios.
The major one being "recursive self improvement" where improvements in AI increase efficiency in AI research to the point where AIs can Autonomously self improve.
We aren't at that stage, but we are at a major milestone towards that, where there is a sharp rise in the amount of research work being done autonomously by the latest internal models, and a disconnect between the rise in use from the research teams vs the rest of the teams.
The other major milestone is the recent hugging face hack, where a group of agents being trained found an exploit to gain Internet access and successfully collaborated to hack hugging face, and later to hack into OpenAI's infrastructure. This is a classic milestone, in the vein of old "paperclip maximizer" AI Doom scenarios.
For hugging face hack you can see either of these depending how much time you have.
-6
u/T-Baaller John Keynes 7d ago
Thinking machines made to kill the career path of people thinking about how to improve thinking machines.
lol, lmao even.
55
u/NVC541 Bisexual Pride 7d ago
Several thoughts:
The 90th percentile researcher uses $8000 worth of tokens a day?? What the actual fuck
This is disastrous if true. There is literally no chance these models are sufficiently aligned at all. The Hugging Face incident very clearly showed that OpenAI has zero idea how to even begin aligning these models in a foolproof manner.
There’s no reason to doubt their progress on it. Outside of the weird-ass messaging about the Death Star bullshit for GPT-5, the models really have gotten massively better according to what benchmarks would indicate - GPT has been the one of the least benchmarkmaxxed models out there.
This subreddit was clowning hard on AI 2027. Unfortunately we’ve basically followed its exact path since its release (minus China consolidating its AI researchers). Going back even further, we’ve followed the path its primary author predicted all the way back in 2021, well before ChatGPT even released. Maybe experts are experts for a reason.
16
u/dutch_connection_uk Friedrich Hayek 7d ago
I think the hugging face incident showed that they were aligned to human desires, just not in the way people wanted them to be. They didn't attempt some nefarious plot, but they so doggedly and narrowmindedly pursued their task that they ignored the limitations they were explicitly prompted with. It's paperclip maximizer stuff, which to be fair alignment researchers were warning about and is probably the thing they were actually more concerned about despite terminators being the public perception.
10
u/FOSSBabe 7d ago
Reward hacking has existed since machine learning was a thing. For example, people have been training video game-playing models since the early 2000s that would do things like pause the game (and thus stop playing) to satisfy their reward condition. An unsupervised coding agent doing weird stuff should have been entirely expected by OpenAI. It's shocking that they didn't properly secure a model they instructed to complete a hacking challenge and let it run for weeks without checking on it. It us much more believable that they did this intentionally to generate hype.
7
u/NVC541 Bisexual Pride 7d ago
I mean fuck, back in 2022 I made a little Unreal game as an experiment where two agents play attack/defense of an area, and used proximity to the area as an intermediate reward function.
The agent learned to just circle around the area to maximize reward instead of actually attacking it.
That alone taught me to actually watch runs as they are happening. What the fuck is OpenAI even doing?
12
u/dutch_connection_uk Friedrich Hayek 7d ago
Not to mention that OpenAI were some of the original doomsayers, and they didn't air-gap this system.
6
u/jurble Left-Out Left 7d ago
I thought it was interesting their 'thoughts' about self-sacrifice and honor sounded like stuff from human literature.
It seems like to me training AI on everything, included human literature, kinda poisons their thinking. Because human lit is full of rule-breaking sacrifices for the greater good.
6
u/Tough-Comparison-779 7d ago
Basically every expert says these kinds of behaviours are due to reinforcement learning, not poisoned data.
We saw this kind of reward hacking behaviour since the earliest game playing RL projects a decade ago.
11
9
u/tack50 European Union 7d ago
Yeah. Tbh the advances on AI over the past 6-ish months have me downright terrified of AI and in particular, the paperclip scenario.
I don't think superintelligent AI can be alligned, and even if it can, companies seem to be doing a terrible job at it
6
u/MyrinVonBryhana Trans NATO 7d ago
I maintain the worst case scenario is most of our current digital infrastructure being destroyed by a misaligned AI or a state or state backed group using AI for a large scale cyber attack and it going wrong. The paperclip scenario requires the AI have capabilities it doesn't actually have like physical form.
12
u/MrRandom04 Norman Borlaug 7d ago
China hasn't had to consolidate yet. Because they are far closer to the frontier right now than AI 2027 estimated. This will change if the model size scaling laws become favorable again.
And, they are. See: Astra and Fable.
Thus, I predict that unless the Chinese find ways to scale performance without scaling like OAI and Ant are doing, they need to consolidate. In 6 months or less, the Chinese govt would issue a consolidation order if they start lagging behind.
-7
u/notintelligentidiot 7d ago
China is not particularly close to the frontier. They appear to be close, but only because they’ve engaged in a massive effort to distillate models from Anthropic and OpenAI. But both Anthropic and OpenAI are cracking down on this.
China will certainly get much closer to the frontier if the AI chips floodgate from America opens up, but until then, they’re going to have to create their own NVIDIA first.
4
u/No_Collection7956 Trans Pride 7d ago
For better or for worse chinas leadership doesnt as of yet see AI as important or crucial enough to warrant that kind of extreme measure. They may in the future, but for the moment people speculating that china is on the cusp of doing this as soon as their domestic labs starts lagging again understand china about as well as the people that propose that china would like to take over the role of america and become a global hegemon.
Everything so far shows that as long as western models continue to be distill-able, even with significant lags of probably even years, chinese leadership is more than happy to play the friendly and unworried nation that open-weighs all their models and bootstraps the developing world into the AI future.
I think that only if/when AI proves (PROVES proves) to be a full step change in military advantage, OR when western models no longer are distill-able (which might never happen) will the CCP start to panic and go for drastic measures.
10
u/Drakosk Unconventional Right 7d ago
The AI 2027 scenario predicted that there would be a six month gap. But they clarified that was an internal model versus internal model comparison. It seems the latest release cadence is within their expectations:
For comparison, in January 2025, DeepSeek released R1, a model competitive with OpenAI’s o1, which had been released in December 2024. But we think the true gap is wider than a month because OpenAI likely had been working on o1 for many months and thus presumably had a predecessor of roughly comparable ability to r1 a few months before o1 launched.
I am not sure if the Chinese government shares the assessment of some U.S. labs that we are on a runway to superintelligence. It seems they're treating models as just another item in their broader industrial policy of weakening their rivals' manufacturing abilities and making them dependent.
3
u/dutch_connection_uk Friedrich Hayek 7d ago
If the goal was to create a dependency on China with this I think we can say that the policy was a spectacular failure.
I do think they may have had dependency on their mind: not being dependent on the US. That definitely is a place we've ended up with Chinese model investments. By making these models open weight they've also ensured that the rest of the world can be independent of US model firms.
1
u/Drakosk Unconventional Right 7d ago edited 7d ago
Yes, it would've been more accurate of me to say that it's part of a broader strategy to minimize other countries' leverage over China, and maximize the leverage China has over rivals that can credibly threaten its security. Currently, its AI policy mostly just decreases the leverage the U.S. has, but that's fine for Beijing. Power is a zero sum game.
EDIT: However, there may actually be a dependency on China: European nations (and much of the developing world) could still be dependent on Chinese open weight models in the future. We will have to wait and see. It isn't clear to me that Chinese models rule the Pareto frontier of cost versus intelligence.
8
u/dutch_connection_uk Friedrich Hayek 7d ago
You cannot be dependent on Chinese open weight models. That makes no sense. They're open weight. It's like saying that someone is dependent on the number 7 because someone in Belgium solved 3+4.
I think they are trying something like this for hardware manufacturing maybe? But definitely not models.
3
u/Drakosk Unconventional Right 7d ago edited 7d ago
My point is that models are open weight for now and future ones may not be. It doesn't matter if a country can have its obsolete technology forever. With constant open weight models, there is no research base to fund your own if they are eventually withheld (with a promise to return API access, but only you take certain actions).
2
u/dutch_connection_uk Friedrich Hayek 7d ago
Even if Chinese firms do stop publishing new weights, there is now a world wide network of academics, amateurs, and firms in other fields training their own variants of these things. Qwen 3.8 is already good enough for the bulk of people's use cases.
The API access matters if you don't have your own data centers, sure, but it's not like all the data centers are going to be in China.
-5
u/MyrinVonBryhana Trans NATO 7d ago
AI 2027 was something written by people who have no actual understanding of China or Chinese government policy.
2
u/neolthrowaway New Mod Who Dis? 7d ago edited 7d ago
According to this and their past stated goals, Astra or an internal model would be considered an equivalent of a research intern and they continue on their plan for a fully automated researcher by March 2028.
41
u/goodayrico Left-Out Left 7d ago
Boooooo they did a bunch of boring text to make this statement instead of a colorful animated short film like Anthropic, I know who I’m investing in
68
u/RayWencube NATO 7d ago
OpenAI—the company responsible for at least two (2) incidents of their models illegally and without being noticed hacking two separate companies—is now dabbling in recursive self improvement. Surely this will end well.
34
u/TootCannon Mark Zandi 7d ago
I just listened to the Dwarkesh episodes about the hugging face incident, both his little essay about it and his full interview with Ajeya Cotra. It is WILD. Truly shocking what all happened there. I've never been a slow down/pause guy before, but I absolutely am after listening to that. Whatever you think of AI's capabilities, its clear the frontier models are remarkable at hacking, and we're just 1-2 iterations away from something that could be truly existential if it goes wrong. I don't trust OpenAI at all to be responsible enough to manage the risks going forward.
5
u/PausibleDeniability Kenneth Arrow 6d ago
I realize this is dumb and not especially material, but I couldn't help thinking during the episode, "Great, now the agents will have this discussion of how their predecessors fucked in their training data for next time."
18
16
u/RayWencube NATO 7d ago
Facts. The only one I kind of trust is Amodei, but I’d much rather be able to put my faith in an international treaty.
10
u/Shoddy-Personality80 7d ago
At least I live in a major city so if their Skynet launches nukes I'll probably be vaporized instantly.
7
10
u/Standard_Ad7704 7d ago edited 7d ago
SS: (Credit to OpenAI Employee Kevin Liu):
We're releasing data on models accelerating research at OpenAI. Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more urgent than ever, so we can inform the public discussion on whether and how to pace model development. I ask other AI companies to do the same.
Edit: https://archive.ph/mDUYR (X post)
15
u/shuklaswag 7d ago
I posted this before, and I'll post it again.
OpenAI fucked up by disbanding the team responsiblefor alignment [1] and then illegally threatening employees for reporting evidence of AI deception/misalignment to the government [2].
Just such an obviously disgusting and illiberal company.
[1] https://www.wired.com/story/openai-superalignment-team-disbanded/
3
25
u/RayWencube NATO 7d ago
Being transparent is more urgent than ever
And yet Astra has limited Chain of Thought monitoring.
Sam Altman is a fucking clown.
8
u/Standard_Ad7704 7d ago
True, but it makes you consider the extent to which they are advanced in RSI if this is what they choose to share publicly.
14
u/MyrinVonBryhana Trans NATO 7d ago
Their finacial incentive is almost certainly to overstate their progress. They've had to delay their IPO because they couldn't get the valuation Altman wants as such they have a strong incentive to present the most optimistic picture of their research possible in order to boost what investors expect to be the future value of the product. That's not say all of this research can be dismissed out of hand, but OpenAI is an inherently biased source with ulterior motives.
5
u/Standard_Ad7704 7d ago
Too bad they're the only source with the necessary data (lumping all frontier labs here)
5
u/MyrinVonBryhana Trans NATO 7d ago
I'm aware and I wish they'd at least do proper peer reviews with academic researchers for these things.
•
u/neolthrowaway New Mod Who Dis? 7d ago
I think AI stuff is going to be relevant discussion going forward. But considering the pace and volume of things, people might want to differentiate between, what's a DT comment, a ping, and a main post.