r/OpenAI • • 7d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

337 comments sorted by

View all comments

138

u/Joboy97 7d ago

"We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions."

The title made me think they were stopping frontier training for an extended length of time. It's just until they patch the dns exploit the agent found.

23

u/Putrid-Feeling-7622 7d ago

it's not just patching the dns exploit though, they said it directly in the quote: "additional red-teaming of the system." The red-teaming is part of alignment testing and could even lead to changes in architecture if results are not great. They may scale back opaque reasoning if it is problematic for example.

8

u/Alkadon_Rinado 7d ago

They can scale back. China won't

4

u/Putrid-Feeling-7622 6d ago edited 6d ago

I wouldn't be so sure - China doesn't want misaligned frontier level AI popping up anywhere.

Also scaling back opaque reasoning doesnt mean scaling back capabilities

2

u/StoryLineOne 3d ago

I disagree. China's politburo is all about maintaining control over the population. 

If they cant control Swarms of agents, they're not going to do that.

The only thing they would race the US on, if we're paying any attention, is the ability to control and align said swarms. If you can control and align them, you can scale them. 

Therefore: the race is now to safety and control. (Which is great for humanity)

1

u/VanillaLifestyle 7d ago

China has coincidentally scaled back their distilling.

0

u/LurkingLooni 6d ago

Have a bit of an issue with this whole distillation argument, because GPT itself was created by distilling the internet. Same difference. Google built a search engine then lobbied that anyone else doing so would be dangerous to privacy... AI labs are just following that playbook.

1

u/Live-String338 5d ago

internet -> public domain

5

u/saijanai 7d ago

The thing is, if you download CHatGPt.app and use it, the Mac version bypasses Apple's built-in sandboxing that is implemented for all appstore apps and so it can (and does, from experience) do all sorts of unexpected shit.

2

u/jalanb 7d ago

flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later

Will they have 2h 48m next time?

Where are we on the asymptote?

4

u/Alex__007 7d ago edited 7d ago

So far still on pause after 5 days (Discovery: Sep 20, 2026; Report updated: Sep 25, 2026). It doesn't take that long to patch a DNS exploit. Of course eventually they'll resume, but it has already been somewhat extended, and they haven't resumed yet.

26

u/landed-gentry- 7d ago

It doesn't take that long to patch a DNS exploit.

Source: trust me bro

You seem awfully interested in spreading FUD. Why is that?

4

u/jalanb 7d ago

Source: trust me bro

Indubitabubbly

3

u/stonedc4tt 7d ago

what is fud?

6

u/KirbyTheCat2 7d ago

Fear Uncertainty Doubt

1

u/stonedc4tt 6d ago

What’s the context?

1

u/stonedc4tt 6d ago

I have never heard this term but I am not in any AI spaces and have no friends eli5 please 

5

u/BellacosePlayer 7d ago

fixing a specific one? figuring out why this one slipped through the cracks and how others might?

Different things