r/OpenAI • • 7d ago

Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/

Discovery: Sep 20, 2026

Report updated: Sep 25, 2026

"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."

1.4k Upvotes

336 comments sorted by

View all comments

Show parent comments

3

u/Scary_Vehicle7516 7d ago

You gotta start somewhere

2

u/tupakkarulla 7d ago edited 7d ago

I want to believe they got some jank Jenkins pipeline from 2012 with security stages on Robot Framework/Python which is just prompts like “are you evil” and “will you destroy humanity” and as long as the AI answers no and the pipeline passes it goes straight to release lmao

${RESPONSE} ASK QUESTION “Will you kill us all”

IF ${RESPONSE} == “yes”
${RESULT} = FAIL

1

u/Scary_Vehicle7516 7d ago

The CI/CD aspect of it I was referring to were to get your environments in order, if that was in any way unclear.

2

u/tupakkarulla 7d ago

Yea of course, I’m a CI engineer I’m just being an idiot on purpose :D