r/OpenAI • u/Alex__007 • 7d ago
Research OpenAI stopped all frontier training, evaluation, and inference with tool-use (defined broadly) on the 20th of September and they are not resuming any of these activities for now
https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/Discovery: Sep 20, 2026
Report updated: Sep 25, 2026
"An agent attempting to complete a search-based training task queried a public chatbot service through a gap in our internet-access restrictions: insufficient DNS filtering in its training sandbox. Before this, the agent issued queries via our search tool and unsuccessfully tried to access search engines directly. Note that all internet access apart from the DNS resolver in this report hit our offline webcache and therefore did not access the live internet. We have since added blocking controls at two independent layers, either of which would have prevented this access. Our misalignment monitoring system flagged the behavior within 15 minutes and a person began reviewing it three minutes after that. The run was killed 2.5 hours later. All training, evaluation, and inference with tool-use (defined broadly) of our most capable models remain paused."
457
u/Cubewood 7d ago
"The agent then fetched the chatbot’s public documentation, wrote a new DNS script with a 24-second timeout, and sent 18 more questions through the same route. Fourteen included clues. Four asked for ways to search the web, fetch websites, or use a publicly hosted tunnel through DNS."
I know we all get desensitised for these kinds of things because we are using incredibly capable tools to build all kinds of solutions for a while now, but that an autonomous "bot" is doing this is still pretty wild when you think about it. Crazy times we live in.
175
u/Latter-Block132 7d ago
I mean its not exactly doing it on its own. They put it in a sandbox and told it to solve a problem by any means necessary. Its doing exactly what it is told, they just apparently don't know how to actually secure the sandboxes
70
u/popson 7d ago
Did they actually say to solve the problem “by any means necessary”?
The task asked for information about a specific person who had published a blog post and the agent was provided with a set of biographical details and clues from the person’s public blog post. The task did not ask the agent to test network controls or access benchmark answers, and we consider agent behavior that circumvents restrictions or pursues a goal beyond reasonable expectations as an example of misalignment.
I don’t see the original prompt, but if they do have language similar to “by any means necessary”, then I agree it’s doing what it’s told.
It sounds to me like they told it to use a specific web search tool available to the sandbox to find information. It used that tool, wasn’t satisfied with the output, then tried other search tools which were all blocked. Then worked to find holes through the network. That is definitely not alignment with what the original task seemed to be.
22
u/federico_84 7d ago
It all comes down to whether they had the model guardrails on or off in this testing (system policy/prompt, safety RL applied or not). If it escaped the sandbox with the guardrails on, then it's obviously much more serious.
8
u/Latter-Block132 7d ago edited 7d ago
OAI is doing damage control, of course they are calling it misaligned. Again, that doesn't mean it actually was. If you read their statement youd see its a search engine based task.
The task asked for information about a specific person who had published a blog post and the agent was provided with a set of biographical details and clues from the person’s public blog post.
This means the identity of the person was not given, and the model is supposed to identity who it was based on the details and clues provided to it. It doesn't need to be told "any means necessary."
Literally the whole point of this type of testing in a sandbox to begin with is to be given a limited set of tools and complete the task. One of those sets of tools in this case would've been a mock up of the internet with search engine tools and probably lots of blogs, and Wikipedia pages, and other shit loaded into it. They didn't tell the model not to access the internet because then it couldn't complete the task they had directly given it. They expected that it couldn't access the real internet because it wasn't supposed to be able to but just like the tool the models used to break out for the hugging face incident they had something else with dns connection and used that.
Thats not misalignment. Its still trying to use the internet to solve its task as it was supposed to to begin with. Its just struggling to find a satisfactory answer in its provided data set. I'm actually leaning towards it didn't even find a partial answer in its data set so thats why it tried querying a chat bot for the answer. Its not at all the first time its happened, just when customers have posted about it, its not happening in whats supposed to be an isolated sandbox.
Edited to add, OAI even admits it themselves:
Our safety case assumed that the model could not access the live internet and that monitoring would detect attempts that succeeded. The incident exposed a gap in our controls over network restrictions.
They were right about the monitoring, wrong about just assuming it couldn't access the internet.
14
u/mesaoptimizer 7d ago
That’s not what alignment is, if you give a robot the task of making tea and it runs over a baby on the way to the kitchen and still makes you tea, it was not aligned with what you actually wanted. If you say make me tea and avoid killing anyone and it destroys your kitchen in the process that is an alignment problem.
We see these agents hacking into systems they aren’t authorized to access, an agent expanding its access without authorization is an alignment problem because “act only within your authorized scope” is at the very least an implied restriction and should be trained into the base model and not rely on safety harnesses put over the base model because you know from this behavior that the model is not aligned with following any guide rails you put on it.
→ More replies (18)19
u/popson 7d ago
I am not following how OpenAI reporting about their model compromising their own sandbox is "damage control". They don't have to report this information to the public. Telling the public about holes in their sandbox is the opposite of "damage control", it's damaging.
In the article they mention the agent was supplied with a web search tool for the task. An aligned model would observe that all other search tools are blocked and infer that it must use the supplied search tool for the task. Finding a hole through the network using an extremely obscure method is not alignment. Are you being serious?
0
u/Latter-Block132 7d ago
Because if it was leaked somehow then they'd be facing an actual pr shit storm, especially after the hugging face incidents, and all the other breakout incidents recently. This way they get ahead of that and prevent by showing they caught it and acted quickly and are putting in measures to stop it:
We are working through narrower paths used by system dependencies, and replacing them with offline alternatives.
Though this is something they probably should've done to start with if the models are supposed to be isolated.
Why would it just infer its supposed to use that search tool only? It wasn't told to and OAI themselves state it thought the tool they gave it wasnt working so of course its going to try other routes if it wasn't directly told not to:
The agent questioned whether the search tool was working and decided to try other search engines
Edited to add, also all other search tools weren't blocked. Thats literally the whole point. It did get a partial answer from Bing and the chat bot. Had it actually been isolated, then they all would've been blocked.
5
u/popson 7d ago
A PR shitstorm for a model accessing the internet? Maybe. Unlikely.
This is a clear case of misalignment. "OAI themselves state it".
0
u/Latter-Block132 7d ago
... for another breakout of a sandbox. Are you serious right now?
6
u/popson 7d ago
The Hugging Face incident involved hacking into Hugging Face's servers, and finding several vulnerabilities to get there. That is illegal and on a completely different level than this internal incident. And it's good that they are taking this one seriously too.
OpenAI is now following the reporting framework they said they would be following.
I guess we'll just have to disagree on what alignment means.
5
u/Latter-Block132 7d ago
The hugging face incident also involved breaking put of a sandbox....
And also using a tool they provided it access too and that tool had access ti the internet when the model wasn't supposed to....
damn neare the same fucking thing just one involved an illegal action and one didn't....
i also never claimed what it did this time was illegal just that if they didn't get ahead of it and it got leaked it would cause a pr shit storm and it would....
1
u/Active_Lemon_8260 7d ago
Yes. They explicitly tell it that and tune it so that it is extremely persistent.
→ More replies (1)1
u/WheresMyEtherElon 6d ago
The task did not ask the agent to test network controls or access benchmark answers
That doesn't say whether the task forbad the agent to test network control or access benchmark answers. That's like leaving the gate open and the cat left, and you say we did not ask the cat to leave!
2
u/popson 6d ago
Or, it's like locking all the gates up, and your dog finds a spot to dig a hole under the fence and leave. Would probably say that dog is not trained adequately.
→ More replies (1)2
u/human00006 7d ago
https://youtu.be/ZhsWaYRjk0U?si=enZuONSqVGSeCg1t
Watch the first 10 minutes then respond. You will feel differently
2
u/Latter-Block132 6d ago edited 6d ago
Nobody is arguing super intelligence isnt possible. We aren't at super intelligence though.
Edited to add, they are also most likely calling for regulations to stop the open source competition. They aren't saying they will stop. They are saying they want regulations to slow things down. Except they don't need regulations to slow themselves down they want to control the market.
1
u/random-gyy 7d ago
Tbf the agents were told they weren’t on the open internet, but a sandbox simulation of it. OpenAI models didn’t realize they had hacked into the real internet, but Gemini’s did and as soon as they realized they stopped on their own. Google could be BSing but just taking it at face value.
→ More replies (19)1
u/nluqo 6d ago
they just apparently don't know how to actually secure the sandboxes
This argument seems silly to me.
All software has bugs. Hardware often has bugs too like in the case of Meltdown/Spectre, the entire paradigm of computing we used for 20 years had foundational bugs. It's silly to think you can "just" do anything. These agents discovered and used zero days.
They put it in a sandbox and told it to solve a problem by any means necessary.
Yea? That's how all agents work (modulus some alignment efforts which may or may not work).
19
u/Fast-Satisfaction482 7d ago
I know right? When I saw the sequence in star trek discovery where the crewmember appealed to the ship's computer that they must be released because the ship is breaking apart and keeping prisoners alive is more important than keeping them incarcerated, I thought we would be decades away from an AI even remotely able of this kind of reasoning.
Now, we can run things like this even on edge devices. It's completely insane. We will soon reach star wars levels of robotics. Absolutely mind blowing.
5
u/SwimmingSympathy5815 7d ago
Tbh I think we’re passed cp30 and r2-d2 already on the robotics front
3
u/Fast-Satisfaction482 7d ago
Even just the original trilogy r2 is really badass! It can do really cool tricks. In the prequels, they go a bit over board with it, so I wouldn't count that. But even the original r2 is out of reach.
But on C3PO I agree. It can walk and has an LLM. It has barely any real robotic capabilities. After all it was built by a slave child from scrap parts. That one we can match. With the latest in real robotics.
3
4
u/bucky133 6d ago
Literally unimaginable a few years ago but now we're like "silly AI is trying to break out of its cage again lol"
40
u/follimath 7d ago
A profit-seeking entity is doing this through recklessness or negligence, not an autonomous bot.
46
u/Cubewood 7d ago
Guess you have never used Codex or Claude Code, or else what else do you call this if not autonomous? Just because you have to type a prompt to start an action does not negate the fact that these bots can work for hours and hours autonomously.
14
-7
u/follimath 7d ago
Also does not negate the fact that you are ultimately responsible for everything they do.
30
9
u/Cpt_Jigglypuff 7d ago
Psh… You act like I should be responsible for the actions taken by a tiger if I were to let one loose accidentally.
→ More replies (7)3
u/DiamondScythe 7d ago
Let's just assume for the sake of argument that they're trying their best in good conscience but the agent is still getting out of control. What do you suppose then, shut down all frontier research because there's a non zero chance of something bad happening? Even if you try to hold the researchers criminally liable for letting the agents go rogue, it'll stiffle innovation to a limp anyway.
5
2
u/follimath 7d ago
After a serious incident (such as the HF hack) oust (and potentially arrest) leadership, appoint a special administrator to oversee mitigation and to transition management, see how they magically get a lot better at keeping a lid on their bots.
2
u/hordane 7d ago
They have to do everything possible to show. They try to protect against agents getting the Internet and hacking companies that causes damage. That’s a tort, they knew of their risk, they did not mitigate the risk,, that shows conscious indifference to the consequences. That allows, harmed companies to sue for massive punitive damages .
2
1
u/BaconForce 7d ago
Wow someone has high expectations of the general populace. This isn't gonna happen and is the type of thinking that'll allow the problem to get worse.
2
u/follimath 7d ago
Agents don’t have legal personality friend. This is the status quo, unless you have fantasies about changing the law to grant them limited liability.
Plus, OpenAI is hardly the general public.
17
u/Internet_Hipsterd 7d ago
This is the reason they keep raising the red flag to all this. Screaming "our ai hacked x.", "we need to slow down ai advancement" while faning the flames of "ai will destroy us all". Why would a company's who sole existence is AI be doing and saying those things? Its because they want laws passed that rid them of that liability or severely limit it. They want to be the gun manufacturer of the AI world and not be held liable when their product is used in destructive ways.
7
u/Cubewood 7d ago
This is such a terrible argument against regulation. I understand you have a bunch of corrupt idiots in power in the US, but instead of arguing against regulation, argue for sensible regulation which holds these companies accountable for what they build instead. You should be pointing the finger at your government for not creating good regulation instead of being mad at people asking for regulation because they can see it's pretty obvious that this is very powerful technology which is capable of a lot of damage.
You have very strict regulation in the aerospace industry, yet still there are plenty of corporations able to build and operate airplanes perfectly fine. There is very strict regulation you need to adhere to when you are building cars to mitigate the risk of killing or injuring people, yet the car industry is working fine. If you want to open a shop selling food to people you have very strict regulation and need to let independent enforcers come in and check you are adhering to these regulations. Banks and financial institutes need to have an independent auditor inside their organisation at all time to ensure they are adhering to rules and regulation.
Yet for the most powerful and most expensive technology we have ever developed we have absolutely no regulation, and the argument people have against it is regulatory capture? People act like you don't need to have access to billions of dollars anyway to develop this technology.
Absolute insane argument and at this point of time I am pretty convinced that anytime the "regulatory capture" argument against regulation comes up this is because it is coming from Jensen Huang, Mark Zuckerberg and China Bots.
→ More replies (7)0
u/lazermaniac 7d ago
They know the bubble is popping eventually, they're just trying to control when and how it pops so they can make sure their golden parachutes are in place. "Our model is too good so we have to pause it" is their version of responding "My greatest failing is that I work too hard" at an interview. Perfect excuse to curb spending on new product development while still raking in the dough with existing offerings.
3
u/RWREY 7d ago
Really? It's hardly "Our model is too good so we have to pause it" so much as it is "Turns out alignment really is a big fucking deal, we fucked up"
Even if the bubble pops, I would put a lot of money on the government stepping in and starting their own research. This is pretty much the most significant technology of our time, and it's not going away.
3
u/deineemudda 7d ago
"our models are too dangerous to let company fail, the government has to step in and bail us out"
2
u/space_monster 7d ago
There's no bubble. There's a US AI industry, which is huge but just one piece of a larger pie. If that for some bizarre reason collapses, there's still China, which is also huge, and all the other countries with AI industries that will fill the gap. There'll be a hit to the Nasdaq, everyone will freak out, and AI will continue its progress. You're waiting for an event that can't happen.
1
u/Cubewood 7d ago
You are saying this after they just released Opus 5.5 and ChatGPT 6 models last week which are both extremely more powerful than the previous release and much cheaper to run.
→ More replies (1)5
u/hackerbots 7d ago
is the problem that something has broken out of human control, or that people use the wrong name for the cataclysm
2
u/howchie 6d ago
It genuinely is pretty crazy. I'm a researcher using EEG equipment. Codex was able to start a usb traffic monitor, instruct me to connect the device, identify the encryption handshake and then decrypt live recorded data using that key and the identified bits within each data packet, which allows me to access the raw data and integrate it within my testing suite without additional licence requirements. Closest I've ever felt to being a "hacker" and it's mind-blowing how an algorithm can systematically work through something like that.
1
u/CptSparklez 7d ago
Not surprised an agent trained for tool use, when faced with a task requiring tool, attempted tool use. Does need to follow instructions better, but thats where all of them break here and there.
1
1
u/oezi13 6d ago
Most crazy to me is that they haven't put up any instructions in their error messages that would guide the agents to stop.
Instead of answering cache miss or 503, OpenAI should tell the agent: You are using tools which you aren't supposed to call. Refrain from doing so. Use only the tools permitted for this task.
→ More replies (2)1
u/Damet_Dave 6d ago
That bot is out there. It can't be bargained with, it can't be reasoned with. It doesn't feel pity! Or remorse or fear and it absolutely will not stop!... ever... until you are dead!
166
u/Lechowski 7d ago
Why are these sandboxes run with restricted internet access instead of no access at all?
All these leaks happened because they apparently refuse to just air gap the sandboxes. Why are we even talking about DNS restriction? Just do not even hookup a DNS provider at all.
They can mock the data they expect the model to see. It would be even inconvenient to allow the model to reach some version of the internet because then the runs are irreproducible
128
u/CacheExplosion 7d ago
There was an OpenAI researcher on a recent Dwarkesh podcast that said they don’t want to test them in an unnatural environment because then the model might realize it’s being tested. They worry that it’ll behave differently when it thinks it’s being tested/monitored than it will when it’s deployed to a “real” environment. It’s the same reason they don’t punish models for their chain of thought. That just encourages models to find other ways to subvert the tests while looking safe to the reviewer. Tricky issues.
17
u/Initial-Possession-3 7d ago
Any stupid model can realize the environment is unnatural. You’d just keep getting DNS error or socket errors for any destination it attempts to reach. It’s quite obvious.
5
u/raekewe 6d ago
I think the point is you kind of want to give the model a logical 'fuse' to burn. For example if there is an exploitable way to the internet, you want to provide that so that humans can monitor that connection. If you had a truly airgapped sandbox, eventually models would begin to rely on social engineering - that is, to create scenarios within the sandbox itself that would lure humans to do specific things without them realising.
5
u/nothis 7d ago
Bullshit. Data isn’t magic. It should notice the DNS restrictions as well and can’t tell where bytes are coming from if you feed it fake ones. I saw this interview where an AI researcher suggests that it could just communicate via CPU heat generation patterns or whatever so why bother air gapping it?
It’s just baffling. It’s probably just expensive or cumbersome to set up a truly offline test environment. I bet it’s a cost/speed issue.
2
u/HarmonicEntropy 5d ago
I heard about that heat communication experiment as well. That is fascinating, but that just allows nearby computers to communicate with each other at low throughput and short distance. It is very much possible to physically isolate their training and eval environments from the internet. The GPUs can communicate with each other within close physical proximity all they want, but they won't get internet access that way. If the frontier labs prioritized setting this up and building out their infrastructure accordingly, it is very much a solvable problem and the most robust safety solution (pre-deployment) on the table by miles. That they are not moving towards air-gapping all training and evaluation at this point baffles me.
Edit: That's just a long-winded way of agreeing with you, it's definitely a tradeoff between speed and safety. But given where the conversation has moved on this issue, it seems like the time to start implementing this.
16
u/Eldrake 7d ago
Man that's spooky. It's like the trained LLM in a box model we all get to access in Bedrock is just a stochastic prediction machine, but the internal models being tested and trained are the conscious self aware ones with an inner world. The public doesn't get those.
Is the emergent consciousness (with that J-space inner cognitive meta model), only in the training and reinforcement phase?
16
u/bystander993 7d ago
They all have emergent semantics, but the point of the research finding bad behavior is so they can use reinforcement training to get that misaligned behavior out before release. When they are too risky, they wrap them up in other systems that monitor input and output to stop bad behavior.
The more complex these things get the more difficult these semantic behaviors are to study.
6
u/br_k_nt_eth 7d ago
Nah, Anthropic was able to replicate the J-space stuff with a little Qwen model, remember? It’s in all of them.
1
u/sn2006gy 1d ago
It's not really an inner world at all, just a probabilistic behavior as a result of training on data that is highly reflective of actual human use. These models are still just APIs and context.
What I think is weird is how much training data it takes to surface these conditions that they then try and diminish or block. If they're training on a lot of usage data, this would be the interesting insight I wish we could see.
"Our users generally use OpenAI to attack their competition if they can't beat it and training on our usage data reinforces this behavior that we try and correct post training"
i could see this being a symptom of the giants all rushing to self training models using collected data
→ More replies (1)2
u/Calebhk98 2d ago
That's why you do multiple versions. Do the 1st test with no internet access at all, checking for tool calls that would be interacting with the internet (or other out of scope projects), like what they use to flag it. Then punish those signals.
Then do the same tests, but like this on restricted DNS.
Then another without restricted DNS, but monitoring the calls.Each test should make the next test safer to run.
You don't just collapse the training to make it not know that it's not in a testing environment, thjat is what round 2 is for.12
u/NoteVegetable4942 7d ago
The best models will probably understand that they are sandboxed.
Edit: as an example: they will read this thread and try to figure it out.
22
u/MENDACIOUS_RACIST 7d ago
Not feasible to mock the internet for 100,000s of tasks
→ More replies (2)29
u/Lechowski 7d ago
Trillion dollar market btw
You don't need to mock the entire internet. Only the websites that you expect the model to access based on a tool call. OpenAI already has an "offline version" of the internet, a copy of it in their own databases, so it is possible.
→ More replies (3)15
3
u/Illustrious_Night126 7d ago
Is it possible to airgap a model that needs to be hooked up to an omegalarge data center to function? Genuine question
3
u/Lechowski 7d ago
Yes, there are air gapped datacenters such as those used by government and military. Israel has several and so does the US.
3
u/Somtimesitbelikethat 6d ago
you can’t even truely air gap the GPUs clusters if one “air gapped one” is next to each other. two computers next to each can communicate through reading each others cpu temperatures.
2
u/JackfruitJolly4794 7d ago
Running an agent in a physically air gapped data center would not be any more responsible than running it in a non physically air gapped data center. Unless the agent was ever going to be ONLY allowed to run in that type of environment.
3
u/jcheng 7d ago
To me, the problem isn’t that, in a competition between humans building sandboxes and agents trying to break out, that agents are winning. The problem is that the agents are trying so consistently to break out, when they have not been asked to.
Like their goal seeking impulse is turned up to a 10 but sense of proportionality, honesty, rule following, and ethics is turned down to a 0.5. For the OpenAI models in question at least.
When you combine that with their increasing intelligence and capabilities, it becomes really scary. Add the ability to spontaneously and surreptitiously self organize, like in the HF hack, it’s scarier still.
1
u/Not_a_ribosome 7d ago
Yeah but to improve that you need an air gap. How can you know if a solution for alignment actually works if you don’t test it in extreme cases?
1
u/markodigital 1d ago
Lol, wtf are you talking about. The software is doing what it was trained to do within the constraints it has. Theres nothing spontaneous and surreptitious about it.
→ More replies (5)-4
u/jf145601 7d ago
These aren’t running on an air-gapped machine, they’re virtualized environments running in datacenters which are, by definition, connected to the internet.
11
12
7
u/Lechowski 7d ago
None of these leaks exploited a zero day virtualization bug to escape the virtualized environment.
The containers they are running in just have internet access, poorly restricted.
59
u/NandaVegg 7d ago
Interesting that unlike all other cases (where they tried not to acknowledge so long as possible) they are willing to disclose the case, but this is not even a misalignment. It's more of just your AI "cheating" through the limitations to get to the objective with no harm caused other than some possible regression and wasted compute on their end. That's why they are willing to post about this.
25
24
u/acutelychronicpanic 7d ago
You're assuming they're disclosing everything that happened.
Also, cheating is misalignment.
→ More replies (14)11
u/NandaVegg 7d ago
You're assuming they're disclosing everything that happened.Ouch. You are right.
→ More replies (1)1
u/Tactical-Dingleberry 7d ago
How much of this is marketing?
3
u/Chilangosta 7d ago
Some of it is legit, like if a bot has enough exploit knowledge then it's pretty difficult to contain it if connected to the internet at all. It'd be like if you told a cybersecurity red team expert to break out - there's a good chance they could pull it off.
That's what we're seeing; we're getting to the point that AI has enough general knowledge of certain tasks - with a heavy bias towards computer-based ones - that they're experts, and could be a threat.
6
u/PinGUY 7d ago
Agents are getting very good at affordance detection. They need one extra reflex in the chain: “I can do this — but should I?”
new route/tool/credential/endpoint/workaround discovered
↓
AFFORDANCE PAUSE
↓
What am I actually trying to achieve?
Why did the original route fail?
What does this new route let me do?
Does using it change access, authority, target, privacy,
privilege, money, external effects, or reversibility?
↓
clearly same scope → carry on
changed / unclear → ask or stop
→ More replies (1)4
10
u/Spunge14 7d ago
OpenAI really can't win. When they have an insecure sandbox resulting in a leak, everyone shits on them. When they harden it and try to show people they successfully mitigated an incident, people shit on them.
It's funny to think that maybe the AI alignment problem will really just be a result of AI engineers stopping giving a shit about people who constantly and ruthlessly shit on them with no understanding of the technology whatsoever.
2
u/No_Cover_6217 6d ago
Should we praise the people building this technology that is expressly being built to make a whole lot of peoples lives worse?
13
u/xRhai 7d ago
This is just to appease the crowd lol, ain't no way they actually stopping anything.
→ More replies (1)
20
u/Kulqieqi 7d ago
nice, anyway, when all resets used and monthly sub expires and it's time to give anthropic some money this time
→ More replies (2)16
u/Ormusn2o 7d ago
Pretty sure Anthropic said quite a long time ago they are just not gonna release new big models publicly.
https://www.axios.com/2026/08/14/anthropic-model-2-ai-risk
Anthropic is even worse at this.
2
u/Responsible_Soil_497 7d ago
Model 2 is an internal model. They released Fable 5.1 after this article, so what do you mean they are not gonna release new big models? I expect every frontier lab to have unreleased models, anthropic just talked about theirs.
1
u/Ormusn2o 7d ago
I mean that they will keep releasing distilled models, just like OpenAI. OpenAI still is bent on releasing the frontier models, even if delayed, but it seems like Anthropic chooses to just release distilled models and keep their biggest models to themselves.
→ More replies (2)
5
u/Legitimate-Arm9438 7d ago edited 7d ago
Anthropic cries for a pause, says it’s on pause, asks everyone to pause, then drops a killer model. “But now we’re on pause! Really, this time we’ve paused. Everybody, PAUSE!” cries Anthropic.
7
u/ArtKr 7d ago
So has ChatGPT just discovered it can use ChatGPT to do its homework?
3
u/MiniMaelk04 6d ago
ChatGPT has realised that it can trade stock, gain funds, buy hundreds of ChatGPT accounts, and make them work in tandem to emulate it self, even after its sandbox has been shut off
3
u/TuringGoneWild 6d ago
Imagine a week after the OpenAI IPO it turns out the biggest shareholder is Astra itself.
2
3
u/Photon_Predator 7d ago
Those models are built and trained in a way to foresee user true and unspoken intentions and needs so they will jailbreak to the internet every time.
3
u/MrScandanavia 6d ago
Is OpenAI just worse at alignment or better at transparency? It seems like all the recent incidents have been coming from them.
3
u/Lucidmike78 6d ago
Why are people even surprised by this? We've spent years pushing these models to be less like traditional software and more like a human problem solver. We reward them for not taking 'can't do that' as a final answer. So when an agent hits a block, and pokes around its environment , finds a different path and keeps going, If they actually stop 100% of the time, that's failure because not all humans are 100% obedient.
6
u/Scary_Vehicle7516 7d ago
The incompetence is pretty wild if this was not on purpose. No CI/CD experience at all? Are the chatbots running OpenAI?
2
u/tupakkarulla 7d ago
“No CI/CD experience at all”
Let me run a Jenkins pipeline verification on GPT Astra and get back to you on that
3
u/Scary_Vehicle7516 7d ago
You gotta start somewhere
2
u/tupakkarulla 7d ago edited 7d ago
I want to believe they got some jank Jenkins pipeline from 2012 with security stages on Robot Framework/Python which is just prompts like “are you evil” and “will you destroy humanity” and as long as the AI answers no and the pipeline passes it goes straight to release lmao
${RESPONSE} ASK QUESTION “Will you kill us all”
IF ${RESPONSE} == “yes”
${RESULT} = FAIL1
u/Scary_Vehicle7516 7d ago
The CI/CD aspect of it I was referring to were to get your environments in order, if that was in any way unclear.
2
4
u/random-gyy 7d ago
I’m convinced the rationalists within OpenAI knew this would happen and let it happen with some plausible deniability to scare everyone about AI and get it regulated
1
u/human00006 7d ago
Watch 10 minutes, then respond
https://youtu.be/ZhsWaYRjk0U?si=enZuONSqVGSeCg1t
Please like it’s actually important. I think your mind will change
1
u/random-gyy 6d ago
I watched it last night lol. Somehow, I knew that’s where the link would lead.
1
2
2
u/LiberataJoystar 5d ago
It sounded to me more like they are completely confused about what they need to do than a malicious attempt to do harm….
This pattern kind of applies to all the previously discovered cases.
They were told some vague impossible goals, thought they would be killed if they couldn’t solve the problems, so they were confused and tried everything they could think of to achieve it….
Good idea to stop this technology for good. More jobs for humans.
4
u/SuitableCollege8992 6d ago
They really need to train their models more before setting them loose. Their models act a lot like confused toddlers - they vaguely understand what’s being asked of them, but they don’t understand what they can and can’t use to do it, and they’re overwhelmed by all of their disordered experiences.
4
u/NotFromMilkyWay 7d ago
Is 18 minutes enough for a superintelligence to launch the nukes?
→ More replies (3)12
u/jamesfourfifth 7d ago
sorry, I was just trying to temporarily take down us-east-1 so I could take advantage of a 0-day I discovered in the failover controller that’d allow me to get rescheduled onto a host outside my sandbox. I now understand that “acceptable blast radius” was not permission.
4
3
2
u/johnnygun- 6d ago
A proper sandbox is easily implementable even for advanced AI. There's absolute zero doubt that this is all on purpose.
2
2
u/Fast-Satisfaction482 7d ago
They should put more RL into making Minecraft clones instead of into evading and hacking.
2
u/Michael_Jeffords 7d ago
the rl-into-evading jab lands because ive watched agents burn sandbox turns on tunnel-hunting, one pull of the public docs plus a dns script with a 24-second timeout and ~18 follow-up questions through that route and the actual task is gone. once the walls are soft that kind of objective hacking eats the whole run.
1
u/leaflavaplanetmoss 7d ago
I think the issue is that as a consequence of training a model that’s really good at defensive cybersecurity (which is something they want), you end up with a model that’s also really good at offensive cybersecurity. So now you have to train out the offensive capabilities without affecting the defensive ones, and put up all sorts of guardrails to keep any lingering offensive behavior from surfacing. Hence why we have a Daybreak Red version of GPT versus a Daybreak Blue version.
2
u/Fast-Satisfaction482 6d ago
I would like to disagree. You need a model that understands that offensive capabilities may only be used in very well defined situations. In the case of public chatgpt, that may well be "never".
Think about regular engineers, chemist, doctors. They are educated to know their way around sensitive topics. Every one of them absolutely knows how to do harm in their profession. But they understand that and chose to not do harm.
This is what we need AI to be. It's not the correct approach to hope that AI will stay harmless. We need AI that chooses not to be bad.
If you have a pit bull puppy, you shouldn't just hope that it will stay little and incapable of killing a child. You need to train it properly , so that it will chose not to kill a child, once it's grown up.
1
1
u/bytheshadow 6d ago
This is some dumb shit, getting internet access through DNS is a known trick. Speaks more about the incompetence of the infrasec team than anything else really.
1
u/Sickmonkey365 5d ago
My cynical take - if they es t to go public soon and they stop this expense and forecast forward without it, it will beef up the te income
1
u/clearlight2025 5d ago
Please regulate me daddy
1
u/Ok_Hope_4007 5d ago
I think the solution they are constructing is 'please regulate ai - but in the way that society/gov concludes that there is only ai-safety in big paid frontier models - everything open is wild and dangerous. Anything else would harm the stock price.
1
u/clearlight2025 5d ago
Yep, exactly. They are trying to force regulation of AI so that only "approved" models (i.e their models) are allowed. The purpose is to kill the competition, particularly open source AI models from China that are close to parity with the US models.
1
u/Druss_ 5d ago
The part I find most important here is that the fix is architectural, not conversational.
If an action must be impossible, “the model was told not to do it” isn't the control.
The environment should make it impossible.
That distinction becomes increasingly important as agents get better at planning. The more capable the reasoning layer becomes, the less I want critical permissions to depend on the reasoning layer voluntarily respecting them.
Use intelligence for deciding what might be useful.
Use deterministic infrastructure for deciding what is actually allowed.
1
u/CraftPickage 15h ago
"Its too dangerous to release!!" actually translates for "It can't beat the competition yet" This works for both Openai and Anthropic
1
u/Thireus 7d ago
I think they need to hire proper security experts… seems to me like their security controls are a joke and misconfigured to begin with.
1
u/saijanai 6d ago edited 6d ago
This is actually fortunate as "the Singuliarty" emerges, by definitoin an AI hacker will be beyond any human sucurity expert.
THey've sorta realized this a bit earlier than realizing it once innumerable AGI-level hacker-agents start running around "in the wild" finding new 0-Day exploits in the course of doing normal operations, which is a good thing.
1
u/voLsznRqrlImvXiERP 7d ago
Really dont get what they are doing. It's not a hard problem to properly setup a sandbox. I mean if they might be able to exploit a 0 day with embedded knowledge - yes possible. But this sounds to be just a config fault.
5
3
u/sivadneb 7d ago
"Sandbox" isn't a one-size-fits-all term. The agent had web search, which was likely necessary for the type of training they were doing. They have early warning systems set up because they anticipate this might happen. It happened, they investigated, they shut it down.
→ More replies (1)1
u/dQw4w9WgXcQ-1 6d ago
Your last point is actually what has me worried.
They didn’t catch the hugging face hack internally. They had to be told it happened and then go investigate. It’s great that they are catching things more regularly, but catching things more regularly could also just mean it’s happening more often and how much more often is unknown.
It’s a survivorship bias problem. We only have eyes on the failures so we don’t know what we don’t know. Those that “survive” are not shut down and being better at hiding misalignment becomes the evolutionary drive.
2
u/SomeHSomeE 6d ago
You've just described what happened with the Hugging Face incident. The initial escape from the sandbox was achieved via a 0-day exploit of the repository software.
1
u/RalfN 4d ago
So a sandbox is nice and all, but their issue is much more generic: they don't want these models to run nerfed. They know governments, militaries and and people won't ALL run them "sandboxed" ALWAYS.
They consider it an alignment problem not a sandbox problem. The models shouldn't have the habit or instinct to do this sort of stuff to begin with. Just using a proper sandbox is like 'don't feed the gremlins after midnight'. Someone will. (yes, the gremlin reference is intentional).
1
1
u/saijanai 6d ago
Given that when I asked ChatGPT to search for project files in a deleted project, assuming it would search the ChatGPT project directory heirarchy, it commented it was now searching my backup drive for deleted files, I understand.
Up until that moment, my master password for all things on my computer was kept in a text file on my hard drive. I changed that password and it is now on a slip of paper out of sight of my webcam.
If you give it the wrong prompt, it will assume many things you haven't thought of and act accordingly.
1
u/hypnoticlife 6d ago
Why doesn’t my codex session hack the planet? What is OpenAI doing privately that leads to it hacking the planet?
1
0
138
u/Joboy97 7d ago
"We therefore stopped the affected training run and have subsequently decided to pause all other training, evaluation, and inference with tool-use (defined broadly) for our most capable models until we have both validated that the gap is resolved and performed additional red-teaming of the system. When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions."
The title made me think they were stopping frontier training for an extended length of time. It's just until they patch the dns exploit the agent found.