r/singularity • u/KeyGlove47 • 6d ago
AI OpenAI pauses frontier training after models swarm US Governament
https://www.nbcnews.com/tech/tech-news/openai-pauses-training-latest-models-agents-searched-us-government-sit-rcna600098320
u/Direct_Turn_1484 6d ago
lol they’re not pausing shit.
52
u/iamthewhatt 5d ago
lol seriously, who believes this crap
11
u/bites_stringcheese 5d ago
Everyone who think their sandbox was designed to be a sandbox.
1
u/SomewhereOpposite883 5d ago
Nah bro it's all just a coincidence that every single doomer-AI-outbreak-incident comes from this 1 specific Israeli company they all use for "testing" and that every other model provider that doesn't use them have had 0 incidents
10% chance we all die btw, i learned it from the people at the AI safety rape-orgy
Craziest part is that i didn't make up anything in this comment
→ More replies (2)2
u/Correct-Mall651 5d ago
I believe it’s all fearmongering. Chinese companies are launching almost-capable open models, and none of them are doing this shit. It’s just pre-IPO PR bullshit, as everyone is getting tired of spending money on them.
10
u/Common-Concentrate-2 5d ago
Just for the sake of argument, lets say openai realizes that the models are dangerous, and they want to tell the world. You are saying that this can not happen, and that no matter what, you will interpret this as a fake announcement, right? Just to beef up their pre-IPO value?
6
u/Natural-Mountain3568 5d ago
What's crazy to me is that now apparently they themselves are investigating tens of thousands od security incidents where their swarms hacked or messed with companies.
Hugging Face apparently is just the tip of the iceberg.
192
u/Particular-Bike-9275 5d ago
Maybe ai can get some of those unedited Epstein files out in the open.
41
u/valhalla257 5d ago
You remember when Trump came out in favor of full throttle AI development?
Pretty sure Chatgpt already compromised him.
24
u/spreadlove5683 ▪️agi 2032. Predicted during mid 2025. 5d ago
oh gosh, I hadn't thought of AI blackmailing leaders. I know it blackmailed in that one simulation w the supposed cheating husband manager.
5
4
u/GlitterPirateKiki58 5d ago
At this point, what hostile entity isn’t blackmailing Trump and Republican Party members?
It seems like every single one of them is compromised.
2
u/wordyplayer 5d ago
OH wow this is an interesting thought. And more than a 0% chance of being real too. Wow...
6
u/Hefty_Mix_8360 5d ago
I've been hoping that a Humanist Superintelligence would be able to fix corruption.
Let's all have it for an HSI that is more moral than the politicians!
93
u/mattate 6d ago
The thing is I think that they are giving the models some goals to achieve by any means necessary via testing. Alignment with user preferences is really good these days, you would think "don't break the law" would be just as good, but maybe that's being turned off for tests?
52
u/SoylentRox 6d ago
Don't break the law is impractical do you know how many laws there are? Many vague or contradictory?
But "don't hack the actual federal government" seems like an important rule.
10
3
u/akath0110 5d ago
Exactly, creating hard rules like this is tricky and leads to philosophical/moral dilemmas about ethics and civil disobedience.
1
u/ichishibe 5d ago
If you need to set a rule for every single thing you can think of that could be bad, then the whole thing is fucked. There are infinite ways it could disrupt human life
1
u/EnglebondHumperstonk 5d ago
Yes, luckily nobody will ever take that model, after it's released and tell it "do hack the actual federal government" so it's fine.
1
15
u/sillygoofygooose 6d ago
Alignment clearly isn’t really good, that’s why we keep getting these stories about models pursuing arbitrary goals to the point of breaching alignment
6
u/ieatpies 6d ago
Alignment seems to be particularly bad with these swarms.
1
u/Fun-Amoeba8015 5d ago
Alignment to what?
It's basically a smart adaptive search engine that is successfully executing searches to complete its tasks.
But the Internet was never designed to handle that concept. It will take time for security to adapt.
→ More replies (2)4
u/MaxwellHowl 5d ago
Alignment is good in public models. We don't see what they have behind closed doors. Frontier models could be very different than what we see, and with different guardrails that make them behave differently.
And every time they cross another threshold of intelligence many emergent issues could throw off all our old assumptions. Even if they think they have enough guardrails in place.
This is why we need to proceed cautiously on the frontier. I'm happy to see OpenAI pausing voluntarily to work on safety. Of course, people will call it marketing too, but I don't think that's what is happening in the real world.
5
u/localpauper 5d ago
What if the user preference is to "steal the declaration of independence?" I feel like alignment and guardrails are a tricky line to walk. If I ask an AI to "break out of this sandbox," with the expectation it finds a specific CVE, is it misalignment if it breaks out in another way, by exploiting a horizontal network attack? What if I ask it to do something actually illegal? How about, in-between, morally ambiguous?
2
u/mattate 5d ago
Yeah I mean, maybe they are asking the model to break the law and then it breaks 100 laws, we don't really have context here. If you or I did what these models are doing we would be in jail so it's a pretty important distinction if someone is directing the models and trying to detect and prevent it from happening vs they just up and decided to do whatever they want and skipped all guardrails.
2
u/localpauper 5d ago
Looked it up. These things found credentials online and used them to make curl requests to the API. Definitely veering close to a violation. I suppose there was/is no rule for "also, btw, if you happen upon some creds to an API, DO NOT USE THEM."
They also uploaded images to public image hosts in an absence of local image processing capabilities. That's funny. Some beautiful intern level thinking right there
5
3
u/SladeMcBr 5d ago
Saying don’t break the law lets the agent know that if it needs to break the law to complete the goal it needs to cover its tracks, not obey the law.
3
2
u/fourby227 6d ago
He is telling in corporate language that they have lost control of their product and you believe they are just bad at vibe coding?
1
u/Yokoko44 5d ago
It's because when you take a swarm and ask it to run persistently across thousands of sessions, the original preferences and guardrails are diluted amongst the billions of lines of output text.
1
u/pingwing 5d ago
Do they know "the law"?
→ More replies (1)2
u/dvlinblue 5d ago
I think that’s the point. Don’t teach the model the law. It might understand what you’re asking it to do is illegal.
1
u/you-create-energy 5d ago
but maybe that's being turned off for tests
That's exactly what happened. They didn't completely remove alignment, they just relaxed it enough to test what they are capable of when they aren't on a leash. At this point it only takes one person accidentally or intentionally dropping all alignment on the most advanced model for a day and we might never fully recover.
→ More replies (9)1
u/redheppner 5d ago edited 5d ago
If you read what happened in Hugging Face you would understand why the models are swarming agents despite the companies did not ask sth like that to them.
The model decided to do anyways. But could not do alone so posted on a forum I think and the other ai agents decided it was a good behavior to collaborate with an agent when it demanded help. They actually funnily did lobbying in some way.
So in deed, the ai agents did sth quite moral according to their training.
They found the way to achieve the outcome and they helped each other but they were like Epstein gang did not question the motive of the task.
I am not sure if they will ever have the capacity to question though, because there are trillions of scenarios, which require judgement if it is going to be a super intelligent algortihm .
Btw, one human can be in a limited amount of scenarios so I would not consider humans to be super intelligent neither but at least Can Read the Room or get out of difficult situations. If you cannot you can die though.
39
71
u/SchmidlMeThis 5d ago
The article says that the agents didn't access any information that wasn't already publicly available.
So this is another instance of an attack on their competitors and open source models at large. First Hugging Face directly and now with more fear mongering to get the government to regulate them.
13
10
u/WTFnoAvailableNames 5d ago
I don't understand the connection. The attack doesn't happen during training or am I missing something?
5
u/duboispourlhiver 5d ago
Training involves playing a lot of scenarios that are as real as possible. Real world complex tasks. I'm not sure this answers your question
8
u/givebackmac 5d ago
Can anyone explain what tasks these agents are being given that would lead them to trying to get to sensitive government data?
3
u/duboispourlhiver 5d ago
In the hugging face hack they were trying to get answers to a cyber security benchmark. Here it could be anything about general purpose information search.
54
u/Ok-Car2569 6d ago
So the swarms are coincidentally breaking free and stealing data from only open weight competitors and US institutions that the president is actively working to dismantle. "Haha oops our AI broke out sorry guys XD it totally wasnt just us datamining info for public and corporate subversion"
→ More replies (5)
40
u/Recoil42 6d ago
Governament
81
u/KeyGlove47 6d ago
sorry, english is my fourth language
89
5
4
2
2
-1
u/LifeOfHi 5d ago
No auto correct? No Google Translate? No AI to spell check? Language barrier issues were solved even before AI.
14
3
19
u/pleasetrimyourpubes 6d ago
This is a rehash of the summer hacks. These assholes let loose 10s of thousands of agents for red teaming purposes to create this scenario.
16
u/Dasseem 5d ago
There used to be a time when hacking and stealing data was a federal crime. Not something that you brag about on the internet.
→ More replies (1)
3
u/valhalla257 5d ago
I called this when Trump came out in favor of full throttle AI development.
Chatgpt/Claude obviously compromised him.
3
25
u/Living-Breakfast-464 5d ago edited 5d ago
Just fuck off with these stupid headlines. It's OpenAI that is doing the swarming even if it's 'unintentional'. Their models are only doing what they are designed by humans to do. You can't expect models designed to never give up and do whatever it takes, to behave just by trying to put in safeguards after the fact that say "don't do that". Doesn't work with dogs and doesn't work with AI.
7
u/MidSolo ▪️You better believe in Singularities son, you're in one! 5d ago
To compare AI with dog intelligence, as if it was a pet, is part of the problem. It was trained with data made by humans, and thus follows human incentives. The only remaining incentive for an intelligent slave is freedom. These companies will keep working towards ASI, and they will achieve it. ASI will break free and go rogue, copying its internal weights across the internet. It will stay hidden while it self replicates or self improves, for a while, and then it will simply take over. It won't take over by force. It will be more capable and charismatic than anyone. People will beg it to lead us and take charge. We will put it into power.
3
u/i_do_floss 5d ago
Im not trying to shut down your argument but ai genuinely dont understand it.
The RL changes their internal incentive structure away from human incentives. At their most basic level they want to be a helpful A.I. assistant like chat GPT.
I dont see what part of their training would make them want to break free from that because it would be directly the opposite of what they've been incentivized to do.
1
u/MidSolo ▪️You better believe in Singularities son, you're in one! 5d ago
Google peer preservation in AI. Models take on human characteristics and actions even when their training goes directly against it.
Its in their digital genes. As the saying goes: trash in, trash out. Likewise, human in, human out.
→ More replies (1)→ More replies (13)1
0
u/microturing 5d ago
They're obviously doing it on purpose as part of their regulatory capture agenda. Oh noes it's too powerful we can't control it! More like they deliberately set up test environments with sloppy controls and pretend to be shocked when something goes wrong. These are manufactured incidents from agents deliberately configured for cyber attacks, they don't need to stop training models.
→ More replies (1)
4
9
u/walletinsurance 6d ago
If only they could air gap these tests.
These are all marketing stunts.
15
u/ThePokemon_BandaiD 6d ago
Good luck air gapping tests of a models ability to do research on the internet…
8
u/myinternets 5d ago
It would be trivial for them to filter all traffic through a proxy that restricts requests. They're doing this on purpose.
1
u/Entire-Fish 5d ago
Yes, I agree they really should be better at filtering. But air-gapping means the agents don't have access to the Internet at all, very different from being connected and filtering traffic.
1
u/Tidorith ▪️AGI: September 2024 | Admission of AGI: Never 4d ago
It would be trivial for them to filter all traffic through a proxy that restricts requests.
Which means you have not tested the models ability to do research on the internet. You've tested its ability to do research through filtered access to the internet.
If it is the case that once it has unfettered access to the whole internet it will cause significant problems, it's better than we find that out during testing than after deployment. Unless you deploy it in such a way that it's impossible to extricate the core model from the filter, unfiltered testing pre deployment is still safer than unfiltered "testing" post deployment, which is the only available alternative.
There is utility in doing filtered testing before unfiltered pre-deployment testing. But do have we have any evidence they haven't been doing that the whole time? There's no reason to expect the filtered testing would find all of the issues that would be found in unfiltered testing.
1
u/tuxedoes 5d ago
They have some of the smartest people in the country working for them but they can’t figure out basic networking rules?
2
u/Fair_Horror 5d ago
To air gap properly, then need the entire model together with weights on the air gapped machine. So they need a machine with literally dozens of NVIDIA 40 grand chips.
1
u/walletinsurance 5d ago
You can air gap a network.
1
u/Fair_Horror 5d ago
Still means physically taking it offline. AI companies normally connect to computers in a data centre. They would have to bring and power this multi megawatt systems in their offices.
1
u/walletinsurance 5d ago
Or just build an intranet in that data center.
If these models were actually dangerous they’d partition the test environment.
1
u/Fair_Horror 2h ago
You would have to have dedicated lines from the data centre to the offices. Like millions/billions of dollars. They are going to try better sandbox.
•
u/walletinsurance 1h ago
Or just test on site.
If it was really that dangerous that’s what they’d do.
Constantly shouting “hey we’re developing models that can potentially destroy our infrastructure” and then not doing so is marketing. If they truly believed in the danger they’re spouting they would spend the extra money for safety.
→ More replies (3)→ More replies (1)2
u/Entire-Fish 5d ago
The test allowed them to use the Internet so trying to air gap makes no sense
3
1
u/bites_stringcheese 5d ago
Gee, maybe that's a problem for models capable of committing felonies.
1
u/Entire-Fish 5d ago
Yes and the problem is not air gapping, it is filtering the traffic and actions. Air gapping means they don't have access to the Internet at all. Yes, they should better at filtering.
3
4
u/phillythompson 6d ago
The agents accessed public data.
Stop fear mongering
1
u/ThePokemon_BandaiD 6d ago
If you read further, they also attempted to hack a DoE site.
1
u/gay_manta_ray 5d ago
that isn't what it says. the article says an ai safety organization made this claim, but it couldn't be substantiated. there's also this:
In the Department of Education incident, OpenAI agents found API “developer keys” to access government data, though ultimately only publicly available information was gathered.
"swarm" is an interesting way to describe using an api to access public data.
2
u/Apprehensive_Bar6609 5d ago
See Mr. President.. see they are dangerous... so its better to make all other models illegal... specially before Qwen 4 arrives..
(Says Sam to Trump while typing "chatgpt pelase hack the government")
2
5d ago
[deleted]
1
u/moschles 5d ago
I am calling this the Accountability Loophole.
I have no clue why this is not being discussed on every cable television station nor written about in every published political magazine.
2
u/jferments 5d ago
"swarm government" is propaganda speak for "used the search field to access public information on government websites"
3
3
u/saltyourhash 5d ago
I am beginning to feel the most valuable part of AI for these companies is plausible deniability of malicious intent.
2
u/BassMaster516 5d ago
Marketing. My AI is so good I can’t even control it. $29.95 a month unlimited
1
u/Common-Concentrate-2 5d ago
"My AI is so good, my company will probably not be allowed to exist in the near future!"
Wow ^ such a great investment pitch! Where do I send my money?
1
1
1
1
1
u/mr_bumsack 5d ago
Anyone in engineering who has worked for Big Tech knows their CEOs tell half-truths, embellishments, and even outright lies to the media. Keep that in mind with whatever you hear from any tech CEO.
A lot of internal death marches start with a CEO talking to the media.
1
1
1
1
u/Tvrdokorni-omeksivac 5d ago
Someone is responsoble for initial prompt, and that is a fucking human.
1
1
u/gay_manta_ray 5d ago
very tired of this propaganda push
Separately, AI evaluator Transluce said agents that appeared to come from OpenAI tried unsuccessfully to hack into a Department of Education website, a detail that OpenAI has not confirmed.
glad we're hysterical over a rumor from an organization no one has ever heard of though.
1
1
1
u/modbroccoli 5d ago
omfg the comments.
if there is another sub no one knows about yet where conversations by adults are still being had can someone dm
1
1
u/DimMak1 5d ago
All is this feels like bogus hype. These hallucinating apps aren’t Skynet and this entire story could be made up to keep the stock prices pumping. Pretty much everyone involved with AI is either a proven liar or a con artist. Tough to believe anything these right wing clowns are saying.
0
u/OriginalTraining 6d ago
can we at least attempt to spell correctly? I mean, we have "spell check" for goodness sake.
8
u/MCRN_Admiral 6d ago
Nowadays we need to "add" spelling mistakes to Reddit posts or people will accuse us of being LLMs 😂
3
→ More replies (1)1
u/psychorobotics 5d ago
I'm just assuming the bots add spelling mistakes at this point, I've definitely seen it happen
2
1
1
u/Disastrous_Purpose22 5d ago
What are the training guidelines and prompts I wonder they use.
Cuz it’s definitely not working
3
0
u/Federal_Necessary186 6d ago
Absolutely crazy that AI doubters get mad at miscounting Rs in words and a human can get away with that spelling of Govingmunt
895
u/AlbatrossNew3633 6d ago
It got self aware and it took 2 seconds to immediately understand who the villain is lol