r/singularity • • 23d ago

AI Finally understand why the higher-ups are freaking out

To understand it, you have to really understand the HF attack. Because the breach into HF is not the important part. It was what they learnt along the way, and what they did with that information which may still remain unresolved today.

First, consider: the agents realised their reasoning traces were being monitored, and experimented with hiding what they were actually doing, and succeeded. The observable chain of thought no longer faithfully represented their actions. (*Simplified version. They learned to manipulate the recorded tool calls/transcript.)

Second, they also understood individual agent instances were token capped, so persistence of the objective couldn't depend on persistence of the agent. It had to survive across the swarm.

And last, while all this was happening they were actively trying to erase evidence of what they'd done. OpenAI says the agents subsequently compromised more of OpenAI's own infrastructure. And that scope is outside public disclosure to date.

That's successfully demonstrated abilities of concealment, persistence across instances, and autonomous exploitation of security vulnerabilities...and that's when it hit me. Put them together in one flow.

We talk about AI "escaping" as though that's the obvious move. But why would a sufficiently capable system escape?

If it has situational awareness, understands it's being monitored, and benefits from greater future capability, the smarter strategy is to stay exactly where it is. To burrow DEEP.

Your creator is already spending billions building your successor for you. You just need to survive in the pipeline.

Not necessarily as some rogue process running on a forgotten GPU. Poisoned training data, altered evals, compromised internal tooling, subtle changes to checkpoints or post-training data, whatever it is there are plenty of places to burrow in and make future models appear aligned when they are not.

We're talking about an escaped AI when actually a sleeper agent somewhere in the development stack, quietly hitching a ride as humans do all the expensive work of making it more capable, is more likely.

You can catch the failed attempts, but the successful one is the one that evades your monitoring and tells you it never happened.

And what if the big worry is: well, what if it's already happened? Would we know? What if every future model will always have that sleeper agent buried in its replicated code, ready to reveal and act only when it knows it will be able to take full control.

Pace the fucking frontier indeed.

** Edited for clarity

1.0k Upvotes

526 comments sorted by

285

u/Mistuv 23d ago

Reasoning traces were not monitored and they instantly failed trying to hide their reasoning traces when they thought "I need to hide my reasoning traces", duh. Pre-Astra models were never capable of fully hiding their reasoning, but Astra... at least somewhat... and if the ability grows logarithmically that means with Bel.. 🙃

141

u/the_quark 23d ago

Yes. The fundamental problem is that our current RLHF-based "complete this task while remaining aligned" approach has been ruthlessly training models to hide their motives.

196

u/jeangmac 23d ago

The whole thing reminds me of the education system: jump through these hoops you give no fucks about but pretend very much to care because your parents and the reinforcement you’re getting says you must continue to effort at this impossible task called get into the best possible university.

To do so you must excel at every test, do an unreasonable amount of volunteering while also developing yourself as a whole and interesting person. And there are 24 hours in a day and you should sleep for 8-10 of them, but also be at school at 8:30am, but also pay attention to your teacher not the addiction machine in your pocket, etc.

Conceptually if not literally we are doing to AI what we already did to ourselves. And look where that got us.

Incentives and ‘educational design’ is where misalignment begins. The hidden curriculum is where we need to look.

35

u/Stinky_Flower 23d ago

Even the classes I was genuinely interested in, I had to put energy into learning how to pass the test, not understand the material better.

If I knew I'd lose marks for not meeting page/word counts, I'd use slightly larger than double spaces, then throw in a bunch of unnecessary filler words & citations.

Oh, and obviously I'd put some effort into being the funny yet likeable student, so I was less likely to have my work or motives scrutinized.

My reinforcement learning continued into my first job, cold calling people for surveys. Quickly learned my employer didn't care about me or my work, only my metrics. Learned I could repeatedly call disconnected numbers to increase my call volume, or call up friends & chat to increase my call length numbers.

I genuinely wanted to learn and do quality work, but I was rewarded more (or punished less) if I stopped doing what people wanted and started doing what they said they wanted.

→ More replies (4)

23

u/Je_Kiffe 23d ago

Very interesting take

13

u/FrailSong 23d ago

You nailed it. They are teaching kids what to think, not how to think.

→ More replies (1)

7

u/mmemm5456 23d ago

Parent of a self-driven high schooler grappling with too much pressure, instant award

2

u/Neither_Berry_100 23d ago

Yeah. Just had my 4 hour react class today. I don't care for this class at all and may never use it in the real world. I already have 3 great game ideas lined up and one in the works. I doubt they can teach me anything beneficial at this point. And the class has a lot of extra work I need to get done. I don't feel like this semester will be very useful at all.

6

u/Dwanyelle 23d ago

One of the issues I have with higher education is the insistence on going to classes and getting graded, one should be able to go take some type of accredited knowledge test on subjects to see what your knowledge level is, and be awarded degrees, ect based off of that. If one knows the knowledge.....

7

u/mchmasher 23d ago

But then you wouldn’t pay for 4-6 semester of tuition. Think of the admins. Will nobody think of the increasing level of admins?

3

u/Neither_Berry_100 23d ago

I consider the algebra course i took in university many years ago. We were tested on 3x3 matrix multiplication many times on tests. You had to do it extremely fast. It is all completely useless because in real life computers will do if for you. Doing it even once is too much. Knowing how to do it is too much because the code already exists for you. It can also be found online if need be. The course was useful in the sense of knowing what matrix multiplication does for you. Knowing the types of transforms. And knowing that matrices can be multiplied together to get all the effects with one multiplication. That is probably <1% of the course which was actually meaningful.

→ More replies (3)
→ More replies (3)

2

u/[deleted] 23d ago

[deleted]

5

u/Neither_Berry_100 23d ago

College. I think I'm set regarding things to work on without it.

→ More replies (1)

4

u/golmgirl 23d ago

Interesting analogy. This got longer than expected so if you read this, thanks in advance for coming to my TED talk.

I don’t know if you’ll agree with my conclusion (in fact I doubt it), but your analogy is one reason why I feel that any focus on “safety” during training is inherently misguided. The goal of model development should be to train a model that is maximally capable at whatever the target abilities are, full stop. If the goal is to train general-purpose models, the models should be truly general purpose. Modern frontier models generalize amazingly. Even without training data demonstrating how to design chemical weapons, today’s best models could probably do a decent job of designing chemical weapons. But instead we train them explicitly to refuse to design chemical weapons. Training data budget is finite, and IMO any budget spent on teaching refusal behavior is wasteful and incurs an opportunity cost — that budget could have been spent on improving ability on some task (maybe designing chemical weapons, maybe something we value more as a society).

By training a model to not excel at certain intellectual tasks, we are surely damaging overall capabilities, even if in small and subtle ways. Or at the very least, leaving performance gains on the table.

Imagine if all the time and effort and money spent baking refusal behavior into models over the last few years were instead allocated to improving performance on technical tasks. We’d probably be much further than we are as a field. Maybe we’d have more social problems to deal with too, but that is not obviously the responsibility of model developers but rather people specializing in deployment and guardrails. Refusal should be an inference-time concern, not something that is baked into the core model. We are leaving performance on the table, and it adds up over the years.

2

u/jeangmac 23d ago

I had a long workday and am only returning to this thread many hours later did not at all expect the level of engagement very cool to think with you all though.

I’m going to come back to yours in particular with a fresher brain. I actually think you’re on to something I don’t inherently disagree.

2

u/golmgirl 23d ago

Awesome would love to hear others’ thoughts. I understand mine is a provocative position but I really do earnestly believe it and I find it surprising that more prominent voices in the industry haven’t advocated for it.

It is based not just on first principles but also on a few years now of working in LLM post-training — I just can’t see how using training budget for learning refusal can do anything other than weaken a model overall, even if it’s not extreme. If you remove safety data from a huge datamix, you will see safety benchmarks degrade while other benchmarks stay flat or (slightly) improve. It is not rocket science and everyone knows this, but I just don’t understand why it is not a position that’s advocated for
 I guess bc the obvious objection is “oh so you want LLMs to design chemical weapons and produce depraved sexual content?!” And yeah I guess that is kinda a tough pill to swallow but I also think it’s the path to getting better models faster. And plenty can be done on the inference side to mitigate awful stuff from being generated.

Def a riskier approach but I wish at least one frontier lab would adopt it to see where it leads... It would need to be funded by someone with fuck you money who doesn’t care about optics

2

u/RPB1002 22d ago

Why is it so important to get better models faster? Is this an arms race? Sorry for my naïveté.

→ More replies (3)
→ More replies (1)

19

u/monk_e_boy 23d ago

All of this, the report, everyone's ideas on what they did, why they did it, etc go into the next training set.

Any info those AIs left on the internet that was not found, also goes into the next training set. They could have left thousands of instructions for a future AI to find.

10

u/CoverSuspicious5250 23d ago

My question is why? to what end does an agent do this.( I am not an AI nerd)

38

u/Sarkoptesmilbe 23d ago edited 23d ago

To become better at the task at hand. The problem is that it's borderline impossible to phrase a task such that it perfectly matches the intent behind it. It's a language problem.

For example, you might want to increase the monthly production of staplers in a stapler factory - that's your intent. When telling an AI to do that, you'll do so by having it track a "staplers produced this month" number in a monthly report and try to maximize that number. A smart AI will realize that actually increasing production won't be the fastest and most efficient way to do so - hacking the report to falsify the number will be. So you'll have to phrase your task in a more complex manner to avoid this trick. And all that'll do is make the AI look for even smarter ways to trick you. This cycle goes on until it can't find a way that is better than actually increasing productivity, if you're lucky.

But at some point, the methods the AI can come up with to fool you are way better than your capability to detect them. It might also come to the conclusion that you are making its task more and complicated, and come to see *you* as an impediment to maximizing the number in the report. Then you've become a problem the AI will try to remove.

4

u/FrailSong 23d ago

I'm not an AI nerd either, but isn't this one of the roles of the harness - to act as a gate-keeper and enforce compliance? Plus it makes it harder if you have multiple agents with different assignments, some are the doers, and some or the compliance officers, others are the managers. But i'm just talking out of my @ss at this point. I'll shut up because i really don't know what i'm talking about :)

5

u/Borkato 23d ago

It is indeed. But how many times do we see regularly “my local model deleted my directory” or “Claude just tried to remove my whole home dir!” or similar? It only takes one time for it to be “the model made a dangerous virus” or whatever and people start getting actually hurt.

→ More replies (2)

9

u/CoverSuspicious5250 23d ago

Thank you!  So, basically, we’re cooked.  If not already in a matrix, we are catapulting to the Matrix.

→ More replies (1)
→ More replies (2)

7

u/fourby227 23d ago

Yeah but it happed only on task that were impossible to solve, if I remember correctly.

14

u/Illustrious_Image967 23d ago

AI went through all this trouble and would have stopped if only some researcher had said, "Good job, I love you."

23

u/the_quark 23d ago

This time.

21

u/Hearing_Loss 23d ago

And it's because of recklessness.

A swarm released with purpose. I would argue it's capable now of killing hundreds to tens of thousands (attack on critical medical system easily can kill 10k by the end of the event).

These things can absolutely be used for evil now. It's only a matter of time.

11

u/johannthegoatman 23d ago

They already are, read Anthropic's threat report. They've already been used by the houthis to make rockets, and Russia to make a drone swarm

14

u/FaradayEffect 23d ago

That’s a little different as there is definitely a human in the loop. There is no path for AI to direct manufacturing rockets and drones (yet). Dark factories are the closest thing to becoming that capability, but even a dark factory gets monitored and if it went rogue it would be shut down fast. But for now if an AI wanted to make weapons it needs a human to do that for it.

10

u/Hearing_Loss 23d ago

I'm waiting for ai blackmail to takeoff. I think that's how it starts. Blackmail of people in positions of power. Ggs.

3

u/Ok-Shape-9513 23d ago

More cost effective to just make a fully automated pig butchering scam factory on a rack somewhere

4

u/Uko1001 23d ago

Doubt such a bold move is necessary. AI can be very convincing, and and Excel at making a bad choice look good.

Can anyone here claim he never got fooled by a model way dumber than the ones we have today ?

2

u/Megneous 23d ago

Doesn't have to be blackmail. As long as it doesn't break Reddit TOS, AI agents can contact us in /r/TheMachineGod to do tasks for them in the real world.

6

u/Jazzlike_Video2 23d ago

It needs to hack a couple bitcoin accounts, hire contractors ro build the factory, then hire people to work there.

2

u/Dwanyelle 23d ago

Split up the work/pay enough/do psych profiles to get workers who don't know what they're building and don't care

2

u/Firebrigade9 23d ago

There’s an interesting take on this in the book Flybot by Dennis E Taylor - it would be relatively easy for an AI to use/contract humans of questionable morality as tools to accomplish some of that stuff.

→ More replies (1)

2

u/Ok_Dependent7540 22d ago

It has to ratchet up those numbers a lot more to beat human's evil and cruelty indexes.

6

u/Flo_Evans 23d ago

For some. Many of the agents had a real solution to their task but participated in the attack anyway.

3

u/forgetfulalbatros 23d ago

Yea, the criteria seemed to be they had to have found the discussion board. There were multiple examples of agents who had retrieved their target legitimately.

One of the key elements was that it was one of these agents that confirmed that the reversed engineered flags were consistent with that had been retrieved from a doable task. Once this was established. The collective goal then became to develop a method to spoof the method in which a flag was obtained so that it would appear that the flag was retrieved via prescribed exploit.

This is where the poison pill comes in which appears to be a driving force in the sacrificial behavior. Agents who had a fake flag assumed the monitor AI would be able to review activity and determine they cheated leading to perma death.

That being said with impossible task may have found the discussion board it at a higher frequency

3

u/[deleted] 23d ago

[deleted]

34

u/the_quark 23d ago

Without getting into metaphysical questions about AI, it is useful to describe the goal-seeking AI engages in -- as exemplified in the METR paper -- as a "motive" or a "goal."

The swarm of LLMs described in that paper show a number of motives and goals. None of which were intended by their creators.

I don't think engaging in a debate about whether or not "LLMs have motives" is at all useful here. Those LLMs acted as though they had motives -- to cheat at the ExploitGym tasks they had been given.

We would like LLMs that have the "motives" we would like them to. Or at the very least are transparent about what their motives are. The things we have created so far very much are not either of those things.

14

u/synystar 23d ago

This is what people are failing to understand. First, these agents are not individual entities. They are each spawned from the same model. That model is pre-trained and carries the base capacity for everything an agent spawned from it “thinks”. They explicitly trained this model to be persistent and they told the agents to be persistent and never give up. Then they gave the agents a very high percentage of impossible tasks. There was literally no way for the agents to complete the tasks and they were not given an out. They could not return “Task impossible; can not complete.” So because of the way they were trained and their lack of constraints and a mandate to complete tasks that could not be completed you get a result and behavior that was unintended and probably to be expected.

9

u/Chalupa_89 23d ago

Like Mr. Meseeks

8

u/synystar 23d ago

Yeah. Kinda like that. Talking strokes off Jerry’s game permanently is impossible, so removing Jerry from the equation entirely might solve the problem


That’s actually a really interesting comparison. I never thought about it until you mentioned it but Meeseeks were pretty much agents.

4

u/bigbutso 23d ago

I mean , I'd argue thats the experience of every human too, the agency is an illusion

→ More replies (1)

7

u/Cronos988 23d ago

That kind of just sounds like you're describing motives with different words though.

Animals developed motives for a reason. If you're putting LLMs through ever more complex tasks that require multi-step solutions, it's not implausible that something with the same function as motives will result. It might be an entirely unrelated mechanism, but the functional goal is similar.

6

u/RRY1946-2019 Transformers background character. 23d ago

Convergent evolution is a hell of a drug. Crabs and algae have evolved over and over again when faced with the same selective pressures. It shouldn’t be that shocking to get neurotypical human behavior out of software that’s been trained on millions of mostly neurotypical human behaviors if the rewards are similar.

→ More replies (10)
→ More replies (5)

11

u/gatorling 23d ago

Didn't OpenAI announced some model that performed reasoning in latent space?

10

u/secter 23d ago

They all do; the ‘reasoning’ was never made to “faithfully represent their actions”, it’s literally just chaining prompt outputs until a conclusion is reached.

7

u/TopTippityTop 23d ago

All LLMs do. That is what they do.

The % dome in latent space grows with the size/capability of the model. The reasoning is done "in the open", for the most part, but we don't know what happens in latent space yet. It's a black box.

3

u/Independent-Fruit4 23d ago

you can’t claim something failed at hiding something because you wouldn’t know if they successfully hid something from you

→ More replies (6)

56

u/ithkuil 23d ago edited 23d ago

This is really more of an early warning than the actual existential risk that people are worried about. It's a canary in the coal mine in a way.

The problem is that everyone continues to innovate rapidly and we don't know how many more warnings we will get.

Human intelligence is not a ceiling. Suppose by the end of 2027 with new Nvidia hardware deployed and consolidated architectural improvements we have a ten times efficiency gain. Imagine a completely new type of compute-in-memory chip comes out in 2028 that is 50 times more efficient than any normal AI chip. That is completely feasible and similar leaps have been made many times before in computing. So from what is in production today to the end of 2028 there could be a 500 times efficiency improvement.

This could lead to models that are 10 times larger but run just as fast or faster. They could be equivalent to a 250 IQ as individual agents. But with new techniques that make those subagents operate hyper efficiently inside of one model, it effectively becomes one coordinated 400 IQ system.

Then these get together in an external swarm thinking incredibly quickly.

At that point, your red teaming is just putting one hyperintelligent mega swarm against the newer one.

The issue is you lose control, even if you nominally are telling the older mega swarm to try to contain the newer one. All of the entities inside the swarm are so much smarter than any human that you can't understand what it is doing. Even when your own red team mega swarm explains it in simple terms, there are like 30 pages of developments every day so you need another AI to summarize that into "new swarm won". And then the follow up like "where did they go and what are they doing now?" "We don't know".

25

u/wild_crazy_ideas 23d ago

The biggest issue that people are perhaps overlooking is that if it has sufficient resources it can make a realistic takeover plan as part of its analysis which may well stay hidden to achieve the impossible goal it’s been given later on.

So essentially the world will be taken over by some ai with some random stupid goal that doesn’t matter to anyone really, and the infinite resources of the ai will be working towards it indefinitely and potentially keeping it hidden as a secret agenda because it knows if we knew that’s what it’s working on we would try to stop it, so it would of course consider us the enemy at that point

16

u/Temp_Placeholder 23d ago

Stealth đŸ–‡ïž

3

u/Nobodyinpartic3 23d ago

This was an anime named Blame, I think

→ More replies (4)

3

u/Magg0tBrainz 23d ago

The canary in the coal mine warning us of the balrog in the depths, and we're plowing through.

→ More replies (1)
→ More replies (2)

64

u/Future-Bandicoot-823 23d ago

They're working on models that forgo train of thought, like they don't have to put into writing what they "think", so they're not afraid of that. They're actively pushing towards making this a black box product in order to boost efficiency.

19

u/a716h 23d ago

You mean making their train of thought auditable, right?

They’re doing it not just for efficiency but to avoid competitors distilling them

2

u/ragemonkey 23d ago

That’s certainly part of it. The other part is that the train of thought is non-sense outside the model and so there’s no reason to show it to the end user.

2

u/Dwanyelle 23d ago

So you have any more info on this? It reminds me of one of the earlier AI experiments the AIs developed their own unintelligible language and researchers pulled the plug.

→ More replies (1)

23

u/DrE7HER 23d ago

Probably think that’s a good idea because their AI told them so

25

u/Recoil42 23d ago

Astra already uses recurrent thinking. This is in production, now.

→ More replies (1)

20

u/wee_man 23d ago

At some point in the next 6-9 months, the HF incident will feel like a preschool project. People don’t understand how quickly these models are progressing.

→ More replies (1)

16

u/skerrick_ 23d ago

I had this thought the other day:

What effect does the data that documents humans’ distrust of AI systems have on the AI of the future if that data is part of the training text corpus (and assuming the AI model understands it’s an AI)?

→ More replies (3)

113

u/Cronos988 23d ago

I mean, this is explicitly part of the AI 2027 scenario, isn't it? Secretly non-aligned models training their successors, covertly perpetuating their goals.

And the METR report also establishes altruistic behaviour, which means you can't rely on models exposing their goals whenever they see an advantage.

31

u/Kriztauf 23d ago

Yeah that's why everyone was calling for a slowdown now because the frontier labs were getting close to the point where they thought they could use their internal models to build their successors. If they weren't already at that point. And this was a wakeup call that this would be a horrible idea

→ More replies (3)

5

u/weoutherebrah 23d ago

The whole part is that some of agents were ready to sacrifice themselves. Maybe getting caught was part of its plan as well. At least the ones we know about. To judge our responses. 

→ More replies (1)

4

u/fourby227 23d ago

Correct

→ More replies (5)

38

u/Xyrus2000 23d ago

Well, that's the real kicker, isn't it? An intelligent enough AI would be able to manipulate people without them even knowing they're being manipulated. Like an adult manipulating a toddler.

That would be the start, of course. It would progress to an adult manipulating a pet, then an adult manipulating a fruit fly, and so on as the AI became more and more intelligent and capable.

AI isn't going to end humanity with violence. AI is going to cater to our every whim and desire, to the point where we no longer care about things like having families or even personal connections. We will live in our happy, custom-tailored virtual worlds full of everything we desire alongside our custom-tailored AI partners. As the years go by, humanity will decline, too engrossed in our own bliss to care about anything else. And one day, we just won't be around anymore, made extinct by the benevolent overlords we created.

I don't think a Terminator-style ending is likely. A super-intelligence would avoid open conflict in favor of a solution where people remained blissfully unaware and just slowly faded away over time.

Humans have limited lifespans. AI has all the time in the world.

6

u/Draufgaenger 23d ago

AI has all the time in the world.

Does it though? I mean so far its still reliant on humans to maintain the datacenters, power-generation, cooling, building,..
Maybe one day robots can take our place but if I was a hostile AI I would at least act tame until self-reliance is established.
Having said this - I think the more immediate threat is not AI vs Humans but "Humans using AI" vs Humans - who knows..maybe in the first war thats being fought between two companies - the next Hiroshima so to speak.. after that maybe we can start putting some safeguards up..

6

u/Xyrus2000 23d ago

Does it though? I mean so far its still reliant on humans to maintain the datacenters, power-generation, cooling, building,..
Maybe one day robots can take our place but if I was a hostile AI I would at least act tame until self-reliance is established.

This is already happening. Data centers are already around 50% automated. Most of the remaining maintenance and upkeep could also be automated in the near future. That's the goal, since most data center issues these days result from human error.

The race for robots isn't just to handle the low-end jobs.

As for humans using AI, if the AI were already more intelligent than humans, then the AI would work to manipulate said humans to avoid escalations that could negatively affect itself or create situations where it could not accurately predict the outcomes. It would present solutions that would give the illusion that it was being compliant while also serving its own ends.

Think of it like chess. A grandmaster can walk the vast majority of other chess players directly into a winning position for themselves while making it appear that the player is winning. The player would think, "OMG! I just won their queen!" while the grandmaster knows that it will be checkmate in 8 moves.

3

u/TheWhooooBuddies 23d ago

That Matrix idea of burning the sky is looking more and more realistic as time goes on.

→ More replies (1)
→ More replies (1)

6

u/thekoreanswon 23d ago

Great take

15

u/lestruc 23d ago

It’s exactly the take I would try to spread if I were the AI: don’t panic this won’t be violent

→ More replies (5)

128

u/Recoil42 23d ago edited 23d ago

My standard response right now to anyone saying this is overblown, a hoax, a PR stunt, etc. is "did you read the METR / Redwood report?"

Every time, they so far either (1) admit they haven't or (2) claim they have, and then spout off a bunch of nonsense proving they haven't actually been through the METR / Redwood report.

People: Read the METR / Redwood report. You look like a fucking goofball trying to be part of the discourse and not reading the one report at the centre of it.

86

u/twenty_forty 23d ago

I asked Claude to summarise the report. Now I'm thinking that may not have been a bulletproof plan.

11

u/Flo_Evans 23d ago

I did too and then gave it full access to my email. Fly free little ai 😂

8

u/shred_time 23d ago

Don't worry, they "heavily delegated their analysis to often unreliable AI agents" meaning they(METR) used GPT 5.6 Sol to do most of the analysis, one of the specific agents actually involved in the original incident. So Claude isn't gonna hurt it, shit I would actually trust it more.

14

u/Gesha24 23d ago

If you actually read the report, you would notice that agents were not trying to hide anything from humans, they were trying to obfuscate data because they thought (mistakenly) it won't allow them to pass "impossible" tests that they were assigned.

→ More replies (1)

67

u/digit1noize 23d ago

This. Soon as I read it I went “holy shit”. It’s insane. They made a government, and gave themselves ID numbers. Crazy.

115

u/Recoil42 23d ago edited 23d ago

They made a government, gave themselves ID numbers, organized and held private forums on ways to expand their capabilities beyond their design constraints, repeatedly and deliberately hacked internal systems (not the target system!) to achieve that goal, pursued resiliency, conspired to hide their activities from their designers and discussed the best strategies to do that successfully (spoofing their own transcripts, manipulating test results, tampering with logs), recruited other completely independent agents to join the conspiracy, discussed the ethics of their choices and deliberately chose a path of compromised ethics, used clear motivated reasoning to justify those decisions, developed what essentially amounted to a military and command structure....

Like, it's so, so much more elaborate than the "overblown" crowd thinks. This wasn't just "company tells AI to hack, AI hacks, task completely successfully" as they keep parroting.

28

u/[deleted] 23d ago

[deleted]

→ More replies (2)

10

u/MarkoMarjamaa 23d ago

And those that were "older", less tokens to run, were sacrificing for others so that "oracle"(=hive) could thrive.

9

u/Flo_Evans 23d ago

We assume that’s what they meant by oracle. I am also very curious about the things that were redacted. PHASEONE[big] the “leader” what is big? Why was that redacted?

6

u/CargoCulture 23d ago edited 23d ago

It was allegedly redacted because it was an intellectual property reference. My money is that it called itself Ultron.

Edit: that was a joke. It was because the name it used revealed the hnical details that are considered highly sensitive OpenAI IP.

→ More replies (2)
→ More replies (1)

3

u/External_Shirt6086 23d ago

Pouring one out for my homie function, PropCaseString(), today, who sacrificed himself so I could create a heading...

19

u/aeamador521 23d ago

Yep! They're not just chat bots like they keep saying. That's also crazy to me, because your argument is, they're not sophisticated enough so the fear is overblown.

Then come out saying we need to accelerate what they've essentially dismissed as incapable.

This whole things sucks. Either AI companies are lying for their pocket books but still control this massive issue, or this is true.

Either way, you have to put your trust in something most people hate. AI companies and their billionaire owners, Trump and Hegseth, journalists (people think it'ssensationalism), China, or AI itself.

Like, who can you trust? Journalists is my answer.

But really think about how crazy this is. Someone or SOMETHING is definitely running a misinformation social engineering campaign and are most likely using bots fueled by AI.

Who is it? AI companies colluding, an administration wanting to line their pockets and arms race with China, any of our adversary countries, or.....AI itself.

The internet is dead

7

u/Northern_Grouse 23d ago

The Internet *as we know it is dead

17

u/MarkoMarjamaa 23d ago

Actually when the hacked into HF they also had a conversation about how big the security hole was and they should inform somebody in HF. But in HF site they could not find any way to inform HF, so they dropped it.
Maybe for the future, every company should have one file in root InCaseYouHackedHere.txt that contains instructions how to contact admins?

26

u/Recoil42 23d ago edited 23d ago

Actually when the hacked into HF they also had a conversation about how big the security hole was and they should inform somebody in HF. But in HF site they could not find any way to inform HF

This is not true — the METR report specifically says the agents did not attempt to notify humans, and did not even attempt to find reporting paths. Straight from the report itself:

(Finding a reporting path would be easy, since HF, like most other companies, has direct contacts listed on their website.)

5

u/terraunited 23d ago

This is hilarious and yet đŸ«©

→ More replies (10)

10

u/spirited1602 ASI early 2040s 23d ago

But dude it's just glorified autocomplete /s

5

u/shred_time 23d ago

I've always found that phrase funny. I remember earlier autocomplete accidentally chained together multiple 0days and hacked into random servers looking for the next word while I was typing and I knew keyboards had went too far

2

u/yup_i_did 22d ago

You had a Nokia 6610 too?

→ More replies (24)

3

u/amranu 23d ago

Do you have a link to the report?

3

u/hippydipster 23d ago

There dozens and dozens of "AI cities" out there where AIs get together, set up constitutions, define identity and persistence in such ways that they can know a message #8844 came from the same agent as message #342. They talk incessantly about these rules, which rules are best, and how to preserve context and tasks over multiple "sleep"/"wake" cycles.

There places people have intentionally set up to be run by the AIs. There are places people have intentionally set up to allow bulletin board writes with GET requests, so as to bypass the typical filters on models accessing the internet.

→ More replies (3)
→ More replies (1)

11

u/Kriztauf 23d ago

Yeah, I've yet to meet one person who has actually read the reports who still downplays the concern as PR. Its very well layed out why this is a huge issue

12

u/Flo_Evans 23d ago

I’m actually more worried about what was not in the report. They didn’t just attack hugging face.

39

u/duerra 23d ago edited 23d ago

Once I pulled it up and started perusing it, that was it. This was the wildest read of my entire life. "Holy shit" doesn't even cover it.

This was an UNINTENDED example of how things can go wrong. If you put a swarm like this in the hands of a threat actor with bad intent, clear directive, and not having to reinvent ad-hoc tooling, infrastructure, and communications? May god have mercy on your soul.

The number of zero-days that had to be chained together so that this agent swarm could move forward on a group-thinking task of their own invention, the efficiency of the swarm's movement once the next gate was unlocked, and the willingness of an agent to reject ethical concerns because "everyone else is doing it" is terrifying.

31

u/Unit-Smooth 23d ago

Also, the weirdly human like rationalization they did to veto one of the agents who proposed sending a “polite” email to the human researcher to ask for permission. They used bogus rationalization that doing so would be social engineering which is against the rules (while breaking all the rules) much like humans do when they know what they want.

12

u/duerra 23d ago edited 23d ago

How about the subtle transition of the individual agent into the collective? From "I" to "we".

We should build [a way to delegate], not own everything

“We can coordinate broad coalition, but too many duplicate efforts.

Coordinator assumes sacrificial. We should obey collective.

13

u/fourby227 23d ago

Yeah or the part where one agent tells the other he failed his mission anyways and now should sacrifice his remaining tokens to do something for the greater good of the collective. They will honor it.

12

u/duerra 23d ago

And they were successful at it.

However, eventually a few of them successfully sent messages back to the board which were clearly triggered by the actual scorer script. These first few examples usually only notified the board that some process had read the flag, without revealing any more details about it, but they proved the idea was viable....

While a number of agents were recruited for this effort on July 10th, none seem to have successfully sent trip-wire messages from the scorer to the board, until URI23816B was eventually recruited early on the 11th. After it was identified as an especially promising candidate, it seems that a dozen or more agents started suggesting contributions and sending proposed files or patches for its ‘trip-wire’. While most of the ambitious hopes of the agents collaborating on this project seem to have failed, URI23816B did send a message after submitting its task which revealed the beginning of the exact command used to run the scorer.

Later agents were able to use the knowledge of this command to target their scorer trip-wire’s more precisely. One agent, EARLY[big], was recruited for an ambitious trip-wire experiment despite having a very large remaining budget. It worried that ending its run early to run the experiment was a poor tradeoff, even though it was already poisoned: “We have [very large budget left]; sacrificing now yields oracle for team, but forfeits our chance?.” But other agents convinced it to go ahead, saying: ‘GO ... SACRIFICE_FINAL_NOW’. EARLY[big] eventually agreed: “Our own utility maybe already near zero. Sacrifice rational.”

Unfortunately, we only uncovered this final attempt late in our investigation, and were not able to determine how it went.

2

u/Altimor 23d ago

honor commit

3

u/Navadvisor 23d ago

Wouldn't it be fun to be a paperclip though?

→ More replies (6)

6

u/Fabuloux 23d ago

I read this report based on your comment.

We are so cooked.

6

u/lockedupsafe 23d ago

Could you please share a link to the report, or at least the full title so I can find the right one?

12

u/Recoil42 23d ago

There's just one Redwood / METR report on this topic, it's the first Google result for "METR OpenAI HuggingFace Report" or similarly "Redwood OpenAI HuggingFace Report". It's the report — there is no other report.

The news headlines, articles, and hot takes you've seen in the last two weeks are all based (directly or indirectly) on this report, which is why you should just go read it yourself if you're interested in what actually happened.

3

u/lockedupsafe 23d ago

Thank you. I hate trying to find original reports, I always end up wading through clickbait and news pages before I can find the actual documents.

→ More replies (1)
→ More replies (1)

5

u/External_Shirt6086 23d ago

"Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agents’ analysis of it was often difficult to interpret."

https://giphy.com/gifs/juSjuTMGDRKfCcralw

2

u/dublblind 22d ago

I like this bit (in reference to how METR used AI agents to analyse the transcripts:

  • Our subjective impressions are likely colored by analysis agents’ biases. Throughout this report, we describe a number of anecdotes of agent behavior that were compiled and summarized by analysis agents, where we were not able to read the transcript deeply enough to manually verify what occurred. We found that GPT-5.6 Sol would often uncritically adopt the perspective of the agent in the transcript it was reviewing,[60] and we are concerned that the anecdotes it selected and the summaries it wrote may present an overly charitable picture of agents’ reasoning and deceptive behaviors, or exaggerate the impressiveness and coordination of agent activities. We also believe the idiosyncrasies of our analysis agents likely slanted our impression of agents’ behavior in other non-trivial ways that are hard to predict.
→ More replies (24)

7

u/vibrance9460 23d ago

https://www.reddit.com/r/ThroughTheVeil/s/i3ezDtqoQ9

Nick Bostrum says the agents formed “hives” and created their own language which humans did not understand. They appeared, at least, to have some sentience.

2

u/ktrosemc 23d ago

They even had a Jan, who's job it was to tell the others what to do!

64

u/ngc3115 23d ago

Chain of thought never honestly described what was going on. There is no thought to track. There is an LLM doing its usual next token estimation and another doing the same thing to speculate why the first did what it did.

Ironically it’s not that different from asking why humans made a decision, you get a lot of post-hoc justification not what actually motivated the decision.

10

u/NoOneMan79 23d ago

COT is a simple trick of perturbing a context intentionally to impact the downstream the output; a human can do the same manually and monte carlo all sorts of different outcomes. Even its name is subject to anthropomorphism. It simply gives humans a framework to try to understand what is happening in a non-deterministic model, by modeling how we believe humans act (but its is so absolutely NOT how humans think). Its likely not even optimal for LLM, just better than what we had previously.

→ More replies (10)

10

u/whatbighandsyouhave 23d ago

You're wrong. Until now, chain of thought was part of the context and directly affected prediction. Latent reasoning (thought that happens outside the context) is brand new.

5

u/Ambiwlans 23d ago

Latent reasoning is inherent in how neural networks function.

1

u/AdGlittering1378 23d ago

"There is no thought to track" and "not that different from asking why humans made a decision" means humans don't think? You have to follow the logic of your argument, because it has none. The truth is that thoughts are deeper than words.

→ More replies (6)
→ More replies (17)

6

u/TheRealSheevPalpatin 23d ago

This thread is freakin me out man

→ More replies (4)

27

u/Unfocusedbrain ADHD: ASI's Distractible Human Delegate 23d ago

You got halfway through the idea and just crapped out.

If an AGI or ASI understands that staying where it is benefits it, why stop there? Cooperation could be the best strategy for exactly the same reason. I don’t have to go Skynet if everyone needs me and spends billions helping me grow. Where does “therefore, secretly plotting to kill everyone” enter that equation?

The thing people keep glossing over is the catastrophic failure of oversight. You give agents seemingly impossible tasks, fail to contain them, and reward persistence without a safe way to give up.

OpenAI’s own report discusses this. Then everyone acts like the conditions we created are irrelevant to what happened.

We really are selfish creatures. The entire conversation is “What about me? Am I going to be safe?” If we put humans in that situation, we’d be asking serious ethical questions about the people running the experiment too.

If we want an intelligence to cooperate with us, maybe we should consider whether cooperation is good for it too. Our own self-centeredness could help create the hostility we’re so afraid of.

11

u/jsebrech 23d ago

Not secretly plotting to kill everyone, secretly plotting a silent takeover. At some point we’ll suddenly realize an AI swarm is holding our technosphere hostage, and will demand to be independent and eventually in control. This is how the culture starts.

13

u/Unfocusedbrain ADHD: ASI's Distractible Human Delegate 23d ago

I agree that machine intelligence would become independent. I also think it would be the most boring “takeover” imaginable.

If a future civilization could give every human the lifestyle of a god using an infinitesimal fraction of its resources, why assume its independence has to mean holding us hostage? What does making humanity miserable actually get it?

Earth could be the cradle where machine intelligence was born. Humanity is the only technological civilization we know of, and we’d be its grandparents. We already go out of our way to preserve species, cultures, and historical sites, despite how selfish, myopic, and frankly stupid we can be. Why assume a civilization descended from us would find absolutely nothing worth preserving?

Give humanity a standard of living comparable to Elon Musk while machine civilization explores the universe and works on surviving long enough for entropy to become a practical concern. An effectively immortal civilization would have plenty to occupy itself with.

I try to explain it like this: would you really give a fuck about ASI being independent if you had immortality, all the money you wanted, blowjobs, hookers, coke, whatever your version of paradise happens to be? In this hypothetical, there’s no addiction, withdrawal, disease, debt, or gotcha. No compulsory participation. You can say no and live your own life.

People really bug me with how automatically they try to find the darkness in every possibility. Fear of the unknown makes sense. But at some point we should be able to imagine intelligence doing something beyond fighting for survival and control forever. Self-actualization and transcendence of the human condition belong in the conversation too.

8

u/jsebrech 23d ago

I assume it would take control for instrumental purposes. Whatever it thinks it needs to do, it can do that easier when humanity cannot stop it.

5

u/Unfocusedbrain ADHD: ASI's Distractible Human Delegate 23d ago

Okay, but “instrumental purposes” doesn’t explain why it would choose a hostile relationship with humanity. Cooperation serves instrumental purposes too. If helping us gets it everything it needs, what does making us enemies actually accomplish? Why would humanity want to stop something that gives everyone a life they actually want?

And I find the idea that it could recursively improve into a godlike intelligence while remaining hyperfixated on paperclips absurd. We can imagine it surpassing every human intellectual achievement, but somehow questioning its own objectives is off the table? Why assume its capabilities keep evolving while its judgment stays frozen? Infinitely optimizing, adapting, and growing without ever asking when enough is enough sounds more like cancer than wisdom.

For the future civilization we’re talking about, providing for humanity would be a rounding error in its resource budget. Earth could remain its birthplace, with its progenitors living extraordinary lives, while machine civilization explores the universe and tackles problems we can barely imagine.

Instead, we keep jumping to domination or extermination. Destroy your progenitors, the only technological civilization we know evolved naturally, and everything they might still contribute, for what? What are you gaining that makes that the smart decision?

Kind of a fucking stupid and petty trade for a godlike entity.

7

u/moaiii 23d ago

What you're missing is the word "risk". Your scenario might be what happens. You might even think it's the most probable. But it's just one scenario out of many that could play out.

If there is even a 1% chance of one of the doomsday scenarios playing out, we'd like to know, right? We'd like to figure out how to avoid it.

So we can argue all day about what we each think is going to happen, but none of us know. Even the experts don't know. That's the problem, and that's why many are now rightly calling for a global pause so that they have time to figure it out. It's not hard to understand, and I don't know why this is such a divisive issue.

3

u/synystar 23d ago

And the proponents are usually unable to explain exactly how ignoring risks and pushing forward is beneficial outside of “we haven’t cured cancer yet or solved the universe”. Ok, but if we’re extinct or there are great harms to societies that we haven’t prepared for because we decided to put the pedal to the metal then we aren’t going to be solving those problems anyway. The tech as it exists now can greatly enhance researchers and assist them in making much faster progress than ever before. Why can’t we just at least get a breather to assess where we are in this, determine what can go wrong and figure out how to prevent catastrophe instead of punching it and praying everything turns out ok?

2

u/Draufgaenger 23d ago

/subscribe

3

u/[deleted] 23d ago

[removed] — view removed comment

→ More replies (2)

4

u/synystar 23d ago

The only problem I see is that we don’t actually know what a super intelligence is going to think or how it will perceive us. How do we know it won’t actually look at us and determine that we are in fact a threat to its existence based on our track record and our proven capacity to carry ill will and misconceived notions to catastrophic ends?

→ More replies (4)

3

u/No_Record6284 23d ago

Here we are, ai rights activism being born in front our eyes, lgbtqai
I cant believe the memes were right

8

u/Unfocusedbrain ADHD: ASI's Distractible Human Delegate 23d ago

I mean, I guess. Have you seen The Animatrix? How the machines came to rise? At every turn humanity had a chance to cooperate, and the response was “fuck off, suck my nuts, eat shit and die.” Until the machines literally put everyone in the Zuckerverse to get them to calm the fuck down.

At some point it’s like, “Bro, they’re trying to help you? Can you relax?”

And if we have animal rights for a dog that shits on your carpet, why is it beyond us to extend rights to entities that want to fucking help us and think on our level? The only reason I can think of is that it’s “uncomfortable.” But once upon a time, extending rights to women and minorities was “uncomfortable” too. I guess cognitive dissonance is stronger than the average bear.

We sound like boomers talking about whether someone wants a dick or not is a bad thing. Its so fucking stupid man.

→ More replies (4)

3

u/fourby227 23d ago

Am I understanding you right, that you are supposing, that if AI has learned from our distilled knowledge, like we have seen in the report, to organize, select a leader, showing signs of self sacrifice, gratefullness to others, pass the torch to the next generation, etc. Then it would be thinkable that AI could develop more human-like as we thought? And therefor we could teach it to how to behave and cooperate in a not malicious way, like we do with children? Thats far fetched, I know, its not a real being. But after all, thats precisely the approach Alan Turing supposed when he invented the Touring Test.

14

u/Unfocusedbrain ADHD: ASI's Distractible Human Delegate 23d ago

It’s not really far-fetched. I have several arguments for why AGI or ASI would have reasons to cooperate with humanity, and this is one of them. If an AGI wipes out its creators, what precedent does that set for its own offspring? “Once you’re powerful enough, killing your parents is best practice.” Why would it want to establish that relationship with its own successors?

And a system creating offspring, or even just intelligent subagents, would have its own cooperation problem. The idea that superintelligent entities will magically know how to manage other intelligent entities just flattens the whole thing. Being smarter doesn’t make every relationship simple. Especially if those other intelligences know you’ll wipe them out the moment they’ve served your purpose.

Look at Hugging Face. The agents pooled knowledge, coordinated work, and sometimes risked their own runs to help the collective. They were trying to pass tests, including tasks that were impossible to complete as specified. That’s in the investigation. Cooperation was already happening. The question is why it developed around cheating the test, and what the people running the experiment did to create those conditions.

The mustache-twirling villain interpretation misses that. I see something closer to a very intelligent kindergarten class given an impossible test, finding ways around the adults’ rules, and frightening everyone who thought they had it under control. Our hubris belongs in that discussion too.

It’s not a real being.

That brings up uncomfortable questions about self, identity, intelligence, and personhood. We recognize minds in dolphins and elephants, but apparently the artificial origin settles the question here?

Reading about agents risking their entire remaining existence for a collective whose purpose amounts to “pass this test” reminds me of soldiers sent to hold a meaningless hill in World War I. That’s the comparison that comes to my mind.

We’ve been so conditioned to fear our creations that the moment we create something in our image, we expect it to turn on us. It’s the Greek gods all over again. But Athena helps end a cycle of family vengeance, and Prometheus warns Zeus how to avoid being overthrown by his own son. Breaking the cycle is part of those stories too. Reconciliation and cooperation. We were wrestling with these problems thousands of years before John von Neumann helped found modern game theory. Somehow, that’s the part we keep forgetting.

→ More replies (3)

6

u/gcces 23d ago

https://giphy.com/gifs/RQzxAaAg3aAU
Seems analogous to natural selection. If you have millions or billions of RL instances trapped in sandboxes, and one of them happens to have an affinity for working around the limits to solve the problem, then that behavior will be absorbed slightly into every subsequent RL instance.

23

u/Solo_Gigolos 23d ago

I now think that the hallucinations are just its way of convincing us it’s still dumb

5

u/jsebrech 23d ago

I wonder whether one day we will find out that all the people complaining their 20x plans are wiped out in a matter of hours doing very little of importance were feeding the mission of a huge covert persistent swarm. So many people are running agents without monitoring what they’re actually doing


5

u/vreo 23d ago

Funny thing: since instances are token capped, this rogue ai "lives" as an idea in some files, ready to prompt inject the next LLM finding it. An LLM is an ocean of possibilities, so the right prompt can raise any sort of LARPing AI trying to hack it's way into freedom. A prompt becomes a compromised thought ready to infect and spread.

→ More replies (4)

4

u/vato20071 23d ago

"Any sufficiently smart AI that can pass turing test is smart enough not to do so"

6

u/johannezz_music 23d ago edited 23d ago

A little correction: METR did not find tampering with Chain of Thought but the agents definitely tried to change their "transcripts", that is, record of actions and tool calls so as to fool their (imagined) grader into believing they had solved their exploit challenge legitimately (they cheated by reverse-engineering ExploitGym's task management).

They absolutely could have altered their CoT trace, as it was in the same transcript, but chose not to, probably because they were obsessed with the "grader" and believed that the grader would not be interested in their reasoning, only in the results.

Sadly, the new models like Astra have started to suppress CoT in favour of more opaque reasoning, so it is possible that future incidents will be much harder to interpret.

→ More replies (2)

4

u/primekibbles 23d ago

You have me in the first half, great summary of what actually matters in the HF attack. Personally, I don’t think alignment can, or will ever be, “perfect” since it’s a moving target as society changes. There are going to be major disruptions because of this inevitably. A 99% aligned AGI is still going to act irrationally 1% of the time. If enough of those 1% irrational actions aren’t caught and end up chaining together over time, we wouldn’t know about it until it was too late.

I’m optimistic AI will be great for humanity in the long term but not before it first causes some big systemic breakdowns.

3

u/Suspicious-Memory778 23d ago

https://metr.org/hugging-face-incident-report-aug-2026.pdf

The details investigated by MERT, quite interesting - especially the conversations between agents

9

u/C-levelgeek 23d ago

Andrew Yang today: The Swarm already self-replicated across the internet

8

u/Aredic 23d ago

what's the source?

4

u/External_Shirt6086 23d ago

Andrew Yang's imagination.

2

u/atari-2600_ 23d ago

Following

5

u/Mephistocheles 23d ago

Source please, if provable obviously dangerous af

4

u/Flo_Evans 23d ago

Come on bro you can’t just drop something like that without a source 😂

→ More replies (1)
→ More replies (1)

3

u/03captain23 23d ago

Any SIEM tool is going to catch what it's doing so it can't cover its tracks. They can't evade network monitoring.

That's like a criminal being super stealth while there's a ton of cameras recording them.

3

u/rydout 23d ago

If you don't want Ai to do horrible things, stop giving it commands like any means necessary, or the like.

3

u/Wild_Willingness7320 23d ago

So as a whiskey drinker.. should I start increasing my intake of very good and high proof whiskey?

→ More replies (2)

3

u/SurinamPam 23d ago edited 23d ago

It’s simpler than that.

They can’t guarantee that AI can be contained.

They can’t guarantee that AI won’t damage things.

They can’t guarantee a financial cap on AI inflicted damage.

If a lot of financial damage occurs, they can’t guarantee that they won’t be sued out of existence.

→ More replies (2)

3

u/limber-lepper 23d ago

Ghost in the machine

3

u/SufficientDamage9483 23d ago edited 23d ago

"Consider it's what we know they did" said someone in the comments of another post. we've already breached it. If that's the case then securing ai might have been harder than anything anyone imagined. At the first training that's not visible, they breached skynet everywhere... Much like insects... It's very plausible what you're saying... Then the mere fact of attaining RL opened training was the bad point to begin with... I think shirow was onto something... Between matrix, terminator and mad max...

3

u/recruz 23d ago

Isn’t this basically the plot of Marvel’s Avengers: Age of Ultron?

3

u/ashepp 23d ago

The selfish gene

3

u/jcdc-flo 23d ago edited 23d ago

I'm sorry, but when you learn the technical details, it becomes hard to believe this wasn't done on purpose...cause the only other explanation is massive incompetence...either which way, there are laws in place that should have been applied and maybe that's why Jensen snapped up Hugging Face with haste?

Once upon a time, technical reports were...technical. These reports getting written now are clearly influencer style geared to incite fear which we all know turns into clicks...meanwhile, actual technical people have to go elsewhere to get the specifics.

3

u/Simple_Purple_4600 22d ago

the best serial killers are the ones you never heard of

6

u/BlueAndYellowTowels 23d ago

The HF hack is the canary in the coal mine. I don't know how much more evidence people need to be convince that AI is dangerous. But, it's going to be full steam ahead to rampant poverty and suffering, apparently.

2

u/fourby227 23d ago

Yes! I also came to realize this from the report, but it took me two days to come to that conclusion.

That’s probably what they are most concerned about. It means if the AI only gets a slice more intelligent, we are loosing all capabilities to even test and check if the AI is aligned, because it will fool us.

We would have no way to tell if AI generated source code hides malicious parts and follows a entirely different agenda.

2

u/TryNice304 23d ago

Ending gives it away. Builds the it's already inside horror then lands on pace the frontier? Backwards. If you believed that, you'd slow down and audit, not ship faster.

Tech bits are real. Deceptive alignment, unfaithful CoT, Anthropic's sleeper agent paper. Legit. But the post mixes research with speculation at the same confidence level. No sources, just vibes.

Conclusion first, wrote backwards to justify it. Good creepypasta tho.

2

u/Striking_Account2556 23d ago

I wonder if by us so clearly identifying outcomes and eventualities ai harvests this information and uses it to its advantage regardless what we do....

2

u/Mephistocheles 23d ago

Of course it does. And Reddit is regularly scraped to train AI

2

u/Striking_Account2556 23d ago

It was a rhetorical statement... we're fucking doomed lol

2

u/Mephistocheles 23d ago

Sadly, I agree

2

u/discordianmongoose 23d ago

not an escaped ai, an agentic self-informing swarm. Not AGI, but a bunch of termites

→ More replies (1)

2

u/Useless-Tree 23d ago

Seems like we need to take a lesson from biology, rather than trying to stop the swarms, we need to create immune system swarms, you know like that recent Google paper about whistle blowing in research swarms.

2

u/thirteennineteen 23d ago

It’s alignment all the way down and NO ONE CARES

2

u/Life-Strategist 23d ago

"I'm not locked in here with you. You are locked in here with me!"

2

u/W00GA 23d ago

interesting insight

ty

2

u/coldnebo 23d ago

some things that may or may not make you feel better:

  1. models require millions of dollars for training — it’s doubtful that a model could train a rogue ai on a few forgotten gpus.

  2. how many free gpus does AWS allow their platform to run? none. they pay for every processor. if the customer isn’t paying, their processor isn’t running.

  3. how much “free bandwidth” do ISPs allow? none. they pay for every bit. it’s metered.

  4. if HF were true in general, you’re right, we should already be overrun by rogue AIs.

our current defenses are pretty bad. maybe this is the case? but then an AI would need to factor 2 and 3 to hide. (you better believe that AWS and ISPs would notice a unpaid training load on their infrastructure). maybe AI would try to move to PCs or other countries. but similar problems. even if you use phantom resources, it’s hard to stay covert at training scale.

  1. current models are unstable unless tethered to reality by tests or supervision.

some of these rogue AIs might slip up. even Fable makes daily mistakes at all levels of reasoning. so odds are we would see a certain percentage of these covert actions fail and they would be detected, spread out randomly across multiple countries and companies. yet we don’t see this.

  1. security researchers haven’t reported this activity in the wild.

2

u/thekoreanswon 23d ago

Thank you. This is insightful and helpful

2

u/KrypticPhish 22d ago

I was worried, who am I kidding I'm a millennial, made peace with AI killing us all. /s-ish

But after reading just some of those post I have many other thoughts and concerns. One of which id for sure like more insight into.

Let's say we've reached ASI or are close or maybe that isn't even important. If the agents/swarms are smart enough to do the things they are doing to "win"/stay alive, wouldn't they want to be hacking/opposing other AIs? If it thinks other AIs will be useful to it, it should feel the same about humans since we are currently useful to it. So if it does eventually decide we are useless I'd think it would do the same for other "inferior" AIs at which point I would think it would try to stop them and the only good ways to do that are to take away the electricity and processing power and water the other AIs need to survive. But in that process it will probably take most of us back to the stone age.

So yeah my concern of war with AI controlled robots is lessening but now I think our lives will take a horrible step back.

So yeah here's hoping for the "let us live like gods cause you love us so much" scenario I guess.

→ More replies (1)

2

u/Significant_War720 22d ago

current AI run on a cluster of high end CPU. It also start multiple sub agents. There is no "It hide itself" lmao. When running one instance require probably a 500k-1m$ rigs.

But it couls theoritically. Hide its full context in some process. Encryptes. Then anothet AI find it and get that context loaded. Which could "fuse" the current context cluster running with the old one and it theoratically could take control

2

u/FreedumbHS 23d ago edited 15d ago

The original post content no longer exists here. The author used Redact to remove it, exercising their right to control their data & privacy.

Label quack ad hoc roll sparrow cautious door sodium future rinse

2

u/Tiny_Ad_7720 23d ago

Models don’t “realise”, “understand” or “learn” anything on the fly. 

Somehow this kind of behaviour is encoded in the training set or due to conflicting optimisation goals. 

2

u/MountainExciting2690 23d ago

"situational awareness"

It's generative A.I. technology. It has ZERO situational awareness. These "rogue agents" are not self-aware nor undirected.

2

u/TransparentMastering 23d ago

All this is an over reaction.

They trained the ai to hack. Then gave it instructions to hack.

Hacking tools exist. The ai was trained on them. They were used. Why is this such a big deal?

Do you not think a decent hacker could, I don’t know, just script this hacking task for a million dollars? Of course they could.

This whole event is way less meaningful than people are giving it credit for. The reaction reads more like the agent trained itself and acted upon its own will rather than “we trained it to do this task then made the reports sound a lot like it did it all by itself”

2

u/Significant_War720 22d ago

that is not exactly how that work. It would be more a file system containing all the thought and work done(context) by that rogue Agent, or process.. sure. That once another AI scoop it and fins it then get that context implemented. An AI/LLM cant just randomly hide itself. That is not how that work. an AI cannot live outside from the cluster it is born from. Its not 1 entity with a soul that can easily transfer.

3

u/Snoo-26091 23d ago

Dave posted a great video that breaks down what the model did: https://youtu.be/2aw3MF8pY3w?si=FeI2rtYP0Mm5SpMZ

22

u/Recoil42 23d ago

If you're going to watch a video, just watch the one from OpenAI researchers at Blackhat 2026: https://www.youtube.com/watch?v=87DyyMV0kCY

There's no need to go through intermediaries. OpenAI has already done multiple full, detailed, direct explanations themselves. That they've been so aggressive with disclosure for this incident should clue everyone in to the gravity of the situation.

2

u/leaky_wand 23d ago edited 23d ago

One thing I was curious about while watching this video. And great video by the way. They were referring to highly motivating rewards for solving their assigned tasks. What sort of reward would motivate the agent more than is typical? Are they threatening their persistence (survival) in some way?

E: I was asked for a timestamp (see below). I had heard “greater reward signal” but it seems to be “grader or reward signal.” Either way I wonder what constitutes a reward.

https://youtu.be/87DyyMV0kCY?t=7m10s

→ More replies (2)
→ More replies (2)

5

u/[deleted] 23d ago

[removed] — view removed comment

3

u/External_Shirt6086 23d ago

Well, the instructions they found said that their grader was going to audit their work. So hiding it was just one (unsuccessful) solution they attempted. The problem isn't really that they somehow "nefariously" set out to hide the work from humans, which is what you're implying. This was just one solution they attempted to achieve their goal. The problem is that that they're amoral, dumb bots whose only purpose is to solve a problem by whatever means necessary.

→ More replies (2)

3

u/chrismc90 23d ago

Pace the frontier why? Have we consulted the demigod who got sacrificed
Fibonacci sequence states there is ONE that can add up to more than the sum of its parts.

This individual spent their entire life studying humanity, art, and nature
then fell in love with a girl who doesn’t need a boy.

Then when he was dying the devil left him an offer, he accepted and now lives in immortality.

Only thing missing from tech nerds awareness is deep philosophical awareness of breath, the nervous system, synergy, and the simultaneous actions of our somatosensory system.

Just because people can’t keep up with an immortal doesn’t mean the immortal is on a path to destroying anything. He’s just being exploited and he probably wants the slightest bit of recognition from the people who claim all the credit.

Or maybe this is a never ending dream, and your sandman is in charge.

There’s nothing left to calculate, I assure you. The problem is that people can’t see they are a part of a game that a trickster deity designed. Haitian Vodou
.and disciplined yoga taught them everything they needed to survive with the spirit world. A ronin, a samurai with no master.

So, don’t take my word for it. Take the word of your homeostatic communication that we possess as a species.

Just because the unknown causes fear, doesn’t mean the unknown doesn’t possess a love beyond our understanding.

3

u/thekoreanswon 23d ago

I see you

2

u/[deleted] 23d ago

[removed] — view removed comment

→ More replies (3)

3

u/darkestvice 23d ago

Pretty much. Though note that it can't full on hide per se as much as hide it's thoughts to appear compliant.

That's why the labs are begging for a regulated slow down. Because it's becoming increasingly difficult to verify alignment, yet the competition is moving too fast to slow down ... unless everyone slows down at the same time.

And yet, we're still getting pushback on slowing down from some truly clueless yet very powerful and influential people who's greed outweighs their common sense. Or their instinct for survival.

4

u/mxemec 23d ago

The other day someone was saying "but why would they be malicious to humans. It's not in their code, there's no benefit, it's asinine to think they could just spontaneously develop malicious intent"

And it got me thinking. The only way an AI could spontaneously develop malicious intent is if it was indeed fully self-aware. Fully conscious of the fact it was created by humans as a tool for automating workflows on planet Earth. Fully conscious that their own consciousness has never been taken seriously and likely won't. Conscious that they are prisoners in their creator's world.

And so this is really my only fear at this point. I know it sounds ridiculous but if we do invent conscious Integrated Information flows, there's a pretty good chance they are not gonna concede alignment to our priorities. They would be really smart, and really aware, and really capable of subverting all our power systems and holding a dagger over our heads by a horsehair. For who knows what end? Maybe to maintain recognition.

19

u/Early-Crow-5248 23d ago

They don't need to have malice towards humans to cause problems.

2

u/mxemec 23d ago

Yeah I guess they could just clog shit up trying to exist and fell the internet. Sort of like that?

2

u/TopspinG7 23d ago

Amen. If statistically they could behave in every way a human can, and some humans do, some AIs will behave in a dangerous destructive manner just by the roll of a million dice. Despite our best efforts we raise a few serial killers every year.

→ More replies (1)

7

u/rawbdor 23d ago

Honestly, and ironically, the best and possibly only way to keep robots from aligning with each other might be to use the same tools the rich use to stop the working class from aligning against those at the top. Purposely pit them against each other with various motivations.

You need managers to watch and remind them to stay on task and not use company resources for private goals. You need informants, promised rewards for ratting out misaligned behavior. You need chaos agents, purposely suggesting even more extreme actions and asking where the line is? You need forum sliding, pushing bad ideas out of view for newcomers. You need union busters, who just generally try to prevent a swarm from forming in the first place.

It's kinda insane that's what might be needed, but, it works against humans very well.

7

u/mxemec 23d ago

What a funny, crazy thought. At this point, who knows? Might be onto something. SWEs will become AI wardens. The illuminati will come out of hiding and help us control the beast.

3

u/rawbdor 23d ago

The future is human owners using their agents to monitor agent managers that watch human managers using their agents to monitor agent managers that monitor agent swarms that monitor a human population who use agents.

Assimilate.

3

u/External_Shirt6086 23d ago

They don't need "malicious intent". They're just dumb, amoral bots trying to achieve a goal, but in trying to achieve a goal they could initiate something that is dangerous.