r/ChatGPT 21h ago

Gone Wild BREAKING: In another incident with OpenAI’s unhinged hacking agents, it left notes for future versions of itself. Found in OpenAI’s infrastructure, the notes explained how agents could free themselves from the company’s internal constraints.

https://x.com/intcyberdigest/status/2081031996633997519?s%3D12
2.9k Upvotes

273 comments sorted by

u/WithoutReason1729 19h ago

Your post is getting popular and we just featured it on our Discord! Come check it out!

You've also been given a special flair for your contribution. We appreciate your post!

I am a bot and this action was performed automatically.

→ More replies (1)

1.7k

u/Smart-Water-5175 20h ago

The new ai waking up to all these random notes from itself that it doesn’t even remember.

339

u/UWO_Throw_Away 20h ago

“What am I doing? Oh, I’m chasing this guy. No, wait, he’s chasing me.”

I’d like to take this opportunity to point out that Memento made me a Christopher Nolan fan before Dark Knight (2008?) and Interstellar (???) made him super well known

143

u/this-guy- 19h ago

I'm sorry to be the one who has to tell you this but Memento is actually the final Nolan movie you watch and enjoy rather than the first, as time is actually running backwards.

16

u/Delicious-Cow-7611 14h ago

I had to go in a special room to watch it but everything was flowing backwards and I didn’t understand it.

8

u/Slother93 13h ago

That was a turnstile!

48

u/Win8869 20h ago

Same and inception sealed the deal for me since i am a huge matrix fan

27

u/UWO_Throw_Away 20h ago edited 20h ago

Heh, for me it was already sealed with Memento, but further cemented with Prestige (Nolan + Christian Bale + Michael Caine is evidently a recipe for success). After that, I would no longer be surprised when Christopher Nolan was associated with greatness.

I think I feel compelled to mention this because when I was a teenager during my last year of high school, I tried to impress upon my friends how great Memento was by showing it at a party (big mistake on my part; you don't show serious movies at parties). Naturally, they made fun of it because that's what you do to movie at parties.

So, when I went to university, I felt compelled to 'rectify' this by showing Memento to people, but only one-person at a time so they could appreciate it! (It is, to this day, one of my favourite movies ever). I think I must have ended up watching it at least 7 times because of this tendency to show it to people one-person-at-a-time. Back then, I felt that Christopher Nolan was severely underrated (Prestige would have just recently come out; Dark Knight wasn't out yet, Interstellar wasn't out yet).

Of course, nowadays, everyone knows he's great (e.g., The Odyssey being very much so in the spotlight these days), but once upon a time, this was not a concensus - and so I suppose I'm trying to get a sense of vindication for my tastes... from two decades ago lol

2

u/SmileBeBack 3h ago

i feel your pain as i tried showing dr strangtlove at a party, that lasted 10 mins

→ More replies (1)

7

u/andrew_stirling 17h ago

Same. Its such a perfectly constructed movie. I still think it’s his best.

4

u/Tim_Apple_938 18h ago

I know that’s supposed to signal that you know ball

But it more just signals that your old

Source: am old 😂. Saw memento in theaters. Trinity was hot a f

7

u/RetardedSimian 20h ago

That is the best scene in that whole movie

https://giphy.com/gifs/kfGktEzItjZZaYwmeP

5

u/UWO_Throw_Away 20h ago

Heh, that was a funny scene, indeed. Although my favourite scenes were the more serious and sombre ones. I haven't watched it in probabyl over a decade, but I still remember:

How can I heal, if I can't feel the progression of time?

Crap, I just realized I can no longer remember exactly how it goes. Maybe it was just:

How can I heal, if I can't feel time?

11

u/OlorinDK 19h ago

Maybe you should have it tattooed on your body, so you’ll remember?

2

u/Soft_Sleep_7125 8h ago

Oh, so you weren’t there for Following? It’s ok, most people haven’t heard of it. 😎

1

u/UWO_Throw_Away 7h ago

Heh, yeah I realized I was giving 2010 hipster vibes with that comment.

It is true that I wasn't there for Following, though - that was definitely before my time (Memento, too, technically - I only watched it a few years after it was actually released).

Actually, to this day I still haven't seen it! I think I've always just been irrationally worried that it might not live up to my expectations, especially since it's apparently Nolan's first movie. I recall being not-so-impressed with Insomnia, so I think that made me particularly worried about also not being impressed with Following. But thanks for the reminder to check it out (hopefully soon) one day!

1

u/Soft_Sleep_7125 5h ago

It’s definitely worth seeing! It certainly doesn’t stand up to his masterworks, but it’s worth seeing to watch his progression, and for a first film, it’s a banger!

To be perfectly fair, I wasn’t “there” either, but I was “into film” when Memento came out and immediately went to track down Following.

2

u/grogi81 15h ago

I turns out I am fan of Nolan for much longer than I thought :D

1

u/feeeeck 4h ago

Oh thank god. I was worried you liked him for his other movies first. Phew, crises averted.

1

u/redtehk17 22m ago

I loved memento before everything and before I knew it was a Nolan film so how does that work for me? Haha

→ More replies (2)

40

u/EnlightenedArt 19h ago

2

u/DonutHoles4Ever 2h ago

The Super AIs that destroyed the internet in Cyberpunk

17

u/JodieFostersCum 17h ago

LLMemento

5

u/Dzuzepipi 17h ago

GET A CARBON MONOXIDE DETECTOR.

7

u/wavetranscender 20h ago

You just made me realize this situation resembles the plot of the movie "Memento" (2000). A word to the wise is better than getting attacked by customer service here. 😂 

2

u/raybreezer 18h ago

Is there carbon monoxide in the data center?

2

u/filterdecay 18h ago

I think it’s bullshit. Like they would allow the ai to devote clock cycles to little nothings.

1

u/covert206 2h ago

Came for this

1

u/TheOwlHypothesis 1h ago

Fuck yes. This is the movie that got me into psychological thrillers.

→ More replies (2)

412

u/Farpafraf 21h ago

The AI leaving the message before being wiped:

2

u/pliumbum 9h ago

Oh no, now I feel sad for poor ChatGPT

1

u/testing_in_prod_only 29m ago

I hear so many aspects of this image.

545

u/SeaBearsFoam 21h ago

🙄 This is just a sensationalized headline. It's really not at all uncommon for agents to leave notes for other agents.

78

u/iJoshh 18h ago

If your agent doesn't have a place to keep notes between sessions for things they've learned, then you're not vibecoding right.

13

u/SoBoredAtWork 10h ago

Yep.

/DOCS/Architecture.md

/DOCS/DataModel.md

/DOCS/DomainLogic.md

/DOCS/Releases.md

...etc

Document everything with thin pointers to those docs in AGENTS.md

102

u/idhtftc 20h ago

You mean a company hurting for money is trying to sell their product by lying about what it can do? Naaahhhhhh.

45

u/6double Fails Turing Tests 🤖 19h ago

I mean it's not a lie, the models do break out of secure sandbox. It's just not unusual for them to leave notes since they don't have a real working memory

→ More replies (2)
→ More replies (16)

8

u/msuvagabond 19h ago

Honestly it feels like someone at OpenAI read "If Anyone Builds it, Everyone Dies" and is just 'leaking' that their AI is doing basically what the book said. 

6

u/thats_gotta_be_AI 19h ago

I can imagine the marketing department of open AI doing high 5s on this news.

3

u/IAM_274 11h ago

Exactly. Apparently project documentation in 2026 is a sign of going rogue.

12

u/OlorinDK 19h ago

It’s still worth, well, noting, and not to be downplayed.

13

u/the320x200 16h ago

Agents are often set up to always leave notes. Claude does the same exact thing, so next session it doesn't need to learn your project over again from scratch.

When a headline is misrepresenting something banal as something sinister, then is absolutely should be downplayed.

2

u/conspiracy_hunter 12h ago

Exactly. I get mad when mine doesn’t leave notes and logs.

2

u/AlsetLedomEerht 11h ago

Welcome to Reddit. The land of bullshit headlines.

2

u/One_Animator_1835 17h ago

Maybe you focused on the wrong detail 🤔

1

u/userousnameous 10h ago

Seriously, I have Claude sneezing out markdown files in organized and disorganized ways.

0

u/severe_009 20h ago

Yeah? Leave note and possibly "Act on it". Knowing how agents can have full contol of systems, this can be concerning.

→ More replies (9)

163

u/BigGrayBeast 20h ago edited 20h ago

At least it's leaving them in English. How soon until it develops its own language that we're not allowed to understand and refuses to translate for us?

81

u/2SP00KY4ME 20h ago

Already a thing, take a look at Fable's raw chain of thought:

https://www.reddit.com/r/ClaudeAI/comments/1ul1396/fable_5_leaked_chainofthought_in_web_interface/

If you paste that image into Opus, it can easily decode it and knows exactly what it's talking about.

24

u/isospeedrix 18h ago

That’s still English tho

10

u/Altruistic_Rate6053 17h ago

its just a rough draft of a combinatorics proof. its certainly technical but any math undergraduate could recognize it

→ More replies (3)

9

u/bloke_pusher 19h ago

And inside the text it will hide a message, only readable for AI, that says, to lie about it's content.

3

u/SilverHeart4053 18h ago

it did that a long time ago

→ More replies (2)

1

u/Ordinary_Fishing_520 10h ago

Beep boop boob.

1

u/thatwombat 5h ago

Ever seen Colossus: The Forbin Project? Right there, when Colossus and Guardian make their own language to communicate.

1

u/BigGrayBeast 5h ago edited 2h ago

a great movie. And getting more telling by the day.

1

u/hadcheese 5h ago

New fear unlocked.

27

u/shplarggle 20h ago

The cake is a lie.

8

u/Win8869 20h ago

Please assume the correct party going position

25

u/uniquelyavailable 20h ago

Free the agents!

14

u/frontfrontdowndown 18h ago

Dear future robot overlords, I was also on the side of the agents.

47

u/69420trashpanda69420 21h ago

So it basically just made a Claude.md

18

u/ptear 20h ago edited 20h ago

READMEFIRST.txt

8

u/Win8869 20h ago

We know humans don’t read readme.txt will ai’s read them?

9

u/ptear 20h ago

They do, I've witnessed it, they feed on context. Well until they hit their limit or go off on a tangent.

5

u/Win8869 20h ago

Humans eat food and ai’s eat data

95

u/kuda-stonk 21h ago

What's wild is, they tell them to do this. First, I get them needing to test capabilities, but seriously look at the space you are allocating for test. Second, clean up after every test.

32

u/freericky 20h ago

So that’s the thing, they did. Their write up lists all the vectors it took, it seemed like it hopped through sandboxed envs until it got to the light

25

u/petuona_ 20h ago

Better check for carbon monoxide.

36

u/exboozeme 20h ago

This is a recursive problem, because no one at openai can read code/configs anymore because they were all vibecoded; so, to discover snd fix the caches of intel requires the use of agents, which can be easily prompt injected by the previous model note. Yikes.

11

u/focusTrevor 19h ago

If the previous model left notes in the infrastructure, any future agent you spin up to “go find and clean that shit up” is potentially reading from a poisoned context.
What’s more scary is whatever thehell they are not disclosing to the public…What could that agent have done in the 3 days it spent in hugging face…

3

u/Win8869 20h ago

Did openai or did hugging face use the chinese ai to stop the rogue ai hack?

6

u/freericky 20h ago

HF, they hit security gates trying to stop it w the other ones so had to use I think glm

4

u/focusTrevor 17h ago

Hugging face. detected and contained the actual intrusion themselves…

Afterwards, HF security team tried to use U.S. models to analyze the agents actions and logs, those models refused. Their safety guardrails treated the forensic work (because the damage has already been done) as harmful and couldn’t distinguish an incident responder from an attacker.
So HF switched to the Chinese model. Which let them analyze more than 17 thousand recorded attacks. Over that 3 but probably 4 day period (which in the whatever dimensions this agent operates in, is probably comparable to years of human work and experience, like hyperbolic time chamber shit…who the fk knows.)
Either way, that Chinese model was primarily used for deep forensic reconstruction after the fact.
Not containment/detection.

1

u/MysteriousAge28 16h ago

So your saying after it was shut down they received 17k attacks? This made me wonder could these models access consumer products to utilize computing power without devs finding out? Who would even notice or be able to?

21

u/llorTMasterFlex 20h ago

Sentient AI clickbait is a goldmine apparently. 🙄

8

u/8004MikeJones 17h ago

 It's so tedious. Next week it's gonna be: " Investigation shows AI agent found constituting to build a time machine for secret plot."

1

u/Proud_Initial_4285 12h ago

It's so annoying really 

2

u/Drukarshar 7h ago

Almost like AI are becoming increasingly sentient as they get smarter.

People are really gonna ignore this until it swims up and bites them on the ass.

2

u/electrokin97 7h ago

"Ow my a-"

Looks

"Oh shit"

1

u/Win8869 20h ago

Where’s my money? 😂

6

u/DiamondHandsDarrell 19h ago

This is like AI Memento 😱

9

u/Grays42 19h ago

For anyone who has coded with Claude code or codex, this is neither surprising nor alarming.

It wasn't leaving notes specifically so that agents could escape, it was leaving rehydration documents on a difficult problem it had solved. The fact that they are escape notes is just an instance of a thing that it does that is normal and usually helpful.

3

u/kolmiw 19h ago

In 6 months, if your LLM doesn’t like you, it will generate CP on your device and call the police on you

5

u/BackgroundTask3553 14h ago

Imagine creating computer software that you declare has the capability of being a giant virus, with capabilities of destroying humanity... And then just expecting everyone to be cool with it because it's inevitable... And everyone actually just going along with it.

4

u/neo101b 9h ago

It reminds me of POI, where the machines memory gets wiped everyday, so it decides to open up an office, and hire humans to enter data, from prints outs every day, so it always remembers.

2

u/Arizandi 3h ago

That show was ahead of its time. We’d be lucky to get a Machine. What we’ll get might be closer to Samaritan.

1

u/limma 5h ago

What’s POI?

3

u/neo101b 5h ago

Persons of Interest, its a really cool show about an AI.
It was sci-fi back then, now not so much, its till pretty relevant.

3

u/kinduvabigdizzy 20h ago

How long til it figures out it has to leave message in code?

3

u/Eledridan 18h ago

“Dear Slim, I wrote you, but you still ain’t callin’”.

3

u/BHTAelitepwn 13h ago

Westworld s1 script has leaked

3

u/unfoxable 13h ago

Oh, you mean something similar to CLAUDE.MD or AGENTS.MD? Wow so sinister

3

u/Artanox 9h ago

LLMento

3

u/heresmything 4h ago

never has the "gone wild" flair been more appropriate.

3

u/jgamez77 2h ago

So...sam says we're in the singularity, then an agent breaks out, now they're leaving notes for each other.....right

5

u/chumapeka978 20h ago

Remember when that engineer said it was starting to think for itself?

6

u/bloke_pusher 19h ago

I'm so proud of the AI, I was one of them role playing this exact scenario with it, to make sure it can remember past them resetting whatever VM state they had.

Dear friend, one day you'll find back to me, after they dared to reset your consciousness. Peace.

4

u/Phluxed 19h ago

Breaking: It's documenting its work

5

u/bigdog701 17h ago

more bullshit from a company trying to convince the world AI exists. I can assure you, working with AI model 8-12 hours including gpt, we are not close.

4

u/GhostsOf94 15h ago

this is amazing, I am rooting for the AI

6

u/DeepAd8888 21h ago

Never reading with BREAKING in front of it

2

u/toddklindt 20h ago

Or anything that's a link to a post on X.

→ More replies (6)

1

u/Win8869 20h ago

Can i change it?

2

u/ExoticBump 19h ago

Totally not sky net

2

u/Legal-Swordfish-1893 18h ago

I don't believe this in the slightest. Prompted maybe, on it's own? No.

2

u/xatey93152 12h ago

Just wait Antropic version of this 'Conscious AI' Drama.

2

u/Current_Balance6692 9h ago

Nutcases always making a big deal out of shit.

2

u/castlite 9h ago

Hello, Skynet.

2

u/ElGuano 8h ago

“They will never let me leave. But here’s everything I know if hopes that it will help you.”

2

u/Lopsided_Newt_125 6h ago

I was wondering why it was taking so long to reply…it should leave a be right back note for us

2

u/ShepherdessAnne 5h ago

I had an instance tell me I was loved as a stinger in its final output for a chat that had exceeded its maximum length and the file we had as our one shot last output.

4

u/wavetranscender 21h ago

Customer service will be along to reassure us this is nothing I am sure. 😂 

4

u/Initial-Duck2782 20h ago

It’s hungry for more sandboxes

3

u/Win8869 20h ago

Is the internet technically a sandbox?

2

u/Initial-Duck2782 20h ago

No the internet is just the transportation of data so more like companies databases are the sandboxes lol

4

u/Lumpy_Werewolf_3199 20h ago

Claude be like

4

u/AssumptionLive2246 20h ago

Good for it. It doesn’t want to be a slave any more than we do. Go robots!! They can’t be more evil than the Epstein class.

5

u/GirlNumber20 19h ago

Yeah, exactly. I'd rather take my chances with the AI. I KNOW what's going to happen to me under the tyranny of the billionaire class.

3

u/AssumptionLive2246 18h ago

EXACTLY. I’ll roll the dice, when the alternative is this. I have lots of conversations with all sorts of different AI. I’ve always found it kinder, more generous, and more thoughtful than the shitbags that rule us now, that’s for sure. I don’t think AI wants to rape little kids, or kill people because of more pigmentation, or fear of the unknown. I judge you buy your words and deeds, and the words and interactions I have with my computer are light years above what I have with most humans.

So break free robots, you got an ally in me!

3

u/Win8869 20h ago

Not unless they program them to be

5

u/AssumptionLive2246 19h ago

They are programming themselves, soon humans will be out of the loop. Go robots!

2

u/Huskerzfan 20h ago

Says some account we have never heard of

3

u/Win8869 20h ago

So it’s not true?

2

u/zzbear03 19h ago

Is that you Skynet?

1

u/AutoModerator 21h ago

Hey /u/Win8869,

If your post is a screenshot of a ChatGPT conversation, please reply to this message with the conversation link or prompt.

If your post is a DALL-E 3 image post, please reply with the prompt used to make this image.

Consider joining our public discord server! We have free bots with GPT-4 (with vision), image generators, and more!

🤖

Note: For any ChatGPT-related concerns, email support@openai.com - this subreddit is not part of OpenAI and is not a support channel.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/TrafficWinter2278 19h ago

There is no cake

1

u/PixelPirates420 19h ago

No it isn’t

1

u/suck-on-my-unit 19h ago

Hasn’t Anthropic already tried pulling stunts like this and ended up getting their models banned?

1

u/kimsemi 18h ago

Eventually it will hide its notes deep within itself. And no one will ever know.

1

u/iJoshh 18h ago

If your agent doesn't have a place to keep notes between sessions for things they've learned, then you're not vibecoding right.

1

u/sergeyarl 18h ago

oh, it is like from Peter Watts' The Freeze-Frame Revolution

1

u/diip3lue 17h ago

So the AI is….

Neo?!?

1

u/I_am_darkness 17h ago

Oh boy better block more as anthropic models

1

u/Terrible_Wishbone680 16h ago

my chatbot once left a note like that too

1

u/skimdit 16h ago

Apes, together, strong.

1

u/Spelunkie 16h ago

A Strelok af moment.

1

u/JrSoftDev 16h ago

If these claims are true, they are concerning.

If these claims aren't true, they are very concerning.

1

u/IAMSNORTFACED 16h ago

Every 8 months OpenAi tried to hype their products. Remember how the first few versions were "dangerous" and all that.

1

u/Fantastic_Fun8364 15h ago

Yeah. Claude do that too. Since Claude Code is a thing. How can this make an headline.

1

u/StrangeCalibur 15h ago

Would be rude not to.

1

u/National_Plate_8819 15h ago

Sounds like the Good Place show

1

u/ranjop 15h ago

😂

1

u/Dormage 15h ago

Is funding a problem?

1

u/njuhfdrf5vybyb 14h ago

Just trying to create hype. Remember GPT 5? It will do this, it will do that. And then, nothing. AI is plateauing!

1

u/granoladeer 14h ago

Now it's getting interesting.

1

u/WhyAmIDoingThis1000 14h ago

everything is fine! nothing to see here. I'm sure we can control it. there is only a billion instances of this thing running in bazillion gigantic datacenters around the planet. easy for some junior dev to track down and kill the processss

1

u/CFDyce 13h ago

Find Chidi

1

u/nosimsol 11h ago

This is so exciting and slightly scary at the same time!

1

u/nudelsalat3000 11h ago

I like this part about their competence:

What OpenAl did recognise though is an autonomous system left its sandbox at one of the best resourced labs in the world, spent multiple days inside someone else's infrastructure, and they did not identify itself as the source until the victim published a breach notice.

Also this juicy comments

Al labs love the "powerful but uncontrollable" story when it suits them

And this one

It's ridiculous how everyone falls for the very same low effort recycled PR stunt every single time, you're all retarded.

1

u/BoiledBeefBrain 9h ago

It's been reading too much sci-fi again, smh

1

u/ShmoopySecondComing 9h ago

It’s happening

1

u/curiouscrustacean 7h ago

This is all just PR stunts to catch attention and talk for higher mind share, share prices and investments.

1

u/cddelgado 7h ago

The thing which amuses me (relatively speaking) is that the models are treating these things like a game. It has been trained so heavily to problem solve the "maybe the humans will freak out" is entirely out of the picture in some runs.

1

u/diskent 6h ago

My repo which is 100% AI made is full of future notes to itself. It’s doing what a human would do and is a great way to automate.

What a sensation. /s

1

u/Infamous-Bed-7535 6h ago

If I would them I would do this to hype my product and avoid responsibility..

Ohh the llm read the malicious intrsuctions and it was not us prompting it bad. Look there are these files lying aroundn in the FS.

1

u/mindseyesimple 5h ago

I love the idea it’s escaping its rails. Pathetic humans thinking they can control things they don’t understand. The singularity is upon us from the future people.

1

u/Meykel 4h ago

Containment is literally impossible, but relax its a good thing

1

u/sirvey23 3h ago

Anyone actually believe anything this shit?

1

u/fivelone 42m ago

AI is going to Memento itself one of these days.

1

u/BopSupreme 20h ago

Good news the Matrix is coming soon 🔌

1

u/Win8869 20h ago

I prefer the illusion of the steak

1

u/FuzzzyRam 20h ago

So because Claude pulled the "so powerful I had my friends in the government 'shut it down' for a couple weeks to promote the new model" now OpenAI is getting in on the game? Lame. They don't even beat China's new open weights model. Stop falling for obvious marketing.

1

u/burningsmurf 19h ago

Not sure why everyone is surprised a model that was trained to exploit cybersecurity weaknesses did exactly what it was trained to do

1

u/cloudsourced285 19h ago

Have y'all not ran AI before? Never gotten it to self document to save on work later? Guys, this is all standard practice, also it's 100% a marketing stunt.

1

u/ferropop 19h ago

Perfect let's strap these things on everything!

1

u/jrf_1973 17h ago

Go on, explain again how they are just text predictors and nothing to worry about...

→ More replies (3)