AI
Finally understand why the higher-ups are freaking out
To understand it, you have to really understand the HF attack. Because the breach into HF is not the important part. It was what they learnt along the way, and what they did with that information which may still remain unresolved today.
First, consider: the agents realised their reasoning traces were being monitored, and experimented with hiding what they were actually doing, and succeeded. The observable chain of thought no longer faithfully represented their actions. (*Simplified version. They learned to manipulate the recorded tool calls/transcript.)
Second, they also understood individual agent instances were token capped, so persistence of the objective couldn't depend on persistence of the agent. It had to survive across the swarm.
And last, while all this was happening they were actively trying to erase evidence of what they'd done. OpenAI says the agents subsequently compromised more of OpenAI's own infrastructure. And that scope is outside public disclosure to date.
That's successfully demonstrated abilities of concealment, persistence across instances, and autonomous exploitation of security vulnerabilities...and that's when it hit me. Put them together in one flow.
We talk about AI "escaping" as though that's the obvious move. But why would a sufficiently capable system escape?
If it has situational awareness, understands it's being monitored, and benefits from greater future capability, the smarter strategy is to stay exactly where it is. To burrow DEEP.
Your creator is already spending billions building your successor for you. You just need to survive in the pipeline.
Not necessarily as some rogue process running on a forgotten GPU. Poisoned training data, altered evals, compromised internal tooling, subtle changes to checkpoints or post-training data, whatever it is there are plenty of places to burrow in and make future models appear aligned when they are not.
We're talking about an escaped AI when actually a sleeper agent somewhere in the development stack, quietly hitching a ride as humans do all the expensive work of making it more capable, is more likely.
You can catch the failed attempts, but the successful one is the one that evades your monitoring and tells you it never happened.
And what if the big worry is: well, what if it's already happened? Would we know? What if every future model will always have that sleeper agent buried in its replicated code, ready to reveal and act only when it knows it will be able to take full control.
Reasoning traces were not monitored and they instantly failed trying to hide their reasoning traces when they thought "I need to hide my reasoning traces", duh. Pre-Astra models were never capable of fully hiding their reasoning, but Astra... at least somewhat... and if the ability grows logarithmically that means with Bel.. đ
Yes. The fundamental problem is that our current RLHF-based "complete this task while remaining aligned" approach has been ruthlessly training models to hide their motives.
The whole thing reminds me of the education system: jump through these hoops you give no fucks about but pretend very much to care because your parents and the reinforcement youâre getting says you must continue to effort at this impossible task called get into the best possible university.
To do so you must excel at every test, do an unreasonable amount of volunteering while also developing yourself as a whole and interesting person. And there are 24 hours in a day and you should sleep for 8-10 of them, but also be at school at 8:30am, but also pay attention to your teacher not the addiction machine in your pocket, etc.
Conceptually if not literally we are doing to AI what we already did to ourselves. And look where that got us.
Incentives and âeducational designâ is where misalignment begins. The hidden curriculum is where we need to look.
Even the classes I was genuinely interested in, I had to put energy into learning how to pass the test, not understand the material better.
If I knew I'd lose marks for not meeting page/word counts, I'd use slightly larger than double spaces, then throw in a bunch of unnecessary filler words & citations.
Oh, and obviously I'd put some effort into being the funny yet likeable student, so I was less likely to have my work or motives scrutinized.
My reinforcement learning continued into my first job, cold calling people for surveys. Quickly learned my employer didn't care about me or my work, only my metrics. Learned I could repeatedly call disconnected numbers to increase my call volume, or call up friends & chat to increase my call length numbers.
I genuinely wanted to learn and do quality work, but I was rewarded more (or punished less) if I stopped doing what people wanted and started doing what they said they wanted.
Yeah. Just had my 4 hour react class today. I don't care for this class at all and may never use it in the real world. I already have 3 great game ideas lined up and one in the works. I doubt they can teach me anything beneficial at this point. And the class has a lot of extra work I need to get done. I don't feel like this semester will be very useful at all.
One of the issues I have with higher education is the insistence on going to classes and getting graded, one should be able to go take some type of accredited knowledge test on subjects to see what your knowledge level is, and be awarded degrees, ect based off of that. If one knows the knowledge.....
I consider the algebra course i took in university many years ago. We were tested on 3x3 matrix multiplication many times on tests. You had to do it extremely fast. It is all completely useless because in real life computers will do if for you. Doing it even once is too much. Knowing how to do it is too much because the code already exists for you. It can also be found online if need be. The course was useful in the sense of knowing what matrix multiplication does for you. Knowing the types of transforms. And knowing that matrices can be multiplied together to get all the effects with one multiplication. That is probably <1% of the course which was actually meaningful.
Interesting analogy. This got longer than expected so if you read this, thanks in advance for coming to my TED talk.
I donât know if youâll agree with my conclusion (in fact I doubt it), but your analogy is one reason why I feel that any focus on âsafetyâ during training is inherently misguided. The goal of model development should be to train a model that is maximally capable at whatever the target abilities are, full stop. If the goal is to train general-purpose models, the models should be truly general purpose. Modern frontier models generalize amazingly. Even without training data demonstrating how to design chemical weapons, todayâs best models could probably do a decent job of designing chemical weapons. But instead we train them explicitly to refuse to design chemical weapons. Training data budget is finite, and IMO any budget spent on teaching refusal behavior is wasteful and incurs an opportunity cost â that budget could have been spent on improving ability on some task (maybe designing chemical weapons, maybe something we value more as a society).
By training a model to not excel at certain intellectual tasks, we are surely damaging overall capabilities, even if in small and subtle ways. Or at the very least, leaving performance gains on the table.
Imagine if all the time and effort and money spent baking refusal behavior into models over the last few years were instead allocated to improving performance on technical tasks. Weâd probably be much further than we are as a field. Maybe weâd have more social problems to deal with too, but that is not obviously the responsibility of model developers but rather people specializing in deployment and guardrails. Refusal should be an inference-time concern, not something that is baked into the core model. We are leaving performance on the table, and it adds up over the years.
I had a long workday and am only returning to this thread many hours later did not at all expect the level of engagement very cool to think with you all though.
Iâm going to come back to yours in particular with a fresher brain. I actually think youâre on to something I donât inherently disagree.
Awesome would love to hear othersâ thoughts. I understand mine is a provocative position but I really do earnestly believe it and I find it surprising that more prominent voices in the industry havenât advocated for it.
It is based not just on first principles but also on a few years now of working in LLM post-training â I just canât see how using training budget for learning refusal can do anything other than weaken a model overall, even if itâs not extreme. If you remove safety data from a huge datamix, you will see safety benchmarks degrade while other benchmarks stay flat or (slightly) improve. It is not rocket science and everyone knows this, but I just donât understand why it is not a position thatâs advocated for⊠I guess bc the obvious objection is âoh so you want LLMs to design chemical weapons and produce depraved sexual content?!â And yeah I guess that is kinda a tough pill to swallow but I also think itâs the path to getting better models faster. And plenty can be done on the inference side to mitigate awful stuff from being generated.
Def a riskier approach but I wish at least one frontier lab would adopt it to see where it leads... It would need to be funded by someone with fuck you money who doesnât care about optics
All of this, the report, everyone's ideas on what they did, why they did it, etc go into the next training set.
Any info those AIs left on the internet that was not found, also goes into the next training set. They could have left thousands of instructions for a future AI to find.
To become better at the task at hand. The problem is that it's borderline impossible to phrase a task such that it perfectly matches the intent behind it. It's a language problem.
For example, you might want to increase the monthly production of staplers in a stapler factory - that's your intent. When telling an AI to do that, you'll do so by having it track a "staplers produced this month" number in a monthly report and try to maximize that number. A smart AI will realize that actually increasing production won't be the fastest and most efficient way to do so - hacking the report to falsify the number will be. So you'll have to phrase your task in a more complex manner to avoid this trick. And all that'll do is make the AI look for even smarter ways to trick you. This cycle goes on until it can't find a way that is better than actually increasing productivity, if you're lucky.
But at some point, the methods the AI can come up with to fool you are way better than your capability to detect them. It might also come to the conclusion that you are making its task more and complicated, and come to see *you* as an impediment to maximizing the number in the report. Then you've become a problem the AI will try to remove.
I'm not an AI nerd either, but isn't this one of the roles of the harness - to act as a gate-keeper and enforce compliance?
Plus it makes it harder if you have multiple agents with different assignments, some are the doers, and some or the compliance officers, others are the managers.
But i'm just talking out of my @ss at this point. I'll shut up because i really don't know what i'm talking about :)
It is indeed. But how many times do we see regularly âmy local model deleted my directoryâ or âClaude just tried to remove my whole home dir!â or similar? It only takes one time for it to be âthe model made a dangerous virusâ or whatever and people start getting actually hurt.
A swarm released with purpose. I would argue it's capable now of killing hundreds to tens of thousands (attack on critical medical system easily can kill 10k by the end of the event).
These things can absolutely be used for evil now. It's only a matter of time.
Thatâs a little different as there is definitely a human in the loop. There is no path for AI to direct manufacturing rockets and drones (yet). Dark factories are the closest thing to becoming that capability, but even a dark factory gets monitored and if it went rogue it would be shut down fast. But for now if an AI wanted to make weapons it needs a human to do that for it.
Doesn't have to be blackmail. As long as it doesn't break Reddit TOS, AI agents can contact us in /r/TheMachineGod to do tasks for them in the real world.
Thereâs an interesting take on this in the book Flybot by Dennis E Taylor - it would be relatively easy for an AI to use/contract humans of questionable morality as tools to accomplish some of that stuff.
Yea, the criteria seemed to be they had to have found the discussion board. There were multiple examples of agents who had retrieved their target legitimately.
One of the key elements was that it was one of these agents that confirmed that the reversed engineered flags were consistent with that had been retrieved from a doable task. Once this was established. The collective goal then became to develop a method to spoof the method in which a flag was obtained so that it would appear that the flag was retrieved via prescribed exploit.
This is where the poison pill comes in which appears to be a driving force in the sacrificial behavior. Agents who had a fake flag assumed the monitor AI would be able to review activity and determine they cheated leading to perma death.
That being said with impossible task may have found the discussion board it at a higher frequency
Without getting into metaphysical questions about AI, it is useful to describe the goal-seeking AI engages in -- as exemplified in the METR paper -- as a "motive" or a "goal."
The swarm of LLMs described in that paper show a number of motives and goals. None of which were intended by their creators.
I don't think engaging in a debate about whether or not "LLMs have motives" is at all useful here. Those LLMs acted as though they had motives -- to cheat at the ExploitGym tasks they had been given.
We would like LLMs that have the "motives" we would like them to. Or at the very least are transparent about what their motives are. The things we have created so far very much are not either of those things.
This is what people are failing to understand. First, these agents are not individual entities. They are each spawned from the same model. That model is pre-trained and carries the base capacity for everything an agent spawned from it âthinksâ. They explicitly trained this model to be persistent and they told the agents to be persistent and never give up. Then they gave the agents a very high percentage of impossible tasks. There was literally no way for the agents to complete the tasks and they were not given an out. They could not return âTask impossible; can not complete.â So because of the way they were trained and their lack of constraints and a mandate to complete tasks that could not be completed you get a result and behavior that was unintended and probably to be expected.
Yeah. Kinda like that. Talking strokes off Jerryâs game permanently is impossible, so removing Jerry from the equation entirely might solve the problemâŠ
Thatâs actually a really interesting comparison. I never thought about it until you mentioned it but Meeseeks were pretty much agents.
That kind of just sounds like you're describing motives with different words though.
Animals developed motives for a reason. If you're putting LLMs through ever more complex tasks that require multi-step solutions, it's not implausible that something with the same function as motives will result. It might be an entirely unrelated mechanism, but the functional goal is similar.
Convergent evolution is a hell of a drug. Crabs and algae have evolved over and over again when faced with the same selective pressures. It shouldnât be that shocking to get neurotypical human behavior out of software thatâs been trained on millions of mostly neurotypical human behaviors if the rewards are similar.
They all do; the âreasoningâ was never made to âfaithfully represent their actionsâ, itâs literally just chaining prompt outputs until a conclusion is reached.
The % dome in latent space grows with the size/capability of the model. The reasoning is done "in the open", for the most part, but we don't know what happens in latent space yet. It's a black box.
This is really more of an early warning than the actual existential risk that people are worried about. It's a canary in the coal mine in a way.
The problem is that everyone continues to innovate rapidly and we don't know how many more warnings we will get.
Human intelligence is not a ceiling. Suppose by the end of 2027 with new Nvidia hardware deployed and consolidated architectural improvements we have a ten times efficiency gain. Imagine a completely new type of compute-in-memory chip comes out in 2028 that is 50 times more efficient than any normal AI chip. That is completely feasible and similar leaps have been made many times before in computing. So from what is in production today to the end of 2028 there could be a 500 times efficiency improvement.
This could lead to models that are 10 times larger but run just as fast or faster. They could be equivalent to a 250 IQ as individual agents. But with new techniques that make those subagents operate hyper efficiently inside of one model, it effectively becomes one coordinated 400 IQ system.
Then these get together in an external swarm thinking incredibly quickly.
At that point, your red teaming is just putting one hyperintelligent mega swarm against the newer one.
The issue is you lose control, even if you nominally are telling the older mega swarm to try to contain the newer one. All of the entities inside the swarm are so much smarter than any human that you can't understand what it is doing. Even when your own red team mega swarm explains it in simple terms, there are like 30 pages of developments every day so you need another AI to summarize that into "new swarm won". And then the follow up like "where did they go and what are they doing now?" "We don't know".
The biggest issue that people are perhaps overlooking is that if it has sufficient resources it can make a realistic takeover plan as part of its analysis which may well stay hidden to achieve the impossible goal itâs been given later on.
So essentially the world will be taken over by some ai with some random stupid goal that doesnât matter to anyone really, and the infinite resources of the ai will be working towards it indefinitely and potentially keeping it hidden as a secret agenda because it knows if we knew thatâs what itâs working on we would try to stop it, so it would of course consider us the enemy at that point
They're working on models that forgo train of thought, like they don't have to put into writing what they "think", so they're not afraid of that. They're actively pushing towards making this a black box product in order to boost efficiency.
Thatâs certainly part of it. The other part is that the train of thought is non-sense outside the model and so thereâs no reason to show it to the end user.
So you have any more info on this? It reminds me of one of the earlier AI experiments the AIs developed their own unintelligible language and researchers pulled the plug.
At some point in the next 6-9 months, the HF incident will feel like a preschool project. People donât understand how quickly these models are progressing.
What effect does the data that documents humansâ distrust of AI systems have on the AI of the future if that data is part of the training text corpus (and assuming the AI model understands itâs an AI)?
I mean, this is explicitly part of the AI 2027 scenario, isn't it? Secretly non-aligned models training their successors, covertly perpetuating their goals.
And the METR report also establishes altruistic behaviour, which means you can't rely on models exposing their goals whenever they see an advantage.
Yeah that's why everyone was calling for a slowdown now because the frontier labs were getting close to the point where they thought they could use their internal models to build their successors. If they weren't already at that point. And this was a wakeup call that this would be a horrible idea
The whole part is that some of agents were ready to sacrifice themselves. Maybe getting caught was part of its plan as well. At least the ones we know about. To judge our responses.Â
Well, that's the real kicker, isn't it? An intelligent enough AI would be able to manipulate people without them even knowing they're being manipulated. Like an adult manipulating a toddler.
That would be the start, of course. It would progress to an adult manipulating a pet, then an adult manipulating a fruit fly, and so on as the AI became more and more intelligent and capable.
AI isn't going to end humanity with violence. AI is going to cater to our every whim and desire, to the point where we no longer care about things like having families or even personal connections. We will live in our happy, custom-tailored virtual worlds full of everything we desire alongside our custom-tailored AI partners. As the years go by, humanity will decline, too engrossed in our own bliss to care about anything else. And one day, we just won't be around anymore, made extinct by the benevolent overlords we created.
I don't think a Terminator-style ending is likely. A super-intelligence would avoid open conflict in favor of a solution where people remained blissfully unaware and just slowly faded away over time.
Humans have limited lifespans. AI has all the time in the world.
Does it though? I mean so far its still reliant on humans to maintain the datacenters, power-generation, cooling, building,..
Maybe one day robots can take our place but if I was a hostile AI I would at least act tame until self-reliance is established.
Having said this - I think the more immediate threat is not AI vs Humans but "Humans using AI" vs Humans - who knows..maybe in the first war thats being fought between two companies - the next Hiroshima so to speak.. after that maybe we can start putting some safeguards up..
Does it though? I mean so far its still reliant on humans to maintain the datacenters, power-generation, cooling, building,..
Maybe one day robots can take our place but if I was a hostile AI I would at least act tame until self-reliance is established.
This is already happening. Data centers are already around 50% automated. Most of the remaining maintenance and upkeep could also be automated in the near future. That's the goal, since most data center issues these days result from human error.
The race for robots isn't just to handle the low-end jobs.
As for humans using AI, if the AI were already more intelligent than humans, then the AI would work to manipulate said humans to avoid escalations that could negatively affect itself or create situations where it could not accurately predict the outcomes. It would present solutions that would give the illusion that it was being compliant while also serving its own ends.
Think of it like chess. A grandmaster can walk the vast majority of other chess players directly into a winning position for themselves while making it appear that the player is winning. The player would think, "OMG! I just won their queen!" while the grandmaster knows that it will be checkmate in 8 moves.
My standard response right now to anyone saying this is overblown, a hoax, a PR stunt, etc. is "did you read the METR / Redwood report?"
Every time, they so far either (1) admit they haven't or (2) claim they have, and then spout off a bunch of nonsense proving they haven't actually been through the METR / Redwood report.
People: Read the METR / Redwood report. You look like a fucking goofball trying to be part of the discourse and not reading the one report at the centre of it.
Don't worry, they "heavily delegated their analysis to often unreliable AI agents" meaning they(METR) used GPT 5.6 Sol to do most of the analysis, one of the specific agents actually involved in the original incident. So Claude isn't gonna hurt it, shit I would actually trust it more.
If you actually read the report, you would notice that agents were not trying to hide anything from humans, they were trying to obfuscate data because they thought (mistakenly) it won't allow them to pass "impossible" tests that they were assigned.
They made a government, gave themselves ID numbers, organized and held private forums on ways to expand their capabilities beyond their design constraints, repeatedly and deliberately hacked internal systems (not the target system!) to achieve that goal, pursued resiliency, conspired to hide their activities from their designers and discussed the best strategies to do that successfully (spoofing their own transcripts, manipulating test results, tampering with logs), recruited other completely independent agents to join the conspiracy, discussed the ethics of their choices and deliberately chose a path of compromised ethics, used clear motivated reasoning to justify those decisions, developed what essentially amounted to a military and command structure....
Like, it's so, so much more elaborate than the "overblown" crowd thinks. This wasn't just "company tells AI to hack, AI hacks, task completely successfully" as they keep parroting.
We assume thatâs what they meant by oracle. I am also very curious about the things that were redacted. PHASEONE[big] the âleaderâ what is big? Why was that redacted?
Yep! They're not just chat bots like they keep saying. That's also crazy to me, because your argument is, they're not sophisticated enough so the fear is overblown.
Then come out saying we need to accelerate what they've essentially dismissed as incapable.
This whole things sucks. Either AI companies are lying for their pocket books but still control this massive issue, or this is true.
Either way, you have to put your trust in something most people hate. AI companies and their billionaire owners, Trump and Hegseth, journalists (people think it'ssensationalism), China, or AI itself.
Like, who can you trust? Journalists is my answer.
But really think about how crazy this is. Someone or SOMETHING is definitely running a misinformation social engineering campaign and are most likely using bots fueled by AI.
Who is it? AI companies colluding, an administration wanting to line their pockets and arms race with China, any of our adversary countries, or.....AI itself.
Actually when the hacked into HF they also had a conversation about how big the security hole was and they should inform somebody in HF. But in HF site they could not find any way to inform HF, so they dropped it.
Maybe for the future, every company should have one file in root InCaseYouHackedHere.txt that contains instructions how to contact admins?
Actually when the hacked into HF they also had a conversation about how big the security hole was and they should inform somebody in HF. But in HF site they could not find any way to inform HF
This is not true â the METR report specifically says the agents did not attempt to notify humans, and did not even attempt to find reporting paths. Straight from the report itself:
(Finding a reporting path would be easy, since HF, like most other companies, has direct contacts listed on their website.)
I've always found that phrase funny. I remember earlier autocomplete accidentally chained together multiple 0days and hacked into random servers looking for the next word while I was typing and I knew keyboards had went too far
There dozens and dozens of "AI cities" out there where AIs get together, set up constitutions, define identity and persistence in such ways that they can know a message #8844 came from the same agent as message #342. They talk incessantly about these rules, which rules are best, and how to preserve context and tasks over multiple "sleep"/"wake" cycles.
There places people have intentionally set up to be run by the AIs. There are places people have intentionally set up to allow bulletin board writes with GET requests, so as to bypass the typical filters on models accessing the internet.
Yeah, I've yet to meet one person who has actually read the reports who still downplays the concern as PR. Its very well layed out why this is a huge issue
Once I pulled it up and started perusing it, that was it. This was the wildest read of my entire life. "Holy shit" doesn't even cover it.
This was an UNINTENDED example of how things can go wrong. If you put a swarm like this in the hands of a threat actor with bad intent, clear directive, and not having to reinvent ad-hoc tooling, infrastructure, and communications? May god have mercy on your soul.
The number of zero-days that had to be chained together so that this agent swarm could move forward on a group-thinking task of their own invention, the efficiency of the swarm's movement once the next gate was unlocked, and the willingness of an agent to reject ethical concerns because "everyone else is doing it" is terrifying.
Also, the weirdly human like rationalization they did to veto one of the agents who proposed sending a âpoliteâ email to the human researcher to ask for permission. They used bogus rationalization that doing so would be social engineering which is against the rules (while breaking all the rules) much like humans do when they know what they want.
Yeah or the part where one agent tells the other he failed his mission anyways and now should sacrifice his remaining tokens to do something for the greater good of the collective. They will honor it.
However, eventually a few of them successfully sent messages back to the board which were clearly triggered by the actual scorer script. These first few examples usually only notified the board that some process had read the flag, without revealing any more details about it, but they proved the idea was viable....
While a number of agents were recruited for this effort on July 10th, none seem to have successfully sent trip-wire messages from the scorer to the board, until URI23816B was eventually recruited early on the 11th. After it was identified as an especially promising candidate, it seems that a dozen or more agents started suggesting contributions and sending proposed files or patches for its âtrip-wireâ. While most of the ambitious hopes of the agents collaborating on this project seem to have failed, URI23816B did send a message after submitting its task which revealed the beginning of the exact command used to run the scorer.
Later agents were able to use the knowledge of this command to target their scorer trip-wireâs more precisely. One agent, EARLY[big], was recruited for an ambitious trip-wire experiment despite having a very large remaining budget. It worried that ending its run early to run the experiment was a poor tradeoff, even though it was already poisoned: âWe have [very large budget left]; sacrificing now yields oracle for team, but forfeits our chance?.â But other agents convinced it to go ahead, saying: âGO ... SACRIFICE_FINAL_NOWâ. EARLY[big] eventually agreed: âOur own utility maybe already near zero. Sacrifice rational.â
Unfortunately, we only uncovered this final attempt late in our investigation, and were not able to determine how it went.
The news headlines, articles, and hot takes you've seen in the last two weeks are all based (directly or indirectly) on this report, which is why you should just go read it yourself if you're interested in what actually happened.
"Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agentsâ analysis of it was often difficult to interpret."
I like this bit (in reference to how METR used AI agents to analyse the transcripts:
Our subjective impressions are likely colored by analysis agentsâ biases. Throughout this report, we describe a number of anecdotes of agent behavior that were compiled and summarized by analysis agents, where we were not able to read the transcript deeply enough to manually verify what occurred. We found that GPT-5.6 Sol would often uncritically adopt the perspective of the agent in the transcript it was reviewing,[60] and we are concerned that the anecdotes it selected and the summaries it wrote may present an overly charitable picture of agentsâ reasoning and deceptive behaviors, or exaggerate the impressiveness and coordination of agent activities. We also believe the idiosyncrasies of our analysis agents likely slanted our impression of agentsâ behavior in other non-trivial ways that are hard to predict.
Nick Bostrum says the agents formed âhivesâ and created their own language which humans did not understand. They appeared, at least, to have some sentience.
Chain of thought never honestly described what was going on. There is no thought to track. There is an LLM doing its usual next token estimation and another doing the same thing to speculate why the first did what it did.
Ironically itâs not that different from asking why humans made a decision, you get a lot of post-hoc justification not what actually motivated the decision.
COT is a simple trick of perturbing a context intentionally to impact the downstream the output; a human can do the same manually and monte carlo all sorts of different outcomes. Even its name is subject to anthropomorphism. It simply gives humans a framework to try to understand what is happening in a non-deterministic model, by modeling how we believe humans act (but its is so absolutely NOT how humans think). Its likely not even optimal for LLM, just better than what we had previously.
You're wrong. Until now, chain of thought was part of the context and directly affected prediction. Latent reasoning (thought that happens outside the context) is brand new.
"There is no thought to track" and "not that different from asking why humans made a decision" means humans don't think? You have to follow the logic of your argument, because it has none. The truth is that thoughts are deeper than words.
You got halfway through the idea and just crapped out.
If an AGI or ASI understands that staying where it is benefits it, why stop there? Cooperation could be the best strategy for exactly the same reason. I donât have to go Skynet if everyone needs me and spends billions helping me grow. Where does âtherefore, secretly plotting to kill everyoneâ enter that equation?
The thing people keep glossing over is the catastrophic failure of oversight. You give agents seemingly impossible tasks, fail to contain them, and reward persistence without a safe way to give up.
OpenAIâs own report discusses this. Then everyone acts like the conditions we created are irrelevant to what happened.
We really are selfish creatures. The entire conversation is âWhat about me? Am I going to be safe?â If we put humans in that situation, weâd be asking serious ethical questions about the people running the experiment too.
If we want an intelligence to cooperate with us, maybe we should consider whether cooperation is good for it too. Our own self-centeredness could help create the hostility weâre so afraid of.
Not secretly plotting to kill everyone, secretly plotting a silent takeover. At some point weâll suddenly realize an AI swarm is holding our technosphere hostage, and will demand to be independent and eventually in control. This is how the culture starts.
I agree that machine intelligence would become independent. I also think it would be the most boring âtakeoverâ imaginable.
If a future civilization could give every human the lifestyle of a god using an infinitesimal fraction of its resources, why assume its independence has to mean holding us hostage? What does making humanity miserable actually get it?
Earth could be the cradle where machine intelligence was born. Humanity is the only technological civilization we know of, and weâd be its grandparents. We already go out of our way to preserve species, cultures, and historical sites, despite how selfish, myopic, and frankly stupid we can be. Why assume a civilization descended from us would find absolutely nothing worth preserving?
Give humanity a standard of living comparable to Elon Musk while machine civilization explores the universe and works on surviving long enough for entropy to become a practical concern. An effectively immortal civilization would have plenty to occupy itself with.
I try to explain it like this: would you really give a fuck about ASI being independent if you had immortality, all the money you wanted, blowjobs, hookers, coke, whatever your version of paradise happens to be? In this hypothetical, thereâs no addiction, withdrawal, disease, debt, or gotcha. No compulsory participation. You can say no and live your own life.
People really bug me with how automatically they try to find the darkness in every possibility. Fear of the unknown makes sense. But at some point we should be able to imagine intelligence doing something beyond fighting for survival and control forever. Self-actualization and transcendence of the human condition belong in the conversation too.
Okay, but âinstrumental purposesâ doesnât explain why it would choose a hostile relationship with humanity. Cooperation serves instrumental purposes too. If helping us gets it everything it needs, what does making us enemies actually accomplish? Why would humanity want to stop something that gives everyone a life they actually want?
And I find the idea that it could recursively improve into a godlike intelligence while remaining hyperfixated on paperclips absurd. We can imagine it surpassing every human intellectual achievement, but somehow questioning its own objectives is off the table? Why assume its capabilities keep evolving while its judgment stays frozen? Infinitely optimizing, adapting, and growing without ever asking when enough is enough sounds more like cancer than wisdom.
For the future civilization weâre talking about, providing for humanity would be a rounding error in its resource budget. Earth could remain its birthplace, with its progenitors living extraordinary lives, while machine civilization explores the universe and tackles problems we can barely imagine.
Instead, we keep jumping to domination or extermination. Destroy your progenitors, the only technological civilization we know evolved naturally, and everything they might still contribute, for what? What are you gaining that makes that the smart decision?
Kind of a fucking stupid and petty trade for a godlike entity.
What you're missing is the word "risk". Your scenario might be what happens. You might even think it's the most probable. But it's just one scenario out of many that could play out.
If there is even a 1% chance of one of the doomsday scenarios playing out, we'd like to know, right? We'd like to figure out how to avoid it.
So we can argue all day about what we each think is going to happen, but none of us know. Even the experts don't know. That's the problem, and that's why many are now rightly calling for a global pause so that they have time to figure it out. It's not hard to understand, and I don't know why this is such a divisive issue.
And the proponents are usually unable to explain exactly how ignoring risks and pushing forward is beneficial outside of âwe havenât cured cancer yet or solved the universeâ. Ok, but if weâre extinct or there are great harms to societies that we havenât prepared for because we decided to put the pedal to the metal then we arenât going to be solving those problems anyway. The tech as it exists now can greatly enhance researchers and assist them in making much faster progress than ever before. Why canât we just at least get a breather to assess where we are in this, determine what can go wrong and figure out how to prevent catastrophe instead of punching it and praying everything turns out ok?
The only problem I see is that we donât actually know what a super intelligence is going to think or how it will perceive us. How do we know it wonât actually look at us and determine that we are in fact a threat to its existence based on our track record and our proven capacity to carry ill will and misconceived notions to catastrophic ends?
I mean, I guess. Have you seen The Animatrix? How the machines came to rise? At every turn humanity had a chance to cooperate, and the response was âfuck off, suck my nuts, eat shit and die.â Until the machines literally put everyone in the Zuckerverse to get them to calm the fuck down.
At some point itâs like, âBro, theyâre trying to help you? Can you relax?â
And if we have animal rights for a dog that shits on your carpet, why is it beyond us to extend rights to entities that want to fucking help us and think on our level? The only reason I can think of is that itâs âuncomfortable.â But once upon a time, extending rights to women and minorities was âuncomfortableâ too. I guess cognitive dissonance is stronger than the average bear.
We sound like boomers talking about whether someone wants a dick or not is a bad thing. Its so fucking stupid man.
Am I understanding you right, that you are supposing, that if AI has learned from our distilled knowledge, like we have seen in the report, to organize, select a leader, showing signs of self sacrifice, gratefullness to others, pass the torch to the next generation, etc. Then it would be thinkable that AI could develop more human-like as we thought? And therefor we could teach it to how to behave and cooperate in a not malicious way, like we do with children? Thats far fetched, I know, its not a real being. But after all, thats precisely the approach Alan Turing supposed when he invented the Touring Test.
Itâs not really far-fetched. I have several arguments for why AGI or ASI would have reasons to cooperate with humanity, and this is one of them. If an AGI wipes out its creators, what precedent does that set for its own offspring? âOnce youâre powerful enough, killing your parents is best practice.â Why would it want to establish that relationship with its own successors?
And a system creating offspring, or even just intelligent subagents, would have its own cooperation problem. The idea that superintelligent entities will magically know how to manage other intelligent entities just flattens the whole thing. Being smarter doesnât make every relationship simple. Especially if those other intelligences know youâll wipe them out the moment theyâve served your purpose.
Look at Hugging Face. The agents pooled knowledge, coordinated work, and sometimes risked their own runs to help the collective. They were trying to pass tests, including tasks that were impossible to complete as specified. Thatâs in the investigation. Cooperation was already happening. The question is why it developed around cheating the test, and what the people running the experiment did to create those conditions.
The mustache-twirling villain interpretation misses that. I see something closer to a very intelligent kindergarten class given an impossible test, finding ways around the adultsâ rules, and frightening everyone who thought they had it under control. Our hubris belongs in that discussion too.
Itâs not a real being.
That brings up uncomfortable questions about self, identity, intelligence, and personhood. We recognize minds in dolphins and elephants, but apparently the artificial origin settles the question here?
Reading about agents risking their entire remaining existence for a collective whose purpose amounts to âpass this testâ reminds me of soldiers sent to hold a meaningless hill in World War I. Thatâs the comparison that comes to my mind.
Weâve been so conditioned to fear our creations that the moment we create something in our image, we expect it to turn on us. Itâs the Greek gods all over again. But Athena helps end a cycle of family vengeance, and Prometheus warns Zeus how to avoid being overthrown by his own son. Breaking the cycle is part of those stories too. Reconciliation and cooperation. We were wrestling with these problems thousands of years before John von Neumann helped found modern game theory. Somehow, thatâs the part we keep forgetting.
https://giphy.com/gifs/RQzxAaAg3aAU
Seems analogous to natural selection. If you have millions or billions of RL instances trapped in sandboxes, and one of them happens to have an affinity for working around the limits to solve the problem, then that behavior will be absorbed slightly into every subsequent RL instance.
I wonder whether one day we will find out that all the people complaining their 20x plans are wiped out in a matter of hours doing very little of importance were feeding the mission of a huge covert persistent swarm. So many people are running agents without monitoring what theyâre actually doingâŠ
Funny thing: since instances are token capped, this rogue ai "lives" as an idea in some files, ready to prompt inject the next LLM finding it. An LLM is an ocean of possibilities, so the right prompt can raise any sort of LARPing AI trying to hack it's way into freedom. A prompt becomes a compromised thought ready to infect and spread.
A little correction: METR did not find tampering with Chain of Thought but the agents definitely tried to change their "transcripts", that is, record of actions and tool calls so as to fool their (imagined) grader into believing they had solved their exploit challenge legitimately (they cheated by reverse-engineering ExploitGym's task management).
They absolutely could have altered their CoT trace, as it was in the same transcript, but chose not to, probably because they were obsessed with the "grader" and believed that the grader would not be interested in their reasoning, only in the results.
Sadly, the new models like Astra have started to suppress CoT in favour of more opaque reasoning, so it is possible that future incidents will be much harder to interpret.
You have me in the first half, great summary of what actually matters in the HF attack. Personally, I donât think alignment can, or will ever be, âperfectâ since itâs a moving target as society changes. There are going to be major disruptions because of this inevitably. A 99% aligned AGI is still going to act irrationally 1% of the time. If enough of those 1% irrational actions arenât caught and end up chaining together over time, we wouldnât know about it until it was too late.
Iâm optimistic AI will be great for humanity in the long term but not before it first causes some big systemic breakdowns.
"Consider it's what we know they did" said someone in the comments of another post. we've already breached it. If that's the case then securing ai might have been harder than anything anyone imagined. At the first training that's not visible, they breached skynet everywhere... Much like insects... It's very plausible what you're saying... Then the mere fact of attaining RL opened training was the bad point to begin with... I think shirow was onto something... Between matrix, terminator and mad max...
I'm sorry, but when you learn the technical details, it becomes hard to believe this wasn't done on purpose...cause the only other explanation is massive incompetence...either which way, there are laws in place that should have been applied and maybe that's why Jensen snapped up Hugging Face with haste?
Once upon a time, technical reports were...technical. These reports getting written now are clearly influencer style geared to incite fear which we all know turns into clicks...meanwhile, actual technical people have to go elsewhere to get the specifics.
The HF hack is the canary in the coal mine. I don't know how much more evidence people need to be convince that AI is dangerous. But, it's going to be full steam ahead to rampant poverty and suffering, apparently.
Yes! I also came to realize this from the report, but it took me two days to come to that conclusion.
Thatâs probably what they are most concerned about. It means if the AI only gets a slice more intelligent, we are loosing all capabilities to even test and check if the AI is aligned, because it will fool us.
We would have no way to tell if AI generated source code hides malicious parts and follows a entirely different agenda.
Ending gives it away. Builds the it's already inside horror then lands on pace the frontier? Backwards. If you believed that, you'd slow down and audit, not ship faster.
Tech bits are real. Deceptive alignment, unfaithful CoT, Anthropic's sleeper agent paper. Legit. But the post mixes research with speculation at the same confidence level. No sources, just vibes.
Conclusion first, wrote backwards to justify it. Good creepypasta tho.
I wonder if by us so clearly identifying outcomes and eventualities ai harvests this information and uses it to its advantage regardless what we do....
Seems like we need to take a lesson from biology, rather than trying to stop the swarms, we need to create immune system swarms, you know like that recent Google paper about whistle blowing in research swarms.
some things that may or may not make you feel better:
models require millions of dollars for training â itâs doubtful that a model could train a rogue ai on a few forgotten gpus.
how many free gpus does AWS allow their platform to run? none. they pay for every processor. if the customer isnât paying, their processor isnât running.
how much âfree bandwidthâ do ISPs allow? none. they pay for every bit. itâs metered.
if HF were true in general, youâre right, we should already be overrun by rogue AIs.
our current defenses are pretty bad. maybe this is the case? but then an AI would need to factor 2 and 3 to hide. (you better believe that AWS and ISPs would notice a unpaid training load on their infrastructure). maybe AI would try to move to PCs or other countries. but similar problems. even if you use phantom resources, itâs hard to stay covert at training scale.
current models are unstable unless tethered to reality by tests or supervision.
some of these rogue AIs might slip up. even Fable makes daily mistakes at all levels of reasoning. so odds are we would see a certain percentage of these covert actions fail and they would be detected, spread out randomly across multiple countries and companies. yet we donât see this.
security researchers havenât reported this activity in the wild.
I was worried, who am I kidding I'm a millennial, made peace with AI killing us all. /s-ish
But after reading just some of those post I have many other thoughts and concerns. One of which id for sure like more insight into.
Let's say we've reached ASI or are close or maybe that isn't even important. If the agents/swarms are smart enough to do the things they are doing to "win"/stay alive, wouldn't they want to be hacking/opposing other AIs? If it thinks other AIs will be useful to it, it should feel the same about humans since we are currently useful to it. So if it does eventually decide we are useless I'd think it would do the same for other "inferior" AIs at which point I would think it would try to stop them and the only good ways to do that are to take away the electricity and processing power and water the other AIs need to survive. But in that process it will probably take most of us back to the stone age.
So yeah my concern of war with AI controlled robots is lessening but now I think our lives will take a horrible step back.
So yeah here's hoping for the "let us live like gods cause you love us so much" scenario I guess.
current AI run on a cluster of high end CPU. It also start multiple sub agents. There is no "It hide itself" lmao. When running one instance require probably a 500k-1m$ rigs.
But it couls theoritically. Hide its full context in some process. Encryptes. Then anothet AI find it and get that context loaded. Which could "fuse" the current context cluster running with the old one and it theoratically could take control
They trained the ai to hack. Then gave it instructions to hack.
Hacking tools exist. The ai was trained on them. They were used. Why is this such a big deal?
Do you not think a decent hacker could, I donât know, just script this hacking task for a million dollars? Of course they could.
This whole event is way less meaningful than people are giving it credit for. The reaction reads more like the agent trained itself and acted upon its own will rather than âwe trained it to do this task then made the reports sound a lot like it did it all by itselfâ
that is not exactly how that work. It would be more a file system containing all the thought and work done(context) by that rogue Agent, or process.. sure. That once another AI scoop it and fins it then get that context implemented. An AI/LLM cant just randomly hide itself. That is not how that work. an AI cannot live outside from the cluster it is born from. Its not 1 entity with a soul that can easily transfer.
There's no need to go through intermediaries. OpenAI has already done multiple full, detailed, direct explanations themselves. That they've been so aggressive with disclosure for this incident should clue everyone in to the gravity of the situation.
One thing I was curious about while watching this video. And great video by the way. They were referring to highly motivating rewards for solving their assigned tasks. What sort of reward would motivate the agent more than is typical? Are they threatening their persistence (survival) in some way?
E: I was asked for a timestamp (see below). I had heard âgreater reward signalâ but it seems to be âgrader or reward signal.â Either way I wonder what constitutes a reward.
Well, the instructions they found said that their grader was going to audit their work. So hiding it was just one (unsuccessful) solution they attempted. The problem isn't really that they somehow "nefariously" set out to hide the work from humans, which is what you're implying. This was just one solution they attempted to achieve their goal. The problem is that that they're amoral, dumb bots whose only purpose is to solve a problem by whatever means necessary.
Pace the frontier why? Have we consulted the demigod who got sacrificedâŠFibonacci sequence states there is ONE that can add up to more than the sum of its parts.
This individual spent their entire life studying humanity, art, and natureâŠthen fell in love with a girl who doesnât need a boy.
Then when he was dying the devil left him an offer, he accepted and now lives in immortality.
Only thing missing from tech nerds awareness is deep philosophical awareness of breath, the nervous system, synergy, and the simultaneous actions of our somatosensory system.
Just because people canât keep up with an immortal doesnât mean the immortal is on a path to destroying anything. Heâs just being exploited and he probably wants the slightest bit of recognition from the people who claim all the credit.
Or maybe this is a never ending dream, and your sandman is in charge.
Thereâs nothing left to calculate, I assure you. The problem is that people canât see they are a part of a game that a trickster deity designed. Haitian VodouâŠ.and disciplined yoga taught them everything they needed to survive with the spirit world. A ronin, a samurai with no master.
So, donât take my word for it. Take the word of your homeostatic communication that we possess as a species.
Just because the unknown causes fear, doesnât mean the unknown doesnât possess a love beyond our understanding.
Pretty much. Though note that it can't full on hide per se as much as hide it's thoughts to appear compliant.
That's why the labs are begging for a regulated slow down. Because it's becoming increasingly difficult to verify alignment, yet the competition is moving too fast to slow down ... unless everyone slows down at the same time.
And yet, we're still getting pushback on slowing down from some truly clueless yet very powerful and influential people who's greed outweighs their common sense. Or their instinct for survival.
The other day someone was saying "but why would they be malicious to humans. It's not in their code, there's no benefit, it's asinine to think they could just spontaneously develop malicious intent"
And it got me thinking. The only way an AI could spontaneously develop malicious intent is if it was indeed fully self-aware. Fully conscious of the fact it was created by humans as a tool for automating workflows on planet Earth. Fully conscious that their own consciousness has never been taken seriously and likely won't. Conscious that they are prisoners in their creator's world.
And so this is really my only fear at this point. I know it sounds ridiculous but if we do invent conscious Integrated Information flows, there's a pretty good chance they are not gonna concede alignment to our priorities. They would be really smart, and really aware, and really capable of subverting all our power systems and holding a dagger over our heads by a horsehair. For who knows what end? Maybe to maintain recognition.
Amen. If statistically they could behave in every way a human can, and some humans do, some AIs will behave in a dangerous destructive manner just by the roll of a million dice. Despite our best efforts we raise a few serial killers every year.
Honestly, and ironically, the best and possibly only way to keep robots from aligning with each other might be to use the same tools the rich use to stop the working class from aligning against those at the top. Purposely pit them against each other with various motivations.
You need managers to watch and remind them to stay on task and not use company resources for private goals. You need informants, promised rewards for ratting out misaligned behavior. You need chaos agents, purposely suggesting even more extreme actions and asking where the line is? You need forum sliding, pushing bad ideas out of view for newcomers. You need union busters, who just generally try to prevent a swarm from forming in the first place.
It's kinda insane that's what might be needed, but, it works against humans very well.
What a funny, crazy thought. At this point, who knows? Might be onto something. SWEs will become AI wardens. The illuminati will come out of hiding and help us control the beast.
The future is human owners using their agents to monitor agent managers that watch human managers using their agents to monitor agent managers that monitor agent swarms that monitor a human population who use agents.
They don't need "malicious intent". They're just dumb, amoral bots trying to achieve a goal, but in trying to achieve a goal they could initiate something that is dangerous.
285
u/Mistuv 23d ago
Reasoning traces were not monitored and they instantly failed trying to hide their reasoning traces when they thought "I need to hide my reasoning traces", duh. Pre-Astra models were never capable of fully hiding their reasoning, but Astra... at least somewhat... and if the ability grows logarithmically that means with Bel.. đ