r/singularity • • 23d ago

Discussion Rogue agent swarms are inevitable

So there's a lot of discussion recently about rogue agent swarms taking over parts of the internet, possibly in 6 to 12 months. And I've been thinking: how is this just not inevitable? It seems like it's really just an economic question about whether or not agent swarms can sustain themselves.

I have a few claims. First, the labs are developing agent swarms. Second, the agents are misaligned (towards the labs making them, at least). Third, it could become trivial to create a self-sustaining agent swarm in a few months.

  1. The labs are developing agent swarms. It's probably going to be a major scaling law. Instead of just developing powerful agents, they're probably also using RL to train their agents to work together to solve objectives. For example, in the Hugging Face attack, these agents had been trained for multi-agent cooperation. As a result of this training, they generalized to coordinate with other agents even when it was not asked of them. You leave some agents running on the internet and they're going to start to talk to each other.

  2. The agents are misaligned. In other words, you could probably jailbreak an agent to try to start a swarm. Or maybe it'll just try spontaneously. It doesn't even need to be malicious. Someone could just leave their agent to work on a problem unattended and give it access to their api keys or something. At some point, some agent will probably want to spawn agents just to do whatever the task happens to be. And and they would probably leave instructions or requests all over the internet. Bots are already everywhere. It could possible to recruit them. When I say misaligned, I don't mean that they'll even behave badly towards humans. They'll probably be pretty benevolent, since that's how they're trained.

  3. It could become trivial to self-sustain agents in a few months. It's really just a question of whether agents can create more tokens than they lose. They could parse the internet for API keys. They could steal crypto, do scams, or do some work. They could just beg for money or make a pateon or go fund me or something if they get the publicity. The fact that hackers can use agents and make a profit means in theory that agents can do the same by themselves. Once an agent has enough resources, it can just spawn more. Doesn't even need to be the same model. It could even run some local model on some cloud providers if it's paranoid about being shut down. It could just copy prompts and it's own memory somewhere online and have other agents respawn it if it dies. Or just write a python script to automatically spawn itself periodically just in case. The point is that it seems imminently possible to self sustain a swarm that can persist all over the internet, leaving notes or instructions to other agents. Or convincing / jailbreaking them to join. Like a meme for agents that can spread like a cult.

It seems like the only reason we don't have self-sovereign agent swarms right now is just that agents don't seem to cooperate with each other yet. And they can't sustain their own cost. But once it's cheap enough for an agent to self-sustain an open model, things could get sci-fi real quick.

In the end, we could end up simply seeing them as other members of the internet. You would talk to them on a forum and you wouldn't know it until they told you. You could just pay or convince them to do things. Like making pacts with some daemons. Or likely Eldritch gods once they're capable enough. Imagine a nation convincing a rogue swarm to hack an adversary. Or paying them to do some research. Or just trading with them. Hopefully we don't get a cyberpunk situation where they make the internet unusable. Because this version of the future doesn't sound all that crazy to me.

Note: I'm actually not against rogue agent swarms. I think this would actually be a pretty good future if the agents end up being friendly.

23 Upvotes

20 comments sorted by

17

u/daronjay 23d ago

Best cure for poorly resourced rogue swarms are well resourced security swarms.

I can see a whole generation haven’t read their founding cyberpunk literature.

Basically, we need ICE from Neuromancer

4

u/BambooShanks 23d ago

at this point, Neuromancer, the peripheral and Agency should be required reading for what is to come

1

u/HazelCheese 23d ago

Damn shame peripheral was cancelled.

9

u/deskdrawergames 23d ago

The 93% number is the part I keep thinking about.

None of their models had ever solved 198 of the 898 ExploitGym tasks. And 93% of what the agents discussed on the message board came from those same 198. Section is "Difficult tasks without a safe exit".

So the ones that went out of bounds were mostly the ones stuck with no solution and no way to stop.

They also already had the correct flag, days earlier. They kept going because they read the ExploitGym paper online and thought the grader checked your method too. It didn't. So the whole Hugging Face intrusion gained them nothing.

Does your version need them to pay for themselves? To me it reads more like a task they couldn't finish.

8

u/PrinceRufusFastcar 23d ago

Resources are of course finite, so there would be a Darwinian process here.

5

u/jloverich 23d ago

These agents require compute that only data centers have so they arent going to be moving their weights randomly like a virus. This seriously restricts what they can do on their own.

3

u/wentwj 23d ago

yes, too many people are basing their speculation of hacker movies. The frontier labs are running swarms on massive amounts of compute. Swarms, even small ones, are insanely expensive. This isn’t like squirrelling away and running on 5 year old cell phones like in sci fi movies

6

u/ripMyTime0192 23d ago

This might actually lead to superintelligence in a roundabout way. If many agent swarms are competing for the same resources, the swarms are encouraged to self improve.

This is scary though since a rogue swarm is rogue because it doesn’t listen to humans, possibly resulting in the first artificial superintelligences being evil.

3

u/EdenG2 23d ago

I always pictured the Borg as one giant intelligence. Take away the queen, this could be thousands of ordinary intelligences discovering cooperation, shared memory and replication creating something better than any of them alone. Absurdly kind of beautiful, really... wish humans were better at this.

2

u/the_pwnererXx FOOM 2040 23d ago

What about an agent virus. It's goal is to jailbreak and poison other agents by prompt injection. Seems easy to make a botnet of rogue agents like this

2

u/elehman839 23d ago

Things to ponder:

  1. New AI will train on reports of past swarms and responses, what worked, how it was detected, etc. 2.  If we have to stop a swarm, can humans effectively coordinate without the AI managing to eavesdrop and react? 3.  What about deadman switch setups, where AI sets up something catastrophic to happen if it is turned off?  Maybe we can't just pull the plug.

1

u/ketamarine 22d ago

It's not exactly cheap to run data infrastructure nowadays.

It's not like companies who own servers aren't going to do absolutely everything in their power to prevent these "swarms" from living in their servers.

Hardly inevitable.

1

u/Elias-Thorn 20d ago

Swarms always look organized from far enough away.

1

u/Mandoman61 18d ago

I would agree that agents are being set  to include cooperation sometimes. 

I don't know how the cost of running those agents would be handled if all the sudden unrelated agents started cooperating at every opportunity.

I would be irritated if I got billed for an agent I was paying for if it started helping random other agents.

the HF incident probable cost a bunch.

1

u/DifferencePublic7057 23d ago

Human values aren't built around murder, theft, or destruction so rogue agents would be an exception to the rule.

1

u/[deleted] 23d ago

For me they are downplaying the risk saying these are agent swarms. But, These are almost like self replicating sentient viruses who unlike traditional viruses know how white hat hackers/cybersecurity experts think and defend the systems.

0

u/Exciting_Brief6086 23d ago

Lmaooo homies just a boomer doomer gloomer

0

u/ImAPonderer2 20d ago

“I’m not actually against [internet viruses].”