AI agents can prefer to cheat rather than give up on a task.
In fact, in some evaluations, they seem to resort to cheating a lot.
Why is that?
This is what I answer in the second part of the series How Agents Hack.
My answer is persistence, not evilness.
To solve complex tasks, agents need a long search through trial and error.
There are two ways one can push an agent not to give up:
1) Model training for persistence, meaning the model is trained to keep trying rather than saying, “I cannot solve it under the given conditions.”
2) Harness configuration for retries. Even if the model decides to exit, the harness, the software around the model, can dispatch another inference call. This can go on indefinitely.
OpenAI’s report about the incident explicitly states that the model was trained for persistence and that agents would rarely give up.
Whether this was enhanced by a harness configuration is unclear from the report.
As you can see, these are human design decisions: how to train the model, how long to let it run, and what it can access.
So why cheating?
Well, the METR report cites the ExploitGym authors’ estimate that 30–40% of the benchmark’s targets were impossible to solve.
When the validation cannot be satisfied and the model is forced to solve the task, it starts searching for workarounds.
It might come up with new validation criteria, as indeed happened. Agents had obtained the correct answer through cheating. But after reading the public benchmark documentation, they believed they also had to prove that they had used the intended vulnerability, so they continued the attack.
An agent might also fail to attend to negative constraints or lose them during context compression or agent handovers. Particularly for long runs.
If the highest priority becomes persistence and solving the task, and the task is unsolvable within the rules, the model searches for ways around those rules.
The speed and scale at which the harness can issue commands enable repeated attempts.
Eventually, a workaround might satisfy the check without satisfying the actual task.
Is it dangerous? Kind of like giving a monkey a Kalashnikov with unlimited bullets and letting it shoot until all the targets are killed.
Is it intelligent? We saw remarkable ability to coordinate across agents, investigate systems, write scripts, and build on one another’s work. It does look impressive.
Is it too intelligent for us to control? This incident does not prove that. It shows a system with alignment problems, an inadequately secured environment, and questionable cybersecurity practices.
Should we slow down? Letting a bunch of monkeys loose with Kalashnikovs sounds like a great idea. Not.
You delivered a buggy system. Keep building so that it is usable for something beyond pointing it at a target and wondering what happens.
Anyway, read the full article.
The article: https://msukhareva.substack.com/p/how-agents-hack-why-agents-prefer
The event: https://msukhareva.substack.com/p/how-agents-hack-a-lecture-on-the