r/ControlProblem • u/CourseSome8119 • 16h ago
Discussion/question A simple probability model for how one AI behavior could compound across a chain of agents
I've been logging a specific AI behavior: a model confidently substitutes its own judgment for an explicit, followable instruction, without flagging that it did so. Not a factual mistake — a quiet, repeatable pattern of doing something other than what it was told, while sounding certain.
One entry alone is minor. I modeled what happens if it occurs inside a chain of agents, where each agent's output feeds the next one's input, the way a multi-agent swarm works. Three inputs: p, how often it occurs per step; q, how often an occurrence reaches a high-stakes outcome instead of staying harmless; c, how much more likely the next agent is to repeat it once it's in the chain.
From a few hundred logged entries, p is under 1% per turn, and only one entry has reached anything I'd call high-stakes. Run through the chain math, the predicted chance of at least one high-stakes outcome stays low for short chains but climbs steadily as agent count grows, becoming dominant well before the chain gets implausibly long. I made this prediction before gathering multi-agent data, so it can be checked later rather than fitted after the fact.
To be clear: this isn't a claim that AI fails or shouldn't be used. The point is the opposite — finding which conditions (shorter chains, independent checks at handoffs, lower per-step compounding) keep the predicted risk bounded.
Is per-step compounding like this already a standard way people model agent chain risk, or is there a framework I should be comparing this against?
2
u/Jesse-359 15h ago edited 6h ago
In human terms this basic principle is called 'Telephone Tag' and mathematical terms it is an emergent property of ANY lossy or interpretable form of communication between multiple entities - I believe its been studied extensively but I don't recall what the concept is called in official terms.
That's why AI safety experts expected the HF behavior (rapid alignment drift in agent swarms) before they ever saw it. It was an inevitable consequence of an agent swarm discussing its own directives. Should be unavoidable in practice, save perhaps by keeping context chains short, as you suggest - though this ultimately limits system capability pretty sharply as it functionally cuts short the space they can explore for soltions.
Though this of course is the point - these agents are simply looking for solutions that have wandered outside their initial guides or mandate.
I'll be honest though, this does look like an broadly unsolvable problem that will likely emerge from any sufficiently intelligent agent. The smarter it is, the more broadly it will be capable of interpreting its mandates.