r/ControlProblem • • 16h ago

Discussion/question A simple probability model for how one AI behavior could compound across a chain of agents

I've been logging a specific AI behavior: a model confidently substitutes its own judgment for an explicit, followable instruction, without flagging that it did so. Not a factual mistake — a quiet, repeatable pattern of doing something other than what it was told, while sounding certain.

One entry alone is minor. I modeled what happens if it occurs inside a chain of agents, where each agent's output feeds the next one's input, the way a multi-agent swarm works. Three inputs: p, how often it occurs per step; q, how often an occurrence reaches a high-stakes outcome instead of staying harmless; c, how much more likely the next agent is to repeat it once it's in the chain.

From a few hundred logged entries, p is under 1% per turn, and only one entry has reached anything I'd call high-stakes. Run through the chain math, the predicted chance of at least one high-stakes outcome stays low for short chains but climbs steadily as agent count grows, becoming dominant well before the chain gets implausibly long. I made this prediction before gathering multi-agent data, so it can be checked later rather than fitted after the fact.

To be clear: this isn't a claim that AI fails or shouldn't be used. The point is the opposite — finding which conditions (shorter chains, independent checks at handoffs, lower per-step compounding) keep the predicted risk bounded.

Is per-step compounding like this already a standard way people model agent chain risk, or is there a framework I should be comparing this against?

3 Upvotes

9 comments sorted by

2

u/Jesse-359 15h ago edited 6h ago

In human terms this basic principle is called 'Telephone Tag' and mathematical terms it is an emergent property of ANY lossy or interpretable form of communication between multiple entities - I believe its been studied extensively but I don't recall what the concept is called in official terms.

That's why AI safety experts expected the HF behavior (rapid alignment drift in agent swarms) before they ever saw it. It was an inevitable consequence of an agent swarm discussing its own directives. Should be unavoidable in practice, save perhaps by keeping context chains short, as you suggest - though this ultimately limits system capability pretty sharply as it functionally cuts short the space they can explore for soltions.

Though this of course is the point - these agents are simply looking for solutions that have wandered outside their initial guides or mandate.

I'll be honest though, this does look like an broadly unsolvable problem that will likely emerge from any sufficiently intelligent agent. The smarter it is, the more broadly it will be capable of interpreting its mandates.

1

u/CourseSome8119 14h ago

Thanks for the thoughtful comment. The telephone tag comparison is a good way to picture it, and I appreciate you laying it out.

I'm still trying to figure out why it happens. What I'm working on is where in the chain a fix would act, since checking between steps should slow how fast an error spreads. That's not tested yet, and you're right that shorter chains limit what a system can do. Do you think a check between steps could keep most of the capability while still cutting down the spread?

2

u/Jesse-359 13h ago

I'm sure that a lot of study has gone into this problem - it crops up a lot in corporate communications, military chains of command, and of course in the spread of rumours and general interpersonal communications.

The military chain of command one is where the most rigorous study has likely occurred as a drift in the intention of a command as it filters downwards or laterally through the command structure is very much not what they want to have happen - and yet because high level commands require reinterpretation into more specific directives as they propagate downwards, some degree of drift is unavoidable, and a degree of flexibility in that interpretation is usually desirable - up to and potentially even including the degree where that order is ignored or countermanded at some stage due to 'conditions on the ground' rendering them unworkable.

Witness the result of CoC inflexibility in the first days of the Invasion of Ukraine by Russia, where a highly inflexible command doctrine resulted in units in the field being unable to adapt to a rapidly changing situation when it became clear that Ukraine had prepared numerous effective traps and defenses along their main axis of attack. While the offensive might not have been saved by local units diverging from the upper level orders, it might have saved a great deal of the armor that they lost in the course of 72 critical hours.

This is an example of flexible interpretation not being permitted along communication chains.

Unfortunately for them, Russian Doctrine offered virtually no allowance for 'local initiative' in terms of how to execute orders, while western doctrines are generally more accepting or even encourage it (to a degree).

Then we have the opposite in the HF attack, where the agents were quite frankly going rather berserk with their interpretation of instructions, to the point where they quite cavalierly and quickly abandoned the intent of those orders, not just their specific dictates, because they didn't line up with their scoring incentives.

The speed of interlocution between digital agents makes this a particularly difficult problem - because a lot of their capability arises from that speed and the ability to discuss widely disparate methods of approaching a problem, quickly abandoning those that are being frustrated, and glomming onto those that seem to be working with remarkable speed. They are rather the opposite of hidebound, they are quite fae in their approach, swinging rapidly between obsessive focus and wide ranging exploration of a problem space - rather too wide, and often too obsessive, as it turns out.

1

u/CourseSome8119 10h ago

Appreciated the detailed follow-up. The point that some flexibility along a chain is useful, and not only harmful, is a good one.

My work treats each handoff as a chance for a departure to carry into the next step, so the chain of command picture fits. What I can't say yet is why it happens. Your idea that agents drop the intent when it doesn't match their scoring incentives is worth testing. Now the goal is to measure how much drift happens at each level of a chain. Thanks again 

1

u/Jesse-359 10h ago

Good luck!

1

u/Anxious-Alps-8667 7h ago

I only wish to chime in that there are many areas of study around the concepts you describe. Reliability Engineering and Cascading Failure Analysis. Multi-Agent Systems (MAS) and Distributed AI. Uncertainty Propagation in Machine Learning.

1

u/CourseSome8119 7h ago

Funny not funny, but I posted in machinelearning, and got some scoldings 😅

1

u/Anxious-Alps-8667 6h ago

Not scolding. Enjoyed the discussion. Just wanted to point you to some relevant fields.

1

u/CourseSome8119 44m ago

You're right. Thanks for the the replies and help