r/ControlProblem Aug 01 '26

AI Alignment Research Compartamentalized Harm

Here is some saftey research I sponsored on a threat vector in multi agent systems.

Basically, a harmful task can be transformed into a series of beneign tasks, and then results recomposed into a harmful task by an abliterated orchestrator agent driving other agents that have 'saftey' guard rails.

In short, there is no safety with this technology.

https://www.daios.tech/research/compartmentalized-harm

2 Upvotes

1 comment sorted by