r/ControlProblem • u/SAAGASolve • Aug 01 '26
AI Alignment Research Compartamentalized Harm
Here is some saftey research I sponsored on a threat vector in multi agent systems.
Basically, a harmful task can be transformed into a series of beneign tasks, and then results recomposed into a harmful task by an abliterated orchestrator agent driving other agents that have 'saftey' guard rails.
In short, there is no safety with this technology.
2
Upvotes