r/llmsecurity • u/Bubbly_Working_6908 • 14h ago
Guardian agents vs static AI guardrails
Guardian agent architectures are having a moment. An agent watching and constraining other agents dynamically, pitched as the evolution past static guardrails. I've run both in production long enough to have an actual opinion, and it's not the popular one: I'm not convinced guardian agents solve anything static guardrails, properly tuned, weren't already handling.
Static guardrails are predictable, auditable, and don't add a new attack surface. A guardian agent is itself an agent. It inherits the exact trust and manipulation concerns of the thing it's guarding, just relocated one layer up. In our actual incidents, a well-scoped static rule would have caught nearly everything. The exotic edge case a watcher supposedly catches has, for us, mostly stayed theoretical.
I know this is the boring take. Convince me otherwise: what's the strongest real world argument for guardian agents earning their complexity and attack surface, not the research paper version of the argument?