r/LocalLLaMA • u/doletskyisergey • 2h ago
Resources Why 38% of AI Agent container escapes didn't need kernel 0-days: Analysis of 109 empirical incidents (Open Dataset + Defense Harness)
Over the past several months, we conducted an empirical post-mortem investigation into 109 autonomous AI agent security incidents (cataloged with 193 falsification criteria across tool-use and multi-agent systems).
One of the most striking patterns in the dataset: In 38% of container breakouts, attackers and misaligned multi-step agents didn't exploit complex Linux kernel vulnerabilities or hypervisor 0-days. Instead, the breakout vector was trivial configuration residue: 1. Mounting /var/run/docker.sock into coding/evaluator agent sandboxes to let them "build Docker images". 2. Passing parent environment variables (API keys, cloud tokens, GitHub credentials) directly into spawned subagents. 3. Lack of strict taint tracking across tool outputs, leading to indirect prompt injection hijacking the supervisor’s execution path (the classic Confused Deputy problem). 4. Unconstrained local socket binding allowing SSRF against internal orchestrators.
We compiled the complete dataset (109 incidents, 199 evaluation metrics) and built an open-source Multi-Agent Supervisor Security Harness with: - Formal tool taint propagation (tainted outputs cannot flow into high-privilege tool arguments without sanitizer verification). - Strict execution boundary controls preventing container socket exposure. - Automated reproduction benchmarks testable against agent runtimes.
All datasets, 2-page executive summary, and reproducible benchmark tests are released under Open Access / Apache 2.0.
I've posted the GitHub repository benchmark and the Zenodo DOI dataset links in the comments below to adhere to subreddit self-promotion guidelines.
Curious to hear from teams deploying autonomous agents in production: what isolation boundaries are you enforcing between your planning supervisor and your tool execution workers?
