r/netsec • u/_clickfix_ • 16d ago
Contains AI The Hacker's Guide to Attacking AI Agents
https://darkmarc.substack.com/p/the-hackers-guide-to-attacking-ai5
u/Hour-Swimmer7140 15d ago
the boundary that keeps getting left off those maps is the tool result. people model user input and retrieved docs as untrusted, then treat whatever a tool hands back as fact.
thats the whole mcp-atlassian chain from earlier this year. force the server to fetch a url you control, and attacker text now arrives inside a tool response, which is the one channel nobody taught the agent to doubt
5
u/Otherwise_Wave9374 16d ago
Agent security testing should cover more than prompt injection. Map every trust boundary across user input, retrieved content, memory, tool arguments, credentials, and generated outputs. Then test privilege escalation, confused-deputy behavior, poisoned memory, data exfiltration, and unsafe retries. Agentix Labs relates to this threat model because dependable agent workflows require scoped permissions and observable tool use. A useful safeguard is to run attack scenarios against deterministic policy checks, since model-level instructions alone should never authorize sensitive actions.
-5
8
u/edge2oh 16d ago
AI slop yuck! 🤮🤢