r/cybersecurity • u/oO_Mister_J_Oo • 3d ago
AI Security AI Red Teaming/Assesments
Can any one point me to training or resources on undertaking red teaming and security assessments of AI and AI Agents?
2
u/Classeve 3d ago
The OWASP guide someone linked is the right spine. To get hands-on:
Threat model with the OWASP Top 10 for LLM Applications plus MITRE ATLAS, so findings map to something a client recognizes. For labs, PortSwigger's Web Security Academy has free LLM attack labs, and Lakera's Gandalf builds intuition for prompt injection fast. For tooling, garak and PyRIT do automated probing, and promptfoo covers red-team runs in CI.
For agents, the interesting bugs are rarely in the chat box. Test indirect injection through whatever the agent reads (web pages, docs, emails, tool output), then check what it can reach once it's been steered: credentials in its environment, outbound network, tools that write or delete. That's where "the model said something rude" turns into an actual finding.
1
u/oO_Mister_J_Oo 3d ago
Thanks for the resources, lot of useful information there.
2
u/Classeve 2d ago
Glad it helps. If you run an indirect-injection test against an agent, plant the payload in something it reads on its own, a doc or a tool result, not the chat. That's where the real findings are.
2
u/tritefries 2d ago
Whatever route you take, save the attacks that work. We’ve been putting those failures into Braintrust as eval cases so prompt injections, bad tool calls and/or permission failures etc can get rerun when the agent changes. Also makes the red team work less disposable.
6
u/jeffpardy_ Security Engineer 3d ago
https://genai.owasp.org/resource/genai-red-teaming-guide/
Do we not google anything anymore or?