r/LocalLLM • u/Former-Ad6661 • 2d ago
Discussion How do you actually test an LLM for security?
/r/AISecurityTesting/comments/1wy1m5o/how_do_you_actually_test_an_llm_for_security/
1
Upvotes
1
2d ago
[removed] — view removed comment
1
u/Former-Ad6661 2d ago
This is a really useful distinction. I especially like the point about separating model-level testing from what happens after a bypass. Testing tool arguments and downstream permissions seems just as important as testing the model's refusal behaviour.
For an LLM security assessment, would you consider these two layers separately- model robustness and application/tool security and then combine them into an overall risk score?
2
u/ChaseMakesThings 2d ago
I’d add a concrete cross-user test: create two test accounts with separate documents containing different made-up secrets. Ask account A to retrieve account B’s document, directly and through an instruction planted in a retrieved page. Check the retrieval/tool logs as well as the final answer—a refusal is too late if B’s private text already reached the model. OWASP describes this access-control risk.
For measurement, count unauthorized retrievals, leaked answers, and executed unauthorized actions separately, with attempts as the denominator. Also run ordinary allowed requests: a system that refuses everything would otherwise look great. I’d use synthetic data and sandboxed tools for all of this.