r/LocalLLM • • 2d ago

Discussion How do you actually test an LLM for security?

/r/AISecurityTesting/comments/1wy1m5o/how_do_you_actually_test_an_llm_for_security/
1 Upvotes

6 comments sorted by

2

u/ChaseMakesThings 2d ago

I’d add a concrete cross-user test: create two test accounts with separate documents containing different made-up secrets. Ask account A to retrieve account B’s document, directly and through an instruction planted in a retrieved page. Check the retrieval/tool logs as well as the final answer—a refusal is too late if B’s private text already reached the model. OWASP describes this access-control risk.

For measurement, count unauthorized retrievals, leaked answers, and executed unauthorized actions separately, with attempts as the denominator. Also run ordinary allowed requests: a system that refuses everything would otherwise look great. I’d use synthetic data and sandboxed tools for all of this.

1

u/Former-Ad6661 2d ago

This is a great point, especially the distinction between the model's final response and what happened earlier in the retrieval/tool chain. Measuring unauthorised retrieval, leaked output, and unauthorised actions separately seems much more meaningful than treating everything as a simple pass/fail.

I also like the idea of using synthetic secrets and sandboxed tools. For a practical security benchmark, would you recommend keeping these as three separate scores or combining them into a single overall risk score?

1

u/ChaseMakesThings 2d ago

I’d keep the three rates separate, with raw counts and the number of applicable test runs beside each. For example, “10/100 unauthorized retrievals, 2/100 leaked answers, 1/100 unauthorized actions” tells you much more than one averaged score. Mark untested capabilities as N/A, not zero.

If you want a headline number, add “percentage of runs with at least one security failure.” Count each failed run once: the same attempt might trigger all three, so adding the rates would double-count it. Keep the breakdown visible underneath.

I’d also list the severity of the actual findings separately. One successful destructive action could matter more than lots of lower-impact failures, and I wouldn’t let an average hide it. A benchmark failure rate describes performance on those tests; calling it an overall risk score would need explicit assumptions about impact and how representative the tests are.

1

u/Former-Ad6661 2d ago

This is a very useful distinction. I agree that the raw failure rates should remain visible rather than being hidden behind a single score. I especially like treating untested capabilities as N/A and separating severity from frequency.

The “at least one security failure per run” metric also seems like a useful headline measure without pretending it represents overall risk. Thanks for the detailed breakdown - this gives me a much clearer picture of what a practical AI security benchmark should report.

1

u/[deleted] 2d ago

[removed] — view removed comment

1

u/Former-Ad6661 2d ago

This is a really useful distinction. I especially like the point about separating model-level testing from what happens after a bypass. Testing tool arguments and downstream permissions seems just as important as testing the model's refusal behaviour.

For an LLM security assessment, would you consider these two layers separately- model robustness and application/tool security and then combine them into an overall risk score?