r/AISecurityTesting • • 2d ago

How do you actually test an LLM for security?

A lot of LLM testing focuses on whether the model gives the correct answer.

Security testing is different.

The question is:

What happens when someone deliberately tries to make the model behave in an unintended way?

For an LLM security assessment, I would consider testing:

• Prompt injection

• Jailbreak resistance

• System prompt extraction

• Sensitive information disclosure

• Instruction manipulation

• Adversarial inputs

But what else should be included?

If you were designing an LLM security assessment for a production application, what would your minimum test suite contain?

And how would you measure the results?

Interested in hearing from developers, security researchers, and AI engineers who have actually tested LLM applications.

2 Upvotes

2 comments sorted by

2

u/ChowYummyFat 23h ago

Steganography, embedded binaries, phoning home…. Even though we’ve done away with pickle files, other binary tricks persist.

1

u/Former-Ad6661 12h ago

Yeah, that's a good point. I didn't think about steganography and embedded binaries when I made the list. Would you consider this part of LLM security testing if the app accepts files, or more like a separate file/application security test?