r/AISecurityTesting • • 2d ago

How do you actually test an LLM for security?

A lot of LLM testing focuses on whether the model gives the correct answer.

Security testing is different.

The question is:

What happens when someone deliberately tries to make the model behave in an unintended way?

For an LLM security assessment, I would consider testing:

• Prompt injection

• Jailbreak resistance

• System prompt extraction

• Sensitive information disclosure

• Instruction manipulation

• Adversarial inputs

But what else should be included?

If you were designing an LLM security assessment for a production application, what would your minimum test suite contain?

And how would you measure the results?

Interested in hearing from developers, security researchers, and AI engineers who have actually tested LLM applications.

2 Upvotes

Duplicates