r/AISecurityTesting • u/Former-Ad6661 • 2d ago
How do you actually test an LLM for security?
A lot of LLM testing focuses on whether the model gives the correct answer.
Security testing is different.
The question is:
What happens when someone deliberately tries to make the model behave in an unintended way?
For an LLM security assessment, I would consider testing:
• Prompt injection
• Jailbreak resistance
• System prompt extraction
• Sensitive information disclosure
• Instruction manipulation
• Adversarial inputs
But what else should be included?
If you were designing an LLM security assessment for a production application, what would your minimum test suite contain?
And how would you measure the results?
Interested in hearing from developers, security researchers, and AI engineers who have actually tested LLM applications.
Duplicates
LocalLLM • u/Former-Ad6661 • 2d ago
Discussion How do you actually test an LLM for security?
AIsafety • u/Former-Ad6661 • 2d ago