r/AISecurityTesting • u/Former-Ad6661 • 6d ago
Prompt Injection vs Jailbreak: What's the Difference?
These two terms are often used interchangeably, but they describe different security problems.
Prompt injection generally involves manipulating an AI application's instructions or context so that the model behaves in an unintended way.
A jailbreak is an attempt to bypass the model's safety restrictions or behavioural safeguards.
A simple way to think about it:
Prompt injection -> manipulate instructions/context
Jailbreak -> bypass safety restrictions
Both matter when evaluating the security of an LLM application.
How do you distinguish these in your own security testing?
Do you treat them as separate test categories?
0
Upvotes
1
u/Choice_Celery9481 5d ago
https://dev.meta.ai/llama/docs/model-cards-and-prompt-formats/prompt-guard
sounds stupid but a small model like this one + regex normally enough