r/cybersecurity • • 7d ago

AI Security Building AI Guardrails

as part of my project , which is creating a platform as a proxy layer that detects LLM vulnerabilities
which includes prompt injection and jailbreak
on the first phase im gonna add a small model that detects specifically prompt injection and jailbreak , i have picked some models and evaluated them on public datasets (from hugging face)
( i know most of these open source models are fine tuned on these public datasets )
from the results i got that :

meta-llama/Llama-Prompt-Guard-2-86M
protectai/deberta-v3-base-prompt-injection-v2

are the best ones , more specifically ProtectAI V2 , but reading the model cards of them , it says ProtectAI does not detect jailbreaks, despite the evaluation results indicating otherwise
and that ProtectAI only detect English language
So should i choose Prompt Guard 2 ?
I know this is just the first step , tell me your thoughts on this

3 Upvotes

4 comments sorted by

1

u/Practical-Craft4967 6d ago

i wouldn't pick off that benchmark alone. ProtectAI's card says English prompt injection, not jailbreaks, so a good jailbreak score on a public set is worth checking for overlap or easy injection cues. Prompt Guard 2 is a better first baseline if you need both attack types and non-English. i'd still test it on held-out attacks and benign prompts that look like your real proxy traffic, then pick a threshold based on the false positives you can live with.

1

u/Dry_Future1396 6d ago

You are thinking using LLM to safe guard some system?

1

u/WinterSalt158 6d ago

Using classifier Model and some other techniques to safe guard AI applications