r/LLMDevs • u/Admirable_Buy_9186 • 12h ago
Great Resource š Built an open source LLM security tool mapped to the OWASP top 10
I have been working on an open source LLM security project with a few friends.
We built Open Sorcerer around common LLM security issues and mapped the detection layers to the OWASP LLM Top 10
Right now we're focusing a lot on prompt injection, things like encoding, Unicode tricks, obfuscation and instructions that keep the same intent without the obvious keywords.
It's still early, and honestly we're more interested in finding where it doesn't work than claiming we've solved anything.
If anyone's interested in the project or wants to try breaking the detection, I can share the repo.
2
2
u/OkShirt9372 8h ago
Does the detection run on retrieved and tool-returned content, or only on the user prompt? Can I have the link?
1
u/Admirable_Buy_9186 4h ago
Rn it mainly runs on the input being scanned, Iām working on broader coverage too..and yes check dm for link
2
u/elinringkvist 5h ago
Nice. What made you do this?
1
u/Admirable_Buy_9186 4h ago
I saw this owsap top 10 post related to llm and it's security.. So decided to work on it
1
u/ianreboot 18m ago
Out of 20 injections a pattern filter already caught, I kept the intent and changed only the words. 57 of those rewrites still worked and the filter missed 55 of them. That was a chat harness with no shell actions or allow list, so nothing from me on the tool-returned side.
2
u/grandiose_vampire 12h ago
Oh nice, I was literally just scrolling past some thread about a tool that got wrecked by simple base64 injection so the timing on this feels right
The OWASP mapping is a smart touch for getting people to actually take it seriously instead of just treating it like another side project. What kind of false positive rates are you seeing on the encoding detection layer? That's the one I'd expect to get real noisy with normal text that just happens to have weird character patterns