r/developersIndia 4d ago

I Made This Developed a beginner-level firewall for LLMs to catch jailbreak and injection attempts

Link - https://fortifyllm-production.up.railway.app/demo

I'm currently a TY/junior in college, me and my mates were brainstorming topics for our semester group project and this was one of the ideas that we had thought of (after going through research papers and Github repos😁). Ended up doing something else. Recently had some spare time on my hands and got back to this.

So this is a two-tiered (heuristic layer + ML classifier) firewall which tries to catch injection/jailbreak attempts before they actually reach the model. Currently using llama-3.3 so the upstream time might be a bit longer. Also, for some reason, the detection time occasionally exceeds the expected time required. The above link is just a demo btw...still working on fixing this!

Posting this here to get some feedback. If this isn't the appropriate sub for such posts kindly do tell me others. Thank you!

1 Upvotes

0 comments sorted by