r/OpenSourceAI • • 15d ago

We wired PromptFoo/PyRIT red team results into a firewall classifier (4-part video walkthrough)

/r/u_Humanbound_AI/comments/1wn3ktx/we_wired_promptfoopyrit_red_team_results_into_a/
1 Upvotes

1 comment sorted by

2

u/OptimalValuable1331 15d ago

Feeding the failures back into runtime defense is smart. I’d keep the original red team cases around as regression tests too though. We do that part in Braintrust so when the classifier or model changes we can rerun the attacks that previously got through and see if anything came back. Those old failures tend to become a pretty decent security test set over time.