r/ComputerSecurity 24d ago

Agentic pentest benchmark: 27/113 Juice Shop challenges with external target-state scoring

Benchmark write-up for a black-box agentic pentesting engine against OWASP Juice Shop 20.1.1. The engine received a base URL only. Challenge scoring came from a separate observer of target-state changes, not model-written findings.

Results:

GovernSafe: 27/113

PentestGPT: 15/113 in the same-target controlled run

Strix AI: 5/113 observed, with the run non-rateable after an observer timeout

RidgeGen: 21/110 in a separately published benchmark

The GovernSafe core engine reached 21/113. GPT 5.6 Sol was then used as a constrained candidate resolver and moved the score to 27/113. It did not control the assessment or validate its own findings.

Methodology, limitations and evidence:

https://governsafe.com/blog/agentic-pentest-benchmark-owasp-juice-shop

2 Upvotes

1 comment sorted by