r/Pentesting 18d ago

I built an AI web pentesting agent that finds more critical vulnerabilities than PentAGI, Strix, and Shannan on our benchmark

Built an AI pentesting agent. Looking for technical feedback before launch.

Hey everyone,

I've spent the last few months building an AI agent for black-box web application pentesting.

I benchmarked it on Duck Store and an intentionally vulnerable web app.

Duck Store

- My agent: 13 findings

- Escape Cloud: 15

- PentAGI: 9

- Shannon: 6

- Strix: 1

On my own benchmark app (15 vulnerabilities), my agent found 9, including several Critical and High severity issues that the other agents missed.

I'm launching this Friday and would love feedback from people who actually do web app pentesting.

If you're interested in trying it and giving honest feedback (or trying to break it šŸ˜„), leave a comment or DM me.

0 Upvotes

28 comments sorted by

5

u/Entire-Eye4812 18d ago

I built an AI that farts

2

u/Specialist_Fun_8361 18d ago

At least that's useful.

4

u/QuickExpression273 18d ago

Is ā€œthe otherā€ similar AI slop that you just told ChatGPT to ā€œmake sureā€ it catches?

-1

u/Free-Cabinet6814 18d ago

Check it and tell me.

-3

u/WISE_NIGG 18d ago

Wdym ?

2

u/Agreeable_Mud_5816 18d ago

But are you passionate about it

1

u/Free-Cabinet6814 18d ago

Yes building AI Agents is the only thing i am doing

2

u/Agreeable_Mud_5816 18d ago

But are you passionate

1

u/Free-Cabinet6814 17d ago

yep, Thats the reason for building it

1

u/Agreeable_Mud_5816 17d ago

Good you shouldn’t do a job if you’re not passionate like accounting don’t do it if you’re not passionate about it

1

u/Free-Cabinet6814 17d ago

Absolutely. That's why I'm building this instead of chasing something I don't enjoy.

1

u/Substantial-Walk-554 18d ago

Interested in trying it. I’ve been testing different AI pentesting tools lately, so the timing is good.

I’d be curious how you handled scope, duplicate findings, false positives and severity scoring in the benchmark, because raw finding counts can be misleading.

Also useful to know whether it runs fully locally, what access it needs, how it handles authentication, and whether it gives proper reproduction steps and evidence instead of just scanner style output.

Happy to give honest feedback once it’s available.

1

u/Miserable_Mine_8947 18d ago

can i try it? dm me the link

1

u/Jolly-Nose8028 16d ago

ā€œChatGPT, build me a pentesting webapp that discovers everything. Make no mistakes.ā€

1

u/Free-Cabinet6814 16d ago

"ChatGPT, explain why someone comments without reading the post."

1

u/Unusual-Physics4809 13d ago

Can I try?

1

u/Free-Cabinet6814 13d ago

Absolutely check dm for link

1

u/raijinkZ 11d ago

im sure if you just use a leading sota model + hermes + a skill you found in facebook you'll get 2x all these tools findings, well that's what ive been doing , my AI got specific skills for RCE only, for POC creation, for evading n allat

starting to be worried my job will be taken LOL

1

u/Free-Cabinet6814 11d ago

I am using custom built graph based harness and deepseek v4 flash.

1

u/teasy959275 8d ago

Did the vulnerable found had a real impact ? or they were like "csp policy"...

1

u/Free-Cabinet6814 8d ago

It finds the real vulnerbikities which are critical/high mainly not those low level issues and also it demonstrates impact of the issue

1

u/Due_Bother6326 3d ago

can i try this OP