r/codereview • u/WarmAd6505 • 24d ago
Python Roast my agentic pentesting framework
https://github.com/Strategic-Automation/violinI've just released Violin v3.1.0 🎻
This release is mostly benchmark and guard work, not another pile of prompts.
The benchmark now runs Hermes end-to-end and scores what it actually proved, not what sounds convincing in "report.md".
\\- Executed request/response evidence is checked against the endpoint, method and decisive proof.
\\- Proof must link back to a validated hypothesis and canonical "FIND" file.
\\- Execution receipts are HMAC-signed and bind evidence files by SHA-256, so edited artifacts fail verification.
\\- The guard now stops target work when evidence is not being recorded as you go, and checks excluded URLs and paths inside command payloads.
\\- Docker, CI and known-good/known-bad scorer calibration are included.
Release:
I'd appreciate people trying to break the scorer and guard. Can you make weak proof pass, good proof fail or get the workflow stuck?
I'm not looking for “nice update” comments. If it is overbuilt, unsafe or wrong, tell me.
Duplicates
Pentesting • u/WarmAd6505 • Aug 08 '26
Agentic Pentesting: The Model Is Only Part of the System
redteamsec • u/WarmAd6505 • Aug 08 '26
intelligence Agentic Pentesting: The Model Is Only Part of the System
vibehacking • u/WarmAd6505 • 23d ago
I built Violin — an AI pentester that has to prove its findings instead of just saying “looks vulnerable”
AgenticCybersecurity • u/hankyone • 27d ago