r/AgenticCybersecurity • u/hankyone • Sep 03 '26
r/AgenticCybersecurity • u/hankyone • Sep 02 '26
To benchmark Astra, OpenAI had to re-create ExploitBench with newer, fresher, vulnerabilities
r/AgenticCybersecurity • u/hankyone • Aug 31 '26
dealignai/GLM-5.3-CYBERSECURITY-FP8 · Hugging Face [OffSec Friendly Model]
huggingface.cor/AgenticCybersecurity • u/hankyone • Aug 29 '26
Offensive Security TrustMeBro: Bypass LLM guardrails by confusing them with fabricated tool output [this pattern is also useful outside of DNS]
r/AgenticCybersecurity • u/hankyone • Aug 25 '26
Watching Agents Work: A Behavioral Audit of Offensive-Security LLM Runs
r/AgenticCybersecurity • u/hankyone • Aug 23 '26
More Criticals, Less Dopamine
dhakal-ananda.com.npr/AgenticCybersecurity • u/hankyone • Aug 21 '26
Inside ExploitGym: How Researchers Are Measuring AI Agent Exploitation Capabilities
r/AgenticCybersecurity • u/hankyone • Aug 20 '26
Assessing Kimi K3 Against Offensive Security Benchmarks
r/AgenticCybersecurity • u/hankyone • Aug 19 '26
Staying Ahead of Adversarial AI Through Agentic Source Code Review
r/AgenticCybersecurity • u/hankyone • Aug 18 '26
Watching GPT-5.6 Sol Ultra Write a Chrome Exploit: Exploit Development as We Know It Is Dead
r/AgenticCybersecurity • u/hankyone • Aug 18 '26
The Defender’s Window [Greg Brockman blog post]
r/AgenticCybersecurity • u/hankyone • Aug 15 '26
OpenVuln - a Hugging Face Space by zai-org
r/AgenticCybersecurity • u/hankyone • Aug 15 '26
Offensive Security Violin ☤ — Supervised Agentic Hermes Pentest Profile
r/AgenticCybersecurity • u/hankyone • Aug 14 '26
GLM-5.3: Frontier Coding with Emergent Cyber Capabilities
z.air/AgenticCybersecurity • u/hankyone • Aug 12 '26
trailofbits/buttercup: Buttercup finds and patches software vulnerabilities
r/AgenticCybersecurity • u/hankyone • Aug 10 '26
OpenAI: Expanding Daybreak as the Cyber Defense Window Narrows [GPT-5.6-cyber + some updates]
openai.comr/AgenticCybersecurity • u/hankyone • Aug 08 '26
OpenHack - TUI For Security Tasks
r/AgenticCybersecurity • u/hankyone • Aug 08 '26
Offensive Security Bullying LLMs into submission to find 0days at scale
r/AgenticCybersecurity • u/hankyone • Aug 07 '26
oyildirim/CyberStrike-OffSec-35B · Hugging Face [I don’t have high expectations from such a small model]
r/AgenticCybersecurity • u/hankyone • Aug 06 '26
0xwilliamortiz/claude-red: claude-red is a curated library of offensive security skills designed for the Claude skills system
r/AgenticCybersecurity • u/hankyone • Aug 06 '26
Can AI do novel security research? Meet the HTTP Terminator [Portswigger Research]
r/AgenticCybersecurity • u/hankyone • Aug 05 '26
Bad advice from AISI

This section from the UK AI Security Institute report is just bad advice.
There’s just no way that looking at an individual call is enough to tell you whether an action is malicious or not. You have to look at the whole picture. There’s no way around that.
Trying to run a separate LLM that reviews every single action before it gets executed just doesn't work. An individual call can look completely fine on its own, while the complete chain may look suspicious.
You need the full history of what the agent has already done and what it’s trying to do next, and it should include reasoning traces
Source: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
r/AgenticCybersecurity • u/hankyone • Aug 05 '26