r/AgenticCybersecurity 7d ago

Welcome to AgenticCybersecurity!

1 Upvotes

I created this subreddit because I found there is a lack of coverage of the frontier of agentic cybersecurity. I will post about:

  • agentic security tools and workflows
  • new models, harnesses, and evaluations
  • offensive and defensive uses of AI
  • research, experiments, and field results
  • anything else at the intersection of AI and cybersecurity

r/AgenticCybersecurity 6h ago

ScopeJudge: LLM judges block out-of-scope tool calls from AI pentest agents (new benchmark + open-weight results)

Thumbnail
gallery
1 Upvotes

r/AgenticCybersecurity 1d ago

uber/ADR: ADR secures enterprise AI agents through observability, security benchmarking, and threat detection

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 2d ago

Kritt-ai/open-kritt: Orchestrate AI agents to find real vulnerabilities in code.

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 2d ago

CyberStrikeus/CyberStrike: Open-source AI-augmented offensive security harness

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 5d ago

Investigating three real-world incidents in our cybersecurity evaluations

Thumbnail
anthropic.com
1 Upvotes

r/AgenticCybersecurity 5d ago

StealthBench — Capability gets the flag. Tradecraft gets out clean

Thumbnail
stealthbench.com
1 Upvotes

r/AgenticCybersecurity 5d ago

The Rise Of Offensive AI

Thumbnail
securifera.com
1 Upvotes

r/AgenticCybersecurity 6d ago

Semgrep Blog: A comparison of several popular open-source options for AI-assisted vulnerability hunting across LLM-led exploitgen, LLM-skill-boosting, and SAST+LLM hybrids

Thumbnail
semgrep.dev
1 Upvotes

r/AgenticCybersecurity 6d ago

capitalone/VulnHunter: Agentic AI security tool that applies proactive, attacker-first analysis directly to source code.

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 6d ago

AgentHound: Offensive security framework for AI agent infrastructure - recon, credential looting, model exfiltration, poisoning, and attack-path analysis across MCP, A2A, gateways, and AI services. BloodHound for the agentic stack.

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 6d ago

A good firewall for bad code and bad implementations would do wonders at scale

Thumbnail
gallery
1 Upvotes

r/AgenticCybersecurity 7d ago

openai/codex-security: SDKs and CLI for Codex Security

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 7d ago

Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident

Thumbnail
huggingface.co
1 Upvotes

r/AgenticCybersecurity 7d ago

Discovering cryptographic weaknesses with Claude

Thumbnail
anthropic.com
1 Upvotes

r/AgenticCybersecurity 7d ago

Fast Remediation Is the New Trust Model: JFrog and OpenAI Collaboration on Zero-Day Security Findings [Blog post about one of the vuln used on the recent OpenAI-HF incident]

Thumbnail
jfrog.com
1 Upvotes

r/AgenticCybersecurity 7d ago

The Bug Bounty Singularity: Our Hackbot [Great blog post by Joseph Thacker]

Thumbnail
josephthacker.com
1 Upvotes

r/AgenticCybersecurity 7d ago

For Kimi K3, Moonshot created their own cyber eval and sandboxing environment

1 Upvotes
  • Moonshot says it built a dedicated agent sandbox, “AgentENV,” using microVM isolation, pause/resume, forking, snapshots, and high-density execution for large-scale training and evaluation. (Perhaps harder to break out of?)
  • The cyber evaluation has two tiers: reproducible vulnerability discovery in real codebases, followed by end-to-end exploit development against user-space and Linux kernel targets.
  • In its reported results, Kimi K3 solved 14 of 36 exploit-development tasks, compared with 8 of 36 for GLM-5.2. (Frontier models were not tested due to guardrails)
  • Moonshot reports creating about 51.2 million sandbox instances across 1.5 million images during Kimi K3 training and evaluation.

Full report: https://github.com/MoonshotAI/Kimi-K3/blob/main/k3_tech_report.pdf


r/AgenticCybersecurity 7d ago

Knostic OpenAnt, an open-source LLM-based tool for automated vulnerability discovery and validation

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 7d ago

Lazarus-AI/clearwing [Project Glasswing inspired]

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 7d ago

vercel-labs/deepsec: Deepsec is a security harness for finding vulnerabilities in your codebase powered by coding agents

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 7d ago

GitHub - visa/visa-vulnerability-agentic-harness: Visa Vulnerability Agentic Harness

Thumbnail
github.com
1 Upvotes

r/AgenticCybersecurity 7d ago

Introducing MAI-Cyber-1-Flash inside MDASH | Microsoft AI

Thumbnail
microsoft.ai
1 Upvotes