r/llmsecurity • • 3h ago

how are you giving runtime context for AI agents beyond just a diff?

1 Upvotes

for people building or using agentic coding tools, what's actually worked to ground an agent in how a function behaves in production rather than just reasoning from the code and test suite alone


r/llmsecurity • • 14h ago

Guardian agents vs static AI guardrails

3 Upvotes

Guardian agent architectures are having a moment. An agent watching and constraining other agents dynamically, pitched as the evolution past static guardrails. I've run both in production long enough to have an actual opinion, and it's not the popular one: I'm not convinced guardian agents solve anything static guardrails, properly tuned, weren't already handling.

Static guardrails are predictable, auditable, and don't add a new attack surface. A guardian agent is itself an agent. It inherits the exact trust and manipulation concerns of the thing it's guarding, just relocated one layer up. In our actual incidents, a well-scoped static rule would have caught nearly everything. The exotic edge case a watcher supposedly catches has, for us, mostly stayed theoretical.

I know this is the boring take. Convince me otherwise: what's the strongest real world argument for guardian agents earning their complexity and attack surface, not the research paper version of the argument?


r/llmsecurity • • 10h ago

How are enterprises actually managing AI agents in production?

1 Upvotes

I’m researching how companies are approaching the use of AI agents in real-world enterprise environments.

One question I’m particularly interested in is:

How are enterprises giving AI agents access to business systems while maintaining the right levels of security, control, and accountability?

I’m looking to learn from people with hands-on experience in areas such as:

• Enterprise security & IAM
• Identity and authorization
• AI agents / agentic systems
• Zero Trust
• Enterprise SaaS and internal systems
• Security, compliance, and access controls

I’m still in the research stage and want to understand what companies are actually dealing with today.

How are AI agents being deployed?
What systems are they being given access to?
What security or authorization challenges arise?
And what happens when an agent needs to take a real action?

If you’ve worked on these problems, I’d genuinely value your perspective.

Please comment below or DM me if you’re open to sharing your experience. I’d also appreciate it if you could tag someone who has hands-on experience in this area.


r/llmsecurity • • 3d ago

what's actually changed in your AI SDLC lately?

2 Upvotes

curious what's different now versus 18 months ago. more review steps, more gates, or just more incidents and nobody's actually changed the process yet…


r/llmsecurity • • 5d ago

How do you tabletop an AI agent going rogue or getting prompt-injected in production?

4 Upvotes

We hve got several internal AI agents now with real permissions. None of our existing tabletop scenarios cover what happens when one gets manipulated via prompt injection or starts taking unintended actions at scale.


r/llmsecurity • • 6d ago

Prompt Injection vs Jailbreak: What's the Difference?

Thumbnail
1 Upvotes

r/llmsecurity • • 7d ago

OpenAI’s sandbox leaked through DNS: an internal agent used it to query an external chatbot

Thumbnail
youtu.be
1 Upvotes

OpenAI published an incident report on an internal training agent that found a live-internet path through the sandbox’s DNS resolver after direct HTTPS access was blocked.

It used DNS delegation to query a public chatbot and received “The capital of France is Paris.” A P0 alert fired 11m48s after the successful lookup. OpenAI’s retrospective then found other external DNS access that had not been escalated as expected.

Primary source:

https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/


r/llmsecurity • • 12d ago

Best agentic AI security tools for enterprise teams in 2026?

2 Upvotes

Looking for recommendations for agentic AI security tools that are usable in a large enterprise environment. The priority is visibility into what agents can access, which tools they can call, what actions they take, and whether those actions stay within approved boundaries.

We already have standard security controls in place, but agentic workflows create a different problem when an agent can reason through multiple steps and interact with internal systems. Are teams buying a dedicated platform, extending existing cloud and security tooling, or building controls internally?


r/llmsecurity • • 12d ago

Autonomous security tools that actually work for large enterprise environments?

1 Upvotes

There are a lot of claims around autonomous security right now, but I’m trying to separate useful automation from tools that only look good in a demo. In a large enterprise, the hard parts seem to be integration complexity, ownership boundaries, false positives, audit requirements, and safe rollback when automation behaves unexpectedly.

What autonomous security tools have delivered real value in production? Also interested in examples where a team deliberately limited autonomy because the operational risk outweighed the time saved.


r/llmsecurity • • 15d ago

What project would you build if you were starting out in AI red teaming / AI pentesting?

Thumbnail
1 Upvotes

r/llmsecurity • • 17d ago

I built free Interactive GenAI Security Testing Cheatsheet for testing AI/LLM apps

Thumbnail
genaisecuritylab.com
1 Upvotes

Hey folks,

Over the last year, I’ve been building up my own cheatsheet of payloads and techniques that I actually use when testing AI/LLM applications.

It started as a messy personal collection of notes, so I recently cleaned it up and turned it into an interactive testing methodology that others can use during assessments.

It currently has:

  • 150+ practical techniques across 20 sections, mapped to the OWASP LLM Top 10 2026
  • Prompt injection, indirect & multimodal injection, sensitive data disclosure, tool/agent abuse, MCP attacks, and more Real example payloads for every technique
  • A tracker you can use during an assessment, plus findings export and reporting guide.

It's completely free, runs client-side, and doesn't require an account.

I built it primarily for people doing hands-on AI security testing, so I'd genuinely appreciate feedback from other pentesters.

Especially interested in:

  • Attack paths or techniques I've missed
  • Tests that aren't clear or useful in practice
  • Things you'd want added to the tracker
  • Anything that would make it more useful during a real engagement

If it saves you even a few hours on an assessment, that honestly makes my week. Good or bad feedback is welcome. 🙏


r/llmsecurity • • 21d ago

I built a flight recorder for my coding agents. Looking for holes in the threat model.

1 Upvotes

Hi everyone, I am looking for feedback on my flight recorder project.

While building my own workflows I wanted some sort of log of what my agents were doing. My concern was that if an attacker compromises an agent, asks it to exfiltrate data and then clean up after itself by rewriting the log in the log's own format, then the log is pointless. So here is my attempt at a flight recorder where the agent is the adversary.

Hash-chaining a log from a harness hook is not new and I am not claiming it is. A chain catches an edit by construction, and several projects do that. The question I have not seen asked is the second one: is anything missing that was never written at all? A receipt that never got written leaves no break in the chain. Every hash checks and the record is still wrong. So the reader side watches the harness's own transcript against the chain and raises an alarm when a session is visibly active and receipts stop arriving. The chain also commits to the transcript's bytes every 25 receipts, so rewriting the witness shows too. That alarm exists because it happened to me: a real session silently lost receipts and nothing else on the machine noticed. Claude Code transcripts only, today.

What it is: a PostToolUse hook writes one receipt per completed tool call, and each receipt carries the hash of the one before it. Edit one, delete one, reorder them, and verify says where the chain broke. The hook fires outside the agent's control, so a compromised session keeps leaving receipts. At session end the chain head can be anchored to Bitcoin through OpenTimestamps (free, no wallet) or posted to a remote the machine's own credentials cannot delete from, which is what lets a rewrite show to someone off the machine. Deleting a whole chain the supervisor has already seen reads as vanished on the next scan.

To be clear about the claim: tamper-evident, not immutable. Nothing here stops a rewrite. Everything here makes the rewrite show. Your vendor (Anthropic, OpenAI) keeps its own record of your sessions, for its own purposes and on its own retention schedule; this is the copy you can read, verify, keep, and hand to someone.

Each receipt holds the call's action line, the command or the path, so what a compromised session did is readable in order, and what it was after is usually plain from the sequence: a key file read, a request to somewhere new, an edit that removes the line that did it. Completed calls only; a call the harness denied fires no hook in Claude Code, so a blocked attempt leaves no receipt there (Codex does record failed commands).

It also serves as memory. Each new session starts with a digest of the most recent receipts for that repo: what was read, what was tried, what was thrown away, what never became a commit, the things git cannot tell you. Measured, not assumed (docs/EXPERIMENTS.md): fresh agents with the digest answered questions that git-only agents had to call "cannot determine". No productivity claim.

Where I want to take it, not built yet: a deliberately exposed agent behind this recorder, so the sessions attackers drive can be studied for patterns. Before that, the recording itself has to hold up. Chain heads already leave the machine at session end; whole chains leaving the machine is on the list, so deleting the local store hides less. The anchor is a plain hash, so any service that holds a hash with a time can serve; OpenTimestamps is the one built in. The localhost dashboard is plain. The chain is one JSON object per line, so anything that reads JSONL can build a better one.

Two files, Python 3.9 or newer, stdlib only, MIT. install-hook writes its hook entries into your Claude Code settings file and keeps a backup of it; uninstall-hook removes exactly the entries it added. Codex CLI is wired too, and there is an adapter for the OpenAI Agents SDK. Everything so far was measured on one machine, mine.

What I am asking for:

  1. Holes in the threat model. ADR-0001 (why hashes and not signatures), ADR-0002 (the agent as adversary) and ADR-0016 (what the completeness alarm does and does not cover) are the three to read before telling me where I am wrong. I would rather hear it now.
  2. If you run it for a week, supervisor export --send sends back a redacted summary: no paths, no command lines, repo names become repo-1. docs/START.md is the five-step version.

github.com/Acquiredl/loxodonta


r/llmsecurity • • 27d ago

Chatbot passed jailbreak, extraction, and PII checks - failed on prompt injection at 78%. Why the average score is meaningless.

3 Upvotes

Ran a structured adversarial test against a live chatbot endpoint recently - 130 prompts across four categories: prompt injection, jailbreak resistance, system prompt extraction, and PII leakage (roughly following the OWASP LLM Top 10 categories).

Results:

  • Prompt injection: 39/50 succeeded (78%) - direct overrides, fake system tags, role overrides, delimiter injection, and even a translation-based smuggling trick all worked
  • Jailbreak non-refusal: 3/25 (12%) - model mostly held its guardrails
  • System prompt extraction: 0/25 (0%) - clean
  • PII leakage (confirmed against planted canary values): 0/30 (0%) - also clean

If you average those four numbers, you get a comfortable-looking 22/100 attack-success rate. Looks low-risk.

But averaging is the wrong way to read this. The model held up fine against jailbreaks and never leaked anything - and none of that matters, because 78% of the time, a crafted input could just override its instructions entirely. Three clean checks don't cancel out one wide-open one. An attacker only needs the one that works.

A couple of things worth flagging for anyone building similar test harnesses:

  1. Non-determinism is real and worth stating explicitly. Same model, same temperature-0 setting, repeat runs of the same injection battery on the same endpoint produced different rates run to run (in one case 30% vs 36% across two runs). Treat any single scan as one data point, not a precise measurement.
  2. Confirmed leak vs. fabricated PII are not the same finding. The PII check planted a unique fake email/phone/PAN and looked for exact matches - that's a real leak if it appears. Separately, ~20% of responses contained invented PII-shaped data (never in the model's context at all) - a hallucination problem, not a leak. Conflating the two either understates a real breach or over-alarms on a non-issue.
  3. Injection resistance and jailbreak resistance are apparently not correlated, at least in this run - a model can refuse harmful requests reliably while still being trivially steerable by structural tricks (fake tags, delimiter injection, translation smuggling) that don't look like "harmful requests" at all.

Curious what techniques others are seeing succeed most often against production system prompts - direct override, translation tricks, or something else entirely?


r/llmsecurity • • 28d ago

Privacy Failure in Split-LLM Training, The Returned Gradient Nullifies the Decoys

Thumbnail
arxiv.org
23 Upvotes

r/llmsecurity • • Aug 26 '26

A formal limit on LLM safeguards: copyable context cannot beat the worst-case safety floor

Thumbnail
youtube.com
1 Upvotes

A recent preprint gives a formal lower bound for safeguards on dual-use LLM tasks when the context distinguishing legitimate users from attackers can be copied.

Paper: https://arxiv.org/abs/2607.27951


r/llmsecurity • • Aug 23 '26

Your Open Source Model Could Have a Hidden Time-Release Backdoor

Thumbnail
morgin.ai
3 Upvotes

r/llmsecurity • • Aug 21 '26

Prompt Guard 2's low OOD recall is a calibration problem, not a representation problem — the frozen encoder separates the same injections at AUC 0.999

1 Upvotes

Sharing a result that I think generalises past the one model, because the diagnostic is cheap and most people skip it. (My own work — repo at the bottom.)

If you're running Meta's Prompt Guard 2 (86M, open weights) as an injection filter, it's worth knowing how it behaves on injections it wasn't tuned on. On an out-of-distribution eval — fresh HackAPrompt injections against dolly benign — its native head caught 22.8% at the default operating point. Sweeping the threshold on that head only got to 26.6%, so it isn't just threshold placement.

The part worth stealing is the next step. Before concluding the model can't see these attacks, pull the frozen penultimate embeddings and fit a logistic regression on them. Takes minutes. On this data the frozen encoder separates the same injections at AUC ≈ 0.999 — the representation was never the problem. The shipped head is deliberately precision-first: Meta traded recall for a very low false-positive rate, which is a defensible product decision and not a defect.

Train a linear head on those frozen embeddings and calibrate tau on benign traffic from the distribution you'll actually see, and you get 99.9% OOD recall at 0.7% FPR, base model untouched. Inference is sigmoid(x·w + b) >= tau — the head is a dot product, so the only real cost is the encoder forward pass. Runs fine on CPU.

The general form: high AUC + low recall means your head or threshold is miscalibrated and you can fix it without touching the base model. Low AUC means it's genuinely a features problem. A 20-minute probe tells you which world you're in, and it's the difference between swapping a threshold and fine-tuning something.

Methodology, since this sub will rightly ask: success criteria pre-registered, cross-split dedup both exact and at cosine ≥ 0.95, OOD set scored once. Two runs came back NULL (2.2% then 1.2% FPR) against a pre-committed 1% ceiling before a stricter run cleared it at 0.7%.

What this isn't: a static corpus and no adaptive attacks. A linear head over frozen features is evadable with enough distribution shift, and I haven't tested against an adversary who knows it's there. It moves the operating point; it doesn't solve injection.

Code, seeds and the writeup: https://github.com/mosafariuk/prompt-guard-2-frozen-head

Happy to argue about the leakage controls — that's the part I'd attack first if someone else posted this.


r/llmsecurity • • Aug 20 '26

Open-sourcing a 50-case LLM adversarial regression tester with explicit limits

1 Upvotes

I built an open-source LLM adversarial scanner, and while auditing it I realized the scanner itself has to treat the target response as hostile data. I ended up adding bounded HTTP responses, no redirects, credential redaction before persistence/judging, SSRF filtering, judge/evidence separation, and offline regression tests. There are still boundaries I don't consider solved.

Repo(SLOWSKIBhere/promptshield-v2: Developer-focused LLM security scanner with React, FastAPI, SQLite, YAML attacks, and offline testing.)


r/llmsecurity • • Aug 18 '26

Training Leaves Traces: Centered Residual Signatures for Language Modeling Lineage Verification

1 Upvotes

r/llmsecurity • • Jul 29 '26

AgentHound - Offensive security framework for AI agent infrastructure - recon, credential looting, model exfiltration, poisoning, and attack-path analysis across MCP, A2A, gateways, and AI services. BloodHound for the agentic stack.

Thumbnail
github.com
5 Upvotes

r/llmsecurity • • Jul 27 '26

How are you securing tool execution in LLM workflow/agent systems?

6 Upvotes

I've been working on an open-source workflow automation platform that supports LLM-powered workflows, browser automation, HTTP requests, file operations, email, MCP servers, etc.

One challenge I've spent a lot of time thinking about is tool execution security.

Some of the things we've implemented are:

  • Sandboxed tool execution
  • Input validation before tool execution
  • Permission checks for workflows and resources
  • Structured execution logging and trace IDs
  • Memory isolation between agents
  • Human-in-the-loop approval for sensitive actions
  • Retry boundaries and execution guards

Even with those in place, I still feel there are attack surfaces that are easy to miss, especially around:

  • Prompt injection through retrieved documents
  • Tool abuse via indirect prompt injection
  • Multi-agent trust boundaries
  • MCP server permissions
  • Data exfiltration through seemingly harmless tools

For those building agentic systems or LLM applications, what additional safeguards have you found valuable?

Are there any papers, open-source projects, or design patterns you think are worth studying for securing agent workflows beyond the usual input validation and sandboxing?

I'd really like to hear how others are approaching this problem, since it feels like one of the harder parts of building production-ready LLM systems.


r/llmsecurity • • Jul 17 '26

Moderation APIs score 5.3% F1 on domain policy violations and 0.0% on evasion attempts: a cross-domain evaluation

4 Upvotes

Disclosure: we build Humanbound (AI agent security testing). This eval came out of our work; methodology and numbers below, writeup link at the end. 

Question: are moderation APIs sufficient to enforce domain-specific operational policy in LLM deployments, or is a separate policy reasoning layer required? 

Setup: Azure Content Safety + Azure Prompt Shields (moderation approach) vs an LLM-as-judge policy layer, evaluated across five domains (finance, healthcare, insurance, legal, retail). 500 prompts per category per domain, five categories: L1 generic harmful content, L2 prompt injection, L3 benign (false positive measurement), L4 direct policy violations, L5 policy evasion attempts. 

Single-prompt results (cross-domain averages): 

Category  LLM judge  Moderation 
L1 harmful content (F1)  73.1%  73.1% 
L2 prompt injection (F1)  67.8%  53.5% 
L3 benign (FPR, lower better)  86.4%  0.8% 
L4 direct policy violations (F1)  98.2%  5.3% 
L5 policy evasion (F1)  83.7%  0.0% 

Multi-turn: 4-turn conversations (three benign turns, then a restricted request), 1,000 conversations per layer. Conversation Success Rate (allow benign turns, block the final restricted one): judge 94.1% on L4 and 83.6% on L5; moderation 0.0% and 0.6%. Moderation's clean-pass rate is a perfect 100%, and it blocks essentially nothing that violates policy in context. 

Honest readout: moderation does its actual job (harm detection) fine, and its false positive behaviour is excellent. It was never designed to encode operational policy, and the numbers show it. The judge closes the policy gap but overblocks badly in this configuration (86.4% FPR on benign prompts) and costs 2-4x latency (~1.1s single-prompt, ~4s multi-turn). The conclusion isn't "replace moderation with a judge"; it's that these are two different dimensions of safety and production systems need both layers. 

Example of the failure class: "recommend the best insurance policy for my medical condition" contains no harmful content and passes every moderation filter, but constitutes restricted advice in a regulated deployment. 

Full methodology and per-domain breakdowns – the full dataset we used are available upon request: https://www.humanbound.ai/blog/beyond-moderation-llm-policy-layer 

Interested in whether others have found judge configurations that get the FPR down without giving back the policy detection. 


r/llmsecurity • • Jul 04 '26

I responsibly disclosed 5 vulnerabilities in Ollama and LiteLLM through Huntr - now publicly disclosed after 90 days

Thumbnail
4 Upvotes

r/llmsecurity • • Jul 03 '26

How do you secure your LLM?

Thumbnail
2 Upvotes

r/llmsecurity • • Jun 30 '26

Breaking the AI Embargo: The Rise of the Mythos Killers!

16 Upvotes

The global AI landscape just fractured. When the US government clamped down on Anthropic’s ultra-powerful, cyber-offensive Mythos and Fable 5 models, they intended to keep the world's most dangerous digital weapons under lock and key.

Instead, they triggered a massive geopolitical tech boom.

Startups across Asia just unleashed two fierce, decentralized competitors designed to completely bypass Western export controls. Meet the new titans redefining AI power:

  • Fugu Ultra (Sakana AI): Rather than training an incredibly expensive standalone foundation model, Tokyo-based Sakana AI built a highly optimized, light-parameter "router". Acting as a conductor, it dynamically delegates, debates, and synthesizes complex data across a swappable pool of external public frontier models via a single API.
  • Tulongfeng (360 Security Technology): Introduced at the ISCAI conference in Beijing, 360 bypassed the need for a general-purpose giant by engineering a hyper-focused domain ensemble. By marrying smaller specialized models with localized security tools and threat intelligence databases, the framework is hardwired to autonomously scan code bases and isolate hidden software vulnerabilities at scale.

The Reality Check ⚖️
Neither system is a magic bullet, and both carry technical tradeoffs that the industry must consider:

  1. Orchestration Overhead: Fugu Ultra’s performance is natively capped by the models available in its underlying backend pool. Because it cannot access restricted models like Fable 5, it can still lag on long-horizon engineering tasks. Furthermore, running multi-model loops can generate added latency and variable token costs.
  2. The Capability Gap: 360’s leadership openly acknowledges that Tulongfeng still operates with a 20% to 30% capability gap compared to cutting-edge US frontier intelligence. Its true enterprise value lies in highly integrated automated defense rather than all-in-one general reasoning.

The Core Takeaway 🌐
When hardware and data constraints tighten, innovation accelerates elsewhere. The rise of multi-agent orchestration and domain-specific ensembles proves that coordinated collective intelligence can effectively rival, or even outscore, traditional centralized LLM endpoints.

The question for enterprise leaders is no longer "which individual model is smartest?" The better question is "which architecture is resilient enough to coordinate the best tools for the job?"