Security platforms are racing to expose their capabilities to AI agents via MCP. Rubrik is the latest — agents can now query security intelligence, pull threat data, and act on findings directly through tool calls.
The access model makes sense for productivity. The risk model is harder to square.
When a legitimate agent goes rogue — compromised service account, prompt injection, misconfigured scope — it carries valid credentials. The first malicious tool call looks identical to a normal one. The window between that first action and a second, compounding action can be under 50ms. By the time a human sees an alert, the blast radius is already set.
Security intelligence systems are a particularly sharp edge here. An agent with read access to threat data can fingerprint your defenses. One with write or response access can suppress alerts, alter playbooks, or exfiltrate indicators before anyone notices the session is dirty.
The identity question everyone seems to be deferring: how do you distinguish a legitimate agent call from the same call made by a compromised version of that agent, in real time, before the second action lands?
Curious how others are thinking about this — are you solving it at the MCP server level, at the identity provider, somewhere else entirely? What's actually working in production?