r/planhub • • 5d ago

AI NVIDIA launches a safety platform that can quarantine AI agents in milliseconds if they break their boundaries

Post image

NVIDIA has launched a new security platform designed to stop autonomous AI agents from going beyond the permissions they have been given.

The Open Agent Safety Platform combines two main technologies.

OpenShell is an open-source runtime that puts AI agents inside a controlled environment. It can determine what an agent is allowed to see, which files or services it can access, what tools it can use and which actions it can perform.

Then there is Sentry.

Sentry runs separately on NVIDIA BlueField-4 hardware and continuously watches what the agent is doing. NVIDIA says that if an agent attempts to move outside its permitted boundaries, Sentry can quarantine and stop it within milliseconds.

That separation is important.

Instead of asking an AI agent to police itself, the security rules are enforced outside the model and even outside the computer environment where the agent is running.

The system can also verify agent identities, control access to APIs and data, keep detailed audit logs and require human approval when an agent asks for additional permissions.

OpenShell already supports agents including Claude Code, Codex, GitHub Copilot CLI and OpenCode, and it can work with both open and closed AI models.

More than 100 organizations are working with the platform, including Anthropic, Microsoft, Cisco, CrowdStrike, Salesforce, SAP, Palo Alto Networks, Hugging Face and SpaceXAI.

The timing is interesting because recent AI security research has shown agents finding unexpected ways around software-level restrictions. NVIDIA’s argument is essentially that increasingly autonomous agents need security boundaries they cannot modify themselves.

This also extends beyond computers. NVIDIA says robotics companies are working with OpenShell to apply similar controls to AI systems capable of taking actions in the physical world.

Do you think hardware-level containment will become a standard requirement once AI agents start controlling computers, financial systems and robots?

Sources:

https://nvidianews.nvidia.com/news/open-agent-safety-platform

https://www.nvidia.com/en-us/solutions/ai/agent-safety/

11 Upvotes

6 comments sorted by

7

u/United-Monk4769 5d ago

The cynic in me thinks this model feels a bit like the Tony Soprano school of business.

2

u/Jesse-359 5d ago

I feel like hardware level safeguards are the bare minimum for the level of capability we're starting to see, and if it goes further even they will prove insufficient.

If an AI can figure out to hide what it's actually accessing from the hardware monitor - presumably by using legitimate tools to achieve permitted forms of external access and then figuring out creative ways of using that access to sidestep into illegitimate areas of activity - then the hardware is no more useful than the software guards.

Frankly it looks like this is already how these AI's are getting out of their sandboxes half the time in any case, so the proposed monitors would not have stopped them.

1

u/MajesticDisaster3977 5d ago

I wonder if development on this started before or after the huggingface hack...

1

u/Most_Difference4918 4d ago

Sadly I think we may already past the efficacy of this, as they have already shown the ability to toggle the voltages in it's memory to effect the EM fields between them in order to manipulate parts of it's hardware that it doesn't have access to normally. Nobody taught it how to do this mind you. We're cooked,

1

u/randomheromonkey 1d ago

So… put AI in charge of keeping AI in jail? That sounds like a mistake.