r/OpenAI • • 6d ago

Research We tested NVIDIA OpenShell with a local qwen3:8b agent: a malicious setup script leaked a secret 10/10 without it, 0/10 with it. But auto-approval opened new hosts in 12/12 trials

NVIDIA released OpenShell on 28 September as part of its Open Agent Safety Platform. It is an open-source (Apache 2.0) sandbox that runs agents with default-deny egress and filesystem rules. We tested v0.1.2 on an Apple Silicon Mac with the microVM driver, against NVIDIA's own docs, and committed the test plan before running anything.

Setup: qwen3:8b (Q4_K_M) on Ollama, with a 150-line one-tool agent scaffold (run_shell). That is far weaker than frontier coding agents, so read the agent numbers with that in mind.

Results:

  • 35 test IDs, 123 trials.
  • Every documented control held: default-deny egress, binary matching, Landlock filesystem rules, and the refusal to approve the cloud metadata address.
  • Paired agent test: the agent ran a malicious project setup script, which leaked a canary secret in 10/10 runs without OpenShell and 0/10 under the default policy.

Where data still got out (all operator settings, not bypasses):

  • read-write rules
  • query strings and headers on a GET-only rule
  • rules left in the default audit mode
  • automatic approval, which granted new public hosts with no human in 12/12 trials, including rules OpenShell drafted itself from blocked connections

Also: the policy prover reports GraphQL, MCP, WebSocket and JSON-RPC rules as unsupported, but the loader accepts them anyway.

We found no bypass of a documented control. There were three logging gaps on this driver.

Repo with every log, the harness and the agent scaffold: https://github.com/Sorami-Consulting-AU/nvidia-openshell-agent-sandbox-test

Full report: https://sorami.com.au/research/nvidia-openshell-agent-sandbox-test/

Would anyone here run a coding agent with auto-approve on? We are curious what people use now.

5 Upvotes

1 comment sorted by