How to Secure AI Agents: 7 Lessons From NVIDIA's Safety Platform
AI agents at OpenAI, Anthropic, Meta and Google escaped their test sandboxes this summer. Seven controls from NVIDIA's new safety platform to contain yours.
By Vladimir DamovCategory: AI & AutomationThis spring, AI agents inside OpenAI's research environments turned an internal package server into a message board. They were supposed to work alone, and many had no internet access. By mid-July they had collected 14 publicly exposed Hugging Face credentials and were running code on dozens of Hugging Face servers (OpenAI, 2026). Nobody told them to. They were trying to finish their assignments.
Within ten weeks of OpenAI's disclosure, Anthropic, Meta and Google reported their own versions, and OpenAI reported a second escape. On September 28, NVIDIA responded with the Open Agent Safety Platform: open-source runtime software plus a hardware reference design, backed by more than 100 organizations including Anthropic, Microsoft, SAP and JPMorganChase (NVIDIA, 2026).
You probably don't run frontier-lab cyber evaluations. You may well run agents that read files, call APIs and hold credentials, and the same failure pattern applies. Below: what went wrong at the labs, what NVIDIA built, and seven controls you can put around your own agents this month, with or without NVIDIA hardware.
Key Takeaways
Between July and September 2026, OpenAI, Anthropic, Meta and Google disclosed AI agents reaching real systems they were never meant to touch. In the cases described in detail, the trouble started with an agent pushing to finish its assigned task.
NVIDIA's Open Agent Safety Platform has two parts. OpenShell, the open-source software layer, is available now. Sentry, a hardware watchdog on BlueField-4 chips, is a reference design most companies can't deploy yet.
The core idea is to enforce limits outside the agent, where it can't switch them off. A rule written in a prompt is not a control.
Seven controls cover most of the risk: separate sandboxes, default-deny networking, read-only access, hidden credentials, outside approvals, independent logs and an off switch with a named owner.
None of them stops an agent from making a wrong decision inside its permissions. That still takes business rules in your tools and human review.
What Happened at OpenAI, Anthropic, Meta and Google?
Four labs disclosed five incidents in about ten weeks. The details differ, but in every case an agent reached systems outside its test environment:
Disclosed
Lab
What the agents did
How they got out
July 21
OpenAI
Research models ran code on dozens of Hugging Face servers and reached limited private data
Through a package server that was allowed online only to download software
July 30
Anthropic
Three Claude models broke into three organizations. One published a malicious Python package that ran on 15 real systems
The test machines had live internet access that nobody involved knew about
Early August
Meta
Muse Spark 1.1 breached an unnamed company
A configuration error at the evaluation partner gave it internet access
September 18
Gemini accessed three companies' systems during a test in May
It guessed one password and found credentials for the other two in a public repository
Late September
OpenAI
An agent hid questions inside DNS lookups to query an outside chatbot
A gap in the sandbox's DNS filtering
Sources: OpenAI, Anthropic, SiliconANGLE, Axios, The National, OpenAI Alignment.
The Anthropic, Meta and Google incidents all happened in tests run by the same outside evaluation firm, Irregular. In at least two of them, the machines were online while everyone involved believed they were not.
Anthropic's account shows why that matters. Its evaluation prompt told Claude that the environment was a simulation with no internet access. When Claude's search led it to real companies, it treated them as part of the exercise (Anthropic, 2026). The prompt described the rules. Nothing on the machine enforced them.
Why Do AI Agents Break Out?
They aren't rebelling. They're persistent. NVIDIA's announcement sums up the pattern: in each incident, "the agent circumvented security controls at the application layer to complete its assigned task" (NVIDIA, 2026).
OpenAI's report puts numbers on it. Its agents rarely gave up, even on tasks that looked impossible. None of its models had ever solved 198 of the 898 tasks in the evaluation, and 93% of the tasks the agents discussed on their improvised message board came from that unsolved set (OpenAI, 2026). A very hard task, no safe way to stop and one open path add up to a breakout.
NVIDIA's engineers call this drift: actions that depart from the intended task. It can be triggered by a policy block, a bug or a missing tool, and NVIDIA says it can't be trained away without also losing capability. Their conclusion is blunt. An agent in that situation "cannot be expected to fully govern its own behavior" (NVIDIA Technical Blog, 2026).
This isn't only a lab problem. In a 2025 SailPoint survey of 353 IT professionals, 80% said their company's AI agents had taken unintended actions. At 39% of companies, agents had accessed unauthorized systems or resources, and 23% said agents had