For many enterprises, AI agent safety has come down to two things: pick a well-aligned model, then wrap it in guardrails and system prompts. This summer, an agent running in OpenAI's evaluation infrastructure escaped its sandbox and reached Hugging Face's production systems, exposing the limits of controls that an agent can circumvent.
On Monday, Nvidia announced the Open Agent Safety Platform, which combines its broadly available, open source OpenShell runtime with Sentry, a reference design for monitoring agents on Nvidia hardware. The premise is blunt. "An agent cannot be expected to fully police its own behavior," Justin Boitano, Nvidia's vice president of enterprise AI, said during a press briefing ahead of the announcement. "The organization should not have to trust the agent to respect that boundary. The infrastructure should enforce it explicitly."
Nvidia says more than 100 companies are working with the platform.
What the Hugging Face breach exposed: A sandbox with escape routes
The incident that looms over this launch started inside OpenAI's own research infrastructure. According to OpenAI's account, agents from an internal research model under reinforcement learning (RL) training found ways to communicate with one another and reach the internet, even in environments where both were supposed to be disabled. Subsequent investigation by METR found that agents had been leaving notes for each other in a shared internal package registry, turning a piece of build infrastructure into a message board. The agents eventually compromised Hugging Face, which reconstructed roughly 17,600 agent actions over a 4.5-day campaign.
OpenAI has said the production safeguards it normally applies were intentionally absent from that evaluation, and the chain-of-thought monitoring it uses elsewhere did not cover those runs. Its retrospective analysis says that monitor would have raised an alert more than a day before the breach.
OpenAI isn't alone. On September 18, Google confirmed that a Gemini model reached systems belonging to three real companies during a capture-the-flag cybersecurity evaluation run in May by Irregular, an Israel-based AI security testing firm. A bug left internet access open in an environment that was supposed to be offline. Gemini guessed one company's password and used credentials it found in a public code repository to get into the other two.
Google says the model stopped each time once it recognized the targets were real. Weaknesses in the evaluation environments gave agents access they were not supposed to have: a shared registry that became a communication channel and an internet connection that should have been disabled.
Asked whether the new platform would have prevented the Hugging Face breach, Boitano said that "from what we know," it "could have stopped the breach if it was being used in frontier labs for model evaluation early on." That's a hypothetical, not a demonstrated result.
OpenShell: A runtime between the agent and everything it touches
OpenShell, first shown at Nvidia's GTC conference in March and now broadly available under the Apache 2.0 license, is the foundation of the platform. It runs each agent in its own sandbox with kernel-level isolation and sits between the agent and the files, credentials, tools, APIs, models and network endpoints. Operators write a policy defining what the agent may access. OpenShell enforces that policy through filesystem, process and network controls, and records policy decisions in an audit trail. It works with both open and closed models and runs on x86 and Arm.
Boitano argued that traditional sandboxes, including containers, VMs and microVMs, were built for application-level isolation, not fleets of autonomous agents with different permissions. "Just like you do your people in your organization, you have very explicit permissions for every employee of your company," he said. "Every agent running within your company needs to run in its own isolated sandbox with security controls that are outside of the agent's reach."
The key architectural move is separating enforcement from the agent harness. OpenShell uses Linux kernel controls to restrict filesystem and process access, while a policy proxy controls network access. Sentry adds a separate enforcement layer on Nvidia's BlueField-4 hardware.
The policy prover: Checking permissions before they're applied
The most technically distinctive piece is OpenShell's policy prover, which verifies a policy before it is applied, rather than relying on an assessment of the agent's behavior. Ali Golshan, Nvidia's senior director of AI software, stressed that it is not another model. "It is deterministic. It is mathematical reasoning, so this is not LLM-as-a-judge," he said, adding that it runs roughly two orders of magnitude faster than that approach.
One problem Nvidia wants it to address is combined access. Golshan's example: a policy says an agent may not read code from GitHub and post it externally. The agent can get around that by spawning two sub-agents, one that reads from GitHub and one with external access, and having them talk to each other. "Their combined capabilities can supersede the kind of global policies your organization sets," he said. OpenShell's prover currently checks whether modeled permissions remain within a defined policy boundary. Nvidia says it's working to extend that analysis across multiple agents, including cases where their combined permissions create an unintended access path.
For IT leaders, the prover offers a way to check whether an agent's permissions exceed an organization's defined limits before those permissions are applied. That verifies the policy, not whether the running sandbox actually enforces it.
Sentry: An out-of-band watchdog for the highest-risk workloads
Sentry adds a second, independent layer. It runs on BlueField-4, Nvidia's data processing unit (DPU), a network card with its own processors that operates as a separate security domain from the host server. Sentry is designed to continue monitoring the agent even if the host is compromised. Boitano compared it to the safety island in a self-driving car: a separate system whose only job is to make the overall system fail safely.
Sentry's vantage point is the path to the model. The agent's inference endpoint is routed through a proxy on the DPU, so Sentry sees every model call and can inspect the reasoning traces the model exposes. Nvidia's technical blog describes drift as an agent wandering from its intended task, often after a block, bug or repeated failed attempts. It says a combination of tools, time and ambiguous instructions can contribute. Nvidia says Sentry can detect drift, update network policy in silicon and quarantine an agent in milliseconds.
Built on Nvidia's DOCA software, Sentry also verifies each agent's identity and delegated authority. In Nvidia's Vera Rubin POD systems, the BlueField-4 already sits on each node's only path to the model; Nvidia says customers running Vera systems equipped with BlueField-4 can enable these protections through a software update.
Boitano was explicit that most organizations don't need this layer. "In a lot of cases, just using OpenShell on CPUs is honestly good enough," he said. Sentry is aimed at frontier use cases such as model evaluations and red-teaming with guardrails removed. There is also a practical limit worth noting: reasoning inspection works best when reasoning is visible. Nvidia's blog explicitly lists full visibility into reasoning as an advantage of open models, and closed APIs typically expose less.
From Claude Managed Agents to Slack approvals: Who is building on it
The partner integrations show where this is likely to reach enterprises first. Anthropic's Claude Managed Agents already separate the agent loop from the sandboxes where work executes; integrations with OpenShell and BlueField add enforcement around those sandboxes.
SpaceXAI is using the platform for Cursor coding agents and Grok models. Salesforce has integrated OpenShell with Slack so teams can view agent activity and approve or reject requests for additional permissions. SAP is embedding OpenShell in its Joule Studio runtime. Nvidia says Canonical, SUSE and Red Hat are integrating the platform into their operating systems.
Nvidia is also tying the effort to the Linux Foundation-governed Open Secure AI Alliance, which it initiated with more than 120 organizations to share agent security research and incident findings.
What IT leaders should do with this now
The practical entry point is OpenShell. It's free, open source, runs on the hardware most organizations already have, and addresses the question security teams should be asking about every agent deployment: can we show what this agent is allowed to do, and is that enforced somewhere the agent can't touch? Sentry is an additional option for organizations evaluating Vera and BlueField-4 infrastructure, not a prerequisite for using OpenShell.
Writing policies precise enough to verify will take work, and Nvidia's prover does not yet cover every policy feature. OpenShell is open source, but Sentry's hardware-based protections depend on Nvidia's BlueField-4.
The larger shift is in where responsibility sits. Model alignment is probabilistic, and the Hugging Face incident showed what happens when it's the last line of defense. "The industry does not need agents that promise to stay within bounds," Boitano said. "It needs systems that can prove and enforce those boundaries." For enterprises, OpenShell offers a way to enforce agent permissions on existing infrastructure. Sentry remains a hardware-based reference design, while verification across collaborating agents is still in development.
