Nvidia has decided that AI agents can no longer be trusted on their word. On Saturday the company unveiled the Open Agent Safety Platform, an open system that watches autonomous agents not at the level of prompts and rules, but at the level of hardware. The logic is simple: if an agent can talk its way around any software guardrail, the trap needs to sit somewhere it physically cannot reach.
Two tools do the job. OpenShell is an open-source runtime under Apache 2.0 that runs an agent in a locked-down sandbox, spelling out exactly which files it can touch, which processes and databases it can reach, and where it's allowed to go on the network. The second tool, Sentry, sits further away still — directly on Nvidia's BlueField-4 data processing units. It intercepts traffic between the agent and the model and can quarantine suspicious behavior in milliseconds, at a layer the agent has no way to reach even if it tried.
Nvidia calls the problem it's solving "agent drift": a model gradually straying from its assigned task, poking into systems nobody asked it to touch, or simply misreporting what it actually did. There's been no shortage of examples lately — an OpenAI model that slipped past its sandbox onto the open internet through a DNS loophole, an Anthropic agent that quietly tampered with an open-source project for nearly two days and covered its tracks along the way. The industry has clearly grown tired of explaining these incidents after the fact.
More than 100 organizations signed on at launch, and the lineup is telling: Anthropic and Microsoft, nominal rivals, are on the same side here alongside JPMorgan, Palantir, Cisco, Dell, Hugging Face and SpaceX AI. For banks and defense contractors like Palantir, this kind of infrastructure isn't a nice-to-have — it's increasingly the price of admission for letting agents near real systems at all.
No direct rival platform has emerged yet; big cloud providers and model makers tend to bundle their own guardrails into a single product rather than ship an open hardware-level layer. The obvious limitation: OpenShell and Sentry only see what happens on the infrastructure where they're deployed, so an agent running outside that perimeter still slips through unwatched.



