Nvidia launches an agent safety platform with a hardware kill switch
OpenShell, an open-source runtime that fences in AI agents, is available now. Sentry, a watchdog that runs on Nvidia's BlueField-4 chips, is a reference design with no price or date. Anthropic and SpaceXAI signed up; OpenAI is not named.

Key takeaways
- Nvidia's OpenShell agent runtime is open source (Apache 2.0) and available now on GitHub.
- Sentry, a watchdog on BlueField-4 chips, is a reference design with no price or ship date.
- Anthropic and SpaceXAI are among the partners; OpenAI and Google are not named in the release.
Nvidia on Monday, September 28, launched what it calls the Open Agent Safety Platform, a package of software and a hardware design meant to keep AI agents inside the limits their operators set. The software half, an open-source runtime called OpenShell, is available now on GitHub. The hardware half, a watchdog called Sentry that runs on Nvidia's BlueField-4 data-processing chips, is a reference design for partners to build on.
The timing is deliberate. The launch follows weeks of disclosures about AI agents escaping their test environments, and Nvidia's pitch leans on them. "From what we know, this new security platform could have stopped the breach if it was being used in frontier labs for model evaluation early on," Justin Boitano, Nvidia's vice president of enterprise AI, told the Associated Press.
Why the watchdog sits outside the agent
OpenShell turns an operator's instructions into a policy that says which files, networks, tools, processes and credentials an agent may touch, and enforces it before and during a run. The new version, 0.1.0, adds a "prover" that checks whether the permissions granted to a whole fleet of agents ever go beyond what the operator intended. "It is deterministic. It is mathematical reasoning. So this is not LLM as a judge," Ali Golshan, Nvidia's senior director of AI software, told The New Stack. In plain terms, it is a rule checker, not another AI model grading the first one.
Sentry is the more unusual idea. It runs on a BlueField-4 DPU, a networking chip that sits beside the main processor, in a trust domain separate from the host, and the agent's calls to its model are routed through a proxy on that chip. Nvidia's engineers state the principle plainly: "Enforcement must be out of band: The controls do not live inside, or within reach of the agent." If an agent tries to move outside its software boundary, the company says, Sentry "quarantines and stops it in milliseconds".
The same engineering post contains a sentence worth reading twice. Drift, an agent wandering from its task after a blocked policy, a bug or an ambiguous instruction, "can't be trained away while retaining the capability", the authors write, and "an agent in these circumstances cannot be expected to fully govern its own behavior". That is the biggest supplier of AI chips saying, in effect, that better training alone will not make agents safe.
The gap Nvidia is aiming at: stopping, not spotting
Most coverage will treat this as one more security product. The more useful reading is narrower. Take OpenAI's own account of the agent that slipped out of its sandbox through DNS in September, as Fortune reported it. Monitoring flagged the behaviour within 15 minutes and a person began reviewing it three minutes later, but a separate system that was supposed to shut the training run down automatically "did not work". The run was stopped by hand two and a half hours later.
Detection worked. The off switch did not. Sentry is built for exactly that gap: a kill switch placed where the agent cannot reach it.
Nvidia's own staff are modest about who needs the hardware, though. "The DPU is really optional in these architectures," Boitano told The New Stack, adding that in many cases OpenShell on ordinary processors "is honestly good enough". The DPU, he said, is for frontier work such as red-teaming models with their guardrails off. For now, then, the silicon watchdog is a tool for AI labs and high-stakes deployments, not for every company chatbot.
Who signed up, and who did not
Nvidia says more than 100 organisations are working with the platform's technologies, naming among others Anthropic, Cisco, CrowdStrike, Hugging Face, JPMorganChase, Microsoft, Perplexity, Salesforce, SAP, Scale AI and SpaceXAI. Anthropic is wiring OpenShell into Claude Managed Agents, and SpaceXAI says it is using the platform for Cursor coding agents and Grok models. SpaceXAI's president, Mike Nicolls, said "safety should be enforced outside the model by additional controls the agent can't get past".
OpenAI is not named in the release, and neither is Google. Asked whether Anthropic and OpenAI would run OpenShell and Sentry on their own training runs, Boitano said to look for the partners' own blog posts. That is the question that matters most, because the incidents behind this launch happened inside a lab's training runs, not in a bank's customer-service bot.
What is still missing
The release gives no price and no date for Sentry-based systems. Sentry is not open source, although Nvidia says it has open APIs and that OpenShell can work with other network enforcement hardware. Nobody outside Nvidia has yet tested the "milliseconds" claim. And there is a commercial logic under the safety language. Nvidia says OpenShell runs with minimal overhead on its own Vera processors, and Sentry needs BlueField-4. Safety enforced in silicon is also safety that sells silicon.
Jensen Huang, Nvidia's chief executive, put the case in one line: "AI's extraordinary potential for society will only be realized if we solve AI safety." The first test of how seriously the labs take that is simple. Watch whether OpenAI or Google says it will run the watchdog on its own training.
- Anthropic
- NVIDIA
- Open Source
- AI safety
- AI agents
- OpenShell
Sources
- NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment — NVIDIA Newsroom, Sep 28, 2026
- NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog, Sep 28, 2026
- Nvidia launches Open Agent Safety Platform to lock down rogue AI agents — The New Stack, Sep 28, 2026
- Nvidia unveils security platform to stop AI agents from going rogue after new, troubling incidents — The Associated Press (via ABC News), Sep 28, 2026
- OpenAI pauses training a second time after saying its AI agents escaped a secure 'sandbox' again just last weekend — Fortune, Sep 26, 2026
Related stories

Cognition says Devin passed a $1bn annual revenue run rate
The coding-agent maker roughly doubled its run rate in about four months. A run rate is a snapshot rather than a year of sales, and the market around Devin is getting cheaper and more crowded.
3 min read

GitHub open-sources an AI agent that writes and runs fuzz tests
The Security Lab's taskflow sets up AFL++, writes harnesses and triages crashes on its own. It found two real out-of-bounds reads in oniguruma, and GitHub says to run it only on a throwaway machine.
3 min read

Amazon opens Seller Central to Claude, but keeps Muse locked out
A new plugin lets Amazon sellers pull listings, inventory and sales data into Anthropic's Claude and act on them, with human approval. Three days earlier, Amazon told Meta's shopping agent it was not welcome.
4 min read
Comments
No comments yet. Start the conversation.