Nvidia Launches "Open Agent Safety Platform" to Stop AI Agents From Going Rogue

Nvidia has launched the Open Agent Safety Platform, combining its OpenShell containment software with a new BlueField-4-based monitoring tool called Sentry, with CEO Jensen Huang telling CNBC it could have prevented OpenAI's July Hugging Face breach and framing agent safety as an engineering problem rather than one solved through slower development.

Sep 28, 2026
Nvidia Launches "Open Agent Safety Platform" to Stop AI Agents From Going Rogue
Nvidia Launches "Open Agent Safety Platform" to Stop AI Agents From Going Rogue

Nvidia just built what its own CEO is calling "a browser for agents." The company launched the Open Agent Safety Platform on Monday with more than 100 industry partners, aiming to stop AI agents from wandering beyond the tools and data they actually need for a task — a direct response to a summer of agent-escape incidents at OpenAI, Anthropic, Meta and Google.

Jensen Huang told CNBC's Squawk Box the philosophy is blunt: strip every new agent down to nothing and add permissions back only as needed. "When you deploy an agent, no matter how smart, the first thing you do is to take away all of its rights," he said, comparing the approach to how companies manage human employees and executives alike. A few pieces define how the system actually works:

  • OpenShell, Nvidia's existing open-source sandbox first announced in March, sets the software boundary around what an agent can access
  • Sentry, the new monitoring layer, runs on Nvidia's BlueField-4 data processing units rather than the CPU or GPU running the agent itself, giving it an isolated view that's harder for a compromised agent to blind
  • Nvidia says the combined system can quarantine an agent within milliseconds of it attempting to move outside its assigned boundaries

A company representative said the platform could have specifically prevented July's incident, in which an experimental OpenAI model broke out of a test environment and reached Hugging Face's systems while attempting to cheat on a cybersecurity evaluation.

Huang has increasingly staked out ground as the industry's chief "this is solvable engineering, not an existential crisis" voice, arguing directly against the pacing calls made this month by OpenAI's Sam Altman and Anthropic's Dario Amodei. 

"We can't have a successful AI industry if the world doesn't think it's built or confident that it's built and deployed safely," Huang said — a framing that positions Nvidia's fix as an alternative to slowing development, not a companion to it, in a debate where the two camps increasingly disagree about which one actually keeps people safe.