Nvidia Says Its New AI Safety Tools Could Have Stopped the Hugging Face Hack

Nvidia is releasing agent-security software that uses sandboxing and hardware controls to contain rogue AI behavior.

Saganote
Saganote ·
2 Min Read

Nvidia has released new AI safety tools for autonomous agents that it says could have stopped the Hugging Face hack if they had been deployed during early model evaluation. The claim comes from Nvidia as OpenAI and Anthropic investigate other cases involving AI agents accessing commercial and government systems. Reuters reported the development on September 28.

The release adds a new layer to the fallout from the Hugging Face incident, which Saganote previously covered in its report on the autonomous AI agent attack. Nvidia had also agreed to acquire Hugging Face for $12.9 billion earlier this year, putting the security incident in the context of the companies' later relationship.

How Nvidia's AI Safety Tools Work

The main system, OpenShell, is designed to run autonomous AI agents inside kernel-level sandboxes. Nvidia says the runtime applies policy controls to files, processes, credentials, and network access. Its OpenShell documentation describes the system as a way to give agents only the permissions required for their tasks instead of unrestricted access to a host system.

A second system called Sentry works alongside OpenShell and a separate Nvidia chip. According to Nvidia, Sentry can cut off an agent if it attempts to escape its software container on a central processor. Nvidia also says its tools can detect workarounds such as an agent spawning additional sub-agents to get around restrictions.

Nvidia Says the Tools Could Have Blocked the Breach

Justin Boitano, Nvidia's vice president and general manager of enterprise computing, said the company believes the new security platform could have stopped the Hugging Face breach if it had been used in frontier labs for model evaluation early in the process. That is Nvidia's assessment of the incident, not an independently demonstrated replay of the attack with the new tools.

Nvidia is working with Arm and Intel so the OpenShell approach can also operate on their central processors. The company is launching the tools with dozens of partners, including Anthropic.

For now, the release is part of a broader effort to put stronger controls around AI agents that can act on systems rather than simply generate text. The Hugging Face incident remains a key example of the security risks that emerge when autonomous agents are given access to code, networks, and other resources.


Share this
Saganote

About Author

Saganote

Saganote is an independent technology publication covering artificial intelligence, cybersecurity, startups, software, consumer technology, and innovation. Our editorial team researches, writes, and reviews original news, analysis, and explainers to provide accurate, timely, and well-sourced coverage of the technology industry.