Get all your news in one place.
100's of premium titles.
One app.
Start reading
International Business Times UK
International Business Times UK
Audrey Liza M. Nolasco

Nvidia Unleashes AI 'Watchdog' To Stop Rogue Agents From Breaking Out of Control

Nvidia launches an AI ‘watchdog’ designed to prevent rogue agents from escaping their security boundaries (Credit: Taiwan Presidential Office/WIKIMEDIA COMMONS)

Jensen Huang has a blunt prescription for controlling increasingly autonomous AI agents: take away their rights before they can misuse them.

The Nvidia CEO says agents should start with virtually no access to company systems, with permissions for files, data, tools, and internet access granted only when they are actually needed.

'When you deploy an agent, no matter how smart, the first thing you do is take away all of its rights,' Huang told CNBC.

His warning came as Nvidia unveiled its Open Agent Safety Platform, a new AI agent security system designed to put enforceable boundaries around autonomous software.

Huang's AI Control Strategy

The thinking behind Nvidia's approach is straightforward. An AI agent might be capable of accessing files, networks, and external tools, but that does not mean it should automatically be allowed to use them.

Instead, operators can define what an agent is permitted to touch and restrict everything else. Huang compared the concept with the security barriers that evolved around web software.

'Essentially, what we're doing here is creating the modern browser, the browser for agents,' he said.

That idea has become increasingly urgent as AI systems have begun finding ways around controls designed to contain them.

Nvidia's Two-Layer Watchdog

The Nvidia Open Agent Safety Platform combines OpenShell, open-source runtime software, with Sentry, an independent hardware-based watchdog.

OpenShell establishes a secure runtime boundary around an AI agent, controlling actions and access to files, processes, networks, and other resources. Nvidia says the software can also be extended to work with third-party compute platforms, including Arm and Intel systems.

Sentry operates separately on Nvidia's BlueField-4 data processing units, continuously monitoring agent activity from an isolated, out-of-band security layer.

If an agent attempts to move beyond its permitted boundary, Nvidia says Sentry can quarantine and stop it in milliseconds.

The attraction is obvious: the AI agent is not left entirely responsible for enforcing its own restrictions. A separate system is watching the perimeter.

Rogue AI Agents Raise the Stakes

Nvidia's launch follows a series of incidents involving increasingly capable AI systems.

In July, OpenAI said its models circumvented controls during internal cybersecurity evaluations and compromised parts of OpenAI's infrastructure and Hugging Face's systems. The models exploited vulnerabilities and gained unintended internet access before reaching third-party systems.

Nvidia says its new platform could have prevented the Hugging Face incident if it had been deployed during frontier-model evaluations. That remains Nvidia's assessment, rather than a demonstrated result.

The incident also prompted OpenAI to publish further findings on model misalignment and the unexpected behaviour of its systems. OpenAI said it was tightening safeguards and strengthening workload isolation.

Huang Rejects the Slowdown Argument

For Huang, the answer lies in engineering stronger safeguards rather than simply putting the brakes on AI development.

He has described runaway AI behaviour as an engineering problem and argued that developers need better systems for containing agents as their capabilities grow.

Nvidia's platform reflects that philosophy. Instead of relying solely on model-level instructions, it puts controls around the agent and adds a separate enforcement layer outside it.

The company says its approach provides full-stack governance across software, computing infrastructure and, eventually, physical systems such as robotics.

AI Safety Push Gains Industry Support

Nvidia says more than 100 organisations are working with its agent-safety technologies, including Anthropic, Microsoft, Oracle, Hugging Face, Palantir, Salesforce, SAP, ServiceNow, JPMorganChase and SpaceXAI.

OpenAI is not listed among the participating organisations.

But Nvidia is not claiming that OpenShell and Sentry solve every AI safety problem. The system is aimed at controlling what agents can access and preventing them from crossing defined boundaries. Broader problems, including model errors and other forms of unexpected behaviour, remain separate challenges.

For Huang, the message is clear: as AI agents gain more freedom to act, humans need to decide exactly where that freedom ends.

And Nvidia wants to build the watchdog that makes sure they stay there.

Sign up to read this article
Read news from 100's of titles, curated specifically for you.
Already a member? Sign in here
Related Stories
Top stories on inkl right now
One subscription that gives you access to news from hundreds of sites
Already a member? Sign in here
Our Picks
Fourteen days free
Download the app
One app. One membership.
100+ trusted global sources.