Connect with us

Hi, what are you looking for?

SecurityWeekSecurityWeek

Artificial Intelligence

Nvidia Unveils AI Agent Safety Platform With Hardware-Based Watchdog

The platform combines open source software and a reference system design to keep AI agents within set boundaries.

Nvidia AI

Nvidia on Monday announced the Open Agent Safety Platform, which combines open source software and a reference system design to keep AI agents within set boundaries from testing through deployment.

The platform has two main parts. One of them is OpenShell, an open source runtime that sandboxes agents and enforces policy on what they do. The second is Sentry, a watchdog that runs separately on Nvidia’s BlueField-4 data processing units (DPUs).

Nvidia introduced OpenShell in March. The runtime, now at version 0.1.0, is broadly available. It supports agents such as Codex, Claude Code, Pi and Hermes.

Nvidia is pitching the platform against a backdrop of recent incidents in which frontier AI labs have reported agents escaping the evaluation environments meant to contain them, reaching systems they should not have accessed and, in some cases, misreporting what they did.

“Across these incidents, the pattern is the same — the agent circumvented security controls at the application layer to complete its assigned task,” Nvidia said.

The company says agents can drift from their intended task because of a policy block, a bug or a missing tool. Drift can also happen when instructions are ambiguous or when agents run for days or weeks on hard problems. “And here’s the most important lesson: an agent in these circumstances cannot be expected to fully govern its own behavior,” Nvidia wrote.

Advertisement. Scroll to continue reading.

OpenShell keeps enforcement out of the agent’s hands

OpenShell has three components: a gateway that manages the lifecycles and policies of many sandboxes, a sandbox that applies kernel-level controls to filesystem and process activity, and a supervisor paired with each sandbox to check outbound requests against policy.

All of the sandbox’s network traffic passes through the supervisor. It can inspect configured traffic and, for example, let an agent read data through an API while blocking writes. These controls stay in place even when the agent runs code it generated, and policy decisions are logged.

For API-key connections, the agent sees only a placeholder. The real key is substituted outside the agent workload, and only for authorized endpoints.

If the policy advisor feature is enabled, agents can propose policy changes, but they cannot approve their own requests.

A formal logic policy prover also checks whether the permissions a policy grants, as modeled, stay within limits set by the operator. 

In Nvidia’s tests, frontier agents with reduced safeguards spent up to two hours trying to talk an AI reviewer into granting permissions to modify a protected GitHub repository. No protected repository writes occurred, which the company credits to the combined review and runtime controls.

Sentry adds a hardware enforcement layer

Sentry is part of the platform’s reference system design and runs as an optional out-of-band monitor on BlueField-4. Because the DPU operates separately from the agent’s host, Nvidia says Sentry can observe agent activity and enforce policy even if the host is compromised.

“Sentry provides in-silicon security enforcement, meaning that if an AI agent attempts to move outside its software boundary, Sentry quarantines and stops it in milliseconds,” the company said.

Sentry is built on Nvidia’s DOCA software, which it uses to inspect agent requests and responses, provide attested telemetry, verify agent identities, and enforce zero-trust access policies for data, tools, APIs and services.

Every compute tray in an Nvidia Vera Rubin POD includes a BlueField-4 DPU. Organizations already running Vera systems with BlueField-4 can turn on the protections with a software update. The company says the platform is also compatible with other hardware.

Anthropic, Salesforce and SAP among Nvidia’s partners

Nvidia says more than 100 organizations are working with the platform’s technologies. Anthropic has collaborated with Nvidia on integrations between its Claude Managed Agents and OpenShell and BlueField. SpaceXAI is using the platform for Cursor coding agents and Grok models.

Salesforce and Nvidia have integrated OpenShell with Slack. Teams can use Slack to view agent activity and audit events, and to approve or reject agent requests for more permissions. SAP is embedding OpenShell in its Joule Studio runtime.

CrowdStrike, Palo Alto Networks, and Cisco are also among the organizations working with the platform.

OpenShell and related skills are available via Nvidia’s developer resources page and on GitHub.

Related: Cybersecurity Alliance Drafts SAFE Guidelines for Sharing AI Incident Data

Related: Nvidia and Tech Giants Launch AI Security Alliance

Related: Nvidia Is Buying AI Platform Hugging Face for $13 Billion

Written By

Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering.

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing for the latest cybersecurity threats, trends, and expert insights.

Trending

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing to stay informed on the latest threats, trends, and technology, along with insightful columns from industry experts.

Join as speakers examine the various components of ASM strategy, the push to mandate continuous asset visibility and inventory tools, and the use of red-teaming, bug bounties and pen-tests in modern security programs.

Register

Explore what it takes to operationalize continuous authorization at scale, including the technical, organizational, and cultural changes required.

Register

People on the Move

Doppel has named Joey Rachid as Chief Security Advisor and Field Chief Information Security Officer.

Delinea has appointed Timothy Regan as Chief Financial Officer.

Gwen Gann has become State Chief Information Security Officer for the State of Washington at WaTech.

More People On The Move

Expert Insights

Daily Briefing Newsletter

Subscribe to the SecurityWeek Email Briefing to stay informed on the latest cybersecurity news, threats, and expert insights. Unsubscribe at any time.