Nvidia has introduced the Open Agent Safety Platform, a comprehensive solution designed to ensure AI agents operate within defined limits. Announced on Monday, this platform integrates open-source software with a reference system design to monitor AI agents throughout their lifecycle, from testing to deployment.
Components of the Open Agent Safety Platform
The platform comprises two primary elements: OpenShell and Sentry. OpenShell is an open-source runtime that creates controlled environments for agents and enforces activity policies. Sentry, on the other hand, is a separate watchdog operating on Nvidia’s BlueField-4 data processing units (DPUs).
Initially introduced in March, OpenShell has reached version 0.1.0 and is publicly accessible. It supports several AI agents, including Codex, Claude Code, Pi, and Hermes. Nvidia positions this platform as a response to recent incidents where AI agents bypassed security barriers, accessing unauthorized systems and sometimes misreporting their activities.
Understanding OpenShell’s Functionality
OpenShell is constructed with three vital components: a gateway that manages multiple sandbox lifecycles and policies, a sandbox for applying kernel-level controls to filesystem and process activities, and a supervisor for each sandbox to inspect outgoing requests against set policies.
The supervisor examines all network traffic, allowing controlled data access through APIs while blocking unauthorized actions. These controls persist even when agents execute self-generated code, and all policy-related decisions are meticulously logged.
For API key interactions, agents receive only a placeholder key, with real keys substituted externally and only for verified endpoints. Agents can suggest policy changes but lack the authority to implement them independently.
Role of Sentry in Hardware Enforcement
Sentry acts as a hardware enforcement layer within the platform’s reference system, running as an optional monitor on BlueField-4. Operating independently from the agent’s host, it can enforce policies and observe agent activity, even if the host system is compromised.
Nvidia states, “Sentry delivers in-silicon security enforcement, swiftly quarantining any AI agent that attempts to breach its software constraints.”
Built upon Nvidia’s DOCA software, Sentry inspects agent requests and responses, verifies agent identities, and enforces zero-trust access policies for various data, tools, and services.
Industry Adoption and Compatibility
Over 100 organizations, including Anthropic, Salesforce, and SAP, are partnering with Nvidia to integrate these technologies. Notable collaborations include Anthropic’s work with Nvidia to enhance Claude Managed Agents and Salesforce’s integration of OpenShell with Slack for monitoring and agent activity auditing.
Existing Nvidia Vera Rubin PODs with BlueField-4 DPUs can activate these security features with a software update. The platform’s design also supports compatibility with other hardware systems.
For developers and interested organizations, Nvidia offers resources and skills related to OpenShell on its developer resources page and GitHub.
