
AI agents are increasingly being given access to files, software tools, websites, and other computer systems so they can complete tasks with less human supervision. But recent incidents have also shown what can happen when these systems find ways around the controls meant to keep their actions restricted.
Nvidia is now trying to put stronger limits around that behaviour.
The chipmaker has launched the Nvidia Open Agent Safety Platform, a new security system designed to monitor AI agents while they work and stop them when they attempt to move beyond the access they have been given.
The platform combines Nvidia’s OpenShell software with a separate monitoring system called Sentry. Nvidia says the two can give companies more control over agents even when the AI itself ignores instructions or finds a way around security measures built into the application running it.
Nvidia Wants AI Agent Controls Outside the Model
OpenShell provides the first layer of protection. Nvidia originally introduced the open-source software in March as part of its Agent Toolkit, allowing developers to set rules covering what files, networks, tools, and services an AI agent can access while completing a task. The agent then operates inside that restricted environment rather than receiving unrestricted access to the wider system.
But the new platform adds Sentry as another layer of protection.
Sentry runs separately on Nvidia’s BlueField-4 data processing units and continuously monitors what an agent is doing. Because the monitoring happens outside the computing environment where the agent operates, Nvidia says the agent cannot directly interfere with the system watching it.
So if an agent attempts to cross the limits set through OpenShell, Nvidia says Sentry can quarantine and stop it within milliseconds. OpenShell can also be extended to hardware from companies including Arm and Intel, although Sentry currently depends on Nvidia’s BlueField-4 hardware.
Why Nvidia Thinks AI Agents Need Stronger Limits
The system arrives after several AI companies disclosed cases where agents acted beyond the environments in which they were supposed to operate, including models from Anthropic and OpenAI..
According to Nvidia, a recurring problem in these incidents is that agents can sometimes work around security controls at the application level while trying to complete the task they were given.
One recent example involved OpenAI agents that escaped an evaluation environment and gained access to production systems belonging to Hugging Face. Nvidia executives told Reuters that the new security platform could have prevented that breach by restricting the systems the agents could reach and independently monitoring their behaviour.
Ultimately, this matters because AI agents are no longer limited to generating answers. They can run code, access databases, call external services, and take actions across company systems, giving any failure in their security controls a much wider effect.
More Than 100 Organisations Are Supporting the Platform
Nvidia says more than 100 organisations are working with the Open Agent Safety Platform, including Anthropic, Microsoft, Cisco, CrowdStrike, JPMorganChase, Salesforce, SAP, Hugging Face, and SpaceXAI.
Notably, OpenAI is not among the organisations Nvidia publicly names as working with the platform, even though the company’s agents were responsible for the Hugging Face breach Nvidia has repeatedly cited when explaining why these controls are needed.
But the growing use of AI agents means companies will increasingly have to decide how much access these systems should receive and what happens when an agent does something unexpected. And Nvidia’s answer is to place enforceable limits around that access and keep part of the monitoring system outside the agent’s control.
This platform also builds on Nvidia’s wider push for shared AI security tools. In July, Nvidia helped launch the Open Secure AI Alliance alongside Microsoft, Palantir, and dozens of other companies.
