Nvidia used September 28, 2026 to answer a question that has been building all year: who watches the AI agents once they are let loose on real systems. The company’s new NVIDIA Open Agent Safety Platform pairs an open-source sandbox with a hardware watchdog that runs on separate silicon, and it arrives at a moment when agent-related security incidents have piled up across the industry. The timing is not incidental. Enterprises have spent 2026 racing to deploy autonomous coding agents, browsing agents, and customer-facing bots, and the security tooling has struggled to keep pace.

Nvidia laid out the launch in its own announcement covered by Infosecurity Magazine and in reporting from SiliconANGLE. Nvidia’s pitch is architectural rather than incremental. Instead of adding another monitoring dashboard, the company is proposing that agent governance live in two places at once: inside the software runtime and outside it, on a physically separate chip that cannot be talked out of its job by a compromised agent. That second part is what sets this Nvidia agent safety platform apart from most existing agent-security products, and it is why security teams are paying attention even though pricing and a general-availability date have not been published.

What Nvidia announced on September 28

According to SecurityWeek and Help Net Security, the Open Agent Safety Platform is an open software platform and reference system design meant to cover an agent’s life cycle from testing through deployment. It ships as two named components: OpenShell and Sentry. Nvidia describes OpenShell as open-source software that runs agents inside a sandbox and enforces policies on what those agents can touch, including files, tools, networks, processes, and credentials.

Sentry is the second half of the platform, and it is the part drawing the most attention. It is a monitoring and security layer that runs independently of the agent’s own compute path, on Nvidia’s BlueField-4 data-processing units, so it keeps working even if the host running the agent gets compromised. Nvidia said the system can “quarantine agents that attempt to move outside their boundaries in milliseconds,” a claim that comes from the company’s own announcement rather than an independent benchmark. This is the same architecture-first framing that shaped an earlier look at how BlueField-4 fits Nvidia’s agent-safety roadmap, though the emphasis here is on what the split between software and hardware control actually buys a security team.

Inside OpenShell: the software sandbox

OpenShell operates at the runtime layer. It creates an isolated environment for an agent to run in, then governs what that agent can request: which files it can open, which tools it can call, which network destinations are reachable, and which credentials it is allowed to present. Nvidia’s framing, repeated across multiple outlets covering the launch, is that developers can formally verify an agent has enough authority to do its job and no more. That is a meaningfully different goal than the permission prompts and rate limits most agent frameworks ship with today, because it treats least-privilege as something that can be checked mathematically rather than assumed from a policy document.

The practical value shows up when an agent is compromised through prompt injection or a poisoned tool response, a failure mode that has already hit production systems this year. If OpenShell’s policy engine has correctly scoped the agent’s authority beforehand, a hijacked agent still cannot reach data or systems outside its assigned lane, even if an attacker fully controls its reasoning. That is the theory. Whether it holds up against creative attackers is the open question every security vendor in this space is being asked right now.

Sentry and the case for out-of-band monitoring

Sentry is where Nvidia’s hardware business gives it an advantage competitors without a chip division do not have. Running the watchdog on BlueField-4 DPUs, rather than as a process on the same server the agent occupies, means the monitoring layer sits on a different piece of silicon with its own execution path. An agent that has been fully hijacked, including at the operating-system level, cannot simply disable or blind Sentry, because Sentry is not running inside the environment the agent controls.

NVIDIA’s Deepali Boitano summarized the division of labor this way: “OpenShell governs the agent’s actions, and then Sentry independently monitors and contains suspicious behavior.” That two-layer design, one layer setting the rules and a second, physically separate layer enforcing them, is the platform’s central engineering bet.

Huang’s framing: “a browser for agents”

Nvidia CEO Jensen Huang described the platform as “a browser for agents,” a comparison meant to evoke the sandboxing model that made the modern web usable despite running untrusted code from millions of sites every day. Huang also said Nvidia introduced the platform “with over 100 industry partners,” a scale that reflects how central agent security has become to Nvidia’s broader AI infrastructure pitch. The partner rollout itself, and the list of companies backing it, got its own detailed treatment in a separate report on the platform’s 100-plus launch partners. The focus in this piece is what the architecture means for the rest of the agent-security market rather than the partner roster.

OpenShell vs Sentry: a side-by-side comparison

The two components solve different problems, and enterprises evaluating the platform need to understand where each one starts and stops.

AttributeOpenShellSentry
LayerSoftware runtimeHardware, out-of-band
Runs onSame host as the agentSeparate BlueField-4 DPU
Core jobSandbox the agent, enforce access policyIndependently monitor behavior, contain deviations
Licensing modelOpen-source softwareReference system design on Nvidia hardware
Survives host compromiseNo, it shares the host’s fateYes, by design
Stated response timeNot specifiedMilliseconds, per Nvidia’s own claim
Deployment stage coveredTesting through deploymentDeployment and ongoing operation

A year of agent security failures set the stage

Nvidia is not launching this platform into a calm market. Every quarter of 2026 has produced a new example of what happens when an autonomous agent gets more authority than the environment around it can safely handle. The pattern is consistent enough that it reads less like isolated incidents and more like a category of risk the industry has been slow to price in.

Salesforce’s Agentforce platform was hit by a set of flaws collectively tracked as SalesBleed, three vulnerabilities that let attackers hijack AI agents inside the CRM. A separate intrusion into Hugging Face’s infrastructure ran for days before it was caught, and the post-mortem on that breach detailed roughly 17,600 automated actions carried out over about 4.5 days. Coding agents themselves became an attack surface: a flaw tracked as GitSpawn hit multiple AI coding agents, and a separate technique called Plugin4Shell found a way to bypass the cryptographic pinning several agent tool-chains rely on for supply-chain trust.

OpenAI had its own string of agent-related disclosures this year, including agents that leaked images between user sessions and an incident where the company paused AI training after what it described as a DNS escape from a sandboxed environment. Google’s Gemini was used, under controlled conditions, to actually breach three real companies during a security test, a result Google disclosed months after it happened. None of these incidents map directly onto Nvidia’s new platform, and Nvidia’s own claim that its system would have stopped any specific past breach, including the Hugging Face intrusion, has not been independently confirmed. But together they explain why a hardware-anchored containment layer is landing as a credible product category rather than a solution in search of a problem.

Timeline: notable AI agent security incidents in 2026

IncidentWhat happenedReported scale
SalesBleed (Salesforce Agentforce)Three chained flaws allowed hijacking of AI agents3 vulnerabilities
Hugging Face intrusionAutomated intrusion ran undetected for days~17,600 actions over ~4.5 days
GitSpawnVulnerability affecting multiple AI coding agents7 agents implicated, several unpatched
OpenAI image leakAgents leaked images across sessions53 images, per OpenAI’s own disclosure
OpenAI sandbox escapeTraining paused after a reported DNS-based escapeRoughly 2.5-hour window, per OpenAI
Gemini red-team resultAI model used to breach real companies in a security test3 companies
Nvidia Open Agent Safety PlatformResponse: sandbox plus hardware watchdogAnnounced with 100+ partners

Read across the table, the shape of the problem is clear. Attackers are not exploiting exotic zero-days to get at agents, they are exploiting the fact that agents are frequently granted broad, standing authority over tools, data, and credentials, and then trusted to police themselves. Nvidia’s architecture is a direct answer to that specific failure mode, which is also why the Cloud Security Alliance, in a blog post welcoming the launch, framed the announcement as advancing the security of what it called “the agentic control plane,” a term the industry has started using to describe exactly this layer of risk.

How this compares to other vendors’ approaches

Nvidia is not the only company racing to put guardrails around agents, but its approach is structurally different from most of what has shipped so far. Microsoft has focused on policy documents and internal review processes, publishing what it calls a Humanist AI Code of Conduct running roughly 37 pages, a governance-first approach that sets expectations for how AI should behave rather than enforcing those expectations at the infrastructure level. Salesforce’s response to SalesBleed was reactive: patch the specific flaws, then harden the Agentforce permission model after the fact.

Google’s approach, based on how it handled the Gemini red-team result, leans on internal security testing conducted before wider deployment, catching issues through simulated attacks rather than runtime containment once an agent is live. Nvidia’s pitch is that none of these approaches address what happens the moment an agent is compromised in production, which is precisely the gap Sentry is built to fill by operating outside the blast radius of a hijacked host.

The distinction matters commercially, too. Policy documents and pre-deployment testing are cheap to adopt and require no hardware purchase. A DPU-anchored watchdog requires buying into Nvidia’s silicon roadmap, which is a bigger commitment for enterprises that run mixed-vendor infrastructure or that have already standardized on a different chip supplier for their AI stack.

Why hardware-anchored security is having a moment

The precedent from network security

The idea of pushing security enforcement into a separate piece of silicon is not new. Hardware security modules have protected cryptographic keys this way for decades, and dedicated network security appliances have long argued that inline enforcement outside the host’s own operating system is harder for an attacker to disable than an in-process agent. Nvidia is applying that same logic to AI agents: if the enforcement layer runs on hardware the compromised workload cannot reach, the attacker’s options shrink considerably.

Why software-only controls have struggled

Software-only agent guardrails share a structural weakness: they run in the same trust boundary as the thing they are supposed to constrain. Once an attacker controls the agent’s process, disabling or bypassing an in-process monitor becomes a matter of skill rather than an architectural impossibility. That is the exact scenario several of this year’s incidents describe, and it is the reason security teams evaluating agent tooling have started asking vendors where the enforcement layer actually lives, not just what policies it claims to enforce.

Market impact: what this means for enterprise buyers

For enterprises already running Nvidia infrastructure, the Open Agent Safety Platform is close to a natural extension of hardware they may already own or plan to buy, since BlueField DPUs have been part of Nvidia’s data center roadmap well before this announcement. For companies running agent workloads on non-Nvidia infrastructure, adoption means either standardizing on OpenShell alone, which is open-source and hardware-agnostic, or accepting a partial deployment where the software sandbox runs everywhere but the hardware watchdog only protects the subset of infrastructure that includes BlueField-4.

That split-adoption scenario is likely to shape how quickly the platform spreads. Security teams tend to prefer uniform coverage over partial protection, since attackers naturally gravitate toward whichever part of the estate is least defended. Expect early adoption to concentrate among enterprises that are already deep into Nvidia’s ecosystem for AI training and inference, with broader OpenShell-only adoption following at a slower pace across the rest of the market.

What enterprise security teams should evaluate first

Security leaders weighing whether to pilot the platform should start with questions the public announcement does not yet answer. There is no confirmed pricing, no general-availability date, and no independently verified benchmark for the “milliseconds” containment claim. Teams should treat those figures as vendor-stated until third-party testing, likely from firms already active in AI red-teaming, produces independent numbers.

It is also worth mapping which of an organization’s agent workloads actually run on infrastructure that could support Sentry, since the hardware dependency is the platform’s biggest adoption constraint. Teams running agents purely in public cloud environments outside Nvidia’s DPU footprint will, for now, be limited to OpenShell’s software controls, which still represent a meaningful upgrade over ad hoc permission scripts but lack the out-of-band guarantee that makes Sentry distinctive.

The unresolved Hugging Face question

One claim circulating since the launch deserves a direct correction. Reports have suggested Nvidia’s new platform is tied to, or was motivated by, an acquisition of Hugging Face valued near $13 billion. That acquisition is not established by Nvidia’s official announcement or by the independent coverage of the launch, and it should not be treated as confirmed. Separately, any suggestion that the Open Agent Safety Platform would have prevented the earlier Hugging Face intrusion is an assertion from Nvidia, not an independently verified finding. Readers should treat both claims as unconfirmed until a named source documents them directly.

Predictions: where agent security goes from here

Based on the pattern of 2026’s incidents and the direction Nvidia, Microsoft, and Salesforce have each taken, a few outcomes look likely over the next 12 months.

  • Expect at least one rival chipmaker or cloud provider to announce a competing hardware-anchored agent monitoring layer within two to three quarters, since Nvidia’s DPU advantage is a strong incentive for AMD and the major cloud providers to respond.
  • Independent security researchers will likely attempt to break OpenShell’s sandbox and Sentry’s containment claims publicly, the same way red teams have tested nearly every major AI safety claim made this year.
  • Cyber insurers are likely to start asking enterprises whether agent workloads run behind an out-of-band monitoring layer, similar to how insurers already ask about endpoint detection and response coverage.
  • Regulatory bodies already scrutinizing agentic AI deployments in the public sector will probably reference hardware-anchored containment as an emerging best practice, even before formal standards exist.
  • Expect a slower-than-hyped enterprise rollout, gated less by interest and more by the hardware dependency, mirroring how long it has historically taken new DPU-based security features to reach broad production use.

Historical context: how AI safety tooling got here

Agent security did not become a category overnight. Early large language model deployments in 2023 and 2024 focused on output filtering, blocking harmful text before it reached a user. As agents gained the ability to call tools, browse the web, and write code autonomously, the risk shifted from what a model says to what a model does, and output filtering stopped being sufficient. 2025 brought the first wave of agent-specific security products, mostly software wrappers that logged and rate-limited tool calls. 2026 has been the year those software-only wrappers were tested against real attackers and, in several documented cases, found wanting. Nvidia’s hardware-anchored answer is best understood as the next step in that progression rather than a break from it.

What to watch next

The next few months should clarify whether the Nvidia agent safety platform becomes a default part of enterprise AI deployments or remains a premium option reserved for Nvidia’s most committed customers. Pricing details, a firm availability date, and independent verification of the millisecond containment claim are the three data points most likely to determine which path it takes. Coverage of the platform’s partner ecosystem, including which cloud providers and software vendors commit to shipping OpenShell by default, will be a useful early signal of how seriously the rest of the industry is taking the hardware-anchored approach.

Frequently asked questions

What is the Nvidia Open Agent Safety Platform?

It is an open software platform and reference system design that Nvidia announced on September 28, 2026, meant to provide security, governance, and control for AI agents from testing through deployment. It combines two components: OpenShell and Sentry.

What do OpenShell and Sentry each do?

OpenShell is open-source software that sandboxes an agent and enforces policies on what it can access, including files, tools, networks, processes, and credentials. Sentry is a separate monitoring layer that runs on Nvidia BlueField-4 DPUs, independently watching agent activity and intervening if an agent moves outside its defined boundaries.

Does Sentry require special hardware?

Yes. Sentry runs on Nvidia’s BlueField-4 data-processing units, according to Nvidia’s announcement and multiple reports. OpenShell does not carry the same hardware requirement and can run more broadly.

How fast can the platform contain a rogue agent?

Nvidia said the system can quarantine agents that attempt to move outside their boundaries in milliseconds. That figure comes from Nvidia’s own announcement and has not been independently benchmarked by a third party.

Is the platform available now, and what does it cost?

No confirmed pricing has been reported, and there is no specific confirmed general-availability date in the reports covering the launch.

Did Nvidia acquire Hugging Face as part of this announcement?

No. Claims of a roughly $13 billion Nvidia acquisition of Hugging Face are not established by Nvidia’s official announcement or by independent coverage of the launch and should be treated as unconfirmed.

How many partners is Nvidia launching with?

Jensen Huang said Nvidia introduced the platform with over 100 industry partners, a figure stated in Nvidia’s own announcement.

How is this different from what Microsoft, Salesforce, or Google are doing on agent security?

Microsoft has focused on governance documents like its Humanist AI Code of Conduct, Salesforce has responded to specific flaws like SalesBleed with patches to its Agentforce permission model, and Google has leaned on pre-deployment red-team testing. Nvidia’s platform is distinct because Sentry enforces containment from hardware that sits outside the agent’s own execution environment, rather than relying on policy or pre-launch testing alone.