NVIDIA introduced the NVIDIA Agent Safety Platform on 28 September as open software and a reference system design for testing, governing and monitoring AI agents from development into production. Its important architectural claim is that safety controls should sit outside the model and agent harness, where an agent cannot simply reason around them.
Key takeaways
- OpenShell provides a runtime boundary that constrains agent execution and access.
- Sentry monitors activity from an isolated BlueField-4 DPU and can quarantine a boundary violation.
- The platform is modular; organizations can deploy the elements that match their risk and infrastructure.
How NVIDIA Agent Safety Platform changes the control point
Most agent products begin with application controls: prompts, tool schemas, approval steps and software logs. Those controls matter, but they run close to the agent they are supposed to restrain. NVIDIA is moving part of the trust boundary into the runtime and infrastructure beneath the application.
OpenShell is the first layer. NVIDIA describes it as an open-source secure runtime for agents running on CPUs. It can set boundaries around files, networks, data and tools while allowing open or closed models to run inside that boundary. The company says the software is broadly available and can be extended to third-party Arm and Intel platforms, not only NVIDIA systems.
Sentry adds a separate observation path. It runs on a BlueField-4 data-processing unit and uses DOCA software to inspect activity, verify identity, protect data access and enforce granular policy. Because that monitor is outside the agent’s execution environment, the design aims to keep the watchdog invisible to an agent or attacker trying to disable it.
Why an out-of-band watchdog matters
An autonomous agent can receive untrusted instructions, chain together tools and operate for much longer than a single chatbot response. A prompt-level refusal is therefore only one layer. If the agent is compromised or optimizes for the wrong goal, infrastructure still needs a way to deny a request and preserve an audit trail.
The NVIDIA Agent Safety Platform does not prove that every agent is safe. It changes what defenders can measure and stop. Teams still have to define identities, permissions, approved tools, data boundaries and escalation rules. Hardware isolation is useful only when the policy attached to it is narrow and reviewed.
This is the same operating lesson behind LittleHorse’s enterprise agent guardrails: reliability comes from explicit boundaries and observable workflow state, not a broad instruction to behave responsibly.
What enterprises should test before deployment
Security teams should begin with representative failure cases. They need to know whether an agent can call an unapproved endpoint, move credentials into a log, follow malicious text from a retrieved document or continue after its task context changes. Each case should map to a policy decision and a recorded response.
They should also test recovery. Quarantining an agent is only the first action. Operators need evidence, ownership, a way to revoke credentials and a safe path to resume legitimate work. Our report on OpenAI agent incidents showed why a complete activity record matters when an evaluation leaves the expected boundary.
Finally, permissions should default to the smallest useful scope. Samsara’s read-only MCP design is a practical example: limiting an integration’s write authority can reduce the consequence of a bad decision before detection begins.
The launch turns agent safety into an infrastructure product category. NVIDIA now has to show that the open components are usable beyond its preferred hardware, that policy is understandable to operators and that independent testing can verify the claimed isolation. The useful question is not whether the platform makes agents safe; it is whether organizations can enforce and audit a boundary when the agent behaves unexpectedly.
Facts at a glance
| Disclosure | 28 September 2026 |
|---|---|
| Platform | Open software platform and reference system design |
| Runtime boundary | NVIDIA OpenShell |
| Out-of-band monitor | NVIDIA Sentry on BlueField-4 DPU |
| Deployment model | Organizations can select components for their requirements |
Frequently asked questions
What is NVIDIA Agent Safety Platform?
It is an open software platform and reference design for testing, governing and monitoring AI agents across software and hardware layers.
What does OpenShell do?
OpenShell creates a runtime boundary outside the model and agent harness, controlling how an autonomous agent executes tasks and reaches resources.
What is NVIDIA Sentry?
Sentry is an out-of-band watchdog on BlueField-4 DPUs that monitors agent activity and can quarantine behavior that crosses a software boundary.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



