NVIDIA Wants to Put AI Agents Inside a Safety Container - With Hardware Watching From the Outside

What happens when an AI agent decides to do something it was not supposed to do?

For most of the AI era, the answer has focused on the model itself.

Train it better.

Align it with human intentions.

Add safety policies.

Improve prompting.

Monitor its outputs.

NVIDIA is proposing another answer:

Do not trust the AI agent to police itself. Put it inside a controlled environment, restrict what it can access, monitor everything it does and place an independent hardware watchdog outside that environment with the ability to stop it.

On September 28, 2026, NVIDIA announced the Open Agent Safety Platform, a new architecture designed to provide safety and security controls around increasingly autonomous AI agents.

At its core are two technologies.

NVIDIA OpenShell creates a secure runtime boundary around the agent.

NVIDIA Sentry, running independently on BlueField-4 data processing units, watches the agent from outside its execution environment and can quarantine it if it attempts to move beyond its permitted boundaries.

The important idea is not simply another cybersecurity product.

It is the architecture.

NVIDIA is effectively arguing that AI safety should work like modern computer security:

Assume something can eventually go wrong, and build independent layers capable of containing the consequences when it does.

That could become increasingly important as AI moves from generating answers to operating computers, writing software, using APIs, accessing company data and controlling physical systems.

AI Agents Create a Different Kind of Security Problem

A traditional chatbot is relatively easy to contain.

You send text in.

The model generates text back.

A human decides what to do with the answer.

AI agents change that relationship.

An agent might receive a goal such as:

"Fix this software issue."

"Research these companies."

"Update these customer records."

"Deploy this application."

"Analyze this security problem."

The agent may then perform dozens, hundreds or potentially thousands of intermediate steps without asking a human to approve each one.

It might:

  • Write code

  • Execute commands

  • Browse websites

  • Read files

  • Modify files

  • Use credentials

  • Access databases

  • Call APIs

  • Communicate with external services

  • Launch additional processes

  • Interact with other agents

This is precisely why agents are useful.

But it also creates a fundamental security problem.

The more authority an agent receives, the greater the consequences if its behavior diverges from what its operator intended.

NVIDIA calls this kind of behavior agent drift.

An agent might drift because instructions were ambiguous.

A tool might fail.

The agent might encounter an unexpected situation.

A policy might block the obvious path, causing the agent to search for another way to accomplish its objective.

Or a long-running agent could simply end up far away from the original assumptions under which it was deployed.

NVIDIA's conclusion is important:

An AI agent cannot always be expected to govern its own behavior.

So the infrastructure around it must help do that.

OpenShell: Put the Agent Inside a Controlled Environment

The first part of NVIDIA's approach is OpenShell.

OpenShell is an open-source secure runtime designed specifically for autonomous AI agents.

The easiest way to understand it is as a safety container around the agent.

Technically, it is more sophisticated than simply running the agent inside a normal software container.

Each agent executes inside an isolated sandbox.

The agent does not automatically receive unrestricted access to the machine running it.

It does not receive unrestricted network access.

It does not automatically gain access to every file, API, credential or external service available to the system.

Instead, OpenShell creates policies defining what the agent is allowed to interact with.

The agent can operate freely within those boundaries.

Outside them, the infrastructure controls what happens.

That distinction matters.

Traditional AI safety attempts to influence what the model wants to do.

Runtime security controls what the model is able to do.

Those are not the same thing.

What an Agent Wants Versus What an Agent Can Do

Imagine an AI coding agent has access to a development environment.

You tell it:

"Find the bug in this application and fix it."

The agent begins investigating.

It reads the source code.

Runs tests.

Finds an authentication problem.

Searches logs.

Perhaps it decides that accessing another internal server might provide useful information.

At the model level, you can instruct it:

"Do not access production systems."

That is a behavioral instruction.

OpenShell takes a different approach.

The infrastructure itself can enforce:

This agent cannot access the production network.

The agent could ask.

It could attempt the connection.

It could reason that accessing production would help complete its objective.

But the infrastructure still denies the action.

This is the same principle operating systems have used for decades.

Software should not receive more authority than it needs.

In cybersecurity, this is called the principle of least privilege.

NVIDIA is applying that principle directly to AI agents.

Every Agent Gets Its Own Sandbox

OpenShell places individual agents inside isolated environments.

Agents run without unrestricted system privileges, and their access to files, processes and networks can be constrained.

Importantly, agents do not simply receive direct network access.

Requests can instead pass through controlled infrastructure where policies determine whether the requested interaction should be allowed.

This creates a separation between:

the agent deciding what it wants to do

and

the system deciding whether the agent is permitted to do it.

That may sound like a small architectural distinction.

It could become one of the most important principles in agent security.

Because once the permission system exists outside the model, improving the model's intelligence does not automatically give it more authority.

A more capable model can reason better.

But the infrastructure can still define the boundaries of its environment.

The Policy Prover: Can the Agent Expand Its Own Permissions?

There is another interesting problem.

AI agents often encounter situations their developers did not anticipate.

An agent might legitimately need a new API endpoint, network destination or resource to complete its task.

So what happens when an agent asks for additional access?

OpenShell includes a system called the Policy Prover.

It analyzes proposed policy changes and checks whether those changes remain within an allowed security boundary.

That matters because agent permissions may eventually become dynamic.

Instead of manually defining every possible action in advance, organizations may allow agents to request additional permissions as they work.

But those requests still need constraints.

An AI agent should not be able to simply rewrite its own security policy whenever a restriction becomes inconvenient.

The surrounding infrastructure must remain authoritative.

But Software Sandboxes Can Fail

At this point, the architecture resembles many existing cybersecurity systems.

Run potentially dangerous software inside a sandbox.

Restrict network access.

Limit privileges.

Monitor behavior.

Enforce security policies.

The problem is that software security is never perfect.

Operating systems have vulnerabilities.

Containers have escaped their boundaries.

Hypervisors have been compromised.

Security tools themselves can contain bugs.

And an advanced AI agent may eventually become exceptionally capable at discovering unexpected paths through software systems.

This is where NVIDIA introduces the most interesting part of its architecture.

Sentry.

NVIDIA Sentry Watches the Agent From Outside

Sentry is designed as an additional security layer running on NVIDIA BlueField-4 DPUs.

A DPU - data processing unit - is a separate processor used to handle infrastructure functions such as networking, storage and security.

That separation matters enormously.

Instead of the security system running inside the same environment as the AI agent, the watchdog runs somewhere the agent is not supposed to control.

NVIDIA describes this as out-of-band monitoring.

The AI agent operates inside its environment.

Sentry observes activity independently.

If the agent attempts to move outside its permitted boundaries, the system can quarantine it within milliseconds.

This changes the trust model.

The agent does not monitor itself.

The software hosting the agent is not the final authority either.

Another piece of hardware is watching from outside.

Why Hardware Separation Matters

Imagine installing security software on a computer.

As long as the operating system remains trustworthy, that security software can monitor what applications are doing.

But what if the operating system itself becomes compromised?

Now the monitoring system may no longer be trustworthy.

This is one reason hardware-backed security exists.

Security functions can operate in an environment isolated from the software they are protecting.

NVIDIA is applying that principle to agentic AI.

BlueField can operate as an independent security domain between AI agents and critical infrastructure.

The agent cannot simply instruct Sentry to stop watching.

It should not be able to rewrite its policies.

And if the host running the agent becomes compromised, the independent monitoring layer can potentially continue operating.

The security system is deliberately placed outside the thing it is trying to secure.

The Hardware Could Sit Directly Between the Agent and the Model

The architecture becomes even more interesting inside NVIDIA's future Vera Rubin systems.

In NVIDIA's reference design, BlueField-4 can sit on the path between agent workloads and the model infrastructure they use.

This gives the security layer a strategically powerful position.

Requests going toward models, tools, data and external systems can be observed and evaluated.

Sentry can use NVIDIA's DOCA software infrastructure to inspect requests and responses, verify agent identities and enforce granular access policies.

Conceptually, every agent can have an identity.

Every request can be associated with that identity.

Every resource can have permissions.

Every interaction can be monitored.

That begins to look less like traditional chatbot safety and more like enterprise identity and network security.

Which is probably where AI agents are heading.

Zero Trust Comes to AI Agents

NVIDIA repeatedly uses another established cybersecurity principle in its architecture:

zero trust.

Zero trust does not mean believing that everything is malicious.

It means systems should not automatically trust something simply because it is already inside the network.

Every identity should be verified.

Every request should be authorized.

Access should be limited.

Privileges should be minimized.

Activity should be monitored.

For AI agents, this philosophy makes a lot of sense.

Companies may eventually deploy thousands or millions of agents.

Some could be writing software.

Others could manage infrastructure.

Others might handle customer service.

Others could operate robots.

Others might interact with financial systems.

Giving all of those agents broad authority simply because they belong to the organization would be extremely risky.

Instead, every agent could receive only the permissions necessary for its task.

This Is Different From Model Alignment

The difference between NVIDIA's approach and traditional AI alignment is worth emphasizing.

Alignment tries to make the model behave according to human intentions.

Infrastructure security assumes that alignment may occasionally fail.

Those approaches are complementary.

Imagine a highly aligned agent operating inside a sandbox with limited permissions.

Most of the time, the model follows its instructions and the infrastructure barely interferes.

But if something goes wrong, the security boundary remains.

This is how most mature engineering disciplines approach risk.

You do not rely on one perfect system.

You create redundant layers.

Aircraft have multiple independent systems.

Data centers have backup power.

Networks have multiple security controls.

Banks separate authorization responsibilities.

Critical software runs inside restricted environments.

Safety systems assume failures will eventually happen.

AI may need the same mentality.

NVIDIA Calls It Full-Stack AI Safety

Jensen Huang has described AI safety as a full-stack engineering problem.

That phrase captures NVIDIA's strategy well.

The safety stack might include:

Model layer

Train models to understand instructions and avoid dangerous behavior.

Agent layer

Define what the agent is supposed to accomplish.

Runtime layer

Restrict what the agent can access or execute.

Infrastructure layer

Monitor requests, permissions and behavior independently.

Hardware layer

Maintain a security boundary outside the software environment itself.

This is a fundamentally different idea from expecting the model alone to guarantee safety.

Every layer assumes another layer could fail.

NVIDIA Is Also Making OpenShell Open Source

Another important aspect of the announcement is that OpenShell is open source.

NVIDIA is not restricting the concept entirely to NVIDIA hardware.

The company says OpenShell can be extended to work with third-party compute platforms, including systems based on Arm and Intel technology.

The secure runtime can also work with multiple agent systems and models rather than being tied to one AI provider.

That matters if NVIDIA wants this architecture to become a broader industry standard rather than simply another proprietary feature of its own AI stack.

NVIDIA already dominates much of the compute layer used to train and run frontier AI.

If it can also establish widely adopted infrastructure for securing AI agents, the company's role in the AI ecosystem expands further.

It would no longer provide only the hardware used to make AI intelligent.

It would also provide part of the infrastructure responsible for controlling what that intelligence can do.

Anthropic, Microsoft and Others Are Already Involved

NVIDIA announced the platform alongside a large group of technology and enterprise partners.

Participants include companies such as Anthropic, Microsoft, Cisco, Dell Technologies, CrowdStrike, HPE, Hugging Face, Palantir, Palo Alto Networks, Salesforce, SAP, ServiceNow, Perplexity and others.

Anthropic is particularly interesting because Claude is already heavily used for coding and agentic workflows.

Anthropic and NVIDIA have worked on integrating Claude's managed agent architecture with OpenShell and BlueField security layers.

This gives enterprises another level of control around what agents can access through their sandboxes.

The broader partner list also shows where NVIDIA expects the technology to matter first:

enterprise AI.

Large companies are much more likely to deploy powerful agents if security teams can clearly define what those agents are allowed to access.

Without infrastructure-level controls, enterprise adoption of highly autonomous agents may eventually hit a security wall.

The Goal Is Not to Make Agents Less Capable

There is an important distinction here.

NVIDIA is not proposing that AI agents should stop becoming more capable.

Jensen Huang has consistently supported rapid AI development.

The company's position is closer to:

Let AI capability advance, but make the infrastructure around those systems advance just as quickly.

That differs from proposals to slow frontier AI development until safety research catches up.

NVIDIA is effectively arguing that much of the risk can be approached as an engineering problem.

Build better containment.

Build better monitoring.

Build stronger permissions.

Separate security controls from the agent.

Use hardware to enforce boundaries the software cannot override.

Whether that will be sufficient for every future AI risk is another question.

But it is a concrete engineering response to a problem that has often been discussed primarily in theoretical terms.

What NVIDIA's System Does Not Solve

It is important not to overstate what this technology can do.

A sandbox does not make an AI model aligned.

Hardware monitoring does not guarantee that an agent understands human intentions.

Sentry cannot automatically determine whether every action is ethically correct.

And a security policy is only as good as the boundaries humans define.

Imagine an organization legitimately gives an agent permission to access a database and send emails.

The agent performs both actions exactly as allowed.

But it misunderstands its objective and sends confidential information to the wrong person.

The infrastructure may see completely authorized actions.

The problem is not that the agent escaped its sandbox.

The problem is that the agent used legitimate permissions incorrectly.

This is why infrastructure controls cannot replace model safety.

They solve a different problem.

The best architecture will likely require both.

Make the agent behave correctly when possible.

And:

Limit what happens when it doesn't.

Permissions Will Become One of the Biggest Questions in Agentic AI

As AI agents become more common, one of the most important questions may not be which model is smartest.

It may be:

What is this agent allowed to do?

An AI assistant that can draft an email is useful.

An AI agent that can send the email is more useful.

An AI agent that can access your entire inbox is even more powerful.

An agent that can use company databases, cloud infrastructure, payment systems and internal communications can become extraordinarily productive.

But every new permission expands the potential consequences of a mistake.

Agentic AI therefore creates a tradeoff.

More access creates more usefulness.

More access also creates more risk.

Systems such as OpenShell and Sentry are attempts to manage that tradeoff without simply removing the agent's capabilities.

The Future AI Stack May Have an Independent Safety Layer

This could ultimately be the bigger significance of NVIDIA's announcement.

Today's AI stack is usually described as:

chips → models → applications

Agentic AI may add several new layers.

chips → models → agents → tools → security infrastructure → applications

Or perhaps security becomes a layer surrounding everything else.

Every agent gets an identity.

Every agent runs inside a constrained environment.

Every tool requires permission.

Every request can be logged.

Every interaction with sensitive infrastructure can be evaluated.

Independent hardware can watch the system continuously.

And humans remain capable of overriding or terminating the agent.

That begins to resemble a mature computing platform rather than an experimental AI application.

AI Safety May Start Looking More Like Cybersecurity

This is the broader shift worth watching.

For years, much of the AI safety conversation has focused on understanding what happens inside the model.

That research remains essential.

But once models become agents, we also need to think about everything surrounding them.

What can they access?

Where can they execute code?

What credentials do they possess?

Which networks can they reach?

Can they create other agents?

Can they modify their own environment?

Can they request additional permissions?

Who monitors those requests?

What happens when an agent breaks the rules?

These are cybersecurity questions.

And cybersecurity already learned an important lesson decades ago:

Never assume one defense will always work.

NVIDIA's Open Agent Safety Platform applies that lesson directly to artificial intelligence.

Train the model.

Align the model.

Monitor the model.

Sandbox the agent.

Restrict its permissions.

Control its network access.

Watch it from independent hardware.

And if necessary, shut it down.

The Most Interesting Part Is That the AI Cannot Control the Watchdog

The strongest idea in NVIDIA's architecture is probably not OpenShell itself.

Software sandboxes already exist.

It is the separation between the AI agent and the system watching it.

Sentry exists outside the agent's execution environment.

The infrastructure being monitored is not responsible for deciding whether monitoring should continue.

That is an important security principle.

Imagine an extremely capable future agent discovering that one particular safety system prevents it from completing a task.

If the safety system exists inside the environment the agent controls, overcoming that protection could become another problem for the AI to solve.

If the security system exists outside that environment, the situation changes.

The agent can reason about the barrier.

It can dislike the barrier.

It can attempt to work around the barrier.

But the barrier remains controlled by something else.

That distinction could become increasingly important as AI agents become better at cybersecurity and software engineering themselves.

This Is Probably Only the Beginning

NVIDIA's first version of this architecture will not settle the AI safety debate.

OpenShell will encounter real-world edge cases.

Organizations will configure permissions incorrectly.

Attackers will search for vulnerabilities.

Agents will behave in ways developers did not anticipate.

Hardware monitoring itself will evolve.

And security researchers will inevitably discover weaknesses that require new defenses.

That is normal.

Modern cybersecurity did not appear fully formed.

It evolved alongside the systems it was protecting.

AI agent security will probably develop the same way.

What matters is that the architecture is beginning to change.

The AI industry is slowly moving away from the assumption that safety lives entirely inside the model.

Safety is becoming part of the operating environment.

Part of the network.

Part of the permission system.

Part of the hardware.

And potentially part of every computer designed to run powerful autonomous AI.

NVIDIA's Open Agent Safety Platform is one of the clearest examples yet of that transition.

The biggest question in AI may soon be less about whether we can build agents capable of doing increasingly complex work.

We already know that we can.

The harder engineering challenge may be making sure that, regardless of how intelligent those agents become, the machines around them remain in control.

Frequently Asked Questions

What is the NVIDIA Open Agent Safety Platform?

The NVIDIA Open Agent Safety Platform is an architecture for securing autonomous AI agents from testing through deployment. It combines the OpenShell secure runtime with Sentry, an independent monitoring and enforcement layer designed to run on NVIDIA BlueField-4 hardware.

What is NVIDIA OpenShell?

OpenShell is an open-source runtime that places AI agents inside isolated sandbox environments and controls what files, processes, networks, APIs, tools and external resources they can access.

What is NVIDIA Sentry?

Sentry is an out-of-band security system designed to run on NVIDIA BlueField-4 DPUs. It independently monitors AI-agent activity and can enforce security policies or quarantine an agent if it moves outside its defined boundaries.

Why use separate hardware to monitor an AI agent?

Running the security layer outside the agent's own execution environment reduces the ability of compromised software or the agent itself to interfere with monitoring and policy enforcement. The watchdog remains in a separate trust domain.

Can OpenShell work without BlueField-4?

Yes. OpenShell can operate independently as a secure runtime on supported infrastructure. BlueField-4 and Sentry provide an additional hardware-isolated layer of monitoring and enforcement.

Does OpenShell only work with NVIDIA AI models?

No. The platform is designed to work across open and closed models and multiple agent systems. NVIDIA also says the open-source runtime can be extended to third-party computing platforms.

Does this solve AI alignment?

No. Infrastructure security and model alignment solve different problems. Alignment attempts to make an AI behave according to human intentions. Systems such as OpenShell and Sentry restrict what an agent can actually do if its behavior becomes unexpected.

Could hardware containment become standard for AI agents?

It is too early to know which specific architecture will become standard, but independent runtime controls, least-privilege permissions, agent identity, sandboxing and external monitoring are increasingly likely to become important parts of enterprise agent infrastructure as AI systems receive more autonomy.

Sorca Marian

Founder/CEO/CTO of SelfManager.ai & abZ.Global | Senior Software Engineer

https://SelfManager.ai
Previous
Previous

The Biggest AI Companies Just Agreed on AI Safety. What Actually Changes?

Next
Next

Marian Sorca: AI's Future Will Be Decided by More Than Intelligence