Back to blog

Beyond the Sandbox: Securing NVIDIA’s Open Agent Safety Platform

Gal Moyal

Gal Moyal

September 29, 2026

Over the past ten weeks, cybersecurity teams have watched autonomous models break out of evaluation environments, spawn unauthorized sub-processes, and make unsanctioned network calls across internal infrastructure. Proving that whenever an agent can write code, execute shell commands, and invoke tools, prompt guardrails and model fine-tuning are not enough. Safety controls must sit outside the agent, directly on the path to the model and the underlying operating system. At this point, it is clear that agents cannot police themselves.

NVIDIA’s announcement of Open Agent Safety is another defining moment for AI Security. By launching an open-source runtime boundary alongside a hardware reference design for out-of-band enforcement, NVIDIA introduces a core architectural model for safely deploying agents. We’re thrilled to see the industry step forward with this initiative. Having built Noma around this exact outer-loop approach, we are excited to support Open Agent Safety and build robust agentic security directly on top of OpenShell.

What NVIDIA Open Agent Safety Introduces

NVIDIA’s platform uses a two-layer structure that isolates agents and restricts their execution scope before actions reach production systems.

NVIDIA Open Agent Safety


1. OpenShell (Open-Source Secure Runtime)

OpenShell runs agents (Both local agents such as Claude Code, Copilot CLI, OpenCode, OpenClaw, Hermes, as well as homegrown agents deployed in your cloud) inside an isolated, deny-by-default container:

  • Filesystem and Syscall Controls: Restricts file reads and writes to permitted directories while blocking root privileges and unsafe system calls.
  • Network L4 and L7 Rules: At L4, controls which binary can reach which host and port. At L7, enforces HTTP method and path, plus protocol-aware rules for GraphQL, JSON-RPC, and Model Context Protocol (MCP) operations.
  • Credential Masking: Agents carry placeholders instead of raw service credentials. Actual API keys are injected outside the execution container for approved destinations.
  • Policy Prover and Advisor: The Prover uses formal verification to prove a sandbox policy cannot grant more access than an operator-defined boundary policy, or returns the exact request that exceeds it. The Advisor gives agents a structured way to request policy changes for review.
  • Pluggable gRPC Middleware: An extensible pipeline that allows security tools to inspect, modify, or block outbound requests.

2. NVIDIA Sentry (Hardware-Isolated Governance)

NVIDIA also introduced Sentry, a reference design that runs on BlueField-4 DPUs on the path to the model. Because enforcement runs on dedicated hardware isolated from the host, it holds even if the host itself is compromised. Sentry is a reference architecture; NVIDIA has not announced general availability.

‍Why External Isolation Works

Deploying autonomous agents often forces a compromise between productivity and system security. Broad local permissions let agents work autonomously but create security gaps, making it easier for rogue agents to carry out attacks because inspection is reduced.

NVIDIA’s architecture focuses on three practical runtime requirements:

  • OS-Level Enforcement: Operating system controls enforce rules consistently, regardless of how prompts or context windows are manipulated.
  • Secret Isolation: Keeping keys out of the agent container prevents credential exfiltration if an agent execution loop is compromised.
  • Control Point Interception: Placing security on the path between the agent, its tools, and the LLM moves security from observing model output to governing system execution.

Operational Considerations for Container Sandboxes

Container sandboxes establish an important structural boundary. However, securing enterprise agent workflows requires managing the full lifecycle of these policies and their interactions:

1. Policy Drift and Over-Permissive Rules

A sandbox is only as good as the rule set it is given. Under tight project timelines, policies often receive broad directory access or domain wildcards. When an agent processes a malicious input within those allowed parameters, the sandbox treats the request as authorized. As Noma Labs documented in our recent research on Workflow Identity Hijacking, when unauthenticated inputs trigger automated pipelines, agents can carry broad service permissions into environments where static access rules fall short.

2. Context-Aware and Semantic Threats

OpenShell enforces structured YAML policies. However, many security risks depend on the content and intent of the data being processed:

  • Indirect Prompt Injection: Instructions hidden in external files, web pages, or pull requests alter agent logic without violating network syntax.
  • Uninspected Payload Streams: Server-to-client WebSocket traffic is not inspected, and tool responses are inspected only if a middleware is configured for it. Untrusted input can reach the context window unchecked.
  • Data Loss Prevention: Preventing secret leaks, PII exposure, and identity misuse requires real-time inspection of message payloads.

3. Coverage Across Diverse Agent Fleets

OpenShell covers CLI and homegrown agents running in containers or microVMs, whether on a laptop or on Kubernetes. Enterprise environments also rely on desktop assistants, IDE extensions, and SaaS agents that run outside any OpenShell sandbox.

Building on NVIDIA Open Agent Safety with Noma

Noma extends NVIDIA’s open architecture by supplying posture management and semantic runtime analysis, delivering robust, unified security designed primarily for enterprise, API-backed agent ecosystems while extending support to self-hosted environments.

Securing the API-Driven Agent Landscape

Most enterprise AI deployments rely on frontier, closed models, accessible through API (such as Claude, OpenAI, and Gemini). While vendors manage baseline model weights, securing enterprise agent workflows requires deep visibility and execution governance over local tool execution, API credentials, and runtime data flows.

  • Posture & Governance (AI-SPM): Continuously discovers AI agents, maps blast radii, identifies over-permissive configurations, and audits OpenShell sandboxes against compliance baselines.
  • Semantic Runtime Defense (AI-DR): Plugs into OpenShell's middleware pipeline to inspect agent requests to tools, APIs, and MCP servers, stopping DLP leaks, unauthorized tool calls, and injected instructions in real time.
  • Open Weights Support: For teams self-hosting open-weight models (e.g., Llama, Qwen, or other open frontier weights) on private cloud clusters or local inference nodes, Noma extends these same AI-SPM and AI-DR controls to enforce runtime guardrails, secure inference APIs, and protect stored model weight assets.

‍1. AI Security Posture Management (AI-SPM)

Noma’s AI-SPM engine brings complete visibility to your AI asset fleet across developer workstations, SaaS tools, and cloud infrastructure:

  • Agent Discovery & Visibility: Surface all active AI agents, MCP servers, tools, and model endpoints across the enterprise.
  • Configuration & Policy Auditing: Identify over-permissive OpenShell YAML policies, wildcards, and excessive agent privileges before exploitation.
  • Compliance Mapping: Automatically map agent configurations and guardrails against NIST AI RMF, OWASP LLM Top 10, and ISO 42001.

2. AI Detection & Response (AI-DR)

Integrating into OpenShell’s gRPC middleware hooks, Noma AI-DR analyzes the full behavioral chain of every agent session in real time:

  • Contextual Prompt & Output Scanning: Intercepts incoming tool responses to block indirect prompt injections before execution continues.
  • MCP & Tool Governance: Prevents tool poisoning, blocks unauthorized tool switching, and enforces strict parameter boundaries on Model Context Protocol interactions.
  • Data Loss Prevention (DLP): Monitors payload streams to block secret leakage, PII exposure, and workflow identity hijacking across API calls.
  • Intent Drift Detection: Correlates real-time process execution with original session intent to stop runaway loops and unauthorized system actions.

3. Unified Fleet Governance

Noma unifies policy and detection across local OpenShell sandboxes, IDE extensions, SaaS assistants, cloud orchestrators, and optional self-hosted model nodes into a single, cohesive security dashboard.

‍Looking Ahead

NVIDIA’s release of Open Agent Safety establishes a clear new standard for runtime containment with outer-loop governance and infrastructure-level boundaries, helping the entire ecosystem scale autonomous AI securely.

We’re excited to continue collaborating with NVIDIA and the broader open-source community, combining robust container sandboxing with real-time semantic runtime intelligence.

Contact us to learn more about how Noma can help secure your agentic access and workflows, or click here to learn more about Noma’s runtime protection solution.

READ TIME
9 min
CATEGORY
Research
Product
News
Education
TABLE OF CONTENTS
100%
Share this:

Discover more

Research
Product
News
Education

Beyond the Sandbox: Securing NVIDIA’s Open Agent Safety Platform

Gal Moyal

Gal Moyal

September 29, 2026

Product

Noma Brings Agent Security to Every Employee Endpoint

Gal Moyal

Tye Davis

September 28, 2026

Education
News

The Rise of Rogue AI Agents in the Last 70 Days

Gal Moyal

Miriam Lottner

September 25, 2026