On 28 September NVIDIA announced the Open Agent Safety Platform: OpenShell, an open-source sandbox runtime for AI agents, and Sentry, a hardware watchdog on BlueField-4 DPUs, with more than a hundred organisations behind it. Since we started Kyvvu we have worked on agent security by focusing on the agent harness: the one place where actions are carried out. It is good to see that NVIDIA and its consortium are now also focusing there. In this post we summarise what NVIDIA released and how it relates to other efforts in the field, including ours.

Some context: why is everyone looking at the harness?

As we have said before, an agent is effectively an LLM coupled with a harness that turns LLM output into actions.

An agent is a model plus a harness

So, the model produces text. The harness turns that text into actions: a file read, a shell command, an HTTP call, an e-mail. Hence, every effect on the outside world goes through the harness. In our view, the model is not trusted: it might hallucinate (leading to possibly harmful actions), and it can be talked into proposing anything by an adversarial party. The model’s behaviour is non-deterministic. The harness, however, is ordinary, deterministic code; it is the place in the agent architecture where we can be sure about what is happening.

The “agent break-out” incidents of the past months should be viewed in this light. We wrote before about the OpenAI–Hugging Face incident, where many stated that “an agent escaped”. But really, an LLM proposed actions that were undesirable, and it seems nobody truly checked the harness’s execution trace. The “break-out” could have been stopped at the level of the harness.

So, what did NVIDIA announce?

On the technical side, cutting through some of the marketing language on safe and secure agents, we end up with two concrete projects:

OpenShell (GitHub, Apache 2.0, documentation) is a runtime that places each agent in a sandbox under OS kernel controls. We wrote about SHADI before; OpenShell is very close to it, with some differences in features as far as we can tell.

Sentry is a reference design for an out-of-band watchdog on NVIDIA BlueField-4 DPUs, built on DOCA. It monitors agent traffic from a trust domain the host cannot reach and can quarantine an agent “in milliseconds”. OpenShell is software you can run today; Sentry would be the hardware enforcement layer underneath it, but it does not appear to be available yet.

NVIDIA named Anthropic, Salesforce, SAP, Scale AI and SpaceXAI as integration partners, Canonical, Red Hat and SUSE as operating-system vendors integrating OpenShell, and over a hundred further organisations, from Cisco, CrowdStrike and Microsoft to Citi and JPMorganChase. The research side is organised as the Open Secure AI Alliance under the Linux Foundation. The platform page has the full list of contributors.

OpenShell in one paragraph

An agent runs inside an OpenShell sandbox as a black box: OpenShell does not know what the agent is, only what the process does. Policy is declarative YAML over four domains.

  1. Filesystem: reads and writes outside declared paths are refused, enforced with Landlock.
  2. Process: privilege escalation and dangerous syscalls are blocked, with seccomp and an unprivileged process identity.
  3. Network: outbound connections are checked by a supervisor that runs outside the sandbox, at the level of binary, destination, method and path.
  4. Providers: credentials never enter the sandbox; the agent holds a placeholder and the supervisor substitutes the real credential only on requests to endpoints the policy allows.

Filesystem and process policy are fixed when the sandbox is created; network policy and credential attachments can be changed on a running sandbox, so an operator can cut an agent’s network or revoke a credential without restarting it. A gateway manages the lifecycle of a fleet of sandboxes on Docker, Podman, microVMs or Kubernetes. Details are in the architecture and policy docs.

For reference, SHADI, AGNTCY’s secure host runtime, is very simmilar: an OS-enforced sandbox around the agent process (Landlock on Linux, Seatbelt on macOS, AppContainer on Windows), with policy for read, write and network access, plus identity and secrets handling. We have run a coding agent under SHADI since September (earlier post). OpenShell adds the supervisor’s per-request network and credential checks and the fleet gateway; Sentry adds a stop from hardware the host cannot block.

What a sandbox sees, and what it does not

A sandbox decides at the level of the process: this path may be read, this syscall may run, this destination may be reached with this credential. It does not know which step the agent is on, what the agent read three steps earlier, or whether the e-mail it is about to send contains what it read from the CRM. A write is a write, whether the model asked for it or a bug caused it. That is by design, and it is what makes the boundary hard to route around: there is no intent for the model to argue with.

The same step, on two paths

Kyvvu builds an Agent Security Kernel: an in-process engine that evaluates every step the harness is about to take against the ordered history of the task so far, before the step runs. In security terms the ASK really is a reference monitor on agent behaviour, in the sense of Anderson’s 1972 report. A reference monitor has three requirements, and the three map onto three different pieces of work:

  • Complete mediation. Every action passes the check, and the check has to see actions as actions, not as text: a typed description of what the harness is about to do. That is what the Agent Action Grammar is for: an open vocabulary of step type, verb and properties that a harness emits, and that lets someone other than us check whether an integration is complete.
  • Verifiability. The engine has to be simple enough to inspect. The SDK is Apache 2.0, the engine is source-available under BSL 1.1, and both are on PyPI. No model is called in the decision path and model output is not parsed; the engine evaluates the action the harness is about to perform against rules the client wrote.
  • Tamper resistance. The thing being monitored must not be able to alter the monitor. This is the property an engine inside the harness cannot provide for itself. Our engine runs in the agent’s process; if the agent’s code can rewrite that process, or the policy file, the check is not defeated but gone. Our architecture docs state this as the trust model: the harness and the engine are trusted, the model is not, and a compromised process is out of scope.

That third property is what an OS-level sandbox provides, and it is why we encourage running a Kyvvu-secured agent in a sandbox like OpenShell or SHADI: Kyvvu decides whether a coding agent may execute code after an approval gate; the sandbox makes sure the code it then executes cannot write to the agent’s own files, to the policy, or to the engine.

So, to summarise: Sentry governs the (hardware) node, OpenShell (or SHADI) the process, and Kyvvu the action.

Where this leaves us

NVIDIA’s platform and SHADI both draw the boundary around the harness, which is where we think it belongs, and both provide what a monitor inside the harness needs underneath it. We have described the agent-security stack as several layers since our first post: alignment at the model side, content guardrails on messages, identity and credentials at the tool side, an action-level engine in the harness (Kyvvu!), and a sandbox around all of it. Each catches a different thing; none is complete on its own. The sandbox layer now has an implementation with a hundred companies behind it, which makes the rest of the stack easier to build.

If you want to see the action layer running, locally and without an account:

pip install kyvvu
kyvvu try

Docs at docs.kyvvu.com. If you run agents under OpenShell or SHADI and want Kyvvu inside the sandbox, we would like to hear from you.

#agent-security #openshell #shadi #runtime-governance #aag