OpenShell is a policy-enforced runtime for autonomous AI agents, a Rust project from NVIDIA that sits between an agent and the machine it runs on. It is for people who have to let agents do real work — read files, install packages, call APIs, use credentials — on hardware and accounts they are answerable for: platform and security engineers running a fleet of agents, and developers who would rather not have a coding agent with shell access be the weakest point on their laptop. You declare what each agent may touch in a policy, and OpenShell enforces that declaration instead of trusting the agent to behave.
What it does
The problem is easy to state and unpleasant to solve. An agent is only useful when it can act on your system, and the usual way to give it that ability is to hand over your shell, your home directory and your API keys. OpenShell keeps the capability and removes the blanket trust.
In practice it gives you:
- A per-agent policy that says which files the agent can reach, which system calls it can make, and where it may connect on the network.
- An isolated sandbox per agent, which by the project's own account a single command spins up.
- Credential handling in which the agent never sees your real secrets.
- Formal verification of a policy change, run before the change takes effect, to establish what the new policy would allow.
- A review gate: access the agent has not been granted yet — a new host, for example — waits for a human to approve it rather than simply being taken.
How it works
Two mechanisms, applied at different moments. At runtime, OpenShell instruments the kernel: every file access, system call and network connection made from inside an agent's sandbox is checked against that agent's policy, and a connection is evaluated before it leaves the sandbox. Enforcement therefore does not depend on the agent's cooperation, on a wrapper library it could route around, or on prompt instructions it might ignore. Secrets stay outside the sandbox, so a prompt injection that talks an agent into exfiltrating a key finds no key to read.
At change time, it uses formal verification. Policies drift — someone needs one more directory, one more outbound host — and the risky moment is usually the edit rather than the steady state. OpenShell checks what an amended policy would permit before applying it, which turns "this diff looks fine" into something closer to a proof obligation. Together with the human-approval step for newly requested access, the intent is that widening an agent's reach becomes a deliberate act with a record, not a side effect of the agent hitting a wall.
Getting started
The repository is Apache-2.0 licensed and the implementation is Rust; there is also a published openshell package on PyPI and full documentation hosted by NVIDIA, which is where the installation path and the policy syntax are written up. The project is at 0.1.x, described as bringing a stable release cadence, new isolation primitives, a larger extension surface and new APIs, with a dedicated 0.1.0 upgrade guide for anyone who started earlier. Read that guide first if you picked the project up before the current line.
When to use it / when not
Reach for it when your agents hold credentials, make outbound calls, or touch a filesystem containing anything you would mind losing, and when you need to be able to say precisely what a given agent was permitted to do. It also fits when several agents run side by side and you want each one's blast radius bounded separately.
It is less of a fit if your agents are already read-only and confined to a disposable container you throw away after each run, since you would be layering policy over an isolation story you already have. Kernel-level enforcement also assumes you control the kernel, which rules out environments where you cannot instrument the host. And the project is young — created in February 2026 and still on a 0.1 line — so expect the API surface to keep moving.
Anyone already letting agents act on production systems, or on a developer machine full of live keys, should take OpenShell seriously. It comes from NVIDIA, it has drawn more than ten thousand stars and over thirteen hundred forks in roughly seven months, and it treats containment as an enforcement problem rather than a prompting one. That last point is the reason to care: policy checked in the kernel and secrets the agent cannot read are properties you can reason about, while asking a model nicely is not. If your agents are experiments in a scratch directory, you can wait for the APIs to settle. If they are not, a young version number is the smaller risk.