TL;DR

OpenShell (NVIDIA/OpenShell) is an Apache-2.0, Rust runtime that runs each AI agent in a sandbox where filesystem access, system calls and every outbound network connection are checked against a declared policy — in the kernel, not in a system prompt. It has 14,567 stars and 1,669 forks as of October 3, 2026, and it is not a weekend project: the repo was created on February 24, 2026, ships tagged stable releases roughly weekly (v0.1.2 landed September 28), and carries SDKs for Python, TypeScript, Go and Rust.

The reason it matters is narrow and real. Agents are useful exactly when they can read your files, install packages, and use your credentials — which is also the moment they become dangerous. Every mitigation most teams actually deploy today is a request: “do not touch production,” “do not push to main.” OpenShell replaces the request with an enforcement boundary. An agent under OpenShell policy does not see your real API keys at all; the supervisor injects them only into requests already bound for an approved endpoint.

Two honest caveats before you install. First, the headline-grabbing hardware half of NVIDIA’s announcement — Sentry — needs a BlueField DPU, so almost nobody reading this will run it; the OpenShell runtime itself needs no NVIDIA hardware and no GPU. Second, this is v0.1.x with 537 open issues, and interfaces marked Experimental can change in a patch release. It is production-shaped, not yet production-proven.

Verdict: if you run more than one or two autonomous agents, this is the first open-source project that treats agent containment as a policy problem with an audit trail instead of a prompt-engineering problem. Start it in audit mode, read the logs for a week, then switch to enforce.

Quick reference

Repogithub.com/NVIDIA/OpenShell
Stars14,567 · 1,669 forks (2026-10-03); created 2026-02-24
LicenceApache-2.0
LanguageRust (CLI, gateway, supervisor, sandbox runtime)
Latest stablev0.1.2 (2026-09-28); weekly cadence, targets Tuesday
Open issues537
PlatformsLinux amd64/arm64 ✅ · macOS Apple Silicon (Docker Desktop) ✅ · Windows WSL 2 ⚠️ experimental
NeedsDocker, Podman, or host virtualization. No GPU, no NVIDIA hardware
Installcurl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh
SDKsuv add openshell · @nvidia/openshell-sdk · go get .../sdk/go · cargo add openshell-sdk
Policy engineOPA for connection-level rules, plus an L7 proxy for HTTP method/path
Docsdocs.nvidia.com/openshell/latest (append .md to any page for clean Markdown)

What OpenShell actually does

Three components, and it is worth keeping them straight because the docs assume you have:

  • Gateway — the control plane. It holds policies, providers (credentials) and sandbox state. You can run it locally via the installer, or on Kubernetes via Helm.
  • Supervisor — sits outside the workload boundary. It owns upstream connections and the real credentials, and it is what mediates TCP and DNS on the agent’s behalf.
  • Sandbox — where the agent runs. Inside it, openshell-sandbox (a static musl binary, staged in automatically) starts and owns the agent’s process tree and observes executable identity, so a rule can say “/usr/local/bin/claude may reach this host” rather than “something in this container may.”

Enforcement happens two ways. At runtime, kernel controls confine which files the agent can touch and which syscalls it can make, and every connection passes a policy check before leaving the sandbox. Separately, when you change a policy, OpenShell runs a prover over the diff to flag newly granted access — reaching a new host with credentials, or calling a new API method — and holds those changes for human review. That second half is the genuinely unusual idea: it is review tooling for the permission grant itself.

Setup: default-deny in about five minutes

Install, then create a sandbox:

curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh
openshell sandbox create --name demo \
  --from registry.example.com/team/agent-tools:1.0 \
  --no-auto-providers

The default sandbox image is a minimal Ubuntu with no agent installed — a detail that trips people up. --no-auto-providers skips the credential-wiring prompt. You land in a shell inside the sandbox:

sandbox@demo:~$

Now the part that sells the project. Try to reach the internet:

curl -s https://api.github.com/zen
curl: (56) Received HTTP code 403 from proxy after CONNECT

Nothing is allowed by default. No --network none to remember, no egress rules you forgot to write. And the denial is not silent — from a second terminal on the host:

openshell logs demo --since 5m --source sandbox
[1775014132.690] [sandbox] [OCSF] [ocsf] NET:OPEN [MED] DENIED
  /usr/bin/curl(64) -> api.github.com:443 [policy:- engine:opa]
  [reason:network connections not allowed by policy]

Note the shape of that line: the destination, the binary that tried it, the engine that decided, and the reason. It is OCSF-formatted, which means it goes into a SIEM without a custom parser.

Granting access, narrowly

Here is where OpenShell separates itself from “run the agent in Docker.” You do not grant “network.” You grant a method class on a host, to a binary:

openshell policy update demo \
  --rule-name github_api \
  --binary /usr/bin/curl \
  --add-endpoint api.github.com:443:read-only:rest:enforce \
  --wait

That endpoint spec is host, port, access preset, protocol, enforcement mode. rest tells the proxy to terminate TLS and inspect each HTTP request; read-only permits GET, HEAD and OPTIONS; enforce blocks everything else. It compiles to this in the policy’s network_policies section:

network_policies:
  github_api:
    endpoints:
      - host: api.github.com
        port: 443
        protocol: rest
        enforcement: enforce
        access: read-only

No restart. Back in the sandbox shell the same GET now works:

curl -s https://api.github.com/zen
# Anything added dilutes everything else.

And a write does not:

curl -s -X POST https://api.github.com/repos/octocat/hello-world/issues \
  -H "Content-Type: application/json" -d '{"title":"oops"}'
{...,"error":"policy_denied",...,"policy":"github_api",
 ...,"rule":"POST /repos/octocat/hello-world/issues",...}

The TCP connection succeeded — api.github.com is permitted — but the proxy read the HTTP method and returned 403. An agent with that policy can read every public repo on GitHub and cannot open an issue, push a commit, or change anything. That distinction is the whole product, and you cannot express it with container networking or a firewall rule.

L7 denials log separately:

[1775014140.412] [sandbox] [OCSF] [ocsf] HTTP:POST [MED] DENIED
  POST http://api.github.com:443/repos/octocat/hello-world/issues
  [policy:github_api engine:l7] [reason:L7_REQUEST deny ...
   reason=POST /repos/octocat/hello-world/issues not permitted by policy]

One operational gotcha worth internalising: policy events are INFO-level regardless of severity. If your log pipeline filters --level warn, you will drop your entire audit trail and conclude nothing is happening.

The workflow that actually works

Do not try to write a correct policy up front. Set enforcement: audit instead of enforce, let the agent run, and read what it reached for:

network_policies:
  pypi:
    endpoints:
      - host: pypi.org
        port: 443
        protocol: rest
        enforcement: audit
        access: read-only

Audit mode logs violations without blocking, so you get a real dependency inventory from a real run instead of a guess. Then flip to enforce. The repo ships examples/policy-advisor and an advisor in the docs for exactly this loop, and examples/sandbox-policy-quickstart/demo.sh runs the whole walkthrough above non-interactively in under a minute.

Agent integration, and the agent-driven angle

OpenShell is built agent-first, including its own development. Two things follow from that.

It ships skills for coding agents:

npx skills add NVIDIA/OpenShell

That installs four skill bundles — openshell-cli, generate-sandbox-policy, debug-inference, debug-openshell-cluster — which teach Claude Code, Codex or OpenCode to drive the CLI and write policies. They work without a source checkout.

And the examples/ directory is unusually practical for a 0.1 release: agent-driven-policy-management, local-inference, jupyter-sandbox, multi-agent-notepad, bring-your-own-container, codex-app-server, governance-interceptor, supervisor-middleware-content-guard, two SPIFFE token-exchange demos, vscode-remote-sandbox.md, aws-s3-sts.md. The “Run Your First Agent” path in the docs drives OpenCode against a free OpenRouter model, so you can evaluate the whole thing without spending anything on inference.

For applications rather than CLI use, the SDKs connect to a gateway (they do not install the CLI), and you should keep SDK and gateway on the same release:

# uv add openshell
import openshell

Community reaction

The launch was loud in a way open-source infrastructure usually is not — AP and CBS both covered it, framing OpenShell as a “sealed workspace with a rule book,” and NVIDIA launched with 100+ ecosystem partners under the Linux Foundation’s Open Secure AI Alliance.

Developer reaction was more interesting than the press. On Hacker News the thread framed as “Nvidia wants to put a watchdog chip next to every AI agent” drew 223 points and 292 comments, and the dominant question was the obvious one: who watches the watchdog? A second HN thread on the formal-methods claim argued about whether policy proving helps in real systems or only inside carefully scoped policy languages — a fair critique, and one the project partly concedes by scoping the prover to policy diffs rather than agent behaviour.

The self-hosting crowd was notably more positive, because the runtime needs nothing NVIDIA sells. From r/LocalLLaMA: “Files, network and tools go behind a real sandbox policy instead of a system prompt. Sentry needs BlueField hardware so I skip that. The runtime itself does not.” That is the correct read.

The strongest adoption signal is the ecosystem that appeared within days: langchain-ai/openshell-deepagent (193 stars) runs a coding agent inside an OpenShell sandbox orchestrated by Deep Agents; TheAiSingularity/hermesclaw sandboxes Nous Research’s Hermes agent; OpenRod/OpenRod builds one-click sandboxes with egress policy and audit on top of it. LangChain shipping an integration this fast is the part I would weight most heavily.

Honest limitations

It is v0.1.x. 537 open issues, and the support policy explicitly says Experimental interfaces may change or be removed in a patch release. Stable interfaces are backward-compatible across patches, but read RFC 0014 before you build a product on it.

Sentry is not for you. The hardware enforcement layer needs a BlueField DPU. Everything in this review is the software runtime, which is the part that is genuinely free and portable.

macOS is “supported” with an asterisk. The CLI and gateway run natively on Apple Silicon, but openshell-sandbox is published only as Linux musl builds for amd64 and arm64 — so on a Mac your agents run inside the Docker Desktop VM. That works; it also means Mac performance and resource overhead are Docker Desktop’s, not OpenShell’s. Windows is WSL 2 and marked experimental.

Kubernetes has a hard prerequisite: your CNI must actually enforce NetworkPolicy. On a cluster whose CNI ignores it, you will deploy the Helm chart, see no errors, and have no network enforcement. Verify this before trusting it.

L7 inspection means TLS termination. For protocol: rest, the proxy terminates TLS to read methods and paths. That is how the method-level control works, and it is a design decision you should understand before pointing it at endpoints handling regulated data.

Telemetry is on by default. It is scoped — NVIDIA documents that it excludes sandbox names, hostnames, file paths, prompts, credentials, provider and model names, and user content — but it is on. Disable with OPENSHELL_TELEMETRY_ENABLED=false on the gateway, server.telemetryEnabled=false for Helm, or compile it out.

Policy writing is real work. Default-deny means every dependency your agent needs is a rule you discover by breaking something. Audit mode and the advisor reduce this to a tractable loop, but budget a day per non-trivial agent, not an hour.

FAQ

Does OpenShell require NVIDIA hardware or a GPU?

No. The OpenShell runtime, gateway and CLI need only Linux, macOS on Apple Silicon, or Windows with WSL 2, plus Docker, Podman or host virtualization. No GPU and no NVIDIA card is required, and GPU passthrough to sandboxes is an optional feature, not a prerequisite. The separate Sentry product from the same announcement does require a BlueField DPU — that is the piece most self-hosters skip.

How is this different from just running my agent in Docker?

Docker gives you an isolation boundary with all-or-nothing networking. OpenShell gives you a policy boundary: rules bind to a specific binary inside the sandbox, outbound traffic is default-deny, and for HTTP endpoints it enforces at the request level, so GET api.github.com can be allowed while POST to the same host is refused with a logged reason. It also keeps real credentials out of the sandbox entirely — the supervisor injects them only into requests already headed to an approved endpoint — and every allow and deny is an OCSF log record. You can approximate the isolation with Docker; you cannot approximate method-level egress control or credential brokering.

Can I use it with Claude Code, Codex or OpenCode?

Yes, and this is the intended use. The documented first-agent walkthrough runs OpenCode against a free OpenRouter model, and examples/codex-app-server and vscode-remote-sandbox.md cover other setups. Because rules target binaries, you scope a policy to /usr/local/bin/claude rather than to the whole container. Running npx skills add NVIDIA/OpenShell also teaches your coding agent to write and debug its own sandbox policies.

Is it production-ready?

For a security boundary, treat v0.1.2 as early. Stable releases ship weekly with conformance, upgrade, API-compatibility and security gates, and NVIDIA maintains the current and previous minor release (N-1) for security and critical reliability fixes — real release engineering. But 537 open issues and patch-level churn in Experimental interfaces mean you should pilot it on agents whose blast radius you already accept, run in audit mode first, export policy events to your SIEM, and not make it the only thing standing between an agent and production.

What does it cost?

Nothing. Apache-2.0, all components open source, no paid tier required for the runtime, gateway, policy prover or SDKs. The costs are your inference provider and the engineering time to author policies.

Should you use it?

Use it if you run multiple autonomous agents, agents with credentials, or agents on a shared machine — especially if your current containment is a paragraph in a system prompt. The combination of default-deny egress, binary-scoped L7 rules, credential brokering and an OCSF audit trail does not exist elsewhere in open source right now, and LangChain shipping an integration inside a week says the ecosystem agrees.

Skip it for now if you run one agent on a throwaway VM, if your cluster’s CNI does not enforce NetworkPolicy, or if you need a security boundary you can certify today rather than pilot. Check back after 1.0.

The pragmatic move: bash examples/sandbox-policy-quickstart/demo.sh takes under a minute and shows you the deny log, the policy diff and the L7 refusal end to end. That is a cheap way to find out whether policy-as-enforcement changes how you think about running agents.

Sources