NVIDIA has launched its Open Agent Safety Platform: OpenShell, an open-source secure runtime that gives agents boundaries and enforces policy as they work, and Sentry, a hardware system that watches an agent at runtime and can quarantine it in milliseconds when it steps outside them. One idea underneath it is worth pulling out on its own.
Prompts, model safeguards and agent frameworks shape what an agent tries to do. Runtime controls decide what it is allowed to do. Those are different things, and only the second one still holds when an agent is confused, manipulated, or simply wrong. We have built around that since the first commit. It is good to see a company with NVIDIA's reach put it in silicon.
So here is where we actually stand against that idea — including the part we got wrong in an earlier draft of this post, which we only caught by going and reading our own source.
The agent should never hold the key
Most agent setups still work one way: paste an API key into an environment variable, hope the model does not leak it, move on. That works until it does not, and when it fails there is no taking it back — the key is in the context window, the transcript, the logs and the agent's memory.
With 1Claw, credentials sit in an HSM-backed vault and the agent gets scoped, audited, revocable access. A step further and the agent does not see the credential at all. Execution Intents let it call an HTTP or GraphQL API through a binding, and the credential is injected server-side. Secrets in prompts become typed placeholders — a tag naming the vault path plus a checksum — and the real value is rehydrated inside the enclave, in a tool call you authorised, to a host on your allowlist. If a model is talked into printing its secrets, there is nothing there to print.
Policy lives outside the agent
An agent cannot be trusted to police itself, so we do not ask it to. Access is decided by a policy engine — Cedar, or OPA Rego — running outside the agent's process. Deny rules win. Time windows, approver roles and spending thresholds are all expressed in the same place.
We also tried to make loosening things by accident hard. Widening a guardrail goes through an approval flow. A preset that would relax a limit tells you exactly what changes before anything is applied. And a new rule can run in shadow mode against real traffic first, so you can see what it would have blocked before it can block anything.
Humans stay in the loop without becoming the bottleneck
Nobody wants to approve every action and nobody should approve none, so approvals are tiered. An agent asks a person about a business action in plain language — refund this customer, post this update — and the server derives the risk tier itself, from the payload and your policy.
That derivation is the load-bearing part. An agent may declare a tier, but the server takes the higher of its own floor and the agent's claim, and tells your operator when the two disagreed. An agent cannot decide that a $500 refund is low risk. The lowest tier can be answered by text message; anything above it needs stronger proof that the right person said yes.
What you can see
The dashboard opens on a live map of your agents, the policies that govern them, the vaults they touch and the systems they call. Threats are ranked by blast radius and each agent carries a behavioural trust score. Every one of those endpoints requires a human principal, so an agent key cannot read any of it — a compromised agent cannot study the organisation it is inside.
The trust score is advice. Nothing in the enforcement path reads it. Which brings us to the correction.
Where we are narrower than we thought
The draft of this post said that nothing suspends an agent automatically, and that we had chosen it that way deliberately. Checking before publishing turned up the opposite: there is a circuit breaker. Ten guardrail denials inside ten minutes and the agent is auto-suspended, with an agent.suspended webhook fired at your org.
The honest version is more interesting than either claim. That breaker is wired into exactly one denial path — an execution binding rejecting a disallowed HTTP method. A host that is not on the allowlist, a path that is not permitted, a spend cap hit over and over: none of those currently feed it. So the mechanism NVIDIA is putting in hardware exists here in software, and it is watching one door out of several.
Sentry quarantines in milliseconds at the hardware boundary. We suspend after ten denials in ten minutes, on one class of denial, in the application. That is the gap, stated plainly, and closing it is now near the top of the list.
Where we are headed
The NVIDIA announcement helped us sort our own thinking. Roughly in order:
- Every denial feeds the breaker. One counter, every guardrail, with the threshold and window configurable per org rather than the current fixed ten-in-ten.
- Graduated response, opt-in. Alert, then require approval for everything, then cut access — shipped in shadow mode first, the same way policy changes are, so you can watch it before it can hurt anything.
- One policy language for the sandbox too. Secrets, actions, and what a runtime can reach on the network or disk should be one set of rules, not three.
- Telemetry you can verify. Logs are only worth your trust in whoever wrote them. Agent activity records tied to a verifiable workload identity, exportable in standard formats.
- Working alongside open runtimes. We want 1Claw to sit under something like OpenShell so credentials and approvals behave the same wherever an agent runs.
- One safety page. What we guarantee against what we merely recommend, with the threat model written out, instead of the current spread across a dozen docs pages.
We will be careful with the word verified. If we say a property is proven, it will be — and the enforcement path is the part we most want someone outside to audit and publish in full, including whatever they find.
Try it
One command:
npx @1claw/cli setupThen ask your assistant to list its secrets. You should see examples/hello, and you are away. To start from the agent's side instead, 1claw agent enroll my-agent --pair prints a fingerprint for a human to confirm, so no key travels through email.
Agents are going to do more work with more access. We would rather the walls were there from the start — and rather say which ones are still missing than imply they all exist. If you want to see something moved up the list, tell us; a lot of this roadmap came straight from people's messages.