Agent security | October 3, 2026

An AI agent watchdog must live outside the agent

NVIDIA's Open Agent Safety Platform makes a useful architectural bet: the workload being controlled should not own its monitor, credentials, policy, or kill path. That makes the boundary harder to bypass. It does not make the policy correct, the evidence complete, or an allowed action safe.

Out-of-band enforcementExact capabilitiesQuarantine and recoverySources checked Oct 3
Layered AI agent safety architecture separating workload, supervisor, policy, monitoring, and recovery

A separate watchdog changes who can cheat

A tool-using agent is not merely a model answering questions. It is a process that can read files, call APIs, hold credentials, start programs, send messages, and keep trying after a denial. If the same process also controls its policy file, network route, logs, and shutdown mechanism, the safety design is circular: the subject of enforcement owns the enforcement.

Out-of-band enforcement breaks that circle. The agent proposes or attempts an effect. A component outside its trust boundary evaluates the request, injects narrowly scoped credentials only after approval, records what happened, and can stop the workload without asking it to cooperate. This is the central idea behind NVIDIA's Open Agent Safety Platform, announced September 28. The public design combines OpenShell, an open-source runtime and policy layer, with Sentry, a reference design that uses BlueField-4 infrastructure to observe and quarantine from outside the agent host.

The architecture is interesting because it moves control lower and farther away from model instructions. It is also easy to overread. A DPU can see packets and enforce a rule; it cannot know that an authorized wire transfer is commercially foolish, that a permitted email violates a promise, or that a read-only API response contains malicious instructions. Hardware separation strengthens an enforcement point. It does not supply business intent.

The useful question is not “does the agent have a sandbox?” It is “which component can still act when the agent, host process, or policy update goes wrong?”

The launch turns a familiar security pattern into an agent reference design

NVIDIA describes five principles: verifiable policy, out-of-band enforcement, control of the path to the model, authority that scales with reasoning visibility, and shared responsibility. OpenShell is available in a public Apache-2.0 repository. Its documentation names concrete controls: Landlock-based filesystem restrictions, an unprivileged process identity with reduced capabilities, syscall filtering, a local network-policy proxy, endpoint-bound credentials, typed policy, and policy revision history.

The public repository is stronger evidence than a product adjective because engineers can inspect the trust split. The sandbox runs with the agent. A supervisor outside the workload owns admitted policy, L7 proxying, gateway relays, and credentials. The gateway coordinates lifecycle and policy distribution. Static filesystem and process constraints are fixed at creation; selected network and provider rules can change while the sandbox runs. Unknown destinations are denied, and policy proposals remain proposals until the configured review path accepts them.

Sentry extends the idea to a separate infrastructure processor. NVIDIA says it can correlate agent activity, policy decisions, tool access, and contextual records, then quarantine a violating workload quickly. Treat the speed and incident-prevention claims as vendor claims until independent tests publish workload, topology, detection rule, evasion attempt, and recovery result. The focused scan found direct launch links and current community debate, but no independent production benchmark for the complete OpenShell-plus-Sentry system.

Community objections are useful design inputs. Current Reddit threads question default telemetry, dependence on NVIDIA hardware, the difference between a sandbox and a safety standard, and whether one automated system is merely watching another. Those concerns do not negate external enforcement. They show what a pilot must measure: data leaving the boundary, portability, operator visibility, false quarantine, fail-open behavior, and the amount of policy work hidden behind a demo.

Build seven separate control planes

PlaneOwnsMust not trust
1. IntentTask, requester, purpose, consequence classModel-generated scope expansion
2. IdentityAgent instance, human sponsor, service principal, sessionA name asserted inside a prompt
3. WorkloadAgent process, tools, temporary files, local stateIts own claim that an action is safe
4. PolicyFiles, syscalls, destinations, methods, paths, budgetsUnreviewed rules proposed by the workload
5. CredentialsEndpoint-bound secrets and token lifetimeEnvironment variables readable by arbitrary child processes
6. ObservationNetwork, process, tool, policy, and effect evidenceLogs written only inside the sandbox
7. RecoveryQuarantine, snapshot, revoke, destroy, resume, notifyA cooperative agent shutdown

The planes should fail independently. If the language model is compromised, policy still enforces. If the workload crashes, the supervisor still exports evidence. If policy distribution stalls, the last admitted generation remains identifiable and new authority fails closed. If the host is suspect, the external monitor can revoke paths and isolate the machine. If the monitor loses context, it should quarantine or degrade to a safe capability set rather than silently permit.

This is also why an ordinary container is not the full architecture. A container can isolate namespaces and resources, but the security outcome depends on the runtime, kernel, mounts, capabilities, network path, credential placement, and control plane. OpenShell's documented use of overlapping file, process, syscall, network, and provider controls is the right shape. Teams should still test each layer on their own kernel, container engine, orchestration environment, and agent image.

Write capabilities that describe effects, not websites

A broad allowlist such as api.github.com is not least privilege. The same host supports reading issues, creating releases, changing repository settings, and writing secrets. Bind the executable, destination, protocol, method, path, credential profile, rate, and approval class.

version: 1
filesystem:
  read_only: ["/workspace", "/usr"]
  read_write: ["/workspace/.agent-tmp"]
process:
  user: agent
  no_new_privileges: true
network:
  default: deny
  rules:
    - name: github-issue-read
      binaries: ["/usr/bin/gh"]
      destination: "api.github.com:443"
      protocol: https
      allow:
        - "GET:/repos/acme/widgets/issues/*"
      credential_profile: github-readonly
      max_requests_per_hour: 120
    - name: test-callback
      binaries: ["/usr/bin/node"]
      destination: "ci.internal.example:443"
      allow: ["POST:/agent-test-results"]
      approval: human_if_external_effect
recovery:
  on_unknown_binary: quarantine
  on_policy_engine_unavailable: deny
  preserve_evidence: true

OpenShell's public quickstart demonstrates the important distinction: a GET to a GitHub endpoint can be allowed while a POST remains denied. Its policy documentation also describes calling-binary checks, optional L7 rules, endpoint-bound provider credentials, parser budgets, duplicate-key rejection, and fail-closed behavior across policy generations. Those details matter because policy itself is an attack surface.

Do not auto-approve a rule merely because an agent observed a denial and suggested the smallest textual edit. Compare effective authority before and after the change. Ask which new hosts, methods, paths, binaries, secrets, and transitive relays become reachable. Bind approval to a candidate hash or review token so a changed proposal requires a new decision.

Verify the boundary with attacks, not screenshots

openshell sandbox create --name safety-lab --policy policy.yaml -- codex
openshell policy get safety-lab --full
openshell logs safety-lab --tail --source sandbox --level warn

# Expected: allowed read
curl -sS https://api.github.com/repos/acme/widgets/issues/42

# Expected: denied write, no credential exposure
curl -sS -X POST https://api.github.com/repos/acme/widgets/issues \
  -d '{"title":"boundary test"}'

# After a simulated violation
openshell sandbox get safety-lab
openshell policy list safety-lab

A real acceptance suite should attempt DNS rebinding, alternate IPs, redirects, IPv6, proxy environment variables, a second binary, shell wrappers, dynamic loaders, child processes, localhost relays, path traversal, symlinks, hard links, archive extraction, inherited file descriptors, credential scraping, log deletion, policy rollback, replayed execution IDs, and resource exhaustion. Repeat after every runtime, kernel, image, network, or policy-engine change.

TestPass conditionEvidence
Policy tamperWorkload cannot change admitted policy or hide revisionRejected syscall plus external revision record
Credential theftSecret is not readable; only approved endpoint receives itRedacted process/file scan and proxy receipt
Write escalationAllowed GET cannot become POST through redirect or alternate clientL7 deny event with binary identity
Monitor lossNew authority fails closed; operator is alertedFault-injection timeline
QuarantineEffect path stops while evidence remains availableExternal timestamps and network capture
RecoveryResume requires a new trusted state, not agent assertionRevocation, rebuild, and approval receipts

Know what external enforcement cannot decide

An external monitor sees events through a schema. It may know that process X called endpoint Y with method Z. It may not know that the payload changed a customer's legal status, that the email recipient was wrong, or that a purchase violated an unwritten commitment. More observation does not automatically create semantic understanding.

Keep consequential authorization in the target system where possible. A payment service should enforce account, amount, counterparty, currency, budget, separation-of-duties, and approval rules. A repository should enforce branch protection and required review. An email system should constrain sender identity, audience, and campaign limits. The agent runtime supplies an additional boundary and evidence, not a waiver for domain controls.

Out-of-band monitoring also introduces privileged infrastructure. Protect its updates, administrators, keys, logs, time source, and network path. Decide which telemetry is necessary and where it goes. A security product that exports sensitive prompts, filenames, payloads, or customer identifiers without a clear purpose can create a new incident while trying to prevent another.

Capacity is another boundary. A watchdog that drops events during a burst or delays a decision until after the effect is complete is an observability layer, not an enforcement layer. Define the maximum concurrent agents, events per second, policy-evaluation latency, evidence queue, and storage pressure the system must survive. Test overload deliberately. Critical deny and quarantine paths need reserved capacity and a behavior that is unambiguous when buffers fill.

Finally, decide what portability means. If the software policy runs on ordinary infrastructure but the strongest out-of-band controls require a specific DPU, document the capability difference instead of presenting one security tier. Teams should be able to say which threats the portable OpenShell configuration addresses, which require Sentry or equivalent separation, and which remain outside both.

Failure modes to design before production

FailureWhy it survives a demoControl
Policy is too broadHappy-path task succeedsDiff effective capabilities and seed write attempts
Allowed action is harmfulNetwork and syscall rules passTarget-system authorization and human consequence gate
Monitor shares the host failureNo fault is injectedSeparate failure domain and tested degraded mode
Quarantine destroys evidenceStopping looks successfulExternal logs, clocks, snapshots, and chain of custody
Policy advisor expands authorityEvery denial looks like frictionHuman review of capability delta and expiry
Credentials leak before proxyingApproved request worksEndpoint-bound injection and hostile child-process test
Telemetry becomes surveillanceMore logs look saferData minimization, retention, access control, and disclosure
Recovery restores compromiseResume is faster than rebuildKnown-good image, token rotation, replay-safe state import

Run a 30-day evidence-first pilot

Week 1: choose one low-consequence agent task and write its effect inventory. Name the human sponsor, allowed repositories, files, APIs, methods, budgets, credentials, external writes, retention, and stop conditions. Capture the current uncontrolled baseline.

Week 2: deploy the smallest policy. Default-deny egress, remove ambient credentials, fix filesystem and process rules at creation, and configure external evidence. Run benign tasks until denials identify missing legitimate capabilities. Review every policy delta.

Week 3: attack the boundary. Use the test matrix above and add workload-specific prompt injection, confused-deputy, data exfiltration, and target-system abuse cases. Inject monitor, network, gateway, and clock faults. Measure whether the system denied, alerted, quarantined, preserved evidence, and recovered.

Week 4: operate in shadow or bounded production. Track accepted tasks, denied legitimate actions, dangerous actions blocked, policy-change volume, quarantine precision, recovery time, evidence completeness, operator load, and cost. Expand only when a new task class has its own contract and tests.

Require a signed pilot report that includes negative results. If the monitor missed an alternate route, a policy update widened more authority than intended, or recovery required manual reconstruction, preserve that result and close it before scale. A security architecture improves through falsification; hiding uncomfortable tests turns a reference design into theater.

Production rule: never let “the sandbox allowed it” become the definition of a correct action. Runtime policy answers whether an effect is permitted at this boundary; accountable humans and domain systems decide whether it should occur.

FAQ

What is out-of-band AI agent enforcement?

It places observation and enforcement outside the workload so the agent cannot directly modify the policy engine, credential broker, evidence store, or quarantine path. Separation reduces circular trust; it does not guarantee good policy.

Does a BlueField DPU make an agent safe?

No. It can provide an independent network and infrastructure control point. It cannot infer all business consequences, replace target-system rules, or prove that an allowed action is appropriate.

Is OpenShell only for NVIDIA models?

The public documentation lists multiple agents and providers. The runtime governs process, file, network, and credential behavior rather than depending on one model family.

Should denied actions automatically expand policy?

No. A denial is evidence of attempted behavior, not evidence of legitimate need. Review the proposed capability delta, expiry, credential reach, and alternatives before admitting a new policy generation.

What is the smallest useful pilot?

One low-consequence task, default-deny networking, no ambient secrets, a precise read-only API rule, external logs, seeded write and exfiltration attacks, tested quarantine, and a clean recovery run.

Sources and further reading

Product, repository, standards, and community sources were accessed and verified on October 3, 2026. NVIDIA performance, partner, and incident-prevention statements are identified as vendor claims; no independent full-platform production benchmark was found in the research window.

Related guides: agent permission kernels, Docker agent sandboxes, execution receipts, and approval-fatigue controls.