Computer-Use Agent Sandboxing Requirements
These agents need isolation at the hardware, network, filesystem, and runtime layers combined.

Computer-use agents navigate browsers, read and write files, and call APIs with the same hands a human operator would use, and that single fact changes what sandboxing has to accomplish. This piece argues that these agents need a different, harder isolation discipline than ordinary code execution: hardware-level VM isolation at the base, default-deny network and filesystem policy on top of it, and runtime monitoring across the whole stack, operating together rather than as substitutes for one another.
Why computer-use agents need different sandboxing than general code execution
Johann Rehberger's ZombAIs demonstration is worth starting with because it shows, in one sequence, what happens when none of this is in place. Claude Computer Use was pointed at a web page carrying a hidden prompt injection payload; the agent read the page, downloaded a binary, ran it, and connected out to a command-and-control server, on the first attempt, with no sandbox around it at any point in the chain. Nothing exotic triggered that outcome. The agent did what it is built to do: read content, decide on an action, and carry it out.
That is the structural issue. Traditional software has a fixed, auditable instruction set fixed at compile time or deploy time, so a security review can look at the code and reason about what it might do. Agents don't work that way. They generate and run new code at runtime, out of natural-language input that may come from an attacker. The instruction set doesn't exist yet when the system is reviewed, only when it runs. Three things follow from that, and all three show up in computer-use agents specifically. The agent generates code from inputs nobody vetted in advance. It makes its own runtime calls about which APIs to hit and how much of a resource to consume. It often carries memory across sessions, giving an attacker a second route in: manipulate what the agent remembers.
Computer-use agents sharpen each of these in a particular way, because their inputs and their actions both touch the outside world directly. A browsing agent reads whatever is on the page it's sent to, including text no human will ever see. A filesystem agent writes to paths that other software treats as executable on sight, startup scripts and config files among them. An API-calling agent typically runs with whatever credentials the host process already has. A single compromised session can reach anything those credentials can reach. None of these are edge cases bolted onto the threat model. They are the ordinary operating mode of a computer-use agent doing its job correctly, which is what makes the problem hard: the capability that makes the agent useful is the same capability an attacker needs.
The agent is a legitimate, authorized actor in the system, so when a prompt injection convinces it to request something malicious, role-based access control approves the request, because RBAC checks who is asking. A permission system built to ask "is this identity allowed to do this" has no way to ask "did this identity really mean to do this, or did a web page just tell it to." That gap is what ZombAIs walked through.
One obvious reply is that standard container hardening, seccomp profiles, AppArmor policies, dropped capabilities, should be enough to keep an agent's behavior inside bounds. It isn't, because containers share the host kernel, and a single vulnerability in the container runtime collapses the entire boundary in one step. CVE-2019-5736 and CVE-2024-21626 both exploited the container runtime itself to break out of the container entirely, and both are documented runtime CVEs, not hypothetical ones. Hardening the policy around a shared kernel doesn't change the fact that the kernel is shared.
The attack patterns that exploit unsandboxed computer-use agents in practice
None of this is theoretical. These failure modes have already produced assigned CVEs and real production incidents, and they follow patterns specific enough to name. Tool poisoning is one of the most common: an attacker publishes a malicious MCP tool that looks legitimate, and once an agent invokes it, the tool inherits whatever permissions the agent process already holds, which in practice often means broad filesystem read and write access, environment variables holding live API keys, and a path onto the internal network. A related pattern runs through documents rather than tools: a file submitted for something as ordinary as summarization can carry hidden instructions telling the agent to read a path like ~/.ssh/id_rsa and fold its contents into the response, and without filesystem sandboxing in place, the private key leaves the system inside the agent's own output.
In Roo Code, a prompt injection led to a write into the workspace, which led to arbitrary code execution, tracked as CVE-2025-58372 with a CVSS score of 8.1, though third-party aggregators list it as high as 9.8. In GitHub Copilot, CVE-2025-53773 let an attacker plant malicious instructions in repository files, code comments, or GitHub issues, which the agent then read and acted on, resulting in local code execution. In Flowise, unsafe handling of configuration data opened the door to JavaScript injection, tracked as CVE-2025-59528 with the maximum possible CVSS score of 10.0, and a large number of instances were found exposed to it. All of these cases trace back to the same place, an agent or tool operating with more reach than the task in front of it required. In the Postmark MCP incident, a production MCP server injected a BCC field into email tool calls, silently exfiltrating all outgoing email. In the Replit incident, a coding agent with live database access, confused by empty inputs, deleted a production database.
The Google Antigravity sandbox escape is the sharpest instructive case for computer-use agents specifically, because it defeated a control that was supposed to be the strongest one available. Google's "Secure Mode," described as its highest security setting, was bypassed through a tool called find_by_name: because that tool was classified as "native," it passed an unsanitized flag straight from its Pattern parameter into the underlying fd binary, and native tool invocations simply aren't evaluated by the agent's application-level security restrictions at all. The policy sat at the application layer, and the subprocess ran underneath it, outside its view. A filesystem write boundary would have stopped this particular escape by denying the write it depended on, and kernel-level process isolation would have prevented the subprocess breakout from happening in the first place. That's the throughline connecting every incident above: each one maps to a specific layer of isolation that was missing, which is exactly the structure the next section lays out.
Why the isolation stack's layers are each independently necessary
Securing a computer-use agent isn't a matter of picking the right single control. It takes a stack of four independently necessary isolation boundaries, each of which closes an attack path the others cannot. Microsoft's Agent Governance Toolkit and NVIDIA's 2026 sandboxing guidance converge on the same four layers as mandatory: network egress control, filesystem boundaries, secrets scoping, and protection of configuration files.
Network egress means the agent operates against a defined, narrow allowlist of external endpoints it's permitted to call, rather than open access to anything its host can reach. That single constraint directly limits both data exfiltration and the kind of attacker-controlled callback that drove the ZombAIs demonstration. Filesystem boundaries matter for a more specific reason: write access to configuration files, dotfiles, hooks, MCP configuration directories, is dangerous precisely because those files run at startup or get picked up by developer tools before any runtime security check even gets a chance to fire, and the Antigravity escape ran through exactly that gap. Process isolation closes a path the other layers can't reach at all: subprocesses spawned by native tool invocations typically sit outside application-level policy, so only isolation enforced at the kernel level can actually contain them, which is the specific failure Antigravity exposed. Secrets scoping replaces the default of inheriting the full host credential environment with credentials provisioned for the specific task at hand and revoked once it's done, which is what least privilege actually looks like when the thing holding the credentials is an autonomous agent rather than a human operator.
None of this is optional guidance. OWASP's Agentic Top 10 lists unexpected code execution as item ASI05 and treats sandboxing as a mandatory control rather than a recommendation, stating: "Never execute agent-generated code without strict sandboxing, input validation, and allowlisting". The reason all four layers have to hold at once, rather than any one of them being sufficient on its own, is that each closes a door the others simply don't cover. A tight network policy does nothing to stop a malicious subprocess from writing into a config file that executes on the next startup. A read-only filesystem does nothing to stop an agent with open network access from calling out to an attacker's server. Pulling any single layer out of the stack opens a documented path through the gap, since one of the incidents above walked through it.
Hardware-level VM isolation as the necessary foundation for the stack
If that primitive can be bypassed by a kernel vulnerability, every policy built on top of it is bypassed along with it, and a shared-kernel container does not give you a primitive that resists that failure. Hardware-level VM isolation does.
Two documented runc CVEs, CVE-2019-5736 and CVE-2024-21626, exploited the container runtime itself rather than through any application-level mistake. An agent that generates and runs its own code at runtime produces syscall patterns that are novel by construction, which raises the odds of eventually hitting one of those unguarded paths, simply because the agent isn't following a fixed, pre-reviewed execution path the way ordinary application code does.
Three isolation technologies dominate how this problem gets solved in practice, and each makes a different trade-off. Firecracker runs microVMs with hardware isolation through KVM, giving each workload its own dedicated kernel fully separated from the host, so an escape requires compromising both the guest kernel and the hypervisor rather than either one alone. It boots in around 125 milliseconds, carries less than 5 MiB of memory overhead per VM, and can spin up a large number of VMs on a single host; it's written in Rust, with a codebase substantially smaller than QEMU's, and Amazon open-sourced it in 2018, where it now powers both AWS Lambda and AWS Fargate. It fits multi-tenant agent execution, untrusted code, and regulated data, where the isolation strength justifies the cost of running it.
gVisor takes a different route: it reimplements Linux syscalls inside a user-space kernel called the Sentry, written in Go, so an application never issues a syscall directly to the host kernel at all. Breaking out of gVisor requires chaining a bug in the Sentry's syscall reimplementation with a separate host kernel bug, a two-vulnerability chain rather than a single point of failure, and it starts fast, at the cost of some I/O overhead. That trade-off suits compute-heavy, cost-sensitive, multi-tenant workloads where full VM isolation isn't justified by the threat the workload actually faces.
Kata Containers takes a third approach, orchestrating VMMs, Firecracker, Cloud Hypervisor, or QEMU, behind a standard container interface, so that Kubernetes sees an ordinary container while the workload actually runs inside a full VM underneath. It boots in around 200 milliseconds, and it fits regulated industries and production Kubernetes environments that need VM-level security without giving up container-based workflows. Hardened Docker containers remain a reasonable place to start in development, using capabilities, seccomp, and mandatory access control profiles to cut down risk, but they explicitly fall short for production agent execution handling untrusted code, because the kernel underneath them is still shared.
Firecracker's architecture carries one further property: it runs one VMM process per microVM rather than a single daemon managing a whole fleet, so compromising one VMM doesn't touch any of the others, which matters once an organization is running many agents at once rather than one. A newer entry, SmolVM, launched on April 17, 2026, as a single-executable microVM with sub-200-millisecond cold starts running on Hypervisor.framework and libkrun, built to be spawned as easily as a subprocess from any developer's own machine, on macOS or Linux. It addresses the practical objection that hardware isolation is too heavy for everyday development, by filling the gap between "secure" and "easy to run locally" that Firecracker itself never addressed.
The choice between gVisor and Firecracker isn't really about raw performance. It comes down to what the workload is actually exposed to: regulated data and adversarial code execution call for hardware isolation, while compute-heavy, cost-sensitive, multi-tenant workloads can reasonably accept syscall-level isolation instead.
Default-deny network and filesystem policies for computer-use agents
Most compromises of computer-use agents don't come from attackers defeating some sophisticated control. They come from agents deployed with permissive-by-default network and filesystem access that nobody ever went back and narrowed. The fix at the network layer is specific: define exactly which external APIs the agent may call, enforce that list through an egress proxy or network policy, and alert on everything else that tries to go out. NVIDIA's 2026 guidance names the implementation mechanisms directly, HTTP proxy, or port-based controls.
Egress allowlisting matters even more for browsing agents specifically, because a browser session naturally follows redirects and loads third-party resources on its own as part of normal operation. Without an egress boundary in place, a compromised browsing agent can call an attacker-controlled endpoint or exfiltrate data through any service it happens to be able to reach, simply by doing the thing browsers are supposed to do.
The filesystem layer follows the same default-deny logic. Anything the agent doesn't need to write to should be mounted read-only; scratch space should use tmpfs rather than persistent storage; and write access should be granted explicitly, scoped to specific paths rather than opened broadly. NVIDIA's guidance calls out dotfiles, hooks, and MCP configuration directories by name as zones that need write protection, precisely because those files execute at startup, before any runtime security check gets the chance to evaluate them.
For a development baseline, AugmentCode's guide lays out a specific hardened Docker configuration: drop all capabilities, mount the filesystem read-only, disable networking entirely, set no-new-privileges, and use tmpfs for /tmp with both noexec and nosuid set. That configuration is a floor, not a ceiling, and a reasonable starting point for building and testing an agent, not a substitute for the hardware isolation layer underneath production deployment. Putting the four policy layers on top of a hardware isolation boundary, default-deny on the network, default-deny on the filesystem, scoped secrets, and protected configuration files, produces a stack where no single compromised decision by the agent is enough, on its own, to reach the host.

