Multi-Tab Orchestration in Long-Running Browser Agents
Multi-tab workflows fail at the infrastructure layer, not the model.

A browser agent that juggles several tabs at once does not fail because it clicks the wrong button or misreads a page. It fails because it forgets what it already knew by the time it needs it. A workflow that logs in on one tab, pulls a token on a second, and submits a form on a third has three separate dependencies stacked on top of each other, and if the session context built on tab one does not travel intact to tab three, the final step executes blind. The agent can still parse the form, still spot the submit button, and still produce a syntactically correct action. It just has no idea what it is supposed to be submitting, or why.
Multi-Tab Browser Agents Fail at the Workflow Level
The instinct when a browser agent stumbles on a multi-step task is to blame the model: not sharp enough at reasoning, not reliable enough at parsing a page, not careful enough in picking the right element to click. That instinct is wrong for exactly the workflows that matter most, the ones that run for minutes or hours and span more than one site. A survey on agent system and harness design makes the underlying claim directly: agentic task quality "emerges from the interaction between model capability, runtime infrastructure, task structure, and evaluation design." Scaling the model does not close the gap on agentic benchmarks, because the gap was never a model problem to begin with. It is a harness problem, and the harness is what carries state from one action to the next.
Consider the three-tab login-extract-submit sequence again. Each tab, taken alone, is a task well within the reach of any competent browser agent: authenticate, read a value, fill a form. The difficulty lives entirely in the seams between them. Tab one establishes a session that tab two depends on. Tab two extracts a token that tab three needs to complete its form. If the session context does not persist across that chain, each tab becomes an isolated, context-free task, and the agent has no way to know it is standing on three steps rather than one. The gap between what a browser agent can do on a single fresh page and what it can do across a long, multi-tab session is where real workflows actually break, and it is a gap in infrastructure, not in capability.
What "persistent state" means for a browser agent running across tabs
"Persistent state" gets used as a catch-all term, which makes it easy to treat as a vague virtue rather than a specific engineering requirement. For a browser agent, it breaks into three layers, each with its own failure mode.
The lowest layer is the browser session itself: login credentials, cookies, auth tokens. Above that sits working context: which tabs are currently open, what each one has already yielded, which actions are still pending. Above both of those is long-term memory: records of past sessions, learned preferences, credentials stored between runs that are entirely separate from each other. This third layer is what separates a long-running agent from one that only ever completes single-session tasks.
The harness survey frames this precisely by decomposing the execution harness into six coupled runtime responsibilities: observation, context, control, action, state, and verification/governance. State sits among those six as a first-class runtime concern, not as something bolted on after the agent's reasoning loop is built. An agent that handles action well, clicking the right things and reading the right fields, but handles state poorly will succeed only on shallow, single-tab tasks and fail precisely on the deeper, multi-tab ones that justify building an agent.
How supervisor-worker topologies distribute state responsibility across tabs
Once state is recognized as the real constraint, the architecture must decide who is responsible for holding it. The supervisor-worker topology has become the dominant answer, because it assigns that responsibility to one component explicitly rather than leaving it implicit across a flat pool of agents.
In a flat multi-agent design, every agent reasons about global state independently. That duplication is expensive on its own, and worse, it creates divergence: two tabs working in parallel can produce contradictory intermediate results with no mechanism to reconcile them before they propagate further. A supervisor avoids this by holding the task decomposition itself, tracking which subtasks are complete, and merging the evidence that workers return. Execution stays distributed across tabs; state stays centralized in one place that can be reasoned about.
WebSwarm illustrates this pattern at research scale. Across 22 baselines on a 13-benchmark suite, Uno-Orchestra reaches 77.0% macro pass@1, roughly 16% above the strongest workflow baseline, at roughly an order of magnitude lower per-query cost. Both examples matter here as illustrations of an architectural pattern, not as endorsements of a particular system.
The counterweight to all of this comes from OrchBench, and it cuts against a very natural instinct. OrchBench finds that good orchestration depends less on how many agents are involved than on whether information is preserved across dependent tasks: additional agents can relieve context pressure, but they also introduce diminishing returns and coordination failures of their own. Adding more workers is not a substitute for getting state management right. The practical consequence is that the supervisor is the state owner in any topology built this way, and its persistence guarantees decide whether the entire workflow can be resumed after an interruption.
Why sleep-and-resume breaks multi-tab workflows without snapshots at the right granularity
Every multi-tab workflow that runs long enough will eventually be suspended, whether by design, by resource constraints, or by simple necessity. Whether it resumes correctly depends entirely on whether the snapshot taken before suspension captured state at the right granularity: tab state, supervisor context, and inter-tab dependencies, all together, at the same moment.
If a snapshot captures only the filesystem, it misses in-memory browser session state. If a snapshot captures browser state but skips the supervisor's task graph, it loses the workflow's place in its own dependency chain. Either partial snapshot produces the same dangerous outcome: a confident but wrong restart, where the agent believes it is continuing a task when it is actually replaying steps it already completed or skipping ones it never finished. A bad snapshot does not usually produce a visible crash; it produces a plausible-looking restart that is simply incorrect, which is what makes the resume failure so hard to catch in testing.
The risk compounds with time. The longer a workflow runs, the more intermediate state it accumulates across tabs, and the more a shallow snapshot has to lose. The correct abstraction is a full process snapshot, covering memory, filesystem, and open connections, taken at a moment when the supervisor has returned to a stable, serializable state between tab operations. That is not something a developer can build reliably as an application-layer pattern. It requires the infrastructure underneath the agent to expose snapshot and restore at the level of the virtual machine or the process itself.
What the infrastructure layer must provide for multi-tab orchestration
That requirement sets the bar for the infrastructure layer directly. Reliable multi-tab orchestration needs VM-level isolation per agent, wake latency measured in under a second, full snapshot-restore across both memory and filesystem, and persistent disk that survives a sleep cycle. Serverless functions and container-based approaches, taken on their own, cannot satisfy all four properties at once.
VM-level isolation matters most when agents act on user-supplied prompts or run AI-generated code against multiple sites in sequence. Firecracker's microVM architecture, which boots to init in under 125 milliseconds with minimal memory overhead and supports high creation rates per host, was built for exactly this kind of workload: high-density, low-latency, isolated execution across many tenants at once.
Wake latency under one second is not a matter of polish. A multi-tab workflow that hits a 30-second cold start every time the agent resumes from sleep will miss any task with a real time constraint attached to it, regardless of how well the agent reasons once it is awake. Without that, every resume is a fresh session in disguise, and any auth-dependent workflow that runs across more than one sleep cycle breaks on exactly that point. Container-based isolation falls short here too, for a related reason: a shared kernel across a multi-tenant agent fleet turns a single kernel-level exploit into a cross-tenant risk, so each agent needs its own guest kernel rather than a namespace carved out of a shared one.
Design patterns for managing shared browser state across a multi-tab agent fleet
Infrastructure that can snapshot and restore correctly only solves half the problem. The agent's own code must also be structured to use that substrate correctly, keeping tab contexts isolated where isolation is needed and shared where sharing is required.
The first pattern is tab-scoped workers paired with a shared supervisor context. Each tab gets its own worker agent holding local state, and the supervisor alone holds the cross-tab dependency graph and is the only component permitted to write to shared state. This keeps workers simple and keeps the one copy of shared truth in a single, auditable place.
The second pattern is serialized checkpoints at task boundaries. The supervisor writes its task graph to persistent storage each time a subtask completes, so a resume after sleep replays from the last completed step. This is the application-layer complement to the VM-level snapshot described above: the infrastructure can restore a process, but the supervisor still needs its own record of exactly where in the task graph it had gotten to.
The third pattern is credential scoping per tab. This prevents a compromised tab from reaching credentials it has no reason to touch, and it makes audit logging per tab practical. The harness survey's inclusion of verification and governance among its six coupled runtime responsibilities applies directly here: each tab action should be verifiable and reversible before the supervisor marks its subtask complete, not after.
The fourth pattern is progressive evidence accumulation. Rather than wait for every tab to finish before it aggregates results, the supervisor should accept partial results as they arrive and revise its task graph incrementally. WebSwarm's progressive recursive delegation follows this logic: evidence returned from one branch can reveal new constraints that reshape how the rest of the search proceeds, and a supervisor that waits for full completion before updating its plan loses that signal.
The failure pattern to avoid cuts across all four: treating the browser session itself as the state store. Sessions expire, get cleared by the sites they belong to, and are not built to be serialized and restored. The agent's own state layer, held by the supervisor and backed by persistent infrastructure, has to be the authoritative record. The browser session is a view into that state, not a substitute for it.
Token and context budget management across multiple tabs
Multi-tab orchestration introduces a context problem that scales with the number of open tabs: a supervisor that loads the full state of every tab into its own context window will exhaust its budget before the workflow produces any real result.
OrchBench's finding is relevant again here. Additional agents help relieve context pressure in some configurations, but that relief is limited and comes with coordination costs attached to it. Adding workers helps most when the working state genuinely exceeds what fits in one context window. Once the state fits, every additional agent is pure coordination overhead with no corresponding benefit. The right approach is for the supervisor to hold a compact task graph rather than raw tab content, with workers summarizing what they find and returning structured evidence instead of raw DOM.
Tool loading follows the same logic. If you load every available tool definition upfront, before a single tab has even opened, you burn context you need later. The better pattern is on-demand loading: fetch a tool's full definition only at the moment a specific tab action actually requires it.
Here is where persistent disk turns from an isolation requirement into a cost-efficiency mechanism. If an agent can read a past session's structured output directly from disk rather than reconstructing it by replaying conversation history, the context window available for the current workflow effectively expands, because none of that history needs to sit in the prompt. The infrastructure decision made for reliability, namely persistent storage that survives sleep, turns out to solve a budget problem that looks, on the surface, like a prompt engineering problem, but it is instead a direct consequence of where state lives.
Security and governance requirements that multi-tab operation introduces
An agent that can read and act across several tabs at once creates a cross-site data flow risk that a single-tab agent never faces, and that risk has to be addressed in the architecture.
Two of the architectural choices already discussed function as the first lines of defense in the meantime. VM-level isolation per agent means a compromised agent cannot reach the filesystem, memory, or browser sessions that belong to another agent running alongside it. Credential scoping means credentials are never stored in a shared location reachable by multiple agents, with each agent's persistent disk holding only what its own workflow needs. Audit logging per tab action, rather than merely per agent session, is what makes a governance investigation possible after an incident occurs, since without it, cross-tab data flows stay opaque to anyone trying to reconstruct what happened.
The same VM isolation and persistent-disk design that makes multi-tab orchestration reliable in the first place is what makes it auditable and governable. Reliability and governance are two outcomes of the same infrastructure decision, made once, at the layer beneath the agent itself.
Sources
- WebSwarm: Recursive Multi-Agent Orchestration for Deep-and-Wide Web Search
- From Question Answering to Task Completion: A Survey on Agent System and Harness Design
- Uno-Orchestra: Parsimonious Agent Routing via Selective Delegation
- OrchBench: Evaluating Multi-Agent Orchestration Plans in Isolation via Deterministic Simulation


