Cua explained: computers, drivers and benchmarks for agents
Cua architecture: code and GUI on one isolated machine
Map the shared guest state and the separate authorization boundary.
What you will learn
- Two ways to use the same guest
- Do not conflate Driver and Sandbox
- Policy sits before dispatch
Before you start
- A disposable target and permitted task
- Basic understanding of host versus guest state
Capture observation, action, policy and final state for one bounded desktop task.
Key takeaways
- Code and GUI operate on shared guest state.
- Host Driver and guest Sandbox are distinct.
- Runtime policy is a technical enforcement point.
Two ways to use the same guest
The sandbox docs describe shell, PTY and Python execution alongside screenshots, accessibility data, clicks and typing. Code and GUI actions share one filesystem, process set and OS state inside a sandbox.
A shell can create a file that the GUI opens, and a GUI download can be inspected in Python. That makes a workflow testable, but also means a wrong action can affect the same guest state used by later steps.
Do not conflate Driver and Sandbox
Driver controls the desktop where its compatible runtime is installed and targeted. Sandbox provides an isolated computer. Putting Driver on the host does not implicitly redirect actions to a guest; installing it inside a guest requires a validated connection.
Local execution and Fleet both expose a Sandbox SDK concept, but provisioning and connection differ. The Fleet’s default guest service is not interchangeable with every Driver command.
Policy sits before dispatch
The permission-policy docs put authorization between callers and native tool implementations. Managed and user policies form ceilings, while modes set default autonomy; a denied call never reaches the tool implementation.
That boundary matters more than a polite prompt telling an agent to avoid a button. Use explicit capabilities, then inspect policy status and the target state after each sensitive action.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Trace one file from shell creation to GUI opening.
- 2
Identify the Driver target and Sandbox guest separately.
- 3
Place policy checks before native action dispatch.
Copy-ready example
agent -> Sandbox SDK -> guest shell + guest GUI
agent -> Driver -> targeted desktop
public tool call -> policy coordinator -> native dispatchFrequently asked questions
Can a sandbox action see files created by its own shell?
Yes. The docs describe shared filesystem and OS state inside the guest.
Does a prompt alone restrict native tools?
Use the runtime permission policy and host controls for enforceable limits.
Sources
- Cua / README.mdSource checked 2026-09-29
- Cua / docs/content/docs/concepts/how-sandboxes-work.mdxSource checked 2026-09-29
- Cua / docs/content/docs/concepts/how-permission-policies-work.mdxSource checked 2026-09-29