Cua explained: computers, drivers and benchmarks for agents
Choose Cua components for your computer-use workload
Use the narrowest component that can produce a verifiable result.
What you will learn
- Map requirements to components
- Compare operational trade-offs
- Pilot with a stop rule
Before you start
- A disposable target and permitted task
- Basic understanding of host versus guest state
Capture observation, action, policy and final state for one bounded desktop task.
Key takeaways
- Driver, Sandbox and Bench answer different questions.
- Provisioning and isolation change cost and risk.
- A pilot needs a clear stop condition.
Map requirements to components
Use Driver when an agent must operate an existing compatible desktop. Use Sandbox when it needs an isolated computer with code and GUI state. Use Bench when the main task is creating and scoring repeatable evaluations.
Fleet adds hosted capacity and lifecycle management, while Lume addresses local Apple Silicon virtual machines. Specialist CUA-S1 research is another choice; its weights and datasets require separate license and scope checks.
Compare operational trade-offs
A host Driver can avoid provisioning a guest but increases exposure to host data. A local guest uses your hardware; Fleet uses remote capacity, credentials and pool cleanup. A full VM may offer OS fidelity at slower startup than a container.
Write down OS requirements, data sensitivity, budget and a testable outcome. Do not pick by a product diagram that omits the task’s actual application and permission boundaries.
Pilot with a stop rule
Approve wider use only after one small task succeeds on the intended target, a forbidden action is denied and cleanup works. If any of these are unmeasured, keep the pilot narrow.
No matched comparison of Cua against another computer-use stack was run here. The decision follows local evidence, not the number of packages in the monorepo.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
List target OS, data boundary and desired outcome.
- 2
Choose one component and one task.
- 3
Require success, denial and cleanup evidence.
Copy-ready example
existing desktop -> Driver
isolated code + GUI -> Sandbox
repeatable score -> Bench
hosted capacity -> Fleet plus cleanup planFrequently asked questions
Should every agent use Fleet?
No. Hosted capacity is relevant when a local or existing target does not meet the workload.
Is local Docker equivalent to a full VM?
No. The sandbox docs distinguish shared-kernel Linux containers from full guest kernels.
Sources
- Cua / README.mdSource checked 2026-09-29
- Cua / docs/content/docs/concepts/how-sandboxes-work.mdxSource checked 2026-09-29
- Cua / docs/content/docs/concepts/how-permission-policies-work.mdxSource checked 2026-09-29