Cua explained: computers, drivers and benchmarks for agents
Cua explained: computers, drivers and benchmarks for agents
Choose the component that matches your computer-use problem before installing a fleet of tools.
What you will learn
- Separate the projects in the monorepo
- Start with a bounded task
- Keep claims scoped
Before you start
- A disposable target and permitted task
- Basic understanding of host versus guest state
Capture observation, action, policy and final state for one bounded desktop task.
Key takeaways
- Cua is several related components, not one installer.
- Driver on a host and an isolated Sandbox have different boundaries.
- A visible action still needs outcome verification.
Separate the projects in the monorepo
Cua Driver lets an agent inspect and operate desktop apps. Sandbox provides an isolated machine with code and GUI access; Fleets provisions hosted capacity. Lume handles local Apple Silicon VMs, CUA-S1 contains specialist model research and Cua Bench defines evaluation tasks.
These components share a computer-use theme but do not automatically configure one another. A Fleet claim is not a Driver installation, and a Driver attached to your personal desktop does not secretly move its actions into a sandbox.
Start with a bounded task
The README suggests using Calculator to compute 6 × 7 and verify 42, or building a simulated benchmark task. For a first trial, choose one app you own and a disposable account or guest. Save the observation and the resulting state, not only the agent’s verbal report.
An agent can use code, APIs and graphical controls in one workflow. The value of a GUI action depends on the target environment and permissions; clicking the right pixel is not evidence that the final application state is correct.
Keep claims scoped
This series reviews a fixed source revision and selected docs. It did not launch a Fleet, execute Driver on a desktop, or reproduce benchmark results. Hosting costs and platform support should be checked for the chosen version before a pilot.
The following chapters cover a safe first test, environment choices, runtime boundaries, source entry points and evaluation. Treat each component’s documentation as its own contract.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Pick Driver, Sandbox or Bench for one specific goal.
- 2
Use a disposable machine and a tiny verification task.
- 3
Record the final application state and component version.
Copy-ready example
agent -> choose Driver or Sandbox SDK
Driver -> targeted desktop
Sandbox -> isolated guest with code + GUI
Bench -> task + evaluator + trajectoryFrequently asked questions
Does a Fleet include a ready Driver?
The sandbox docs say Driver needs its own compatible installation and integration in the guest.
Is CUA-S1 required?
No. The README invites users to bring their own agent and model.
Sources
- Cua / README.mdSource checked 2026-09-29
- Cua / docs/content/docs/concepts/how-sandboxes-work.mdxSource checked 2026-09-29