Cua explained: computers, drivers and benchmarks for agents
Evaluate a computer-use agent with Cua Bench
Define task success from the environment, not from the agent’s own report.
What you will learn
- Create a verifiable task
- Separate agent and environment cost
- Avoid borrowed benchmark claims
Before you start
- A disposable target and permitted task
- Basic understanding of host versus guest state
Capture observation, action, policy and final state for one bounded desktop task.
Key takeaways
- Task success should come from observed state.
- Warm and cold environments have different costs.
- A reference task does not establish general agent performance.
Create a verifiable task
Cua Bench supports simulated tasks as well as heavier app environments. Its README points to a first task that can run without a VM, Docker or model API key; the reference solution and evaluator are the starting checks.
Write an initial state, allowed actions and a final-state predicate. An agent saying it completed a form is not the same as the evaluator finding the intended values in the application.
Separate agent and environment cost
Record task setup, execution time, model tokens, sandbox startup and any Fleet capacity cost independently. Reusing a warm pool changes latency and price, so do not compare it directly with a cold local launch.
Capture trajectories and failed cases, including timeouts and mis-targeted clicks. A reward of 1.0 on a reference solution validates that task’s evaluator path, not a model’s general ability.
Avoid borrowed benchmark claims
The repository contains model research and benchmark components, but this series did not run them. A published score needs its dataset, agent revision, environment and evaluation procedure.
For adoption, test tasks matching your actual applications and permission boundaries. Keep missing measurements blank instead of filling them with a project chart.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Build one simulated task and run its reference solution.
- 2
Evaluate an agent on the same final-state predicate.
- 3
Report failures, latency and cost by component.
Copy-ready example
task,agent_revision,environment,reference_reward,agent_reward,startup_ms,model_cost,verdict
calculator,,,,,,,not-runFrequently asked questions
Can I start without a model API key?
The README describes a simulated first Bench task that does not need one.
Was a Cua benchmark reproduced here?
No. Only the evaluation design and source were reviewed.
Sources
- Cua / README.mdSource checked 2026-09-29
- Cua / docs/content/docs/concepts/what-is-cua-bench.mdxSource checked 2026-09-29
- Cua / docs/content/docs/concepts/how-sandboxes-work.mdxSource checked 2026-09-29