Cua explained: computers, drivers and benchmarks for agents
Build a Cua task receipt that proves what the agent changed
Capture observation, action, policy and final state for one bounded desktop task.
What you will learn
- Define the task contract
- Capture the action sequence
- Finish with cleanup and limits
Before you start
- A disposable target and permitted task
- Basic understanding of host versus guest state
Capture observation, action, policy and final state for one bounded desktop task.
Key takeaways
- An agent narrative is not a state check.
- Policy and target belong in the receipt.
- Cleanup is part of task completion.
Define the task contract
Choose a disposable app and describe initial state, allowed operations and a final condition. A Calculator example is intentionally small; a later form task should include field-level checks rather than a vague “done” message.
Select host Driver or a guest Sandbox explicitly. Record OS, image or runtime version, policy snapshot and which credentials are present.
Capture the action sequence
Before each action save the observation used by the agent, then the tool name, arguments and response. If a screenshot is involved, store it with privacy controls. The resulting trajectory should distinguish observation from assumption.
After the final action, inspect the app itself or use a task evaluator. Keep any failed click and retry in order so another reader can understand how the state changed.
Finish with cleanup and limits
Export required artifacts before deleting a guest. For Fleet, verify both claim release and pool capacity removal if the trial is over. Record anything you could not verify, including external services contacted.
This receipt format is a reader exercise proposed here; Cua does not promise this exact report file. We did not run the task for the article.
Decision guide
| Criterion | Option A | Option B |
|---|---|---|
| Best when | You need predictable behavior and easy auditing | You need adaptive optimization and have reliable telemetry |
| Main risk | May leave performance on the table | Can become difficult to explain or debug |
Implementation steps
- 1
Write initial state, allowed actions and success predicate.
- 2
Log observation and tool result for each step.
- 3
Verify final state and resource cleanup.
Copy-ready example
{"task":"calculator-6x7","target":"disposable-guest","policy":"record-hash","observed_display":null,"expected_display":"42","cleanup":"pending"}Frequently asked questions
Does Cua automatically create this exact receipt?
No. It is a proposed review artifact using documented components.
What if the agent reaches the right answer without using the app?
Mark the task incomplete if app-state interaction is part of the requirement.
Sources
- Cua / README.mdSource checked 2026-09-29
- Cua / docs/content/docs/concepts/how-sandboxes-work.mdxSource checked 2026-09-29
- Cua / docs/content/docs/concepts/what-is-cua-bench.mdxSource checked 2026-09-29